# Spice.ai OSS > A portable SQL query and AI compute engine, written in Rust, for data-grounded apps and agents. - [Spice.ai OSS](/) ## cookbook A collection of guides and samples to help you build data-grounded AI apps and agents with Spice.ai Open-Source. Find ready-to-use examples for data acceleration, AI agents, LLM memory, and more. - [🧑‍🍳 Spice.ai OSS Cookbook](/cookbook): A collection of guides and samples to help you build data-grounded AI apps and agents with Spice.ai Open-Source. Find ready-to-use examples for data acceleration, AI agents, LLM memory, and more. ## docs Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven applications and AI agents. Learn about data federation, acceleration, RAG, and building intelligent apps. - [Spice.ai Open Source](/docs): Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven applications and AI agents. Learn about data federation, acceleration, RAG, and building intelligent apps. ### next Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven applications and AI agents. Learn about data federation, acceleration, RAG, and building intelligent apps. - [Spice.ai Open Source](/docs/next): Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven applications and AI agents. Learn about data federation, acceleration, RAG, and building intelligent apps. - [Tags](/docs/next/tags) - [One doc tagged with "Acknowledgements"](/docs/next/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [3 docs tagged with "ADBC"](/docs/next/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/next/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Argo CD"](/docs/next/tags/argocd): Argo CD declarative GitOps continuous delivery for Kubernetes. - [2 docs tagged with "Arrow"](/docs/next/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/next/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/next/tags/auth): Authentication and authorization mechanisms. - [2 docs tagged with "Authentication"](/docs/next/tags/authentication): User authentication methods and security protocols. - [5 docs tagged with "Azure"](/docs/next/tags/azure): Microsoft Azure cloud services and integrations. - [2 docs tagged with "Blob Storage"](/docs/next/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Caching"](/docs/next/tags/caching): Data caching strategies and performance optimization. - [13 docs tagged with "Catalogs"](/docs/next/tags/catalogs): Data catalog connectors and metadata management. - [3 docs tagged with "Cayenne"](/docs/next/tags/cayenne): Cayenne (Vortex) data accelerator built on Vortex columnar format. - [2 docs tagged with "CLI"](/docs/next/tags/cli): Spice command-line interface commands and usage. - [5 docs tagged with "Component Metrics"](/docs/next/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/next/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/next/tags/configuration): System configuration files and runtime settings. - [2 docs tagged with "Cosmos DB"](/docs/next/tags/cosmosdb): Azure Cosmos DB (NoSQL / Core SQL) data connector. - [One doc tagged with "C#"](/docs/next/tags/csharp): C# programming language topics. - [8 docs tagged with "Data Accelerators"](/docs/next/tags/data-accelerators): Data acceleration engines for high-performance query execution. - [48 docs tagged with "Data Connectors"](/docs/next/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/next/tags/data-lake): Data lake architectures and storage solutions. - [3 docs tagged with "Databricks"](/docs/next/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/next/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debezium"](/docs/next/tags/debezium): Debezium change data capture integration. - [One doc tagged with "Debugging"](/docs/next/tags/debugging): Debugging techniques and troubleshooting tools. - [3 docs tagged with "Delta Lake"](/docs/next/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/next/tags/dependencies): Third-party dependencies and software requirements. - [8 docs tagged with "Deployment"](/docs/next/tags/deployment): Production deployment using Docker and Kubernetes. - [2 docs tagged with "Docker"](/docs/next/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with ".NET"](/docs/next/tags/dotnet): .NET framework topics. - [One doc tagged with "Dremio"](/docs/next/tags/dremio): Dremio data lake engine integration. - [2 docs tagged with "DuckDB"](/docs/next/tags/duckdb): DuckDB embedded analytical database integration. - [2 docs tagged with "DuckLake"](/docs/next/tags/ducklake): DuckLake catalog integration. - [2 docs tagged with "DynamoDB"](/docs/next/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [2 docs tagged with "Elasticsearch"](/docs/next/tags/elasticsearch): Elasticsearch data connector and vector engine integration. - [7 docs tagged with "Embeddings"](/docs/next/tags/embeddings): Vector embeddings and semantic similarity operations. - [6 docs tagged with "Features"](/docs/next/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/next/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "File"](/docs/next/tags/file): File-based data connectors and local file access. - [One doc tagged with "Flux"](/docs/next/tags/flux): Flux CD GitOps toolkit for Kubernetes. - [One doc tagged with "Flux CD"](/docs/next/tags/fluxcd): Flux CD GitOps toolkit for Kubernetes. - [2 docs tagged with "Functions"](/docs/next/tags/functions): User-defined SQL scalar functions and remote function endpoints. - [One doc tagged with "Getting Started"](/docs/next/tags/getting-started): Installation guides and quickstart tutorials. - [3 docs tagged with "GitHub"](/docs/next/tags/github): GitHub repository data connector and integration. - [3 docs tagged with "GitOps"](/docs/next/tags/gitops): GitOps continuous delivery patterns for Kubernetes. - [2 docs tagged with "Glue"](/docs/next/tags/glue): AWS Glue ETL service and data catalog integration. - [One doc tagged with "Go"](/docs/next/tags/go): Go programming language topics. - [One doc tagged with "Golang"](/docs/next/tags/golang): Go (Golang) programming language topics. - [One doc tagged with "GraphQL"](/docs/next/tags/graphql): GraphQL API data connectors and query language support. - [3 docs tagged with "Helm"](/docs/next/tags/helm): Helm package manager for deploying Spice.ai on Kubernetes. - [One doc tagged with "HTTPS"](/docs/next/tags/https): HTTPS data connectors and secure web API access. - [2 docs tagged with "Hugging Face"](/docs/next/tags/huggingface): Hugging Face model hub and transformer model integration. - [2 docs tagged with "Iceberg"](/docs/next/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/next/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/next/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "Java"](/docs/next/tags/java): Java programming language topics. - [One doc tagged with "JavaScript"](/docs/next/tags/javascript): JavaScript programming language topics. - [One doc tagged with "JDBC"](/docs/next/tags/jdbc): Java Database Connectivity driver configuration. - [One doc tagged with "Kafka"](/docs/next/tags/kafka): Apache Kafka event streaming platform integration. - [6 docs tagged with "Kubernetes"](/docs/next/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "libSQL"](/docs/next/tags/libsql): libSQL database topics. - [One doc tagged with "Local"](/docs/next/tags/local): Local model execution and on-premise services. - [One doc tagged with "Logging"](/docs/next/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/next/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/next/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/next/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/next/tags/memory): In-memory data connectors and volatile storage. - [16 docs tagged with "Models"](/docs/next/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MongoDB"](/docs/next/tags/mongodb): MongoDB NoSQL document database integration. - [One doc tagged with "MSSQL"](/docs/next/tags/mssql): Microsoft SQL Server database integration. - [3 docs tagged with "MySQL"](/docs/next/tags/mysql): MySQL relational database integration. - [One doc tagged with "Node.js"](/docs/next/tags/nodejs): Node.js runtime environment topics. - [3 docs tagged with "NoSQL"](/docs/next/tags/nosql): NoSQL database systems and document stores. - [29 docs tagged with "Observability"](/docs/next/tags/observability): Monitoring, tracing, and metrics for system visibility. - [One doc tagged with "ODBC"](/docs/next/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/next/tags/open-source): Open source licenses and community contributions. - [3 docs tagged with "OpenAI"](/docs/next/tags/openai): OpenAI API integration and GPT model access. - [2 docs tagged with "Oracle"](/docs/next/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/next/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/next/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/next/tags/parameters): Configuration parameters and runtime options. - [2 docs tagged with "Performance"](/docs/next/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/next/tags/persistence): Data persistence and durable storage mechanisms. - [4 docs tagged with "Postgres"](/docs/next/tags/postgres): PostgreSQL database integration and configuration. - [One doc tagged with "Python"](/docs/next/tags/python): Python programming language topics. - [3 docs tagged with "Query"](/docs/next/tags/query): SQL query execution, parameterized queries, and prepared statements. - [4 docs tagged with "Reference"](/docs/next/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/next/tags/relational): Relational database systems and SQL operations. - [3 docs tagged with "Runtime"](/docs/next/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Rust"](/docs/next/tags/rust): Rust programming language topics. - [2 docs tagged with "S3"](/docs/next/tags/s3): Amazon S3 object storage integration. - [One doc tagged with "S3 Express"](/docs/next/tags/s3-express): Amazon S3 Express One Zone storage. - [One doc tagged with "Sandbox"](/docs/next/tags/sandbox): Sandbox environments and isolated testing. - [One doc tagged with "ScyllaDB"](/docs/next/tags/scylladb): ScyllaDB database integration. - [6 docs tagged with "SDK"](/docs/next/tags/sdk): Software Development Kit topics and usage. - [10 docs tagged with "Search"](/docs/next/tags/search): Vector search, semantic search, and ranking capabilities. - [2 docs tagged with "Security"](/docs/next/tags/security): Security features and data protection mechanisms. - [One doc tagged with "Snowflake"](/docs/next/tags/snowflake): Snowflake cloud data warehouse integration. - [7 docs tagged with "SpiceAI"](/docs/next/tags/spiceai): SpiceAI cloud platform and managed services. - [5 docs tagged with "Spicepod"](/docs/next/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [6 docs tagged with "SQL"](/docs/next/tags/sql): SQL query language support and database operations. - [5 docs tagged with "Tools"](/docs/next/tags/tools): Development tools and utility integrations. - [2 docs tagged with "Tracing"](/docs/next/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/next/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Turso"](/docs/next/tags/turso): Turso data accelerator integration. - [2 docs tagged with "UDF"](/docs/next/tags/udf): User-defined functions registered with the SQL engine. - [2 docs tagged with "Unity Catalog"](/docs/next/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [One doc tagged with "Views"](/docs/next/tags/views): Virtual views and data transformation layers. - [One doc tagged with "Vortex"](/docs/next/tags/vortex): Vortex columnar file format and storage engine. - [10 docs tagged with "Write"](/docs/next/tags/write): Data connectors and catalogs that support write operations. - [One doc tagged with "YAML"](/docs/next/tags/yaml): YAML configuration syntax and file formats. - [One doc tagged with "Zipkin"](/docs/next/tags/zipkin): Zipkin distributed tracing system integration. - [Open Source Acknowledgements](/docs/next/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: - [Spice.ai API Reference](/docs/next/api): Spice.ai API reference including HTTP REST API, Arrow Flight SQL, JDBC, ODBC, ADBC connectors, authentication, and TLS configuration. - [ADBC: Arrow Database Connectivity](/docs/next/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/next/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/next/api/auth): Authentication documentation - [Generate Package](/docs/next/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/next/api/HTTP/get-catalogs): Returns a list of all registered catalogs (data sources). Catalogs provide metadata about schemas and tables available from external data sources. - [Get Iceberg API config](/docs/next/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/next/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/next/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [List Models](/docs/next/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [Check if a namespace exists.](/docs/next/api/HTTP/get-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Spicepods](/docs/next/api/HTTP/get-spicepods): Get a list of spicepods and their summary details. - [Check Runtime Status](/docs/next/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Get a table.](/docs/next/api/HTTP/get-table): This endpoint returns the table if it exists, otherwise it returns a 404 Not Found response. - [List Workers](/docs/next/api/HTTP/get-workers): Returns a list of all registered workers in the runtime. Workers are configurable processing units that can perform tasks like load balancing between models or implementing fallback strategies. - [Check Namespace exists](/docs/next/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [Check if a table exists.](/docs/next/api/HTTP/head-table): This endpoint returns a 200 OK response if the table exists, otherwise it returns a 404 Not Found response. - [list_tables](/docs/next/api/HTTP/list-tables): list_tables - [List Tools](/docs/next/api/HTTP/list-tools): Returns a list of all available tools in the Spice runtime. Tools provide reusable functionality that can be invoked programmatically or by AI agents. - [Send a Model Context Protocol message](/docs/next/api/HTTP/mcp-message): Send a JSON-RPC message to the Spice MCP server using the MCP Streamable HTTP transport. The response is either a single JSON-RPC response (`application/json`) or an SSE stream (`text/event-stream`), selected via the `Accept` header. Session continuity is carried via the `Mcp-Session-Id` header. - [Open an MCP server-to-client SSE stream](/docs/next/api/HTTP/mcp-stream): Open a long-lived server-to-client SSE stream for the current MCP session as defined by the Streamable HTTP transport. The `Mcp-Session-Id` header must identify an existing session created via `POST /v1/mcp`. - [Terminate an MCP Streamable HTTP session](/docs/next/api/HTTP/mcp-terminate-session): Terminate the MCP session identified by the `Mcp-Session-Id` header. Subsequent requests bearing the same session id will receive `404 Not Found`. - [Update Refresh SQL](/docs/next/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Create Chat Completion](/docs/next/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/next/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/next/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Text-to-SQL (NSQL)](/docs/next/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [post_responses](/docs/next/api/HTTP/post-responses): post_responses - [Search](/docs/next/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/next/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/next/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [Run Tool](/docs/next/api/HTTP/run-tool): Execute a specific tool by name. The request body schema and response format are defined by each individual tool's specification. Use `GET /v1/tools` to discover available tools and their parameter schemas. - [runtime](/docs/next/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/next/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/next/api/odbc): ODBC API Documentation - [API Overview](/docs/next/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/next/api/tls): Encryption in transit with TLS documentation - [Spice.ai CLI Reference](/docs/next/cli): Complete CLI reference for Spice.ai including commands to create, manage Spicepods, run queries, and interact with the Spice runtime. - [Spice.ai OSS CLI command reference](/docs/next/cli/reference): Spice CLI command reference - [add](/docs/next/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/next/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/next/cli/reference/chat): spice chat CLI documentation - [completions](/docs/next/cli/reference/completions): Generate shell completions for the Spice CLI. - [connect](/docs/next/cli/reference/connect): Enroll this host with Spice Cloud for remote management (Cloud Connect). - [dataset](/docs/next/cli/reference/dataset): Add or configure dataset entries in spicepod.yaml. - [datasets](/docs/next/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [feedback](/docs/next/cli/reference/feedback): Open the Spice.ai community Slack in the default browser to share feedback. - [init](/docs/next/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/next/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/next/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/next/cli/reference/models): Lists models loaded by the Spice runtime - [nsql](/docs/next/cli/reference/nsql): spice nsql CLI documentation - [pods](/docs/next/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [query](/docs/next/cli/reference/query): Submit an async query or start an interactive async query REPL against the Spice runtime's distributed query engine. - [refresh](/docs/next/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/next/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/next/cli/reference/search): Performs embeddings-based searches across search-configured datasets. - [spiced](/docs/next/cli/reference/spiced): Command-line reference for the spiced runtime binary — flags, defaults, and differences from spice run. - [sql](/docs/next/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/next/cli/reference/status): Spice runtime status - [trace](/docs/next/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/next/cli/reference/upgrade): Upgrades the Spice CLI and runtime to the latest or specified version - [validate](/docs/next/cli/reference/validate): Validate a spicepod.yaml without starting the runtime. - [version](/docs/next/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/next/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels - [Clients and Tools](/docs/next/clients): Client and tools for connecting to Spice - [DBeaver](/docs/next/clients/dbeaver): Configure DBeaver to query Spice via JDBC - [JetBrains DataGrip](/docs/next/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Microsoft Power BI Connector](/docs/next/clients/powerbi): Use Microsoft Power BI to access, visualize and analyze Spice datasets. - [Apache Superset](/docs/next/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/next/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. - [How Spice Compares](/docs/next/comparison): Compare Spice.ai with data platforms (Databricks, Snowflake), query engines (Trino, Dremio, ClickHouse), vector databases (Turbopuffer, LanceDB), search engines (Elasticsearch), and AI frameworks (LangChain, LlamaIndex, Ollama). - [Spice.ai Runtime Components](/docs/next/components): Configure Spice.ai runtime components including data connectors, data accelerators, catalog connectors, model providers, embedding models, and secret stores. - [Catalog Connectors](/docs/next/components/catalogs): Connect to external catalog providers like Unity Catalog, Databricks, Iceberg, AWS Glue, Snowflake, ADBC, PostgreSQL, MySQL, MSSQL, Oracle, and more for federated SQL query in Spice. - [ADBC Catalog Connector](/docs/next/components/catalogs/adbc): Connect to databases via ADBC for automatic schema and table discovery. - [Databricks Catalog Connector](/docs/next/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [DuckLake Catalog Connector](/docs/next/components/catalogs/ducklake): Connect to a DuckLake catalog for federated SQL query. - [Glue Catalog Connector](/docs/next/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/next/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Microsoft SQL Server Catalog Connector](/docs/next/components/catalogs/mssql): Connect to a Microsoft SQL Server database as a catalog provider for federated SQL query. - [MySQL Catalog Connector](/docs/next/components/catalogs/mysql): Connect to a MySQL database as a catalog provider for federated SQL query. - [Oracle Catalog Connector](/docs/next/components/catalogs/oracle): Connect to an Oracle database as a catalog provider for federated SQL query. - [PostgreSQL Catalog Connector](/docs/next/components/catalogs/postgres): Connect to a PostgreSQL database as a catalog provider for federated SQL query. - [Snowflake Catalog Connector](/docs/next/components/catalogs/snowflake): Connect to a Snowflake database as a catalog provider for federated SQL query. - [Spice.ai Catalog Connector](/docs/next/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/next/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Unity Catalog Catalog Connector Deployment Guide](/docs/next/components/catalogs/unity-catalog/deployment): Operating guide for the Unity Catalog catalog connector in production: workspace authentication, table-type filtering, effective-permissions flow, and observability. - [Data Accelerators](/docs/next/components/data-accelerators): Data acceleration engines for local materialization and query acceleration in Spice - [In-Memory Arrow Data Accelerator](/docs/next/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [Arrow Data Accelerator Deployment Guide](/docs/next/components/data-accelerators/arrow/deployment): Operating guide for the Arrow (in-memory) data accelerator in production: memory sizing, indexes, and observability. - [Spice Cayenne Data Accelerator](/docs/next/components/data-accelerators/cayenne): Spice Cayenne Data Accelerator (Vortex) Documentation - [Cayenne Data Accelerator Deployment Guide](/docs/next/components/data-accelerators/cayenne/deployment): Operating guide for Spice Cayenne in production: footer and segment caches, S3 Express, metastore durability, and observability. - [Spice Cayenne Performance Tuning](/docs/next/components/data-accelerators/cayenne/performance): Tuning Spice Cayenne: self-tuning and goal-driven SLOs, cache sizing, compression strategy, and file-size tuning. - [DuckDB Data Accelerator](/docs/next/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [DuckDB Data Accelerator Deployment Guide](/docs/next/components/data-accelerators/duckdb/deployment): Operating guide for the DuckDB data accelerator in production: memory vs file mode, checkpointing, spill, pool sizing, and observability. - [PostgreSQL Data Accelerator](/docs/next/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [PostgreSQL Data Accelerator Deployment Guide](/docs/next/components/data-accelerators/postgres/deployment): Operating guide for the PostgreSQL data accelerator in production: authentication, connection pooling, and observability. - [SQLite Data Accelerator](/docs/next/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [SQLite Data Accelerator Deployment Guide](/docs/next/components/data-accelerators/sqlite/deployment): Operating guide for the SQLite data accelerator in production: file mode, busy timeout, pool, and observability. - [Turso Data Accelerator](/docs/next/components/data-accelerators/turso): Turso (libSQL) Data Accelerator Documentation - [Data Connectors](/docs/next/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/next/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ADBC Data Connector](/docs/next/components/data-connectors/adbc): ADBC Data Connector Documentation - [ClickHouse Data Connector](/docs/next/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Azure Cosmos DB Data Connector](/docs/next/components/data-connectors/cosmosdb): Query Azure Cosmos DB (NoSQL / Core SQL API) containers as SQL tables in Spice. Read-only scan with schema inferred from a sample of documents. - [Azure Cosmos DB Data Connector Deployment Guide](/docs/next/components/data-connectors/cosmosdb/deployment): Operating guide for the Azure Cosmos DB data connector in production: authentication, RU sizing, resilience, metrics, and observability. - [Databricks Data Connector](/docs/next/components/data-connectors/databricks): Databricks Data Connector Documentation - [Databricks Deployment Guide](/docs/next/components/data-connectors/databricks/deployment): Operating guide for the Databricks connector in production: resilience controls, Unity Catalog behavior, metrics, and observability. - [Debezium Data Connector](/docs/next/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/next/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Delta Lake Data Connector Deployment Guide](/docs/next/components/data-connectors/delta-lake/deployment): Operating guide for the Delta Lake data connector in production: object store auth, metadata caching, metrics, and observability. - [Dremio Data Connector](/docs/next/components/data-connectors/dremio): Dremio Data Connector Documentation - [Dremio Data Connector Deployment Guide](/docs/next/components/data-connectors/dremio/deployment): Operating guide for the Dremio data connector in production: authentication, Flight SQL transport, and observability. - [DuckDB Data Connector](/docs/next/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DuckDB Data Connector Deployment Guide](/docs/next/components/data-connectors/duckdb/deployment): Operating guide for the DuckDB data connector in production: file vs memory mode, durability, and observability. - [DuckLake Data Connector](/docs/next/components/data-connectors/ducklake): DuckLake Data Connector Documentation - [DynamoDB Data Connector](/docs/next/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [DynamoDB Data Connector Deployment Guide](/docs/next/components/data-connectors/dynamodb/deployment): Operating guide for the DynamoDB data connector in production: IAM, streams, checkpointing, lag behavior, and observability. - [Elasticsearch Data Connector](/docs/next/components/data-connectors/elasticsearch): Query Elasticsearch indexes as SQL tables in Spice, including kNN vector search, full-text search, and hybrid search. - [Elasticsearch Data Connector Deployment Guide](/docs/next/components/data-connectors/elasticsearch/deployment): Operating guide for the Elasticsearch data connector in production: authentication, TLS, resilience, and operational tuning. - [File Data Connector](/docs/next/components/data-connectors/file): File Data Connector Documentation - [File Data Connector Deployment Guide](/docs/next/components/data-connectors/file/deployment): Operating guide for the File data connector in production: permissions, formats, performance, and observability. - [Flight SQL Data Connector](/docs/next/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/next/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GCS Data Connector](/docs/next/components/data-connectors/gcs): GCS (Google Cloud Storage) Data Connector Documentation - [GitHub Data Connector](/docs/next/components/data-connectors/github): GitHub Data Connector Documentation - [GitHub Data Connector Deployment Guide](/docs/next/components/data-connectors/github/deployment): Operating guide for the GitHub data connector in production: PATs, rate limits, pagination, and observability. - [Glue Data Connector](/docs/next/components/data-connectors/glue): Connect to and query tables in an AWS Glue Data Catalog - [GraphQL Data Connector](/docs/next/components/data-connectors/graphql): GraphQL Data Connector Documentation - [GraphQL Data Connector Deployment Guide](/docs/next/components/data-connectors/graphql/deployment): Operating guide for the GraphQL data connector in production: authentication, pagination, rate limits, and observability. - [HTTP(s) Data Connector](/docs/next/components/data-connectors/https): HTTP(s) Data Connector Documentation - [HTTP(s) Data Connector Deployment Guide](/docs/next/components/data-connectors/https/deployment): Operating guide for the HTTP(s) data connector in production: authentication, rate control, retries, and observability. - [Iceberg Data Connector](/docs/next/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/next/components/data-connectors/imap): IMAP Data Connector Documentation - [Kafka Data Connector](/docs/next/components/data-connectors/kafka): Kafka Data Connector Documentation - [Localpod Data Connector](/docs/next/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/next/components/data-connectors/memory): Memory Data Connector Documentation - [MongoDB Data Connector](/docs/next/components/data-connectors/mongodb): MongoDB Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/next/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [Microsoft SQL Server Performance](/docs/next/components/data-connectors/mssql/performance): Performance tuning for the Microsoft SQL Server connector: TopK / ORDER BY ... LIMIT pushdown and NULL ordering. - [MySQL Data Connector](/docs/next/components/data-connectors/mysql): MySQL Data Connector Documentation - [MySQL Data Connector Deployment Guide](/docs/next/components/data-connectors/mysql/deployment): Operating guide for the MySQL data connector in production: authentication, connection pooling, TLS, metrics, and observability. - [NFS Data Connector](/docs/next/components/data-connectors/nfs): NFS Data Connector Documentation - [ODBC Data Connector](/docs/next/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/next/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/next/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [PostgreSQL Data Connector Deployment Guide](/docs/next/components/data-connectors/postgres/deployment): Operating guide for the PostgreSQL data connector in production: authentication, connection pooling, TLS, metrics, and observability. - [Amazon Redshift Data Connector](/docs/next/components/data-connectors/redshift): Connect to Amazon Redshift using the PostgreSQL connector in Spice. - [S3 Data Connector](/docs/next/components/data-connectors/s3): S3 Data Connector Documentation - [S3 Data Connector Deployment Guide](/docs/next/components/data-connectors/s3/deployment): Operating guide for the S3 data connector in production: IAM, credential chains, file formats, metrics, and observability. - [ScyllaDB Data Connector](/docs/next/components/data-connectors/scylladb): ScyllaDB Data Connector Documentation - [ScyllaDB Performance](/docs/next/components/data-connectors/scylladb/performance): Performance considerations for the ScyllaDB connector: partition/clustering key filters, acceleration, and datacenter locality. - [SharePoint Data Connector](/docs/next/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [SMB Data Connector](/docs/next/components/data-connectors/smb): SMB Data Connector Documentation - [Snowflake Data Connector](/docs/next/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/next/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/next/components/data-connectors/spiceai): Federated SQL across Spice runtimes — Spice Cloud Platform datasets and self-hosted Spice instances (cluster-sidecar pattern). - [Spice.ai Data Connector Deployment Guide](/docs/next/components/data-connectors/spiceai/deployment): Operating guide for the Spice.ai connector in production: API keys, Flight endpoints, message sizing, sidecar topology, and observability. - [Embedding Models](/docs/next/components/embeddings): Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. - [Azure OpenAI Embedding Models](/docs/next/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/next/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/next/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [Google AI Embedding Models](/docs/next/components/embeddings/google): To use a hosted Google AI embedding model, specify the google path in the from field of your configuration. - [HuggingFace Text Embedding Models](/docs/next/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Hugging Face Embedding Deployment Guide](/docs/next/components/embeddings/huggingface/deployment): Operating guide for Hugging Face embeddings in production: tokens, download cache, pooling, device selection, and observability. - [Local Filesystem Embedding Models](/docs/next/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [Local Embedding Deployment Guide](/docs/next/components/embeddings/local/deployment): Operating guide for filesystem-loaded embedding models in production: formats, pooling, device selection, and observability. - [Model2Vec Embedding Models](/docs/next/components/embeddings/model2vec): Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning. - [OpenAI (or Compatible) Embedding Models](/docs/next/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [OpenAI Embedding Deployment Guide](/docs/next/components/embeddings/openai/deployment): Operating guide for the OpenAI embedding provider in production: API keys, usage tiers, batching, retries, metrics, and observability. - [Model Providers](/docs/next/components/models): Overview of supported model providers for LLMs in Spice. - [Anthropic Models](/docs/next/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/next/components/models/azure): Instructions for using Azure OpenAI models - [Amazon Bedrock Models](/docs/next/components/models/bedrock): How to use Amazon Bedrock models with Spice. - [Databricks Model Provider](/docs/next/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/next/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [Filesystem Model Deployment Guide](/docs/next/components/models/filesystem/deployment): Operating guide for filesystem-loaded models in production: formats, device selection, memory footprint, and observability. - [Google AI Models](/docs/next/components/models/google): Instructions for using language models hosted on Google AI with Spice. - [HuggingFace](/docs/next/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [Hugging Face Model Deployment Guide](/docs/next/components/models/huggingface/deployment): Operating guide for the Hugging Face model in production: tokens, download cache, device selection, local inference footprint, and observability. - [OpenAI (or Compatible) Language Models](/docs/next/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [OpenAI Model Deployment Guide](/docs/next/components/models/openai/deployment): Operating guide for the OpenAI model in production: API keys, usage tiers, rate limiting, Responses API, metrics, and observability. - [Perplexity Models (Deprecated)](/docs/next/components/models/perplexity): Perplexity model support is no longer supported in Spice. - [Spice Cloud Platform](/docs/next/components/models/spiceai): Instructions for using language models served by the Spice.ai Cloud Platform with Spice. - [xAI Models](/docs/next/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/next/components/secret-stores): Configure secret stores to manage sensitive data like passwords, tokens, and API keys. - [AWS Secrets Manager Secret Store](/docs/next/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Azure Key Vault Secret Store](/docs/next/components/secret-stores/azure-keyvault): Azure Key Vault Secret Store Documentation - [Environment Secret Store](/docs/next/components/secret-stores/env): Environment Variables Secret Store Documentation - [HashiCorp Vault Secret Store](/docs/next/components/secret-stores/hashicorp-vault): HashiCorp Vault Secret Store Documentation - [Keyring Secret Store](/docs/next/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/next/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/next/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/next/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Streamable HTTP transport. - [Web Search Tool (Deprecated)](/docs/next/components/tools/websearch): The websearch tool is no longer supported in Spice. - [Vector Engines](/docs/next/components/vectors): Configure vector engines for efficient embedding storage and similarity search in Spice. - [DuckDB Vector Engine](/docs/next/components/vectors/duckdb): Use DuckDB as a vector engine in Spice for HNSW-based vector search via the DuckDB VSS extension. - [Elasticsearch Vector Engine](/docs/next/components/vectors/elasticsearch): Use Elasticsearch as a vector engine in Spice for kNN vector search, full-text search, and hybrid search. - [Amazon S3 Vectors Engine](/docs/next/components/vectors/s3_vectors): Amazon S3 Vectors Engine Documentation - [Spice.ai Deployment Guide](/docs/next/deployment): Deploy Spice.ai in your environment using Docker, Kubernetes, AWS, Azure, or the Spice Cloud Platform. Learn about sidecar, microservice, tiered, and cluster deployment architectures. - [Deployment Architectures](/docs/next/deployment/architectures): Explore Spice deployment architectures including sidecar, microservice, tiered, sharded, and cluster configurations. - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/next/deployment/architectures/cluster): Deploying Spice as a cluster - [Cluster-Sidecar Deployment](/docs/next/deployment/architectures/cluster-sidecar): Deploy Spice with application-local sidecars for localhost query, search, and inference, backed by a centralized cluster for ingest, acceleration, and distributed query. - [Cloud Hosted](/docs/next/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Microservice Deployment (Single or Multiple Replicas)](/docs/next/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/next/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/next/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/next/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/next/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [AWS Integrations](/docs/next/deployment/aws/integrations): Complete guide to Spice.ai integrations with Amazon Web Services, including data connectors, AI models, vector stores, and secret management. - [Azure Deployment Options](/docs/next/deployment/azure): Guide to deploying Spice.ai applications on Microsoft Azure - [Azure Integrations](/docs/next/deployment/azure/integrations): Spice.ai integrations with Microsoft Azure, including data connectors, AI models, embeddings, and authentication. - [CI/CD Deployment](/docs/next/deployment/ci-cd): Deploy Spice.ai applications using continuous integration and delivery pipelines, including Helm, Kubernetes GitOps with Argo CD or Flux, GitHub Actions, and the Spice Cloud deploy action. - [Spice Cloud Platform Deployment](/docs/next/deployment/cloud): Guide to deploying data and AI applications using the managed Spice Cloud Platform - [Docker](/docs/next/deployment/docker): Run Spice.ai as a Docker container. - [Docker Sandbox Guide - v1.3.0](/docs/next/deployment/docker/sandbox): Migrating to v1.3.0 - [Google Cloud Deployment Options](/docs/next/deployment/gcp): Guide to deploying Spice.ai applications on Google Cloud Platform (GCP). - [GCP Integrations](/docs/next/deployment/gcp/integrations): Spice.ai integrations with Google Cloud Platform, including data connectors, AI models, embeddings, and authentication. - [Kubernetes Deployment](/docs/next/deployment/kubernetes): Deploy Spice.ai on Kubernetes using Helm, Argo CD, or Flux. - [Kubernetes - Argo CD](/docs/next/deployment/kubernetes/argocd): Deploy Spice.ai on self-hosted Kubernetes using Argo CD and the Spice Helm chart. - [Kubernetes - Flux](/docs/next/deployment/kubernetes/flux): Deploy Spice.ai on Kubernetes using Flux CD and the Spice Helm chart. - [Kubernetes - Helm](/docs/next/deployment/kubernetes/helm): Deploy Spice.ai in Kubernetes using Helm. - [Read/Write Separation](/docs/next/deployment/read-write-separation): Separate write/ingest workloads (cluster) from read workloads (application sidecars, agents) using shared snapshots and live query delegation. - [Spice.ai FAQ](/docs/next/faq): Answers to frequently asked questions about Spice.ai including features, use cases, differences from Trino/Presto/Dremio, federated queries, caching, and AI capabilities. - [Spice.ai Features](/docs/next/features): Explore Spice.ai features including data federation, data acceleration, caching, search, LLM integration, embeddings, observability, and more for building data-driven AI applications. - [Caching](/docs/next/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/next/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Debezium (CDC over Kafka)](/docs/next/features/cdc/debezium): Consume Debezium change events from Kafka into a Spice-accelerated dataset for sources without a native Spice CDC path. - [Debezium Push Ingest (CDC without Kafka)](/docs/next/features/cdc/debezium-ingest): Stream Debezium change events directly into a Spice-accelerated dataset over HTTP — any Debezium source plugin, no Kafka bus required. - [DynamoDB Streams (Native CDC)](/docs/next/features/cdc/dynamodb-streams): Stream INSERT, UPDATE, and DELETE events from Amazon DynamoDB directly into a Spice-accelerated dataset using DynamoDB Streams. - [MongoDB Change Streams (Native CDC)](/docs/next/features/cdc/mongodb-streams): Stream insert, update, replace, and delete events from MongoDB directly into a Spice-accelerated dataset using MongoDB Change Streams. - [MySQL Binlog Replication (Native CDC)](/docs/next/features/cdc/mysql-replication): Stream INSERT, UPDATE, and DELETE events from MySQL directly into a Spice-accelerated dataset using native binary log (binlog) replication. - [PostgreSQL Logical Replication (Native CDC)](/docs/next/features/cdc/postgres-replication): Stream INSERT, UPDATE, and DELETE events from PostgreSQL directly into a Spice-accelerated dataset using native logical replication. - [Data Acceleration](/docs/next/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/next/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/next/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Hash Index for Arrow Acceleration](/docs/next/features/data-acceleration/hash-index): Learn how to use hash indexes for O(1) point lookups on Arrow-accelerated datasets. - [Indexes](/docs/next/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/next/features/data-acceleration/partitioning): Partition accelerated datasets to make filtered queries faster by reading only the relevant partitions. - [Refresh Modes](/docs/next/features/data-acceleration/refresh-modes): Refresh modes for accelerated datasets in Spice. - [Append Refresh Mode](/docs/next/features/data-acceleration/refresh-modes/append): Incrementally append new rows to an accelerated dataset. - [Caching Refresh Mode](/docs/next/features/data-acceleration/refresh-modes/caching): Learn how to use caching refresh mode for HTTP-based datasets - [Changes Refresh Mode](/docs/next/features/data-acceleration/refresh-modes/changes): Apply incremental inserts, updates, and deletes via Change Data Capture. - [Full Refresh Mode](/docs/next/features/data-acceleration/refresh-modes/full): Replace the entire accelerated dataset on each refresh. - [Snapshot Refresh Mode](/docs/next/features/data-acceleration/refresh-modes/snapshot): Reload acceleration data exclusively from the snapshot store. - [Snapshots](/docs/next/features/data-acceleration/snapshots): Bootstrap file-mode accelerations from managed snapshots to eliminate cold starts. - [Data Ingestion](/docs/next/features/data-ingestion): Learn how to ingest data in Spice. - [Distributed Query](/docs/next/features/distributed-query): Learn how to run Spice in distributed mode for larger scale queries, including the async queries API. - [Embedding Datasets](/docs/next/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Functions](/docs/next/features/functions): Define custom scalar and table SQL functions inline (SQL tier) or by calling remote HTTP services (Remote tier), automatically exposed as SQL functions and LLM tools. - [Large Language Models](/docs/next/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models (Deprecated)](/docs/next/features/large-language-models/evals): Language model evals are no longer supported in Spice. - [Model Context Protocol (MCP)](/docs/next/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/next/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/next/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/next/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/next/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/next/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/next/features/machine-learning-models): Deprecated in vNext: Support for loading and serving traditional machine learning (ONNX) models for inference was removed in vNext, along with the /v1/predict and /v1/models//predict prediction endpoints. See the v2.1 docs for documentation of this feature. - [Observability & Monitoring](/docs/next/features/observability): Monitor Spice with Prometheus metrics, OpenTelemetry, and distributed tracing. - [Component Metrics](/docs/next/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/next/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Parameterized Queries](/docs/next/features/query-federation/parameterized-queries): Learn how to use prepared statements and parameterized queries in Spice for improved security and performance. - [URL Tables](/docs/next/features/query-federation/url-tables): Query object store files directly using URLs without pre-registering datasets - [Search Functionality](/docs/next/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/next/features/search/full-text): Learn how Spice can perform full text search - [Multi-Vector Search](/docs/next/features/search/multi-vector): Embed list-of-strings columns as a column of vectors and use ColBERT-style late-interaction search in Spice. - [Reranking](/docs/next/features/search/rerank): Rerank search results using dedicated reranker models or LLM-as-reranker for improved relevance. - [Vector-Based Search](/docs/next/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/next/features/semantic-model): Attach descriptions and metadata to datasets, views, and columns in Spice so LLMs, SQL functions, and humans share the same understanding of your data. - [Tool Registry](/docs/next/features/tool-registry): Reduce per-turn token cost and improve LLM tool selection accuracy by replacing individual tool definitions with searchable tool_search and tool_invoke meta-tools backed by hybrid full-text, keyword, schema, and vector search. - [Views](/docs/next/features/views): Documentation for defining Views in Spice - [Web Search](/docs/next/features/web-search): Learn how Spice can perform web search - [Workers](/docs/next/features/workers): Configure workers in the Spice runtime to coordinate interactions between LLMs and tools, with load-balancing, round-robin, and fallback strategies. - [Getting Started with Spice.ai OSS](/docs/next/getting-started): Get started with Spice.ai in 5 minutes. Install the CLI, connect to datasets, run SQL queries, and use AI models with OpenAI-compatible APIs. - [Community Data](/docs/next/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/next/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/next/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. - [Install Spice.ai OSS](/docs/next/installation): Install Spice.ai OSS on macOS, Linux, Windows, or WSL using the install script, Homebrew, PowerShell, or direct download from GitHub releases. - [Building Intelligent AI Applications with Spice.ai](/docs/next/intelligent-applications): Learn how to build intelligent, data-driven AI applications and agents with Spice.ai. Explore patterns for RAG, LLM integration, and real-time AI inference. - [Monitoring](/docs/next/monitoring): Monitor Spice.ai deployments with Datadog, Grafana, New Relic, Prometheus, and Zipkin integrations. - [Datadog](/docs/next/monitoring/datadog): Monitoring Spice with Datadog - [Grafana & Prometheus](/docs/next/monitoring/grafana): Monitoring Spice instances with Grafana & Prometheus - [New Relic](/docs/next/monitoring/new-relic): Monitoring Spice with New Relic - [Spice Cloud Platform](/docs/next/monitoring/spice-cloud): Connect a self-hosted Spice runtime to the Spice Cloud Platform to centralize task history and runtime observability across deployments. - [Zipkin Integration](/docs/next/monitoring/zipkin): Learn how to integrate Spice with Zipkin tracing. - [Spice.ai OSS Reference Docs](/docs/next/reference): Reference documentation for Spice.ai including API reference, CLI commands, Spicepod configuration syntax, SQL reference, and data type specifications. - [Cron Schedules](/docs/next/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/next/reference/datatypes): Spice uses Apache Arrow data types internally, providing consistent type handling across different data sources and accelerators. This section documents how Arrow types map to specific accelerators and object store formats. - [Accelerator Data Types](/docs/next/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/next/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Spice Runtime Distributions](/docs/next/reference/distributions): Distribution variants of the Spice runtime for different use cases and deployment scenarios, including data-only, GPU-accelerated, NAS, and allocator variants. - [Duration](/docs/next/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/next/reference/file_format): File-based data connectors — including s3//, file//, sftp://, and others — support multiple structured and document file formats. This page details the format-specific parameters available for each. - [Managing Memory Usage](/docs/next/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice deployments. - [Models Grade Report](/docs/next/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [Performance Tuning](/docs/next/reference/performance-tuning): Comprehensive guide to optimizing query performance, acceleration, and resource utilization in Spice deployments. - [YAML syntax for Spicepod manifests](/docs/next/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/next/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/next/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/next/reference/spicepod/embeddings): Embeddings YAML reference - [Evals (Deprecated)](/docs/next/reference/spicepod/evals): The evals Spicepod component is no longer supported in Spice. - [Functions (User-Defined Functions)](/docs/next/reference/spicepod/functions): User-defined functions YAML reference - [Reserved Keywords](/docs/next/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/next/reference/spicepod/models): Models YAML reference - [Runtime](/docs/next/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/next/reference/spicepod/tools): Tools YAML reference - [Views](/docs/next/reference/spicepod/views): Views YAML reference - [Workers](/docs/next/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/next/reference/sql): Complete SQL reference for Spice.ai including SELECT syntax, subqueries, DML statements, aggregate functions, AI functions, JSON operators, and search capabilities. - [Aggregate Functions](/docs/next/reference/sql/aggregate_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [AI Functions](/docs/next/reference/sql/ai): AI functions in Spice provide direct integration with large language models (LLMs) and embedding models within SQL queries. These functions process text through configured model providers and return generated responses or vector embeddings. - [DML (Data Manipulation Language)](/docs/next/reference/sql/dml): Data Manipulation Language (DML) statements for inserting and modifying data in Spice. - [Explain](/docs/next/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/next/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [JSON Functions and Operators](/docs/next/reference/sql/json): Reference for JSON functions and operators in Spice SQL - [Operators](/docs/next/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/next/reference/sql/prepared_statements): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Scalar Functions](/docs/next/reference/sql/scalar_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [Search in SQL](/docs/next/reference/sql/search): Reference for search functions and filtering in Spice SQL. - [SELECT](/docs/next/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/next/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/next/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/next/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. - [SDKs](/docs/next/sdks): Connect to Spice using official SDKs - [Dotnet SDK](/docs/next/sdks/dotnet): Connect to Spice using the Dotnet SDK - [Go SDK](/docs/next/sdks/golang): Connect to Spice using the Go SDK - [Java SDK](/docs/next/sdks/java): Connect to Spice using the Java SDK - [JavaScript SDK](/docs/next/sdks/javascript): Connect to Spice using the JavaScript SDK - [Python SDK](/docs/next/sdks/python): Connect to Spice using the Python SDK - [Rust SDK](/docs/next/sdks/rust): Connect to Spice using the Rust SDK - [Troubleshooting Spice](/docs/next/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Spice.ai Use Cases](/docs/next/use-cases): Discover how to use Spice.ai for data federation, reverse-ETL, database CDN, enterprise search, RAG, and building AI-powered applications and agents. - [AI Applications and Agents](/docs/next/use-cases/ai): AI Applications and Agents - [Agentic AI Applications and Agents](/docs/next/use-cases/ai/agentic-apps): Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. - [Edge-Enabled AI Applications and Agents](/docs/next/use-cases/ai/edge-ai): Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases. - [Federated MCP Client for Distributed Tool Ecosystems](/docs/next/use-cases/ai/federated-mcp-server): Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, improving threat analysis. - [Multi-Tenant AI Agents](/docs/next/use-cases/ai/multi-tenant-agents): Deploy AI agents across many SaaS tenants with strict isolation and no per-tenant ETL pipelines. - [Object-Store Based SQL Query, Search, and LLM Inference Engine](/docs/next/use-cases/ai/object-store-ai-engine): Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. - [Real-Time Decision-Making for Intelligent Applications](/docs/next/use-cases/ai/real-time-decision-making): Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. - [Tool-Augmented AI with Model Context Protocol Server](/docs/next/use-cases/ai/tool-calling-ai): Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. - [Caching](/docs/next/use-cases/caching): Caching use cases with Spice.ai, including write-through, read-through, SQL, S3, and HTTP caching. - [HTTP Cache](/docs/next/use-cases/caching/http-cache): Use Spice.ai to cache HTTP API responses locally, reducing API call frequency and providing fast, SQL-queryable access to external API data. - [Read-Through Cache](/docs/next/use-cases/caching/read-through-cache): Use Spice.ai as a read-through cache with the SQL results cache for federated data sources and HTTP APIs. - [S3 Cache](/docs/next/use-cases/caching/s3-cache): Use Spice.ai to cache S3 and object store data locally for fast, repeatable SQL queries without re-reading from remote storage. - [SQL and Database Cache](/docs/next/use-cases/caching/sql-database-cache): Use Spice.ai to cache SQL database tables and query results locally for low-latency access and reduced load on upstream databases. - [Write-Through Cache](/docs/next/use-cases/caching/write-through-cache): Use Spice.ai as a write-through cache that writes data to both a local accelerator and the upstream data source. - [Data Federation, Acceleration, and SQL Query](/docs/next/use-cases/data): Data Federation, Acceleration, and SQL Query - [Application Resilience and Performance Optimization](/docs/next/use-cases/data/application-resilence-and-acceleration): Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance. - [Data Mesh for Unified Data Access](/docs/next/use-cases/data/data-mesh): Give domain teams decentralized, real-time data ownership and access across disparate sources through a unified SQL interface. - [Database CDN for Enhanced Performance](/docs/next/use-cases/data/database-cdn): Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience. - [ETL-free Workflows and Data Migrations](/docs/next/use-cases/data/etl-free-workflows): Federate legacy and modern data systems without ETL for faster migrations, lower overhead, and zero application downtime. - [Object-Store Data Engine](/docs/next/use-cases/data/object-store-data-engine): Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses. - [Reverse-ETL for Operational Workflows](/docs/next/use-cases/data/reverse-etl): Serve enriched data from warehouses and data lakes to operational systems and applications, eliminating complex pipelines. - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/next/use-cases/rag): Use Spice for Retrieval-Augmented-Generation (RAG) - [RAG for Contextual Applications](/docs/next/use-cases/rag/applications): Build context-rich AI applications using Spice for Retrieval-Augmented Generation (RAG). - [Retrieval-Augmented Generation for AI-Powered Reporting](/docs/next/use-cases/rag/reporting): Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. - [Search & Retrieval](/docs/next/use-cases/search): Search & Retrieval - [Simplifying Real-Time Data Collection and Search](/docs/next/use-cases/search/data-collection-and-search): Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic. - [Enterprise Search and Retrieval](/docs/next/use-cases/search/enterprise-search): Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities. - [Object-Store Native Search Engine](/docs/next/use-cases/search/object-store-search-engine): Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search. ### v1.10 Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Spice.ai Open Source](/docs/v1.10): Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Tags](/docs/v1.10/tags) - [One doc tagged with "Acknowledgements"](/docs/v1.10/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [One doc tagged with "ADBC"](/docs/v1.10/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/v1.10/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Arrow"](/docs/v1.10/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/v1.10/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/v1.10/tags/auth): Authentication and authorization mechanisms. - [One doc tagged with "Authentication"](/docs/v1.10/tags/authentication): User authentication methods and security protocols. - [2 docs tagged with "Azure"](/docs/v1.10/tags/azure): Microsoft Azure cloud services and integrations. - [One doc tagged with "Blob Storage"](/docs/v1.10/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Caching"](/docs/v1.10/tags/caching): Data caching strategies and performance optimization. - [5 docs tagged with "Catalogs"](/docs/v1.10/tags/catalogs): Data catalog connectors and metadata management. - [One doc tagged with "Cayenne"](/docs/v1.10/tags/cayenne): Cayenne (Vortex) data accelerator built on Vortex columnar format. - [2 docs tagged with "CLI"](/docs/v1.10/tags/cli): Spice command-line interface commands and usage. - [5 docs tagged with "Component Metrics"](/docs/v1.10/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/v1.10/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/v1.10/tags/configuration): System configuration files and runtime settings. - [One doc tagged with "Data Accelerators"](/docs/v1.10/tags/data-accelerators): Data acceleration engines for high-performance query execution. - [19 docs tagged with "Data Connectors"](/docs/v1.10/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/v1.10/tags/data-lake): Data lake architectures and storage solutions. - [2 docs tagged with "Databricks"](/docs/v1.10/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/v1.10/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debezium"](/docs/v1.10/tags/debezium): Debezium change data capture integration. - [One doc tagged with "Debugging"](/docs/v1.10/tags/debugging): Debugging techniques and troubleshooting tools. - [2 docs tagged with "Delta Lake"](/docs/v1.10/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/v1.10/tags/dependencies): Third-party dependencies and software requirements. - [3 docs tagged with "Deployment"](/docs/v1.10/tags/deployment): Production deployment using Docker and Kubernetes. - [2 docs tagged with "Docker"](/docs/v1.10/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with "DynamoDB"](/docs/v1.10/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [3 docs tagged with "Embeddings"](/docs/v1.10/tags/embeddings): Vector embeddings and semantic similarity operations. - [One doc tagged with "Evaluation"](/docs/v1.10/tags/evaluation): Model evaluation metrics and assessment frameworks. - [2 docs tagged with "Features"](/docs/v1.10/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/v1.10/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "Getting Started"](/docs/v1.10/tags/getting-started): Installation guides and quickstart tutorials. - [One doc tagged with "GitHub"](/docs/v1.10/tags/github): GitHub repository data connector and integration. - [2 docs tagged with "Glue"](/docs/v1.10/tags/glue): AWS Glue ETL service and data catalog integration. - [2 docs tagged with "Iceberg"](/docs/v1.10/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/v1.10/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/v1.10/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "JDBC"](/docs/v1.10/tags/jdbc): Java Database Connectivity driver configuration. - [One doc tagged with "Kafka"](/docs/v1.10/tags/kafka): Apache Kafka event streaming platform integration. - [2 docs tagged with "Kubernetes"](/docs/v1.10/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "Logging"](/docs/v1.10/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/v1.10/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/v1.10/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/v1.10/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/v1.10/tags/memory): In-memory data connectors and volatile storage. - [13 docs tagged with "Models"](/docs/v1.10/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MongoDB"](/docs/v1.10/tags/mongodb): MongoDB NoSQL document database integration. - [One doc tagged with "MySQL"](/docs/v1.10/tags/mysql): MySQL relational database integration. - [One doc tagged with "NoSQL"](/docs/v1.10/tags/nosql): NoSQL database systems and document stores. - [One doc tagged with "Observability"](/docs/v1.10/tags/observability): Monitoring, tracing, and metrics for system visibility. - [One doc tagged with "ODBC"](/docs/v1.10/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/v1.10/tags/open-source): Open source licenses and community contributions. - [One doc tagged with "OpenAI"](/docs/v1.10/tags/openai): OpenAI API integration and GPT model access. - [One doc tagged with "Oracle"](/docs/v1.10/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/v1.10/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/v1.10/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/v1.10/tags/parameters): Configuration parameters and runtime options. - [2 docs tagged with "Performance"](/docs/v1.10/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/v1.10/tags/persistence): Data persistence and durable storage mechanisms. - [4 docs tagged with "Reference"](/docs/v1.10/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/v1.10/tags/relational): Relational database systems and SQL operations. - [One doc tagged with "Runtime"](/docs/v1.10/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Sandbox"](/docs/v1.10/tags/sandbox): Sandbox environments and isolated testing. - [5 docs tagged with "Search"](/docs/v1.10/tags/search): Vector search, semantic search, and ranking capabilities. - [One doc tagged with "Security"](/docs/v1.10/tags/security): Security features and data protection mechanisms. - [2 docs tagged with "SpiceAI"](/docs/v1.10/tags/spiceai): SpiceAI cloud platform and managed services. - [5 docs tagged with "Spicepod"](/docs/v1.10/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [2 docs tagged with "SQL"](/docs/v1.10/tags/sql): SQL query language support and database operations. - [3 docs tagged with "Tools"](/docs/v1.10/tags/tools): Development tools and utility integrations. - [2 docs tagged with "Tracing"](/docs/v1.10/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/v1.10/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Unity Catalog"](/docs/v1.10/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [One doc tagged with "Views"](/docs/v1.10/tags/views): Virtual views and data transformation layers. - [One doc tagged with "Vortex"](/docs/v1.10/tags/vortex): Vortex columnar file format and storage engine. - [6 docs tagged with "Write"](/docs/v1.10/tags/write): Data connectors and catalogs that support write operations. - [One doc tagged with "YAML"](/docs/v1.10/tags/yaml): YAML configuration syntax and file formats. - [One doc tagged with "Zipkin"](/docs/v1.10/tags/zipkin): Zipkin distributed tracing system integration. - [Open Source Acknowledgements](/docs/v1.10/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: - [API](/docs/v1.10/api): API - [ADBC: Arrow Database Connectivity](/docs/v1.10/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/v1.10/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/v1.10/api/auth): Authentication documentation - [Generate Package](/docs/v1.10/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/v1.10/api/HTTP/get-catalogs): List Catalogs - [Get Iceberg API config](/docs/v1.10/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/v1.10/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/v1.10/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [ML Prediction](/docs/v1.10/api/HTTP/get-model-predict): Make a ML prediction using a specific model. - [List Models](/docs/v1.10/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [List Spicepods](/docs/v1.10/api/HTTP/get-spicepods): Get a list of spicepods and their details. In CSV format, it will return a summarised form. - [Check Runtime Status](/docs/v1.10/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Check Namespace exists](/docs/v1.10/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Evals](/docs/v1.10/api/HTTP/list): Return all evals available to run in the runtime. - [Send message to MCP server](/docs/v1.10/api/HTTP/mcp-event): Send message to the MCP endoint, for a given session. - [Establish an MCP SSE Connection](/docs/v1.10/api/HTTP/operation-id): Initiates a Server-Sent Events (SSE) connection using the Model Context Protocol (MCP) to interact with Spice tools. - [Update Refresh SQL](/docs/v1.10/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Run Tool](/docs/v1.10/api/HTTP/post): The request body and JSON response formats match the tool’s specification. - [Batch ML Predictions](/docs/v1.10/api/HTTP/post-batch-predict): Perform a batch of ML predictions, using multiple models, in one request. This is useful for ensembling or A/B testing different models. - [Create Chat Completion](/docs/v1.10/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/v1.10/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/v1.10/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Run Eval](/docs/v1.10/api/HTTP/post-eval): Evaluate a model against a eval spice specification - [Text-to-SQL (NSQL)](/docs/v1.10/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [Search](/docs/v1.10/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/v1.10/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/v1.10/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [runtime](/docs/v1.10/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/v1.10/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/v1.10/api/odbc): ODBC API Documentation - [API Overview](/docs/v1.10/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/v1.10/api/tls): Encryption in transit with TLS documentation - [Spice.ai OSS CLI documentation](/docs/v1.10/cli): Detailed documentation on the Spice.ai OSS CLI - [Spice.ai OSS CLI command reference](/docs/v1.10/cli/reference): Spice CLI command reference - [add](/docs/v1.10/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/v1.10/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/v1.10/cli/reference/chat): spice chat CLI documentation - [Completion](/docs/v1.10/cli/reference/completion): Generate the autocompletion script for spice for the specified shell. - [connect](/docs/v1.10/cli/reference/connect): Connect to an app on the Spice.ai Cloud Platform. - [dataset](/docs/v1.10/cli/reference/dataset): Configure a Spice dataset. - [datasets](/docs/v1.10/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [init](/docs/v1.10/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/v1.10/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/v1.10/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/v1.10/cli/reference/models): Lists models loaded by the Spice runtime - [pods](/docs/v1.10/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [refresh](/docs/v1.10/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/v1.10/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/v1.10/cli/reference/search): Performs embeddings-based searches across search configured datasets. Note: Search requires the ai feature to be installed. - [sql](/docs/v1.10/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/v1.10/cli/reference/status): Spice runtime status - [trace](/docs/v1.10/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/v1.10/cli/reference/upgrade): Upgrades the Spice CLI & Runtime to the latest release - [version](/docs/v1.10/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/v1.10/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels - [Clients and Tools](/docs/v1.10/clients): Client and tools - [DBeaver](/docs/v1.10/clients/dbeaver): Configure DBeaver to query Spice via JDBC - [JetBrains DataGrip](/docs/v1.10/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Microsoft Power BI Connector](/docs/v1.10/clients/powerbi): Use Microsoft Power BI to access, visualize and analyze Spice datasets. - [Apache Superset](/docs/v1.10/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/v1.10/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. - [Runtime Components](/docs/v1.10/components): Runtime components' - [Catalog Connectors](/docs/v1.10/components/catalogs) - [Databricks Catalog Connector](/docs/v1.10/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [Glue Catalog Connector](/docs/v1.10/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/v1.10/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Spice.ai Catalog Connector](/docs/v1.10/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/v1.10/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Data Accelerators](/docs/v1.10/components/data-accelerators) - [In-Memory Arrow Data Accelerator](/docs/v1.10/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [Cayenne Data Accelerator](/docs/v1.10/components/data-accelerators/cayenne): Cayenne Data Accelerator (Vortex) Documentation - [DuckDB Data Accelerator](/docs/v1.10/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [PostgreSQL Data Accelerator](/docs/v1.10/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [SQLite Data Accelerator](/docs/v1.10/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [Data Connectors](/docs/v1.10/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/v1.10/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ClickHouse Data Connector](/docs/v1.10/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Databricks Data Connector](/docs/v1.10/components/data-connectors/databricks): Databricks Data Connector Documentation - [Debezium Data Connector](/docs/v1.10/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/v1.10/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Dremio Data Connector](/docs/v1.10/components/data-connectors/dremio): Dremio Data Connector Documentation - [DuckDB Data Connector](/docs/v1.10/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DynamoDB Data Connector](/docs/v1.10/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [File Data Connector](/docs/v1.10/components/data-connectors/file): File Data Connector Documentation - [Flight SQL Data Connector](/docs/v1.10/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/v1.10/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GitHub Data Connector](/docs/v1.10/components/data-connectors/github): GitHub Data Connector Documentation - [Glue Data Connector](/docs/v1.10/components/data-connectors/glue): Glue Data Connector Documentation - [GraphQL Data Connector](/docs/v1.10/components/data-connectors/graphql): GraphQL Data Connector Documentation - [HTTP(s) Data Connector](/docs/v1.10/components/data-connectors/https): HTTP(s) Data Connector Documentation - [Iceberg Data Connector](/docs/v1.10/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/v1.10/components/data-connectors/imap): IMAP Data Connector Documentation - [Kafka Data Connector](/docs/v1.10/components/data-connectors/kafka): Kafka Data Connector Documentation - [Localpod Data Connector](/docs/v1.10/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/v1.10/components/data-connectors/memory): Memory Data Connector Documentation - [MongoDB Data Connector](/docs/v1.10/components/data-connectors/mongodb): MongoDB Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/v1.10/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [MySQL Data Connector](/docs/v1.10/components/data-connectors/mysql): MySQL Data Connector Documentation - [ODBC Data Connector](/docs/v1.10/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/v1.10/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/v1.10/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [Amazon Redshift Data Connector](/docs/v1.10/components/data-connectors/redshift): Connect to Amazon Redshift using the PostgreSQL connector in Spice. - [S3 Data Connector](/docs/v1.10/components/data-connectors/s3): S3 Data Connector Documentation - [SharePoint Data Connector](/docs/v1.10/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [Snowflake Data Connector](/docs/v1.10/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/v1.10/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/v1.10/components/data-connectors/spiceai): Spice.ai Data Connector Documentation - [Embedding Models](/docs/v1.10/components/embeddings): Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. - [Azure OpenAI Embedding Models](/docs/v1.10/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/v1.10/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/v1.10/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [HuggingFace Text Embedding Models](/docs/v1.10/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Local Filesystem Embedding Models](/docs/v1.10/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [Model2Vec Embedding Models](/docs/v1.10/components/embeddings/model2vec): Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning. - [OpenAI (or Compatible) Embedding Models](/docs/v1.10/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [Model Providers](/docs/v1.10/components/models): Overview of supported model providers for ML and LLMs in Spice. - [Anthropic Models](/docs/v1.10/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/v1.10/components/models/azure): Instructions for using Azure OpenAI models - [Amazon Bedrock Models](/docs/v1.10/components/models/bedrock): How to use Amazon Bedrock models with Spice. - [Databricks Model Provider](/docs/v1.10/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/v1.10/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [HuggingFace](/docs/v1.10/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [OpenAI (or Compatible) Language Models](/docs/v1.10/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [Perplexity Models](/docs/v1.10/components/models/perplexity): Instructions for using language models hosted on Perplexity with Spice. - [Spice Cloud Platform](/docs/v1.10/components/models/spiceai): Instructions for using models hosted on the Spice Cloud Platform with Spice. - [xAI Models](/docs/v1.10/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/v1.10/components/secret-stores) - [AWS Secrets Manager Secret Store](/docs/v1.10/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Environment Secret Store](/docs/v1.10/components/secret-stores/env): Environment Variables Secret Store Documentation - [Keyring Secret Store](/docs/v1.10/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/v1.10/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/v1.10/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/v1.10/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Server-Sent Events (SSE) protocol. - [Web Search Tool](/docs/v1.10/components/tools/websearch): The Web Search Tool enables Spice models to search the web for information. The tool is available through the websearch tool, and backed by different search engines. - [Vector Engines](/docs/v1.10/components/vectors) - [Amazon S3 Vectors Engine](/docs/v1.10/components/vectors/s3_vectors): Amazon S3 Vectors Engine Documentation - [Views](/docs/v1.10/components/views): Documentation for defining Views in Spice - [Workers Overview](/docs/v1.10/components/workers): Detailed documentation for workers in the Spice runtime. - [Deployment](/docs/v1.10/deployment): Deploy Spice.ai in your environment - [Deployment Architectures](/docs/v1.10/deployment/architectures): Spice.ai Open Source Deployment architectures - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/v1.10/deployment/architectures/cluster): Deploying Spice as a cluster - [Cloud Hosted](/docs/v1.10/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Microservice Deployment (Single or Multiple Replicas)](/docs/v1.10/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/v1.10/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/v1.10/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/v1.10/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/v1.10/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [AWS Integrations](/docs/v1.10/deployment/aws/integrations): Complete guide to Spice.ai integrations with Amazon Web Services, including data connectors, AI models, vector stores, and secret management. - [Spice Cloud Platform Deployment](/docs/v1.10/deployment/cloud): Guide to deploying data and AI applications using the managed Spice Cloud Platform - [Docker - Kubernetes](/docs/v1.10/deployment/docker): Running Spice.ai as Docker container - [Docker Sandbox Guide - v1.3.0](/docs/v1.10/deployment/docker/sandbox): Migrating to v1.3.0 - [Helm - Kubernetes](/docs/v1.10/deployment/kubernetes): Deploy Spice.ai in Kubernetes using Helm. - [Frequently Asked Questions](/docs/v1.10/faq): Get answers to common questions about Spice.ai, including its features, differences from other tools, and use cases - [Features](/docs/v1.10/features): Features - [Caching](/docs/v1.10/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/v1.10/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Data Acceleration](/docs/v1.10/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/v1.10/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/v1.10/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Indexes](/docs/v1.10/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/v1.10/features/data-acceleration/partitioning): Partitioning for accelerated datasets - [Caching Refresh Mode](/docs/v1.10/features/data-acceleration/refresh-modes/caching): Learn how to use caching refresh mode for HTTP-based datasets - [Snapshots](/docs/v1.10/features/data-acceleration/snapshots): Bootstrap file-mode accelerations from managed snapshots to eliminate cold starts. - [Data Ingestion](/docs/v1.10/features/data-ingestion): Learn how to ingest data in Spice. - [Distributed Query](/docs/v1.10/features/distributed-query): Learn how to run Spice in distributed mode for larger scale queries. - [Embedding Datasets](/docs/v1.10/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Large Language Models](/docs/v1.10/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models](/docs/v1.10/features/large-language-models/evals): Learn how Spice evaluates, tracks, compares, and improves language model performance for specific tasks - [Model Context Protocol (MCP)](/docs/v1.10/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/v1.10/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/v1.10/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/v1.10/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/v1.10/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/v1.10/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/v1.10/features/machine-learning-models): Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform. - [Observability & Monitoring](/docs/v1.10/features/observability): Monitor Spice with Prometheus metrics, OpenTelemetry, and distributed tracing. - [Component Metrics](/docs/v1.10/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/v1.10/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Search Functionality](/docs/v1.10/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/v1.10/features/search/full-text): Learn how Spice can perform full text search - [Vector-Based Search](/docs/v1.10/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/v1.10/features/semantic-model): Learn how to define and use semantic data models with Spice. - [Web Search](/docs/v1.10/features/web-search): Learn how Spice can perform web search - [Getting started with Spice.ai OSS](/docs/v1.10/getting-started): Get started with Spice in 5 minutes - [Community Data](/docs/v1.10/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/v1.10/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/v1.10/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. - [Spice.ai OSS Installation](/docs/v1.10/installation): Instructions for installing Spice.ai OSS - [Intelligent Applications](/docs/v1.10/intelligent-applications): Building intelligent data and AI-driven applications with Spice.ai - [Monitoring](/docs/v1.10/monitoring): Monitoring Spice.ai deployments - [Datadog](/docs/v1.10/monitoring/datadog): Monitoring Spice with Datadog - [Grafana & Prometheus](/docs/v1.10/monitoring/grafana): Monitoring Spice instances with Grafana & Prometheus - [Zipkin Integration](/docs/v1.10/monitoring/zipkin): Learn how to integrate Spice with Zipkin tracing. - [Spice.ai OSS Reference Docs](/docs/v1.10/reference): Reference documentation on the Spice API, CLI and Pod manifest syntax. - [Cron Schedules](/docs/v1.10/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/v1.10/reference/datatypes) - [Accelerator Data Types](/docs/v1.10/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/v1.10/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Duration](/docs/v1.10/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/v1.10/reference/file_format): Spice currently supports CSV, JSON, and Parquet data file-formats for data connectors that can read files from a file system or cloud object storage (i.e. s3//, file://, etc.). Support for Iceberg and other file-formats are on the roadmap. - [Managing Memory Usage](/docs/v1.10/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice.ai Open Source deployments. - [Models Grade Report](/docs/v1.10/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [YAML syntax for Spicepod manifests](/docs/v1.10/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/v1.10/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/v1.10/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/v1.10/reference/spicepod/embeddings): Embeddings YAML reference - [evals](/docs/v1.10/reference/spicepod/evals): Evaluations YAML reference - [Reserved Keywords](/docs/v1.10/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/v1.10/reference/spicepod/models): Models YAML reference - [Runtime](/docs/v1.10/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/v1.10/reference/spicepod/tools): Tools YAML reference - [Views](/docs/v1.10/reference/spicepod/views): Views YAML reference - [Workers](/docs/v1.10/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/v1.10/reference/sql): This section provides a comprehensive reference for SQL support in Spice.ai, including syntax, data types, operators, functions, and system features. The reference is organized by topic for ease of navigation. - [Aggregate Functions](/docs/v1.10/reference/sql/aggregate_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. Note, when using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [AI Functions](/docs/v1.10/reference/sql/ai): AI functions in Spice provide direct integration with large language models (LLMs) and embedding models within SQL queries. These functions process text through configured model providers and return generated responses or vector embeddings. - [DML (Data Manipulation Language)](/docs/v1.10/reference/sql/dml): Data Manipulation Language (DML) statements for inserting and modifying data in Spice. - [Explain](/docs/v1.10/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/v1.10/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [JSON Functions and Operators](/docs/v1.10/reference/sql/json): Reference for JSON functions and operators in Spice SQL - [Operators](/docs/v1.10/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/v1.10/reference/sql/prepared_statements): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Scalar Functions](/docs/v1.10/reference/sql/scalar_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. Note, when using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [Search in SQL](/docs/v1.10/reference/sql/search): Reference for search functions and filtering in Spice SQL. - [SELECT](/docs/v1.10/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/v1.10/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/v1.10/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/v1.10/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. - [SDKs](/docs/v1.10/sdks): Connect to spice, using official Spice SDKs - [Dotnet SDK for Spice.ai](/docs/v1.10/sdks/dotnet): Connect to Spice using Spice Dotnet SDK - [Go SDK](/docs/v1.10/sdks/golang): Connect to spice using spice go SDK - [Java SDK](/docs/v1.10/sdks/java): Connect to Spice using Spice Java SDK - [JavaScript SDK](/docs/v1.10/sdks/javascript): Connect to spice using Spice.js SDK - [Python SDK](/docs/v1.10/sdks/python): Connect to spice using spice python SDK - [Rust SDK](/docs/v1.10/sdks/rust): Connect to spice using spice rust SDK - [Troubleshooting Spice](/docs/v1.10/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Spice.ai Use Cases](/docs/v1.10/use-cases): Use cases for Spice.ai' - [AI Applications and Agents](/docs/v1.10/use-cases/ai): AI Applications and Agents - [Agentic AI Applications and Agents](/docs/v1.10/use-cases/ai/agentic-apps): Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. - [Edge-Enabled AI Applications and Agents](/docs/v1.10/use-cases/ai/edge-ai): Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases. - [Federated MCP Client for Distributed Tool Ecosystems](/docs/v1.10/use-cases/ai/federated-mcp-server): Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, enhancing threat analysis. - [Object-Store Based SQL Query, Search, and LLM Inference Engine](/docs/v1.10/use-cases/ai/object-store-ai-engine): Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. - [Real-Time Decision-Making for Intelligent Applications](/docs/v1.10/use-cases/ai/real-time-decision-making): Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. - [Tool-Augmented AI with Model Context Protocol Server](/docs/v1.10/use-cases/ai/tool-calling-ai): Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. - [Data Federation, Acceleration, and SQL Query](/docs/v1.10/use-cases/data): Data Federation, Acceleration, and SQL Query - [Application Resilience and Performance Optimization](/docs/v1.10/use-cases/data/application-resilence-and-acceleration): Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance. - [Data Mesh for Unified Data Access](/docs/v1.10/use-cases/data/data-mesh): Spice.ai enables unified data access across disparate sources for health-tech applications, fostering a data mesh architecture. - [Database CDN for Enhanced Performance](/docs/v1.10/use-cases/data/database-cdn): Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience. - [ETL-free Workflows and Data Migrations](/docs/v1.10/use-cases/data/etl-free-workflows): Spice.ai enables data migrations and workflows without ETL by federating legacy and modern systems for seamless transitions. - [Object-Store Data Engine](/docs/v1.10/use-cases/data/object-store-data-engine): Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses. - [Reverse-ETL for Operational Workflows](/docs/v1.10/use-cases/data/reverse-etl): Spice.ai serves enriched data from warehouses to operational systems for real-time actions, eliminating complex ETL pipelines. - [Retrieval-Augmented-Generation (RAG)](/docs/v1.10/use-cases/rag): Retrieval-Augmented-Generation (RAG) - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/v1.10/use-cases/rag/applications): Use Spice for Retrieval-Augmented-Generation (RAG) - [Retrieval-Augmented Generation for AI-Powered Reporting](/docs/v1.10/use-cases/rag/reporting): Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. - [Search & Retrieval](/docs/v1.10/use-cases/search): Search & Retrieval - [Simplifying Real-Time Data Collection and Search](/docs/v1.10/use-cases/search/data-collection-and-search): Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic. - [Enterprise Search and Retrieval](/docs/v1.10/use-cases/search/enterprise-search): Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities. - [Object-Store Native Search Engine](/docs/v1.10/use-cases/search/object-store-search-engine): Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search. ### v1.11 Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven applications and AI agents. Learn about data federation, acceleration, RAG, and building intelligent apps. - [Spice.ai Open Source](/docs/v1.11): Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven applications and AI agents. Learn about data federation, acceleration, RAG, and building intelligent apps. - [Tags](/docs/v1.11/tags) - [One doc tagged with "Acknowledgements"](/docs/v1.11/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [2 docs tagged with "ADBC"](/docs/v1.11/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/v1.11/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Arrow"](/docs/v1.11/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/v1.11/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/v1.11/tags/auth): Authentication and authorization mechanisms. - [One doc tagged with "Authentication"](/docs/v1.11/tags/authentication): User authentication methods and security protocols. - [3 docs tagged with "Azure"](/docs/v1.11/tags/azure): Microsoft Azure cloud services and integrations. - [One doc tagged with "Blob Storage"](/docs/v1.11/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Caching"](/docs/v1.11/tags/caching): Data caching strategies and performance optimization. - [5 docs tagged with "Catalogs"](/docs/v1.11/tags/catalogs): Data catalog connectors and metadata management. - [One doc tagged with "Cayenne"](/docs/v1.11/tags/cayenne): Cayenne (Vortex) data accelerator built on Vortex columnar format. - [2 docs tagged with "CLI"](/docs/v1.11/tags/cli): Spice command-line interface commands and usage. - [5 docs tagged with "Component Metrics"](/docs/v1.11/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/v1.11/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/v1.11/tags/configuration): System configuration files and runtime settings. - [One doc tagged with "C#"](/docs/v1.11/tags/csharp): C# programming language topics. - [2 docs tagged with "Data Accelerators"](/docs/v1.11/tags/data-accelerators): Data acceleration engines for high-performance query execution. - [20 docs tagged with "Data Connectors"](/docs/v1.11/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/v1.11/tags/data-lake): Data lake architectures and storage solutions. - [2 docs tagged with "Databricks"](/docs/v1.11/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/v1.11/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debezium"](/docs/v1.11/tags/debezium): Debezium change data capture integration. - [One doc tagged with "Debugging"](/docs/v1.11/tags/debugging): Debugging techniques and troubleshooting tools. - [2 docs tagged with "Delta Lake"](/docs/v1.11/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/v1.11/tags/dependencies): Third-party dependencies and software requirements. - [3 docs tagged with "Deployment"](/docs/v1.11/tags/deployment): Production deployment using Docker and Kubernetes. - [2 docs tagged with "Docker"](/docs/v1.11/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with ".NET"](/docs/v1.11/tags/dotnet): .NET framework topics. - [One doc tagged with "DynamoDB"](/docs/v1.11/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [3 docs tagged with "Embeddings"](/docs/v1.11/tags/embeddings): Vector embeddings and semantic similarity operations. - [One doc tagged with "Evaluation"](/docs/v1.11/tags/evaluation): Model evaluation metrics and assessment frameworks. - [5 docs tagged with "Features"](/docs/v1.11/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/v1.11/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "Getting Started"](/docs/v1.11/tags/getting-started): Installation guides and quickstart tutorials. - [One doc tagged with "GitHub"](/docs/v1.11/tags/github): GitHub repository data connector and integration. - [2 docs tagged with "Glue"](/docs/v1.11/tags/glue): AWS Glue ETL service and data catalog integration. - [One doc tagged with "Go"](/docs/v1.11/tags/go): Go programming language topics. - [One doc tagged with "Golang"](/docs/v1.11/tags/golang): Go (Golang) programming language topics. - [2 docs tagged with "Iceberg"](/docs/v1.11/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/v1.11/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/v1.11/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "Java"](/docs/v1.11/tags/java): Java programming language topics. - [One doc tagged with "JavaScript"](/docs/v1.11/tags/javascript): JavaScript programming language topics. - [One doc tagged with "JDBC"](/docs/v1.11/tags/jdbc): Java Database Connectivity driver configuration. - [One doc tagged with "Kafka"](/docs/v1.11/tags/kafka): Apache Kafka event streaming platform integration. - [2 docs tagged with "Kubernetes"](/docs/v1.11/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "libSQL"](/docs/v1.11/tags/libsql): libSQL database topics. - [One doc tagged with "Logging"](/docs/v1.11/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/v1.11/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/v1.11/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/v1.11/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/v1.11/tags/memory): In-memory data connectors and volatile storage. - [13 docs tagged with "Models"](/docs/v1.11/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MongoDB"](/docs/v1.11/tags/mongodb): MongoDB NoSQL document database integration. - [One doc tagged with "MySQL"](/docs/v1.11/tags/mysql): MySQL relational database integration. - [One doc tagged with "Node.js"](/docs/v1.11/tags/nodejs): Node.js runtime environment topics. - [One doc tagged with "NoSQL"](/docs/v1.11/tags/nosql): NoSQL database systems and document stores. - [One doc tagged with "Observability"](/docs/v1.11/tags/observability): Monitoring, tracing, and metrics for system visibility. - [One doc tagged with "ODBC"](/docs/v1.11/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/v1.11/tags/open-source): Open source licenses and community contributions. - [One doc tagged with "OpenAI"](/docs/v1.11/tags/openai): OpenAI API integration and GPT model access. - [One doc tagged with "Oracle"](/docs/v1.11/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/v1.11/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/v1.11/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/v1.11/tags/parameters): Configuration parameters and runtime options. - [2 docs tagged with "Performance"](/docs/v1.11/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/v1.11/tags/persistence): Data persistence and durable storage mechanisms. - [One doc tagged with "Python"](/docs/v1.11/tags/python): Python programming language topics. - [3 docs tagged with "Query"](/docs/v1.11/tags/query): SQL query execution, parameterized queries, and prepared statements. - [4 docs tagged with "Reference"](/docs/v1.11/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/v1.11/tags/relational): Relational database systems and SQL operations. - [One doc tagged with "Runtime"](/docs/v1.11/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Rust"](/docs/v1.11/tags/rust): Rust programming language topics. - [One doc tagged with "S3"](/docs/v1.11/tags/s3): Amazon S3 object storage integration. - [One doc tagged with "S3 Express"](/docs/v1.11/tags/s3-express): Amazon S3 Express One Zone storage. - [One doc tagged with "Sandbox"](/docs/v1.11/tags/sandbox): Sandbox environments and isolated testing. - [One doc tagged with "ScyllaDB"](/docs/v1.11/tags/scylladb): ScyllaDB database integration. - [6 docs tagged with "SDK"](/docs/v1.11/tags/sdk): Software Development Kit topics and usage. - [5 docs tagged with "Search"](/docs/v1.11/tags/search): Vector search, semantic search, and ranking capabilities. - [2 docs tagged with "Security"](/docs/v1.11/tags/security): Security features and data protection mechanisms. - [2 docs tagged with "SpiceAI"](/docs/v1.11/tags/spiceai): SpiceAI cloud platform and managed services. - [5 docs tagged with "Spicepod"](/docs/v1.11/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [5 docs tagged with "SQL"](/docs/v1.11/tags/sql): SQL query language support and database operations. - [3 docs tagged with "Tools"](/docs/v1.11/tags/tools): Development tools and utility integrations. - [2 docs tagged with "Tracing"](/docs/v1.11/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/v1.11/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Turso"](/docs/v1.11/tags/turso): Turso data accelerator integration. - [One doc tagged with "Unity Catalog"](/docs/v1.11/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [One doc tagged with "Views"](/docs/v1.11/tags/views): Virtual views and data transformation layers. - [One doc tagged with "Vortex"](/docs/v1.11/tags/vortex): Vortex columnar file format and storage engine. - [6 docs tagged with "Write"](/docs/v1.11/tags/write): Data connectors and catalogs that support write operations. - [One doc tagged with "YAML"](/docs/v1.11/tags/yaml): YAML configuration syntax and file formats. - [One doc tagged with "Zipkin"](/docs/v1.11/tags/zipkin): Zipkin distributed tracing system integration. - [Open Source Acknowledgements](/docs/v1.11/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: - [Spice.ai API Reference](/docs/v1.11/api): Spice.ai API reference including HTTP REST API, Arrow Flight SQL, JDBC, ODBC, ADBC connectors, authentication, and TLS configuration. - [ADBC: Arrow Database Connectivity](/docs/v1.11/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/v1.11/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/v1.11/api/auth): Authentication documentation - [Generate Package](/docs/v1.11/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/v1.11/api/HTTP/get-catalogs): List Catalogs - [Get Iceberg API config](/docs/v1.11/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/v1.11/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/v1.11/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [ML Prediction](/docs/v1.11/api/HTTP/get-model-predict): Make a ML prediction using a specific model. - [List Models](/docs/v1.11/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [List Spicepods](/docs/v1.11/api/HTTP/get-spicepods): Get a list of spicepods and their summary details. - [Check Runtime Status](/docs/v1.11/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Check Namespace exists](/docs/v1.11/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Evals](/docs/v1.11/api/HTTP/list): Return all evals available to run in the runtime. - [Send message to MCP server](/docs/v1.11/api/HTTP/mcp-event): Send message to the MCP endoint, for a given session. - [Establish an MCP SSE Connection](/docs/v1.11/api/HTTP/operation-id): Initiates a Server-Sent Events (SSE) connection using the Model Context Protocol (MCP) to interact with Spice tools. - [Update Refresh SQL](/docs/v1.11/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Run Tool](/docs/v1.11/api/HTTP/post): The request body and JSON response formats match the tool’s specification. - [Batch ML Predictions](/docs/v1.11/api/HTTP/post-batch-predict): Perform a batch of ML predictions, using multiple models, in one request. This is useful for ensembling or A/B testing different models. - [Create Chat Completion](/docs/v1.11/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/v1.11/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/v1.11/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Run Eval](/docs/v1.11/api/HTTP/post-eval): Evaluate a model against a eval spice specification - [Text-to-SQL (NSQL)](/docs/v1.11/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [Search](/docs/v1.11/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/v1.11/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/v1.11/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [runtime](/docs/v1.11/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/v1.11/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/v1.11/api/odbc): ODBC API Documentation - [API Overview](/docs/v1.11/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/v1.11/api/tls): Encryption in transit with TLS documentation - [Spice.ai CLI Reference](/docs/v1.11/cli): Complete CLI reference for Spice.ai including commands to create, manage Spicepods, run queries, and interact with the Spice runtime. - [Spice.ai OSS CLI command reference](/docs/v1.11/cli/reference): Spice CLI command reference - [add](/docs/v1.11/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/v1.11/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/v1.11/cli/reference/chat): spice chat CLI documentation - [Completion](/docs/v1.11/cli/reference/completion): Generate the autocompletion script for spice for the specified shell. - [connect](/docs/v1.11/cli/reference/connect): Connect to an app on the Spice.ai Cloud Platform. - [dataset](/docs/v1.11/cli/reference/dataset): Configure a Spice dataset. - [datasets](/docs/v1.11/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [init](/docs/v1.11/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/v1.11/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/v1.11/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/v1.11/cli/reference/models): Lists models loaded by the Spice runtime - [pods](/docs/v1.11/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [query](/docs/v1.11/cli/reference/query): Submit an async query or start an interactive async query REPL against the Spice runtime's distributed query engine. - [refresh](/docs/v1.11/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/v1.11/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/v1.11/cli/reference/search): Performs embeddings-based searches across search configured datasets. Note: Search requires the ai feature to be installed. - [sql](/docs/v1.11/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/v1.11/cli/reference/status): Spice runtime status - [trace](/docs/v1.11/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/v1.11/cli/reference/upgrade): Upgrades the Spice CLI & Runtime to the latest release - [version](/docs/v1.11/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/v1.11/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels - [Clients and Tools](/docs/v1.11/clients): Client and tools for connecting to Spice - [DBeaver](/docs/v1.11/clients/dbeaver): Configure DBeaver to query Spice via JDBC - [JetBrains DataGrip](/docs/v1.11/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Microsoft Power BI Connector](/docs/v1.11/clients/powerbi): Use Microsoft Power BI to access, visualize and analyze Spice datasets. - [Apache Superset](/docs/v1.11/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/v1.11/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. - [How Spice Compares](/docs/v1.11/comparison): Compare Spice.ai with data platforms (Databricks, Snowflake), query engines (Trino, Dremio, ClickHouse), vector databases (Turbopuffer, LanceDB), search engines (Elasticsearch), and AI frameworks (LangChain, LlamaIndex, Ollama). - [Spice.ai Runtime Components](/docs/v1.11/components): Configure Spice.ai runtime components including data connectors, data accelerators, catalog connectors, model providers, embedding models, and secret stores. - [Catalog Connectors](/docs/v1.11/components/catalogs): Connect to external catalog providers like Unity Catalog, Databricks, Iceberg, and AWS Glue for federated SQL query in Spice. - [Databricks Catalog Connector](/docs/v1.11/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [Glue Catalog Connector](/docs/v1.11/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/v1.11/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Spice.ai Catalog Connector](/docs/v1.11/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/v1.11/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Data Accelerators](/docs/v1.11/components/data-accelerators): Data acceleration engines for local materialization and query acceleration in Spice - [In-Memory Arrow Data Accelerator](/docs/v1.11/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [Spice Cayenne Data Accelerator](/docs/v1.11/components/data-accelerators/cayenne): Spice Cayenne Data Accelerator (Vortex) Documentation - [DuckDB Data Accelerator](/docs/v1.11/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [PostgreSQL Data Accelerator](/docs/v1.11/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [SQLite Data Accelerator](/docs/v1.11/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [Turso Data Accelerator](/docs/v1.11/components/data-accelerators/turso): Turso (libSQL) Data Accelerator Documentation - [Data Connectors](/docs/v1.11/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/v1.11/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ClickHouse Data Connector](/docs/v1.11/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Databricks Data Connector](/docs/v1.11/components/data-connectors/databricks): Databricks Data Connector Documentation - [Debezium Data Connector](/docs/v1.11/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/v1.11/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Dremio Data Connector](/docs/v1.11/components/data-connectors/dremio): Dremio Data Connector Documentation - [DuckDB Data Connector](/docs/v1.11/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DynamoDB Data Connector](/docs/v1.11/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [File Data Connector](/docs/v1.11/components/data-connectors/file): File Data Connector Documentation - [Flight SQL Data Connector](/docs/v1.11/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/v1.11/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GitHub Data Connector](/docs/v1.11/components/data-connectors/github): GitHub Data Connector Documentation - [Glue Data Connector](/docs/v1.11/components/data-connectors/glue): Glue Data Connector Documentation - [GraphQL Data Connector](/docs/v1.11/components/data-connectors/graphql): GraphQL Data Connector Documentation - [HTTP(s) Data Connector](/docs/v1.11/components/data-connectors/https): HTTP(s) Data Connector Documentation - [Iceberg Data Connector](/docs/v1.11/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/v1.11/components/data-connectors/imap): IMAP Data Connector Documentation - [Kafka Data Connector](/docs/v1.11/components/data-connectors/kafka): Kafka Data Connector Documentation - [Localpod Data Connector](/docs/v1.11/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/v1.11/components/data-connectors/memory): Memory Data Connector Documentation - [MongoDB Data Connector](/docs/v1.11/components/data-connectors/mongodb): MongoDB Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/v1.11/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [MySQL Data Connector](/docs/v1.11/components/data-connectors/mysql): MySQL Data Connector Documentation - [NFS Data Connector](/docs/v1.11/components/data-connectors/nfs): NFS Data Connector Documentation - [ODBC Data Connector](/docs/v1.11/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/v1.11/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/v1.11/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [Amazon Redshift Data Connector](/docs/v1.11/components/data-connectors/redshift): Connect to Amazon Redshift using the PostgreSQL connector in Spice. - [S3 Data Connector](/docs/v1.11/components/data-connectors/s3): S3 Data Connector Documentation - [ScyllaDB Data Connector](/docs/v1.11/components/data-connectors/scylladb): ScyllaDB Data Connector Documentation - [SharePoint Data Connector](/docs/v1.11/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [SMB Data Connector](/docs/v1.11/components/data-connectors/smb): SMB Data Connector Documentation - [Snowflake Data Connector](/docs/v1.11/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/v1.11/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/v1.11/components/data-connectors/spiceai): Spice.ai Data Connector Documentation - [Embedding Models](/docs/v1.11/components/embeddings): Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. - [Azure OpenAI Embedding Models](/docs/v1.11/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/v1.11/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/v1.11/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [Google AI Embedding Models](/docs/v1.11/components/embeddings/google): To use a hosted Google AI embedding model, specify the google path in the from field of your configuration. - [HuggingFace Text Embedding Models](/docs/v1.11/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Local Filesystem Embedding Models](/docs/v1.11/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [Model2Vec Embedding Models](/docs/v1.11/components/embeddings/model2vec): Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning. - [OpenAI (or Compatible) Embedding Models](/docs/v1.11/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [Model Providers](/docs/v1.11/components/models): Overview of supported model providers for ML and LLMs in Spice. - [Anthropic Models](/docs/v1.11/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/v1.11/components/models/azure): Instructions for using Azure OpenAI models - [Amazon Bedrock Models](/docs/v1.11/components/models/bedrock): How to use Amazon Bedrock models with Spice. - [Databricks Model Provider](/docs/v1.11/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/v1.11/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [Google AI Models](/docs/v1.11/components/models/google): Instructions for using language models hosted on Google AI with Spice. - [HuggingFace](/docs/v1.11/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [OpenAI (or Compatible) Language Models](/docs/v1.11/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [Perplexity Models](/docs/v1.11/components/models/perplexity): Instructions for using language models hosted on Perplexity with Spice. - [Spice Cloud Platform](/docs/v1.11/components/models/spiceai): Instructions for using models hosted on the Spice Cloud Platform with Spice. - [xAI Models](/docs/v1.11/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/v1.11/components/secret-stores): Configure secret stores to manage sensitive data like passwords, tokens, and API keys. - [AWS Secrets Manager Secret Store](/docs/v1.11/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Environment Secret Store](/docs/v1.11/components/secret-stores/env): Environment Variables Secret Store Documentation - [Keyring Secret Store](/docs/v1.11/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/v1.11/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/v1.11/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/v1.11/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Server-Sent Events (SSE) protocol. - [Web Search Tool](/docs/v1.11/components/tools/websearch): Configure web search tools for LLMs to search the internet using Perplexity and other search engines. - [Vector Engines](/docs/v1.11/components/vectors): Configure vector engines for efficient embedding storage and similarity search in Spice. - [Amazon S3 Vectors Engine](/docs/v1.11/components/vectors/s3_vectors): Amazon S3 Vectors Engine Documentation - [Views](/docs/v1.11/components/views): Documentation for defining Views in Spice - [Workers Overview](/docs/v1.11/components/workers): Detailed documentation for workers in the Spice runtime. - [Spice.ai Deployment Guide](/docs/v1.11/deployment): Deploy Spice.ai in your environment using Docker, Kubernetes, AWS, or the Spice Cloud Platform. Learn about sidecar, microservice, tiered, and cluster deployment architectures. - [Deployment Architectures](/docs/v1.11/deployment/architectures): Explore Spice deployment architectures including sidecar, microservice, tiered, sharded, and cluster configurations. - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/v1.11/deployment/architectures/cluster): Deploying Spice as a cluster - [Cloud Hosted](/docs/v1.11/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Hybrid Deployment](/docs/v1.11/deployment/architectures/hybrid): Deploying Spice with sidecar caching backed by a centralized cluster for acceleration, distributed query, and ingestion. - [Microservice Deployment (Single or Multiple Replicas)](/docs/v1.11/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/v1.11/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/v1.11/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/v1.11/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/v1.11/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [AWS Integrations](/docs/v1.11/deployment/aws/integrations): Complete guide to Spice.ai integrations with Amazon Web Services, including data connectors, AI models, vector stores, and secret management. - [Spice Cloud Platform Deployment](/docs/v1.11/deployment/cloud): Guide to deploying data and AI applications using the managed Spice Cloud Platform - [Docker - Kubernetes](/docs/v1.11/deployment/docker): Running Spice.ai as Docker container - [Docker Sandbox Guide - v1.3.0](/docs/v1.11/deployment/docker/sandbox): Migrating to v1.3.0 - [Helm - Kubernetes](/docs/v1.11/deployment/kubernetes): Deploy Spice.ai in Kubernetes using Helm. - [Spice.ai FAQ](/docs/v1.11/faq): Answers to frequently asked questions about Spice.ai including features, use cases, differences from Trino/Presto/Dremio, federated queries, caching, and AI capabilities. - [Spice.ai Features](/docs/v1.11/features): Explore Spice.ai features including data federation, data acceleration, caching, search, LLM integration, embeddings, observability, and more for building data-driven AI applications. - [Caching](/docs/v1.11/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/v1.11/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Data Acceleration](/docs/v1.11/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/v1.11/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/v1.11/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Hash Index for Arrow Acceleration](/docs/v1.11/features/data-acceleration/hash-index): Learn how to use hash indexes for O(1) point lookups on Arrow-accelerated datasets. - [Indexes](/docs/v1.11/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/v1.11/features/data-acceleration/partitioning): Partitioning for accelerated datasets - [Caching Refresh Mode](/docs/v1.11/features/data-acceleration/refresh-modes/caching): Learn how to use caching refresh mode for HTTP-based datasets - [Snapshots](/docs/v1.11/features/data-acceleration/snapshots): Bootstrap file-mode accelerations from managed snapshots to eliminate cold starts. - [Data Ingestion](/docs/v1.11/features/data-ingestion): Learn how to ingest data in Spice. - [Distributed Query](/docs/v1.11/features/distributed-query): Learn how to run Spice in distributed mode for larger scale queries. - [Embedding Datasets](/docs/v1.11/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Large Language Models](/docs/v1.11/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models](/docs/v1.11/features/large-language-models/evals): Learn how Spice evaluates, tracks, compares, and improves language model performance for specific tasks - [Model Context Protocol (MCP)](/docs/v1.11/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/v1.11/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/v1.11/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/v1.11/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/v1.11/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/v1.11/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/v1.11/features/machine-learning-models): Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform. - [Observability & Monitoring](/docs/v1.11/features/observability): Monitor Spice with Prometheus metrics, OpenTelemetry, and distributed tracing. - [Component Metrics](/docs/v1.11/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/v1.11/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Parameterized Queries](/docs/v1.11/features/query-federation/parameterized-queries): Learn how to use prepared statements and parameterized queries in Spice for improved security and performance. - [URL Tables](/docs/v1.11/features/query-federation/url-tables): Query object store files directly using URLs without pre-registering datasets - [Search Functionality](/docs/v1.11/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/v1.11/features/search/full-text): Learn how Spice can perform full text search - [Vector-Based Search](/docs/v1.11/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/v1.11/features/semantic-model): Learn how to define and use semantic data models with Spice. - [Web Search](/docs/v1.11/features/web-search): Learn how Spice can perform web search - [Getting Started with Spice.ai OSS](/docs/v1.11/getting-started): Get started with Spice.ai in 5 minutes. Install the CLI, connect to datasets, run SQL queries, and use AI models with OpenAI-compatible APIs. - [Community Data](/docs/v1.11/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/v1.11/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/v1.11/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. - [Install Spice.ai OSS](/docs/v1.11/installation): Install Spice.ai OSS on macOS, Linux, Windows, or WSL using the install script, Homebrew, PowerShell, or direct download from GitHub releases. - [Building Intelligent AI Applications with Spice.ai](/docs/v1.11/intelligent-applications): Learn how to build intelligent, data-driven AI applications and agents with Spice.ai. Explore patterns for RAG, LLM integration, and real-time AI inference. - [Monitoring](/docs/v1.11/monitoring): Monitor Spice.ai deployments with Datadog, Grafana, Prometheus, and Zipkin integrations. - [Datadog](/docs/v1.11/monitoring/datadog): Monitoring Spice with Datadog - [Grafana & Prometheus](/docs/v1.11/monitoring/grafana): Monitoring Spice instances with Grafana & Prometheus - [Zipkin Integration](/docs/v1.11/monitoring/zipkin): Learn how to integrate Spice with Zipkin tracing. - [Spice.ai OSS Reference Docs](/docs/v1.11/reference): Reference documentation for Spice.ai including API reference, CLI commands, Spicepod configuration syntax, SQL reference, and data type specifications. - [Cron Schedules](/docs/v1.11/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/v1.11/reference/datatypes): Spice uses Apache Arrow data types internally, providing consistent type handling across different data sources and accelerators. This section documents how Arrow types map to specific accelerators and object store formats. - [Accelerator Data Types](/docs/v1.11/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/v1.11/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Spice Runtime Distributions](/docs/v1.11/reference/distributions): Distribution variants of the Spice runtime for different use cases and deployment scenarios, including data-only, GPU-accelerated, NAS, and allocator variants. - [Duration](/docs/v1.11/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/v1.11/reference/file_format): File-based data connectors — including s3//, file//, sftp://, and others — support multiple structured and document file formats. This page details the format-specific parameters available for each. - [Managing Memory Usage](/docs/v1.11/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice deployments. - [Models Grade Report](/docs/v1.11/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [Performance Tuning](/docs/v1.11/reference/performance-tuning): Comprehensive guide to optimizing query performance, acceleration, and resource utilization in Spice deployments. - [YAML syntax for Spicepod manifests](/docs/v1.11/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/v1.11/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/v1.11/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/v1.11/reference/spicepod/embeddings): Embeddings YAML reference - [evals](/docs/v1.11/reference/spicepod/evals): Evaluations YAML reference - [Reserved Keywords](/docs/v1.11/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/v1.11/reference/spicepod/models): Models YAML reference - [Runtime](/docs/v1.11/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/v1.11/reference/spicepod/tools): Tools YAML reference - [Views](/docs/v1.11/reference/spicepod/views): Views YAML reference - [Workers](/docs/v1.11/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/v1.11/reference/sql): Complete SQL reference for Spice.ai including SELECT syntax, subqueries, DML statements, aggregate functions, AI functions, JSON operators, and search capabilities. - [Aggregate Functions](/docs/v1.11/reference/sql/aggregate_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [AI Functions](/docs/v1.11/reference/sql/ai): AI functions in Spice provide direct integration with large language models (LLMs) and embedding models within SQL queries. These functions process text through configured model providers and return generated responses or vector embeddings. - [DML (Data Manipulation Language)](/docs/v1.11/reference/sql/dml): Data Manipulation Language (DML) statements for inserting and modifying data in Spice. - [Explain](/docs/v1.11/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/v1.11/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [JSON Functions and Operators](/docs/v1.11/reference/sql/json): Reference for JSON functions and operators in Spice SQL - [Operators](/docs/v1.11/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/v1.11/reference/sql/prepared_statements): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Scalar Functions](/docs/v1.11/reference/sql/scalar_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [Search in SQL](/docs/v1.11/reference/sql/search): Reference for search functions and filtering in Spice SQL. - [SELECT](/docs/v1.11/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/v1.11/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/v1.11/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/v1.11/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. - [SDKs](/docs/v1.11/sdks): Connect to Spice using official SDKs - [Dotnet SDK](/docs/v1.11/sdks/dotnet): Connect to Spice using the Dotnet SDK - [Go SDK](/docs/v1.11/sdks/golang): Connect to Spice using the Go SDK - [Java SDK](/docs/v1.11/sdks/java): Connect to Spice using the Java SDK - [JavaScript SDK](/docs/v1.11/sdks/javascript): Connect to Spice using the JavaScript SDK - [Python SDK](/docs/v1.11/sdks/python): Connect to Spice using the Python SDK - [Rust SDK](/docs/v1.11/sdks/rust): Connect to Spice using the Rust SDK - [Troubleshooting Spice](/docs/v1.11/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Spice.ai Use Cases](/docs/v1.11/use-cases): Discover how to use Spice.ai for data federation, reverse-ETL, database CDN, enterprise search, RAG, and building AI-powered applications and agents. - [AI Applications and Agents](/docs/v1.11/use-cases/ai): AI Applications and Agents - [Agentic AI Applications and Agents](/docs/v1.11/use-cases/ai/agentic-apps): Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. - [Edge-Enabled AI Applications and Agents](/docs/v1.11/use-cases/ai/edge-ai): Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases. - [Federated MCP Client for Distributed Tool Ecosystems](/docs/v1.11/use-cases/ai/federated-mcp-server): Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, improving threat analysis. - [Object-Store Based SQL Query, Search, and LLM Inference Engine](/docs/v1.11/use-cases/ai/object-store-ai-engine): Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. - [Real-Time Decision-Making for Intelligent Applications](/docs/v1.11/use-cases/ai/real-time-decision-making): Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. - [Tool-Augmented AI with Model Context Protocol Server](/docs/v1.11/use-cases/ai/tool-calling-ai): Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. - [Data Federation, Acceleration, and SQL Query](/docs/v1.11/use-cases/data): Data Federation, Acceleration, and SQL Query - [Application Resilience and Performance Optimization](/docs/v1.11/use-cases/data/application-resilence-and-acceleration): Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance. - [Data Mesh for Unified Data Access](/docs/v1.11/use-cases/data/data-mesh): Spice.ai enables unified data access across disparate sources for health-tech applications, fostering a data mesh architecture. - [Database CDN for Enhanced Performance](/docs/v1.11/use-cases/data/database-cdn): Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience. - [ETL-free Workflows and Data Migrations](/docs/v1.11/use-cases/data/etl-free-workflows): Spice.ai enables data migrations and workflows without ETL by federating legacy and modern systems for seamless transitions. - [Object-Store Data Engine](/docs/v1.11/use-cases/data/object-store-data-engine): Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses. - [Reverse-ETL for Operational Workflows](/docs/v1.11/use-cases/data/reverse-etl): Spice.ai serves enriched data from warehouses to operational systems for real-time actions, eliminating complex ETL pipelines. - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/v1.11/use-cases/rag): Use Spice for Retrieval-Augmented-Generation (RAG) - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/v1.11/use-cases/rag/applications): Use Spice for Retrieval-Augmented-Generation (RAG) - [Retrieval-Augmented Generation for AI-Powered Reporting](/docs/v1.11/use-cases/rag/reporting): Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. - [Search & Retrieval](/docs/v1.11/use-cases/search): Search & Retrieval - [Simplifying Real-Time Data Collection and Search](/docs/v1.11/use-cases/search/data-collection-and-search): Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic. - [Enterprise Search and Retrieval](/docs/v1.11/use-cases/search/enterprise-search): Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities. - [Object-Store Native Search Engine](/docs/v1.11/use-cases/search/object-store-search-engine): Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search. ### v1.5 Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Spice.ai Open Source](/docs/v1.5): Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Tags](/docs/v1.5/tags) - [One doc tagged with "Acknowledgements"](/docs/v1.5/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [One doc tagged with "ADBC"](/docs/v1.5/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/v1.5/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Arrow"](/docs/v1.5/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/v1.5/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/v1.5/tags/auth): Authentication and authorization mechanisms. - [One doc tagged with "Authentication"](/docs/v1.5/tags/authentication): User authentication methods and security protocols. - [2 docs tagged with "Azure"](/docs/v1.5/tags/azure): Microsoft Azure cloud services and integrations. - [One doc tagged with "Blob Storage"](/docs/v1.5/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Cache Control"](/docs/v1.5/tags/cache-control-policy): Cache control headers and directives. - [One doc tagged with "Caching"](/docs/v1.5/tags/caching): Data caching strategies and performance optimization. - [5 docs tagged with "Catalogs"](/docs/v1.5/tags/catalogs): Data catalog connectors and metadata management. - [2 docs tagged with "CLI"](/docs/v1.5/tags/cli): Spice command-line interface commands and usage. - [2 docs tagged with "Component Metrics"](/docs/v1.5/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/v1.5/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/v1.5/tags/configuration): System configuration files and runtime settings. - [One doc tagged with "Data Connector"](/docs/v1.5/tags/data-connector): Data connector tools and integrations. - [13 docs tagged with "Data Connectors"](/docs/v1.5/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/v1.5/tags/data-lake): Data lake architectures and storage solutions. - [2 docs tagged with "Databricks"](/docs/v1.5/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/v1.5/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debugging"](/docs/v1.5/tags/debugging): Debugging techniques and troubleshooting tools. - [2 docs tagged with "Delta Lake"](/docs/v1.5/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/v1.5/tags/dependencies): Third-party dependencies and software requirements. - [3 docs tagged with "Deployment"](/docs/v1.5/tags/deployment): Production deployment using Docker and Kubernetes. - [2 docs tagged with "Docker"](/docs/v1.5/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with "DynamoDB"](/docs/v1.5/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [3 docs tagged with "Embeddings"](/docs/v1.5/tags/embeddings): Vector embeddings and semantic similarity operations. - [One doc tagged with "Evaluation"](/docs/v1.5/tags/evaluation): Model evaluation metrics and assessment frameworks. - [One doc tagged with "Features"](/docs/v1.5/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/v1.5/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "Getting Started"](/docs/v1.5/tags/getting-started): Installation guides and quickstart tutorials. - [One doc tagged with "GitHub"](/docs/v1.5/tags/github): GitHub repository data connector and integration. - [One doc tagged with "Glue"](/docs/v1.5/tags/glue): AWS Glue ETL service and data catalog integration. - [One doc tagged with "Iceberg"](/docs/v1.5/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/v1.5/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/v1.5/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "Introduction"](/docs/v1.5/tags/introduction): Introductory guides and getting started. - [One doc tagged with "JDBC"](/docs/v1.5/tags/jdbc): Java Database Connectivity driver configuration. - [2 docs tagged with "Kubernetes"](/docs/v1.5/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "Logging"](/docs/v1.5/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/v1.5/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/v1.5/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/v1.5/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/v1.5/tags/memory): In-memory data connectors and volatile storage. - [12 docs tagged with "Models"](/docs/v1.5/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MySQL"](/docs/v1.5/tags/mysql): MySQL relational database integration. - [One doc tagged with "NoSQL"](/docs/v1.5/tags/nosql): NoSQL database systems and document stores. - [One doc tagged with "ODBC"](/docs/v1.5/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/v1.5/tags/open-source): Open source licenses and community contributions. - [One doc tagged with "OpenAI"](/docs/v1.5/tags/openai): OpenAI API integration and GPT model access. - [One doc tagged with "Oracle"](/docs/v1.5/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/v1.5/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/v1.5/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/v1.5/tags/parameters): Configuration parameters and runtime options. - [One doc tagged with "Performance"](/docs/v1.5/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/v1.5/tags/persistence): Data persistence and durable storage mechanisms. - [3 docs tagged with "Reference"](/docs/v1.5/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/v1.5/tags/relational): Relational database systems and SQL operations. - [One doc tagged with "Results Caching"](/docs/v1.5/tags/results-caching-feature): Query results caching configuration. - [One doc tagged with "Runtime"](/docs/v1.5/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Sandbox"](/docs/v1.5/tags/sandbox): Sandbox environments and isolated testing. - [4 docs tagged with "Search"](/docs/v1.5/tags/search): Vector search, semantic search, and ranking capabilities. - [One doc tagged with "Security"](/docs/v1.5/tags/security): Security features and data protection mechanisms. - [2 docs tagged with "SpiceAI"](/docs/v1.5/tags/spiceai): SpiceAI cloud platform and managed services. - [3 docs tagged with "Spicepod"](/docs/v1.5/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [One doc tagged with "Spicepods"](/docs/v1.5/tags/spicepods): Spicepod configuration and usage. - [2 docs tagged with "SQL"](/docs/v1.5/tags/sql): SQL query language support and database operations. - [3 docs tagged with "Tools"](/docs/v1.5/tags/tools): Development tools and utility integrations. - [One doc tagged with "Tracing"](/docs/v1.5/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/v1.5/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Unity Catalog"](/docs/v1.5/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [One doc tagged with "YAML"](/docs/v1.5/tags/yaml): YAML configuration syntax and file formats. - [Open Source Acknowledgements](/docs/v1.5/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: - [API](/docs/v1.5/api): API - [ADBC: Arrow Database Connectivity](/docs/v1.5/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/v1.5/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/v1.5/api/auth): Authentication documentation - [Generate Package](/docs/v1.5/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/v1.5/api/HTTP/get-catalogs): List Catalogs - [Get Iceberg API config](/docs/v1.5/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/v1.5/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/v1.5/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [ML Prediction](/docs/v1.5/api/HTTP/get-model-predict): Make a ML prediction using a specific model. - [List Models](/docs/v1.5/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [List Spicepods](/docs/v1.5/api/HTTP/get-spicepods): Get a list of spicepods and their details. In CSV format, it will return a summarised form. - [Check Runtime Status](/docs/v1.5/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Check Namespace exists](/docs/v1.5/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Evals](/docs/v1.5/api/HTTP/list): Return all evals available to run in the runtime. - [Send message to MCP server](/docs/v1.5/api/HTTP/mcp-event): Send message to the MCP endoint, for a given session. - [Establish an MCP SSE Connection](/docs/v1.5/api/HTTP/operation-id): Initiates a Server-Sent Events (SSE) connection using the Model Context Protocol (MCP) to interact with Spice tools. - [Update Refresh SQL](/docs/v1.5/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Run Tool](/docs/v1.5/api/HTTP/post): The request body and JSON response formats match the tool’s specification. - [Batch ML Predictions](/docs/v1.5/api/HTTP/post-batch-predict): Perform a batch of ML predictions, using multiple models, in one request. This is useful for ensembling or A/B testing different models. - [Create Chat Completion](/docs/v1.5/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/v1.5/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/v1.5/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Run Eval](/docs/v1.5/api/HTTP/post-eval): Evaluate a model against a eval spice specification - [Text-to-SQL (NSQL)](/docs/v1.5/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [Search](/docs/v1.5/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/v1.5/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/v1.5/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [runtime](/docs/v1.5/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/v1.5/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/v1.5/api/odbc): ODBC API Documentation - [API Overview](/docs/v1.5/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/v1.5/api/tls): Encryption in transit with TLS documentation - [Spice.ai OSS CLI documentation](/docs/v1.5/cli): Detailed documentation on the Spice.ai OSS CLI - [Spice.ai OSS CLI command reference](/docs/v1.5/cli/reference): Spice CLI command reference - [add](/docs/v1.5/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/v1.5/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/v1.5/cli/reference/chat): spice chat CLI documentation - [Completion](/docs/v1.5/cli/reference/completion): Generate the autocompletion script for spice for the specified shell. - [connect](/docs/v1.5/cli/reference/connect): Connect to an app on the Spice.ai Cloud Platform. - [dataset](/docs/v1.5/cli/reference/dataset): Configure a Spice dataset. - [datasets](/docs/v1.5/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [init](/docs/v1.5/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/v1.5/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/v1.5/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/v1.5/cli/reference/models): Lists models loaded by the Spice runtime - [pods](/docs/v1.5/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [refresh](/docs/v1.5/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/v1.5/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/v1.5/cli/reference/search): Performs embeddings-based searches across search configured datasets. Note: Search requires the ai feature to be installed. - [sql](/docs/v1.5/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/v1.5/cli/reference/status): Spice runtime status - [trace](/docs/v1.5/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/v1.5/cli/reference/upgrade): Upgrades the Spice CLI & Runtime to the latest release - [version](/docs/v1.5/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/v1.5/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels - [Clients and Tools](/docs/v1.5/clients): Client and tools - [Datadog](/docs/v1.5/clients/Datadog): Monitoring Spice with Datadog - [DBeaver](/docs/v1.5/clients/DBeaver): Configure DBeaver to query Spice via JDBC - [Grafana & Prometheus](/docs/v1.5/clients/grafana): Monitoring Spice instances with Grafana & Prometheus - [JetBrains DataGrip](/docs/v1.5/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Apache Superset](/docs/v1.5/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/v1.5/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. - [Runtime Components](/docs/v1.5/components): Runtime components' - [Catalog Connectors](/docs/v1.5/components/catalogs) - [Databricks Catalog Connector](/docs/v1.5/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [Glue Catalog Connector](/docs/v1.5/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/v1.5/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Spice.ai Catalog Connector](/docs/v1.5/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/v1.5/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Data Accelerators](/docs/v1.5/components/data-accelerators) - [In-Memory Arrow Data Accelerator](/docs/v1.5/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [DuckDB Data Accelerator](/docs/v1.5/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [PostgreSQL Data Accelerator](/docs/v1.5/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [SQLite Data Accelerator](/docs/v1.5/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [Data Connectors](/docs/v1.5/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/v1.5/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ClickHouse Data Connector](/docs/v1.5/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Databricks Data Connector](/docs/v1.5/components/data-connectors/databricks): Databricks Data Connector Documentation - [Debezium Data Connector](/docs/v1.5/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/v1.5/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Dremio Data Connector](/docs/v1.5/components/data-connectors/dremio): Dremio Data Connector Documentation - [DuckDB Data Connector](/docs/v1.5/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DynamoDB Data Connector](/docs/v1.5/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [File Data Connector](/docs/v1.5/components/data-connectors/file): File Data Connector Documentation - [Flight SQL Data Connector](/docs/v1.5/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/v1.5/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GitHub Data Connector](/docs/v1.5/components/data-connectors/github): GitHub Data Connector Documentation - [Glue Data Connector](/docs/v1.5/components/data-connectors/glue): Glue Data Connector Documentation - [GraphQL Data Connector](/docs/v1.5/components/data-connectors/graphql): GraphQL Data Connector Documentation - [HTTP(s) Data Connector](/docs/v1.5/components/data-connectors/https): HTTP(s) Data Connector Documentation - [Iceberg Data Connector](/docs/v1.5/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/v1.5/components/data-connectors/imap): IMAP Data Connector Documentation - [Localpod Data Connector](/docs/v1.5/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/v1.5/components/data-connectors/memory): Memory Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/v1.5/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [MySQL Data Connector](/docs/v1.5/components/data-connectors/mysql): MySQL Data Connector Documentation - [ODBC Data Connector](/docs/v1.5/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/v1.5/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/v1.5/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [S3 Data Connector](/docs/v1.5/components/data-connectors/s3): S3 Data Connector Documentation - [SharePoint Data Connector](/docs/v1.5/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [Snowflake Data Connector](/docs/v1.5/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/v1.5/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/v1.5/components/data-connectors/spiceai): Spice.ai Data Connector Documentation - [Embedding Models](/docs/v1.5/components/embeddings) - [Azure OpenAI Embedding Models](/docs/v1.5/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/v1.5/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/v1.5/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [HuggingFace Text Embedding Models](/docs/v1.5/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Local Filesystem Embedding Models](/docs/v1.5/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [OpenAI (or Compatible) Embedding Models](/docs/v1.5/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [Model Providers](/docs/v1.5/components/models): Overview of supported model providers for ML and LLMs in Spice. - [Anthropic Models](/docs/v1.5/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/v1.5/components/models/azure): Instructions for using Azure OpenAI models - [Databricks Model Provider](/docs/v1.5/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/v1.5/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [HuggingFace](/docs/v1.5/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [OpenAI (or Compatible) Language Models](/docs/v1.5/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [Perplexity Models](/docs/v1.5/components/models/perplexity): Instructions for using language models hosted on Perplexity with Spice. - [Spice Cloud Platform](/docs/v1.5/components/models/spiceai): Instructions for using models hosted on the Spice Cloud Platform with Spice. - [xAI Models](/docs/v1.5/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/v1.5/components/secret-stores) - [AWS Secrets Manager Secret Store](/docs/v1.5/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Environment Secret Store](/docs/v1.5/components/secret-stores/env): Environment Variables Secret Store Documentation - [Keyring Secret Store](/docs/v1.5/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/v1.5/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/v1.5/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/v1.5/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Server-Sent Events (SSE) protocol. - [Web Search Tool](/docs/v1.5/components/tools/websearch): The Web Search Tool enables Spice models to search the web for information. The tool is available through the websearch tool, and backed by different search engines. - [Vector Engines](/docs/v1.5/components/vectors) - [AWS S3 Vectors Engine](/docs/v1.5/components/vectors/s3_vectors): AWS S3 Vectors Engine Documentation - [Views](/docs/v1.5/components/views): Documentation for defining Views in Spice - [Workers Overview](/docs/v1.5/components/workers): Detailed documentation for workers in the Spice runtime. - [Deployment](/docs/v1.5/deployment): Deploy Spice.ai in your environment - [Deployment Architectures](/docs/v1.5/deployment/architectures): Spice.ai Open Source Deployment architectures - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/v1.5/deployment/architectures/cluster): Deploying Spice as a cluster - [Cloud Hosted](/docs/v1.5/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Microservice Deployment (Single or Multiple Replicas)](/docs/v1.5/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/v1.5/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/v1.5/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/v1.5/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/v1.5/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [Spice Cloud Platform Deployment](/docs/v1.5/deployment/cloud): Guide to deploying data and AI applications using the managed Spice Cloud Platform - [Docker - Kubernetes](/docs/v1.5/deployment/docker): Running Spice.ai as Docker container - [Docker Sandbox Guide - v1.3.0](/docs/v1.5/deployment/docker/sandbox): Migrating to v1.3.0 - [Helm - Kubernetes](/docs/v1.5/deployment/kubernetes): Deploy Spice.ai in Kubernetes using Helm. - [Frequently Asked Questions](/docs/v1.5/faq): Get answers to common questions about Spice.ai, including its features, differences from other tools, and use cases - [Features](/docs/v1.5/features): Features - [Caching](/docs/v1.5/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/v1.5/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Data Acceleration](/docs/v1.5/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/v1.5/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/v1.5/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Indexes](/docs/v1.5/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/v1.5/features/data-acceleration/partitioning): Partitioning for accelerated datasets - [Data Ingestion](/docs/v1.5/features/data-ingestion): Learn how to ingest data in Spice. - [Embedding Datasets](/docs/v1.5/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Large Language Models](/docs/v1.5/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models](/docs/v1.5/features/large-language-models/evals): Learn how Spice evaluates, tracks, compares, and improves language model performance for specific tasks - [Model Context Protocol (MCP)](/docs/v1.5/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/v1.5/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/v1.5/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/v1.5/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/v1.5/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/v1.5/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/v1.5/features/machine-learning-models): Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform. - [Observability & Monitoring](/docs/v1.5/features/observability): Learn how to use Spice telemetry. - [Component Metrics](/docs/v1.5/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/v1.5/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Search Functionality](/docs/v1.5/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/v1.5/features/search/full-text): Learn how Spice can perform full text search - [Vector-Based Search](/docs/v1.5/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/v1.5/features/semantic-model): Learn how to define and use semantic data models with Spice. - [Getting started with Spice.ai OSS](/docs/v1.5/getting-started): Get started with Spice in 5 minutes - [Community Data](/docs/v1.5/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/v1.5/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/v1.5/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. - [Spice.ai OSS Installation](/docs/v1.5/installation): Instructions for installing Spice.ai OSS - [Intelligent Applications](/docs/v1.5/intelligent-applications): Building intelligent data and AI-driven applications with Spice.ai - [Spice.ai OSS Reference Docs](/docs/v1.5/reference): Reference documentation on the Spice API, CLI and Pod manifest syntax. - [Cron Schedules](/docs/v1.5/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/v1.5/reference/datatypes) - [Accelerator Data Types](/docs/v1.5/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/v1.5/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Duration](/docs/v1.5/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/v1.5/reference/file_format): Spice currently supports CSV and Parquet data file-formats for data connectors that can read files from a file system or cloud object storage (i.e. s3//, file://, etc.). Support for Iceberg and other file-formats are on the roadmap. - [Managing Memory Usage](/docs/v1.5/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice.ai Open Source deployments. - [Models Grade Report](/docs/v1.5/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [YAML syntax for Spicepod manifests](/docs/v1.5/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/v1.5/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/v1.5/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/v1.5/reference/spicepod/embeddings): Embeddings YAML reference - [evals](/docs/v1.5/reference/spicepod/evals): Evaluations YAML reference - [Reserved Keywords](/docs/v1.5/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/v1.5/reference/spicepod/models): Models YAML reference - [Runtime](/docs/v1.5/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/v1.5/reference/spicepod/tools): Tools YAML reference - [Workers](/docs/v1.5/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/v1.5/reference/sql) - [EXPLAIN](/docs/v1.5/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/v1.5/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Operators](/docs/v1.5/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/v1.5/reference/sql/prepared_statements): Positional Arguments - [SELECT](/docs/v1.5/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/v1.5/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/v1.5/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/v1.5/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. - [Timestamps](/docs/v1.5/reference/timestamp): In Spice all timestamps are represented as an integer value denoting the number of seconds that have passed since the Unix epoch in UTC time. The Unix epoch is defined as 1970-01-01T0000Z. - [SDKs](/docs/v1.5/sdks): Connect to spice, using official Spice SDKs - [Dotnet SDK for Spice.ai](/docs/v1.5/sdks/dotnet): Connect to Spice using Spice Dotnet SDK - [Go SDK](/docs/v1.5/sdks/golang): Connect to spice using spice go SDK - [Java SDK](/docs/v1.5/sdks/java): Connect to Spice using Spice Java SDK - [JavaScript SDK](/docs/v1.5/sdks/javascript): Connect to spice using Spice.js SDK - [Python SDK](/docs/v1.5/sdks/python): Connect to spice using spice python SDK - [Rust SDK](/docs/v1.5/sdks/rust): Connect to spice using spice rust SDK - [Troubleshooting Spice](/docs/v1.5/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Use Cases](/docs/v1.5/use-cases): Examples and use cases for Spice' - [Using Spice.ai for Agentic AI Applications](/docs/v1.5/use-cases/agentic-apps): Build intelligent autonomous agents that act contextually by grounding AI models in secure, full-knowledge datasets with fast, iterative feedback loops. - [Spice for SQL Query Mesh/Federation](/docs/v1.5/use-cases/data-mesh): Spice for SQL Query Mesh/Federation - [Spice as a CDN for Databases](/docs/v1.5/use-cases/database-cdn): Use Spice as a CDN for Databases - [Spice for Enterprise Search](/docs/v1.5/use-cases/enterprise-search): Use Spice for Enterprise Search - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/v1.5/use-cases/rag): Use Spice for Retrieval-Augmented-Generation (RAG) ### v1.6 Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Spice.ai Open Source](/docs/v1.6): Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Tags](/docs/v1.6/tags) - [One doc tagged with "Acknowledgements"](/docs/v1.6/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [One doc tagged with "ADBC"](/docs/v1.6/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/v1.6/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Arrow"](/docs/v1.6/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/v1.6/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/v1.6/tags/auth): Authentication and authorization mechanisms. - [One doc tagged with "Authentication"](/docs/v1.6/tags/authentication): User authentication methods and security protocols. - [2 docs tagged with "Azure"](/docs/v1.6/tags/azure): Microsoft Azure cloud services and integrations. - [One doc tagged with "Blob Storage"](/docs/v1.6/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Caching"](/docs/v1.6/tags/caching): Data caching strategies and performance optimization. - [5 docs tagged with "Catalogs"](/docs/v1.6/tags/catalogs): Data catalog connectors and metadata management. - [2 docs tagged with "CLI"](/docs/v1.6/tags/cli): Spice command-line interface commands and usage. - [2 docs tagged with "Component Metrics"](/docs/v1.6/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/v1.6/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/v1.6/tags/configuration): System configuration files and runtime settings. - [15 docs tagged with "Data Connectors"](/docs/v1.6/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/v1.6/tags/data-lake): Data lake architectures and storage solutions. - [2 docs tagged with "Databricks"](/docs/v1.6/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/v1.6/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debugging"](/docs/v1.6/tags/debugging): Debugging techniques and troubleshooting tools. - [2 docs tagged with "Delta Lake"](/docs/v1.6/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/v1.6/tags/dependencies): Third-party dependencies and software requirements. - [3 docs tagged with "Deployment"](/docs/v1.6/tags/deployment): Production deployment using Docker and Kubernetes. - [2 docs tagged with "Docker"](/docs/v1.6/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with "DynamoDB"](/docs/v1.6/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [3 docs tagged with "Embeddings"](/docs/v1.6/tags/embeddings): Vector embeddings and semantic similarity operations. - [One doc tagged with "Evaluation"](/docs/v1.6/tags/evaluation): Model evaluation metrics and assessment frameworks. - [One doc tagged with "Features"](/docs/v1.6/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/v1.6/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "Getting Started"](/docs/v1.6/tags/getting-started): Installation guides and quickstart tutorials. - [One doc tagged with "GitHub"](/docs/v1.6/tags/github): GitHub repository data connector and integration. - [One doc tagged with "Glue"](/docs/v1.6/tags/glue): AWS Glue ETL service and data catalog integration. - [One doc tagged with "Iceberg"](/docs/v1.6/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/v1.6/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/v1.6/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "JDBC"](/docs/v1.6/tags/jdbc): Java Database Connectivity driver configuration. - [2 docs tagged with "Kubernetes"](/docs/v1.6/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "Logging"](/docs/v1.6/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/v1.6/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/v1.6/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/v1.6/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/v1.6/tags/memory): In-memory data connectors and volatile storage. - [12 docs tagged with "Models"](/docs/v1.6/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MongoDB"](/docs/v1.6/tags/mongodb): MongoDB NoSQL document database integration. - [One doc tagged with "MySQL"](/docs/v1.6/tags/mysql): MySQL relational database integration. - [One doc tagged with "NoSQL"](/docs/v1.6/tags/nosql): NoSQL database systems and document stores. - [One doc tagged with "ODBC"](/docs/v1.6/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/v1.6/tags/open-source): Open source licenses and community contributions. - [One doc tagged with "OpenAI"](/docs/v1.6/tags/openai): OpenAI API integration and GPT model access. - [One doc tagged with "Oracle"](/docs/v1.6/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/v1.6/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/v1.6/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/v1.6/tags/parameters): Configuration parameters and runtime options. - [One doc tagged with "Performance"](/docs/v1.6/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/v1.6/tags/persistence): Data persistence and durable storage mechanisms. - [3 docs tagged with "Reference"](/docs/v1.6/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/v1.6/tags/relational): Relational database systems and SQL operations. - [One doc tagged with "Runtime"](/docs/v1.6/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Sandbox"](/docs/v1.6/tags/sandbox): Sandbox environments and isolated testing. - [4 docs tagged with "Search"](/docs/v1.6/tags/search): Vector search, semantic search, and ranking capabilities. - [One doc tagged with "Security"](/docs/v1.6/tags/security): Security features and data protection mechanisms. - [2 docs tagged with "SpiceAI"](/docs/v1.6/tags/spiceai): SpiceAI cloud platform and managed services. - [4 docs tagged with "Spicepod"](/docs/v1.6/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [2 docs tagged with "SQL"](/docs/v1.6/tags/sql): SQL query language support and database operations. - [3 docs tagged with "Tools"](/docs/v1.6/tags/tools): Development tools and utility integrations. - [One doc tagged with "Tracing"](/docs/v1.6/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/v1.6/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Unity Catalog"](/docs/v1.6/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [One doc tagged with "YAML"](/docs/v1.6/tags/yaml): YAML configuration syntax and file formats. - [Open Source Acknowledgements](/docs/v1.6/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: - [API](/docs/v1.6/api): API - [ADBC: Arrow Database Connectivity](/docs/v1.6/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/v1.6/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/v1.6/api/auth): Authentication documentation - [Generate Package](/docs/v1.6/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/v1.6/api/HTTP/get-catalogs): List Catalogs - [Get Iceberg API config](/docs/v1.6/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/v1.6/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/v1.6/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [ML Prediction](/docs/v1.6/api/HTTP/get-model-predict): Make a ML prediction using a specific model. - [List Models](/docs/v1.6/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [List Spicepods](/docs/v1.6/api/HTTP/get-spicepods): Get a list of spicepods and their details. In CSV format, it will return a summarised form. - [Check Runtime Status](/docs/v1.6/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Check Namespace exists](/docs/v1.6/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Evals](/docs/v1.6/api/HTTP/list): Return all evals available to run in the runtime. - [Send message to MCP server](/docs/v1.6/api/HTTP/mcp-event): Send message to the MCP endoint, for a given session. - [Establish an MCP SSE Connection](/docs/v1.6/api/HTTP/operation-id): Initiates a Server-Sent Events (SSE) connection using the Model Context Protocol (MCP) to interact with Spice tools. - [Update Refresh SQL](/docs/v1.6/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Run Tool](/docs/v1.6/api/HTTP/post): The request body and JSON response formats match the tool’s specification. - [Batch ML Predictions](/docs/v1.6/api/HTTP/post-batch-predict): Perform a batch of ML predictions, using multiple models, in one request. This is useful for ensembling or A/B testing different models. - [Create Chat Completion](/docs/v1.6/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/v1.6/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/v1.6/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Run Eval](/docs/v1.6/api/HTTP/post-eval): Evaluate a model against a eval spice specification - [Text-to-SQL (NSQL)](/docs/v1.6/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [Search](/docs/v1.6/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/v1.6/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/v1.6/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [runtime](/docs/v1.6/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/v1.6/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/v1.6/api/odbc): ODBC API Documentation - [API Overview](/docs/v1.6/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/v1.6/api/tls): Encryption in transit with TLS documentation - [Spice.ai OSS CLI documentation](/docs/v1.6/cli): Detailed documentation on the Spice.ai OSS CLI - [Spice.ai OSS CLI command reference](/docs/v1.6/cli/reference): Spice CLI command reference - [add](/docs/v1.6/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/v1.6/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/v1.6/cli/reference/chat): spice chat CLI documentation - [Completion](/docs/v1.6/cli/reference/completion): Generate the autocompletion script for spice for the specified shell. - [connect](/docs/v1.6/cli/reference/connect): Connect to an app on the Spice.ai Cloud Platform. - [dataset](/docs/v1.6/cli/reference/dataset): Configure a Spice dataset. - [datasets](/docs/v1.6/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [init](/docs/v1.6/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/v1.6/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/v1.6/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/v1.6/cli/reference/models): Lists models loaded by the Spice runtime - [pods](/docs/v1.6/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [refresh](/docs/v1.6/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/v1.6/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/v1.6/cli/reference/search): Performs embeddings-based searches across search configured datasets. Note: Search requires the ai feature to be installed. - [sql](/docs/v1.6/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/v1.6/cli/reference/status): Spice runtime status - [trace](/docs/v1.6/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/v1.6/cli/reference/upgrade): Upgrades the Spice CLI & Runtime to the latest release - [version](/docs/v1.6/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/v1.6/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels - [Clients and Tools](/docs/v1.6/clients): Client and tools - [Datadog](/docs/v1.6/clients/datadog): Monitoring Spice with Datadog - [DBeaver](/docs/v1.6/clients/dbeaver): Configure DBeaver to query Spice via JDBC - [Grafana & Prometheus](/docs/v1.6/clients/grafana): Monitoring Spice instances with Grafana & Prometheus - [JetBrains DataGrip](/docs/v1.6/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Microsoft Power BI Connector](/docs/v1.6/clients/powerbi): Use Microsoft Power BI to access, visualize and analyze Spice datasets. - [Apache Superset](/docs/v1.6/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/v1.6/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. - [Runtime Components](/docs/v1.6/components): Runtime components' - [Catalog Connectors](/docs/v1.6/components/catalogs) - [Databricks Catalog Connector](/docs/v1.6/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [Glue Catalog Connector](/docs/v1.6/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/v1.6/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Spice.ai Catalog Connector](/docs/v1.6/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/v1.6/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Data Accelerators](/docs/v1.6/components/data-accelerators) - [In-Memory Arrow Data Accelerator](/docs/v1.6/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [DuckDB Data Accelerator](/docs/v1.6/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [PostgreSQL Data Accelerator](/docs/v1.6/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [SQLite Data Accelerator](/docs/v1.6/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [Data Connectors](/docs/v1.6/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/v1.6/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ClickHouse Data Connector](/docs/v1.6/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Databricks Data Connector](/docs/v1.6/components/data-connectors/databricks): Databricks Data Connector Documentation - [Debezium Data Connector](/docs/v1.6/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/v1.6/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Dremio Data Connector](/docs/v1.6/components/data-connectors/dremio): Dremio Data Connector Documentation - [DuckDB Data Connector](/docs/v1.6/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DynamoDB Data Connector](/docs/v1.6/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [File Data Connector](/docs/v1.6/components/data-connectors/file): File Data Connector Documentation - [Flight SQL Data Connector](/docs/v1.6/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/v1.6/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GitHub Data Connector](/docs/v1.6/components/data-connectors/github): GitHub Data Connector Documentation - [Glue Data Connector](/docs/v1.6/components/data-connectors/glue): Glue Data Connector Documentation - [GraphQL Data Connector](/docs/v1.6/components/data-connectors/graphql): GraphQL Data Connector Documentation - [HTTP(s) Data Connector](/docs/v1.6/components/data-connectors/https): HTTP(s) Data Connector Documentation - [Iceberg Data Connector](/docs/v1.6/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/v1.6/components/data-connectors/imap): IMAP Data Connector Documentation - [Kafka Data Connector](/docs/v1.6/components/data-connectors/kafka): Kafka Data Connector Documentation - [Localpod Data Connector](/docs/v1.6/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/v1.6/components/data-connectors/memory): Memory Data Connector Documentation - [MongoDB Data Connector](/docs/v1.6/components/data-connectors/mongodb): MongoDB Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/v1.6/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [MySQL Data Connector](/docs/v1.6/components/data-connectors/mysql): MySQL Data Connector Documentation - [ODBC Data Connector](/docs/v1.6/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/v1.6/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/v1.6/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [Amazon Redshift Data Connector](/docs/v1.6/components/data-connectors/redshift): Connect to Amazon Redshift using the PostgreSQL connector in Spice. - [S3 Data Connector](/docs/v1.6/components/data-connectors/s3): S3 Data Connector Documentation - [SharePoint Data Connector](/docs/v1.6/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [Snowflake Data Connector](/docs/v1.6/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/v1.6/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/v1.6/components/data-connectors/spiceai): Spice.ai Data Connector Documentation - [Embedding Models](/docs/v1.6/components/embeddings): Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. - [Azure OpenAI Embedding Models](/docs/v1.6/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/v1.6/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/v1.6/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [HuggingFace Text Embedding Models](/docs/v1.6/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Local Filesystem Embedding Models](/docs/v1.6/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [Model2Vec Embedding Models](/docs/v1.6/components/embeddings/model2vec): Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning. - [OpenAI (or Compatible) Embedding Models](/docs/v1.6/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [Model Providers](/docs/v1.6/components/models): Overview of supported model providers for ML and LLMs in Spice. - [Anthropic Models](/docs/v1.6/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/v1.6/components/models/azure): Instructions for using Azure OpenAI models - [Amazon Bedrock Models](/docs/v1.6/components/models/bedrock): How to use Amazon Bedrock models with Spice. - [Databricks Model Provider](/docs/v1.6/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/v1.6/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [HuggingFace](/docs/v1.6/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [OpenAI (or Compatible) Language Models](/docs/v1.6/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [Perplexity Models](/docs/v1.6/components/models/perplexity): Instructions for using language models hosted on Perplexity with Spice. - [Spice Cloud Platform](/docs/v1.6/components/models/spiceai): Instructions for using models hosted on the Spice Cloud Platform with Spice. - [xAI Models](/docs/v1.6/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/v1.6/components/secret-stores) - [AWS Secrets Manager Secret Store](/docs/v1.6/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Environment Secret Store](/docs/v1.6/components/secret-stores/env): Environment Variables Secret Store Documentation - [Keyring Secret Store](/docs/v1.6/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/v1.6/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/v1.6/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/v1.6/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Server-Sent Events (SSE) protocol. - [Web Search Tool](/docs/v1.6/components/tools/websearch): The Web Search Tool enables Spice models to search the web for information. The tool is available through the websearch tool, and backed by different search engines. - [Vector Engines](/docs/v1.6/components/vectors) - [Amazon S3 Vectors Engine](/docs/v1.6/components/vectors/s3_vectors): Amazon S3 Vectors Engine Documentation - [Views](/docs/v1.6/components/views): Documentation for defining Views in Spice - [Workers Overview](/docs/v1.6/components/workers): Detailed documentation for workers in the Spice runtime. - [Deployment](/docs/v1.6/deployment): Deploy Spice.ai in your environment - [Deployment Architectures](/docs/v1.6/deployment/architectures): Spice.ai Open Source Deployment architectures - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/v1.6/deployment/architectures/cluster): Deploying Spice as a cluster - [Cloud Hosted](/docs/v1.6/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Microservice Deployment (Single or Multiple Replicas)](/docs/v1.6/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/v1.6/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/v1.6/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/v1.6/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/v1.6/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [Spice Cloud Platform Deployment](/docs/v1.6/deployment/cloud): Guide to deploying data and AI applications using the managed Spice Cloud Platform - [Docker - Kubernetes](/docs/v1.6/deployment/docker): Running Spice.ai as Docker container - [Docker Sandbox Guide - v1.3.0](/docs/v1.6/deployment/docker/sandbox): Migrating to v1.3.0 - [Helm - Kubernetes](/docs/v1.6/deployment/kubernetes): Deploy Spice.ai in Kubernetes using Helm. - [Frequently Asked Questions](/docs/v1.6/faq): Get answers to common questions about Spice.ai, including its features, differences from other tools, and use cases - [Features](/docs/v1.6/features): Features - [Caching](/docs/v1.6/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/v1.6/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Data Acceleration](/docs/v1.6/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/v1.6/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/v1.6/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Indexes](/docs/v1.6/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/v1.6/features/data-acceleration/partitioning): Partitioning for accelerated datasets - [Data Ingestion](/docs/v1.6/features/data-ingestion): Learn how to ingest data in Spice. - [Embedding Datasets](/docs/v1.6/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Large Language Models](/docs/v1.6/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models](/docs/v1.6/features/large-language-models/evals): Learn how Spice evaluates, tracks, compares, and improves language model performance for specific tasks - [Model Context Protocol (MCP)](/docs/v1.6/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/v1.6/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/v1.6/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/v1.6/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/v1.6/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/v1.6/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/v1.6/features/machine-learning-models): Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform. - [Observability & Monitoring](/docs/v1.6/features/observability): Learn how to use Spice telemetry. - [Component Metrics](/docs/v1.6/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/v1.6/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Search Functionality](/docs/v1.6/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/v1.6/features/search/full-text): Learn how Spice can perform full text search - [Vector-Based Search](/docs/v1.6/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/v1.6/features/semantic-model): Learn how to define and use semantic data models with Spice. - [Getting started with Spice.ai OSS](/docs/v1.6/getting-started): Get started with Spice in 5 minutes - [Community Data](/docs/v1.6/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/v1.6/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/v1.6/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. - [Spice.ai OSS Installation](/docs/v1.6/installation): Instructions for installing Spice.ai OSS - [Intelligent Applications](/docs/v1.6/intelligent-applications): Building intelligent data and AI-driven applications with Spice.ai - [Spice.ai OSS Reference Docs](/docs/v1.6/reference): Reference documentation on the Spice API, CLI and Pod manifest syntax. - [Cron Schedules](/docs/v1.6/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/v1.6/reference/datatypes) - [Accelerator Data Types](/docs/v1.6/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/v1.6/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Duration](/docs/v1.6/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/v1.6/reference/file_format): Spice currently supports CSV and Parquet data file-formats for data connectors that can read files from a file system or cloud object storage (i.e. s3//, file://, etc.). Support for Iceberg and other file-formats are on the roadmap. - [Managing Memory Usage](/docs/v1.6/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice.ai Open Source deployments. - [Models Grade Report](/docs/v1.6/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [YAML syntax for Spicepod manifests](/docs/v1.6/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/v1.6/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/v1.6/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/v1.6/reference/spicepod/embeddings): Embeddings YAML reference - [evals](/docs/v1.6/reference/spicepod/evals): Evaluations YAML reference - [Reserved Keywords](/docs/v1.6/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/v1.6/reference/spicepod/models): Models YAML reference - [Runtime](/docs/v1.6/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/v1.6/reference/spicepod/tools): Tools YAML reference - [Workers](/docs/v1.6/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/v1.6/reference/sql): This section provides a comprehensive reference for SQL support in Spice.ai, including syntax, data types, operators, functions, and system features. The reference is organized by topic for ease of navigation. - [Explain](/docs/v1.6/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/v1.6/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Operators](/docs/v1.6/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/v1.6/reference/sql/prepared_statements): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Scalar Functions](/docs/v1.6/reference/sql/scalar_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Search in SQL](/docs/v1.6/reference/sql/search): Reference for search functions and filtering in Spice SQL. - [SELECT](/docs/v1.6/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/v1.6/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/v1.6/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/v1.6/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. - [SDKs](/docs/v1.6/sdks): Connect to spice, using official Spice SDKs - [Dotnet SDK for Spice.ai](/docs/v1.6/sdks/dotnet): Connect to Spice using Spice Dotnet SDK - [Go SDK](/docs/v1.6/sdks/golang): Connect to spice using spice go SDK - [Java SDK](/docs/v1.6/sdks/java): Connect to Spice using Spice Java SDK - [JavaScript SDK](/docs/v1.6/sdks/javascript): Connect to spice using Spice.js SDK - [Python SDK](/docs/v1.6/sdks/python): Connect to spice using spice python SDK - [Rust SDK](/docs/v1.6/sdks/rust): Connect to spice using spice rust SDK - [Troubleshooting Spice](/docs/v1.6/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Spice.ai Use Cases](/docs/v1.6/use-cases): Use cases for Spice.ai' - [AI Applications and Agents](/docs/v1.6/use-cases/ai): AI Applications and Agents - [Agentic AI Applications and Agents](/docs/v1.6/use-cases/ai/agentic-apps): Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. - [Edge-Enabled AI Applications and Agents](/docs/v1.6/use-cases/ai/edge-ai): Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases. - [Federated MCP Client for Distributed Tool Ecosystems](/docs/v1.6/use-cases/ai/federated-mcp-server): Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, enhancing threat analysis. - [Object-Store Based SQL Query, Search, and LLM Inference Engine](/docs/v1.6/use-cases/ai/object-store-ai-engine): Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. - [Real-Time Decision-Making for Intelligent Applications](/docs/v1.6/use-cases/ai/real-time-decision-making): Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. - [Tool-Augmented AI with Model Context Protocol Server](/docs/v1.6/use-cases/ai/tool-calling-ai): Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. - [Data Federation, Acceleration, and SQL Query](/docs/v1.6/use-cases/data): Data Federation, Acceleration, and SQL Query - [Application Resilience and Performance Optimization](/docs/v1.6/use-cases/data/application-resilence-and-acceleration): Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance. - [Data Mesh for Unified Data Access](/docs/v1.6/use-cases/data/data-mesh): Spice.ai enables unified data access across disparate sources for health-tech applications, fostering a data mesh architecture. - [Database CDN for Enhanced Performance](/docs/v1.6/use-cases/data/database-cdn): Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience. - [ETL-free Workflows and Data Migrations](/docs/v1.6/use-cases/data/etl-free-workflows): Spice.ai enables data migrations and workflows without ETL by federating legacy and modern systems for seamless transitions. - [Object-Store Data Engine](/docs/v1.6/use-cases/data/object-store-data-engine): Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses. - [Reverse-ETL for Operational Workflows](/docs/v1.6/use-cases/data/reverse-etl): Spice.ai serves enriched data from warehouses to operational systems for real-time actions, eliminating complex ETL pipelines. - [Retrieval-Augmented-Generation (RAG)](/docs/v1.6/use-cases/rag): Retrieval-Augmented-Generation (RAG) - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/v1.6/use-cases/rag/applications): Use Spice for Retrieval-Augmented-Generation (RAG) - [Retrieval-Augmented Generation for AI-Powered Reporting](/docs/v1.6/use-cases/rag/reporting): Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. - [Search & Retrieval](/docs/v1.6/use-cases/search): Search & Retrieval - [Simplifying Real-Time Data Collection and Search](/docs/v1.6/use-cases/search/data-collection-and-search): Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic. - [Enterprise Search and Retrieval](/docs/v1.6/use-cases/search/enterprise-search): Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities. - [Object-Store Native Search Engine](/docs/v1.6/use-cases/search/object-store-search-engine): Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search. ### v1.7 Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Spice.ai Open Source](/docs/v1.7): Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Tags](/docs/v1.7/tags) - [One doc tagged with "Acknowledgements"](/docs/v1.7/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [One doc tagged with "ADBC"](/docs/v1.7/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/v1.7/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Arrow"](/docs/v1.7/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/v1.7/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/v1.7/tags/auth): Authentication and authorization mechanisms. - [One doc tagged with "Authentication"](/docs/v1.7/tags/authentication): User authentication methods and security protocols. - [2 docs tagged with "Azure"](/docs/v1.7/tags/azure): Microsoft Azure cloud services and integrations. - [One doc tagged with "Blob Storage"](/docs/v1.7/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Caching"](/docs/v1.7/tags/caching): Data caching strategies and performance optimization. - [5 docs tagged with "Catalogs"](/docs/v1.7/tags/catalogs): Data catalog connectors and metadata management. - [2 docs tagged with "CLI"](/docs/v1.7/tags/cli): Spice command-line interface commands and usage. - [4 docs tagged with "Component Metrics"](/docs/v1.7/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/v1.7/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/v1.7/tags/configuration): System configuration files and runtime settings. - [17 docs tagged with "Data Connectors"](/docs/v1.7/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/v1.7/tags/data-lake): Data lake architectures and storage solutions. - [2 docs tagged with "Databricks"](/docs/v1.7/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/v1.7/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debezium"](/docs/v1.7/tags/debezium): Debezium change data capture integration. - [One doc tagged with "Debugging"](/docs/v1.7/tags/debugging): Debugging techniques and troubleshooting tools. - [2 docs tagged with "Delta Lake"](/docs/v1.7/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/v1.7/tags/dependencies): Third-party dependencies and software requirements. - [3 docs tagged with "Deployment"](/docs/v1.7/tags/deployment): Production deployment using Docker and Kubernetes. - [2 docs tagged with "Docker"](/docs/v1.7/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with "DynamoDB"](/docs/v1.7/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [3 docs tagged with "Embeddings"](/docs/v1.7/tags/embeddings): Vector embeddings and semantic similarity operations. - [One doc tagged with "Evaluation"](/docs/v1.7/tags/evaluation): Model evaluation metrics and assessment frameworks. - [One doc tagged with "Features"](/docs/v1.7/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/v1.7/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "Getting Started"](/docs/v1.7/tags/getting-started): Installation guides and quickstart tutorials. - [One doc tagged with "GitHub"](/docs/v1.7/tags/github): GitHub repository data connector and integration. - [One doc tagged with "Glue"](/docs/v1.7/tags/glue): AWS Glue ETL service and data catalog integration. - [One doc tagged with "Iceberg"](/docs/v1.7/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/v1.7/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/v1.7/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "JDBC"](/docs/v1.7/tags/jdbc): Java Database Connectivity driver configuration. - [One doc tagged with "Kafka"](/docs/v1.7/tags/kafka): Apache Kafka event streaming platform integration. - [2 docs tagged with "Kubernetes"](/docs/v1.7/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "Logging"](/docs/v1.7/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/v1.7/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/v1.7/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/v1.7/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/v1.7/tags/memory): In-memory data connectors and volatile storage. - [13 docs tagged with "Models"](/docs/v1.7/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MongoDB"](/docs/v1.7/tags/mongodb): MongoDB NoSQL document database integration. - [One doc tagged with "MySQL"](/docs/v1.7/tags/mysql): MySQL relational database integration. - [One doc tagged with "NoSQL"](/docs/v1.7/tags/nosql): NoSQL database systems and document stores. - [One doc tagged with "Observability"](/docs/v1.7/tags/observability): Monitoring, tracing, and metrics for system visibility. - [One doc tagged with "ODBC"](/docs/v1.7/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/v1.7/tags/open-source): Open source licenses and community contributions. - [One doc tagged with "OpenAI"](/docs/v1.7/tags/openai): OpenAI API integration and GPT model access. - [One doc tagged with "Oracle"](/docs/v1.7/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/v1.7/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/v1.7/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/v1.7/tags/parameters): Configuration parameters and runtime options. - [One doc tagged with "Performance"](/docs/v1.7/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/v1.7/tags/persistence): Data persistence and durable storage mechanisms. - [3 docs tagged with "Reference"](/docs/v1.7/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/v1.7/tags/relational): Relational database systems and SQL operations. - [One doc tagged with "Runtime"](/docs/v1.7/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Sandbox"](/docs/v1.7/tags/sandbox): Sandbox environments and isolated testing. - [5 docs tagged with "Search"](/docs/v1.7/tags/search): Vector search, semantic search, and ranking capabilities. - [One doc tagged with "Security"](/docs/v1.7/tags/security): Security features and data protection mechanisms. - [2 docs tagged with "SpiceAI"](/docs/v1.7/tags/spiceai): SpiceAI cloud platform and managed services. - [4 docs tagged with "Spicepod"](/docs/v1.7/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [2 docs tagged with "SQL"](/docs/v1.7/tags/sql): SQL query language support and database operations. - [3 docs tagged with "Tools"](/docs/v1.7/tags/tools): Development tools and utility integrations. - [2 docs tagged with "Tracing"](/docs/v1.7/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/v1.7/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Unity Catalog"](/docs/v1.7/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [One doc tagged with "YAML"](/docs/v1.7/tags/yaml): YAML configuration syntax and file formats. - [One doc tagged with "Zipkin"](/docs/v1.7/tags/zipkin): Zipkin distributed tracing system integration. - [Open Source Acknowledgements](/docs/v1.7/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: - [API](/docs/v1.7/api): API - [ADBC: Arrow Database Connectivity](/docs/v1.7/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/v1.7/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/v1.7/api/auth): Authentication documentation - [Generate Package](/docs/v1.7/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/v1.7/api/HTTP/get-catalogs): List Catalogs - [Get Iceberg API config](/docs/v1.7/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/v1.7/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/v1.7/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [ML Prediction](/docs/v1.7/api/HTTP/get-model-predict): Make a ML prediction using a specific model. - [List Models](/docs/v1.7/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [List Spicepods](/docs/v1.7/api/HTTP/get-spicepods): Get a list of spicepods and their details. In CSV format, it will return a summarised form. - [Check Runtime Status](/docs/v1.7/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Check Namespace exists](/docs/v1.7/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Evals](/docs/v1.7/api/HTTP/list): Return all evals available to run in the runtime. - [Send message to MCP server](/docs/v1.7/api/HTTP/mcp-event): Send message to the MCP endoint, for a given session. - [Establish an MCP SSE Connection](/docs/v1.7/api/HTTP/operation-id): Initiates a Server-Sent Events (SSE) connection using the Model Context Protocol (MCP) to interact with Spice tools. - [Update Refresh SQL](/docs/v1.7/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Run Tool](/docs/v1.7/api/HTTP/post): The request body and JSON response formats match the tool’s specification. - [Batch ML Predictions](/docs/v1.7/api/HTTP/post-batch-predict): Perform a batch of ML predictions, using multiple models, in one request. This is useful for ensembling or A/B testing different models. - [Create Chat Completion](/docs/v1.7/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/v1.7/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/v1.7/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Run Eval](/docs/v1.7/api/HTTP/post-eval): Evaluate a model against a eval spice specification - [Text-to-SQL (NSQL)](/docs/v1.7/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [Search](/docs/v1.7/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/v1.7/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/v1.7/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [runtime](/docs/v1.7/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/v1.7/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/v1.7/api/odbc): ODBC API Documentation - [API Overview](/docs/v1.7/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/v1.7/api/tls): Encryption in transit with TLS documentation - [Spice.ai OSS CLI documentation](/docs/v1.7/cli): Detailed documentation on the Spice.ai OSS CLI - [Spice.ai OSS CLI command reference](/docs/v1.7/cli/reference): Spice CLI command reference - [add](/docs/v1.7/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/v1.7/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/v1.7/cli/reference/chat): spice chat CLI documentation - [Completion](/docs/v1.7/cli/reference/completion): Generate the autocompletion script for spice for the specified shell. - [connect](/docs/v1.7/cli/reference/connect): Connect to an app on the Spice.ai Cloud Platform. - [dataset](/docs/v1.7/cli/reference/dataset): Configure a Spice dataset. - [datasets](/docs/v1.7/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [init](/docs/v1.7/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/v1.7/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/v1.7/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/v1.7/cli/reference/models): Lists models loaded by the Spice runtime - [pods](/docs/v1.7/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [refresh](/docs/v1.7/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/v1.7/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/v1.7/cli/reference/search): Performs embeddings-based searches across search configured datasets. Note: Search requires the ai feature to be installed. - [sql](/docs/v1.7/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/v1.7/cli/reference/status): Spice runtime status - [trace](/docs/v1.7/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/v1.7/cli/reference/upgrade): Upgrades the Spice CLI & Runtime to the latest release - [version](/docs/v1.7/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/v1.7/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels - [Clients and Tools](/docs/v1.7/clients): Client and tools - [DBeaver](/docs/v1.7/clients/dbeaver): Configure DBeaver to query Spice via JDBC - [JetBrains DataGrip](/docs/v1.7/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Microsoft Power BI Connector](/docs/v1.7/clients/powerbi): Use Microsoft Power BI to access, visualize and analyze Spice datasets. - [Apache Superset](/docs/v1.7/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/v1.7/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. - [Runtime Components](/docs/v1.7/components): Runtime components' - [Catalog Connectors](/docs/v1.7/components/catalogs) - [Databricks Catalog Connector](/docs/v1.7/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [Glue Catalog Connector](/docs/v1.7/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/v1.7/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Spice.ai Catalog Connector](/docs/v1.7/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/v1.7/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Data Accelerators](/docs/v1.7/components/data-accelerators) - [In-Memory Arrow Data Accelerator](/docs/v1.7/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [DuckDB Data Accelerator](/docs/v1.7/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [PostgreSQL Data Accelerator](/docs/v1.7/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [SQLite Data Accelerator](/docs/v1.7/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [Data Connectors](/docs/v1.7/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/v1.7/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ClickHouse Data Connector](/docs/v1.7/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Databricks Data Connector](/docs/v1.7/components/data-connectors/databricks): Databricks Data Connector Documentation - [Debezium Data Connector](/docs/v1.7/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/v1.7/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Dremio Data Connector](/docs/v1.7/components/data-connectors/dremio): Dremio Data Connector Documentation - [DuckDB Data Connector](/docs/v1.7/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DynamoDB Data Connector](/docs/v1.7/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [File Data Connector](/docs/v1.7/components/data-connectors/file): File Data Connector Documentation - [Flight SQL Data Connector](/docs/v1.7/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/v1.7/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GitHub Data Connector](/docs/v1.7/components/data-connectors/github): GitHub Data Connector Documentation - [Glue Data Connector](/docs/v1.7/components/data-connectors/glue): Glue Data Connector Documentation - [GraphQL Data Connector](/docs/v1.7/components/data-connectors/graphql): GraphQL Data Connector Documentation - [HTTP(s) Data Connector](/docs/v1.7/components/data-connectors/https): HTTP(s) Data Connector Documentation - [Iceberg Data Connector](/docs/v1.7/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/v1.7/components/data-connectors/imap): IMAP Data Connector Documentation - [Kafka Data Connector](/docs/v1.7/components/data-connectors/kafka): Kafka Data Connector Documentation - [Localpod Data Connector](/docs/v1.7/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/v1.7/components/data-connectors/memory): Memory Data Connector Documentation - [MongoDB Data Connector](/docs/v1.7/components/data-connectors/mongodb): MongoDB Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/v1.7/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [MySQL Data Connector](/docs/v1.7/components/data-connectors/mysql): MySQL Data Connector Documentation - [ODBC Data Connector](/docs/v1.7/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/v1.7/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/v1.7/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [Amazon Redshift Data Connector](/docs/v1.7/components/data-connectors/redshift): Connect to Amazon Redshift using the PostgreSQL connector in Spice. - [S3 Data Connector](/docs/v1.7/components/data-connectors/s3): S3 Data Connector Documentation - [SharePoint Data Connector](/docs/v1.7/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [Snowflake Data Connector](/docs/v1.7/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/v1.7/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/v1.7/components/data-connectors/spiceai): Spice.ai Data Connector Documentation - [Embedding Models](/docs/v1.7/components/embeddings): Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. - [Azure OpenAI Embedding Models](/docs/v1.7/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/v1.7/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/v1.7/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [HuggingFace Text Embedding Models](/docs/v1.7/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Local Filesystem Embedding Models](/docs/v1.7/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [Model2Vec Embedding Models](/docs/v1.7/components/embeddings/model2vec): Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning. - [OpenAI (or Compatible) Embedding Models](/docs/v1.7/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [Model Providers](/docs/v1.7/components/models): Overview of supported model providers for ML and LLMs in Spice. - [Anthropic Models](/docs/v1.7/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/v1.7/components/models/azure): Instructions for using Azure OpenAI models - [Amazon Bedrock Models](/docs/v1.7/components/models/bedrock): How to use Amazon Bedrock models with Spice. - [Databricks Model Provider](/docs/v1.7/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/v1.7/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [HuggingFace](/docs/v1.7/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [OpenAI (or Compatible) Language Models](/docs/v1.7/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [Perplexity Models](/docs/v1.7/components/models/perplexity): Instructions for using language models hosted on Perplexity with Spice. - [Spice Cloud Platform](/docs/v1.7/components/models/spiceai): Instructions for using models hosted on the Spice Cloud Platform with Spice. - [xAI Models](/docs/v1.7/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/v1.7/components/secret-stores) - [AWS Secrets Manager Secret Store](/docs/v1.7/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Environment Secret Store](/docs/v1.7/components/secret-stores/env): Environment Variables Secret Store Documentation - [Keyring Secret Store](/docs/v1.7/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/v1.7/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/v1.7/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/v1.7/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Server-Sent Events (SSE) protocol. - [Web Search Tool](/docs/v1.7/components/tools/websearch): The Web Search Tool enables Spice models to search the web for information. The tool is available through the websearch tool, and backed by different search engines. - [Vector Engines](/docs/v1.7/components/vectors) - [Amazon S3 Vectors Engine](/docs/v1.7/components/vectors/s3_vectors): Amazon S3 Vectors Engine Documentation - [Views](/docs/v1.7/components/views): Documentation for defining Views in Spice - [Workers Overview](/docs/v1.7/components/workers): Detailed documentation for workers in the Spice runtime. - [Deployment](/docs/v1.7/deployment): Deploy Spice.ai in your environment - [Deployment Architectures](/docs/v1.7/deployment/architectures): Spice.ai Open Source Deployment architectures - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/v1.7/deployment/architectures/cluster): Deploying Spice as a cluster - [Cloud Hosted](/docs/v1.7/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Microservice Deployment (Single or Multiple Replicas)](/docs/v1.7/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/v1.7/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/v1.7/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/v1.7/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/v1.7/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [Spice Cloud Platform Deployment](/docs/v1.7/deployment/cloud): Guide to deploying data and AI applications using the managed Spice Cloud Platform - [Docker - Kubernetes](/docs/v1.7/deployment/docker): Running Spice.ai as Docker container - [Docker Sandbox Guide - v1.3.0](/docs/v1.7/deployment/docker/sandbox): Migrating to v1.3.0 - [Helm - Kubernetes](/docs/v1.7/deployment/kubernetes): Deploy Spice.ai in Kubernetes using Helm. - [Frequently Asked Questions](/docs/v1.7/faq): Get answers to common questions about Spice.ai, including its features, differences from other tools, and use cases - [Features](/docs/v1.7/features): Features - [Caching](/docs/v1.7/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/v1.7/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Data Acceleration](/docs/v1.7/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/v1.7/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/v1.7/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Indexes](/docs/v1.7/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/v1.7/features/data-acceleration/partitioning): Partitioning for accelerated datasets - [Data Ingestion](/docs/v1.7/features/data-ingestion): Learn how to ingest data in Spice. - [Embedding Datasets](/docs/v1.7/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Large Language Models](/docs/v1.7/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models](/docs/v1.7/features/large-language-models/evals): Learn how Spice evaluates, tracks, compares, and improves language model performance for specific tasks - [Model Context Protocol (MCP)](/docs/v1.7/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/v1.7/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/v1.7/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/v1.7/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/v1.7/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/v1.7/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/v1.7/features/machine-learning-models): Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform. - [Observability & Monitoring](/docs/v1.7/features/observability): Learn how to use Spice telemetry. - [Component Metrics](/docs/v1.7/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/v1.7/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Search Functionality](/docs/v1.7/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/v1.7/features/search/full-text): Learn how Spice can perform full text search - [Vector-Based Search](/docs/v1.7/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/v1.7/features/semantic-model): Learn how to define and use semantic data models with Spice. - [Web Search](/docs/v1.7/features/web-search): Learn how Spice can perform web search - [Getting started with Spice.ai OSS](/docs/v1.7/getting-started): Get started with Spice in 5 minutes - [Community Data](/docs/v1.7/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/v1.7/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/v1.7/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. - [Spice.ai OSS Installation](/docs/v1.7/installation): Instructions for installing Spice.ai OSS - [Intelligent Applications](/docs/v1.7/intelligent-applications): Building intelligent data and AI-driven applications with Spice.ai - [Monitoring](/docs/v1.7/monitoring): Monitoring Spice.ai deployments - [Datadog](/docs/v1.7/monitoring/datadog): Monitoring Spice with Datadog - [Grafana & Prometheus](/docs/v1.7/monitoring/grafana): Monitoring Spice instances with Grafana & Prometheus - [Zipkin Integration](/docs/v1.7/monitoring/zipkin): Learn how to integrate Spice with Zipkin tracing. - [Spice.ai OSS Reference Docs](/docs/v1.7/reference): Reference documentation on the Spice API, CLI and Pod manifest syntax. - [Cron Schedules](/docs/v1.7/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/v1.7/reference/datatypes) - [Accelerator Data Types](/docs/v1.7/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/v1.7/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Duration](/docs/v1.7/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/v1.7/reference/file_format): Spice currently supports CSV, JSON, and Parquet data file-formats for data connectors that can read files from a file system or cloud object storage (i.e. s3//, file://, etc.). Support for Iceberg and other file-formats are on the roadmap. - [Managing Memory Usage](/docs/v1.7/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice.ai Open Source deployments. - [Models Grade Report](/docs/v1.7/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [YAML syntax for Spicepod manifests](/docs/v1.7/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/v1.7/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/v1.7/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/v1.7/reference/spicepod/embeddings): Embeddings YAML reference - [evals](/docs/v1.7/reference/spicepod/evals): Evaluations YAML reference - [Reserved Keywords](/docs/v1.7/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/v1.7/reference/spicepod/models): Models YAML reference - [Runtime](/docs/v1.7/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/v1.7/reference/spicepod/tools): Tools YAML reference - [Workers](/docs/v1.7/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/v1.7/reference/sql): This section provides a comprehensive reference for SQL support in Spice.ai, including syntax, data types, operators, functions, and system features. The reference is organized by topic for ease of navigation. - [Aggregate Functions](/docs/v1.7/reference/sql/aggregate_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. Note, when using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [Explain](/docs/v1.7/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/v1.7/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [JSON Functions and Operators](/docs/v1.7/reference/sql/json): Reference for JSON functions and operators in Spice SQL - [Operators](/docs/v1.7/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/v1.7/reference/sql/prepared_statements): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Scalar Functions](/docs/v1.7/reference/sql/scalar_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. Note, when using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [Search in SQL](/docs/v1.7/reference/sql/search): Reference for search functions and filtering in Spice SQL. - [SELECT](/docs/v1.7/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/v1.7/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/v1.7/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/v1.7/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. - [SDKs](/docs/v1.7/sdks): Connect to spice, using official Spice SDKs - [Dotnet SDK for Spice.ai](/docs/v1.7/sdks/dotnet): Connect to Spice using Spice Dotnet SDK - [Go SDK](/docs/v1.7/sdks/golang): Connect to spice using spice go SDK - [Java SDK](/docs/v1.7/sdks/java): Connect to Spice using Spice Java SDK - [JavaScript SDK](/docs/v1.7/sdks/javascript): Connect to spice using Spice.js SDK - [Python SDK](/docs/v1.7/sdks/python): Connect to spice using spice python SDK - [Rust SDK](/docs/v1.7/sdks/rust): Connect to spice using spice rust SDK - [Troubleshooting Spice](/docs/v1.7/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Spice.ai Use Cases](/docs/v1.7/use-cases): Use cases for Spice.ai' - [AI Applications and Agents](/docs/v1.7/use-cases/ai): AI Applications and Agents - [Agentic AI Applications and Agents](/docs/v1.7/use-cases/ai/agentic-apps): Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. - [Edge-Enabled AI Applications and Agents](/docs/v1.7/use-cases/ai/edge-ai): Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases. - [Federated MCP Client for Distributed Tool Ecosystems](/docs/v1.7/use-cases/ai/federated-mcp-server): Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, enhancing threat analysis. - [Object-Store Based SQL Query, Search, and LLM Inference Engine](/docs/v1.7/use-cases/ai/object-store-ai-engine): Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. - [Real-Time Decision-Making for Intelligent Applications](/docs/v1.7/use-cases/ai/real-time-decision-making): Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. - [Tool-Augmented AI with Model Context Protocol Server](/docs/v1.7/use-cases/ai/tool-calling-ai): Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. - [Data Federation, Acceleration, and SQL Query](/docs/v1.7/use-cases/data): Data Federation, Acceleration, and SQL Query - [Application Resilience and Performance Optimization](/docs/v1.7/use-cases/data/application-resilence-and-acceleration): Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance. - [Data Mesh for Unified Data Access](/docs/v1.7/use-cases/data/data-mesh): Spice.ai enables unified data access across disparate sources for health-tech applications, fostering a data mesh architecture. - [Database CDN for Enhanced Performance](/docs/v1.7/use-cases/data/database-cdn): Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience. - [ETL-free Workflows and Data Migrations](/docs/v1.7/use-cases/data/etl-free-workflows): Spice.ai enables data migrations and workflows without ETL by federating legacy and modern systems for seamless transitions. - [Object-Store Data Engine](/docs/v1.7/use-cases/data/object-store-data-engine): Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses. - [Reverse-ETL for Operational Workflows](/docs/v1.7/use-cases/data/reverse-etl): Spice.ai serves enriched data from warehouses to operational systems for real-time actions, eliminating complex ETL pipelines. - [Retrieval-Augmented-Generation (RAG)](/docs/v1.7/use-cases/rag): Retrieval-Augmented-Generation (RAG) - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/v1.7/use-cases/rag/applications): Use Spice for Retrieval-Augmented-Generation (RAG) - [Retrieval-Augmented Generation for AI-Powered Reporting](/docs/v1.7/use-cases/rag/reporting): Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. - [Search & Retrieval](/docs/v1.7/use-cases/search): Search & Retrieval - [Simplifying Real-Time Data Collection and Search](/docs/v1.7/use-cases/search/data-collection-and-search): Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic. - [Enterprise Search and Retrieval](/docs/v1.7/use-cases/search/enterprise-search): Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities. - [Object-Store Native Search Engine](/docs/v1.7/use-cases/search/object-store-search-engine): Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search. ### v1.8 Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Spice.ai Open Source](/docs/v1.8): Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Tags](/docs/v1.8/tags) - [One doc tagged with "Acknowledgements"](/docs/v1.8/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [One doc tagged with "ADBC"](/docs/v1.8/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/v1.8/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Arrow"](/docs/v1.8/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/v1.8/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/v1.8/tags/auth): Authentication and authorization mechanisms. - [One doc tagged with "Authentication"](/docs/v1.8/tags/authentication): User authentication methods and security protocols. - [2 docs tagged with "Azure"](/docs/v1.8/tags/azure): Microsoft Azure cloud services and integrations. - [One doc tagged with "Blob Storage"](/docs/v1.8/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Caching"](/docs/v1.8/tags/caching): Data caching strategies and performance optimization. - [5 docs tagged with "Catalogs"](/docs/v1.8/tags/catalogs): Data catalog connectors and metadata management. - [2 docs tagged with "CLI"](/docs/v1.8/tags/cli): Spice command-line interface commands and usage. - [4 docs tagged with "Component Metrics"](/docs/v1.8/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/v1.8/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/v1.8/tags/configuration): System configuration files and runtime settings. - [19 docs tagged with "Data Connectors"](/docs/v1.8/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/v1.8/tags/data-lake): Data lake architectures and storage solutions. - [2 docs tagged with "Databricks"](/docs/v1.8/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/v1.8/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debezium"](/docs/v1.8/tags/debezium): Debezium change data capture integration. - [One doc tagged with "Debugging"](/docs/v1.8/tags/debugging): Debugging techniques and troubleshooting tools. - [2 docs tagged with "Delta Lake"](/docs/v1.8/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/v1.8/tags/dependencies): Third-party dependencies and software requirements. - [3 docs tagged with "Deployment"](/docs/v1.8/tags/deployment): Production deployment using Docker and Kubernetes. - [2 docs tagged with "Docker"](/docs/v1.8/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with "DynamoDB"](/docs/v1.8/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [3 docs tagged with "Embeddings"](/docs/v1.8/tags/embeddings): Vector embeddings and semantic similarity operations. - [One doc tagged with "Evaluation"](/docs/v1.8/tags/evaluation): Model evaluation metrics and assessment frameworks. - [2 docs tagged with "Features"](/docs/v1.8/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/v1.8/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "Getting Started"](/docs/v1.8/tags/getting-started): Installation guides and quickstart tutorials. - [One doc tagged with "GitHub"](/docs/v1.8/tags/github): GitHub repository data connector and integration. - [2 docs tagged with "Glue"](/docs/v1.8/tags/glue): AWS Glue ETL service and data catalog integration. - [2 docs tagged with "Iceberg"](/docs/v1.8/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/v1.8/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/v1.8/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "JDBC"](/docs/v1.8/tags/jdbc): Java Database Connectivity driver configuration. - [One doc tagged with "Kafka"](/docs/v1.8/tags/kafka): Apache Kafka event streaming platform integration. - [2 docs tagged with "Kubernetes"](/docs/v1.8/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "Logging"](/docs/v1.8/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/v1.8/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/v1.8/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/v1.8/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/v1.8/tags/memory): In-memory data connectors and volatile storage. - [13 docs tagged with "Models"](/docs/v1.8/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MongoDB"](/docs/v1.8/tags/mongodb): MongoDB NoSQL document database integration. - [One doc tagged with "MySQL"](/docs/v1.8/tags/mysql): MySQL relational database integration. - [One doc tagged with "NoSQL"](/docs/v1.8/tags/nosql): NoSQL database systems and document stores. - [One doc tagged with "Observability"](/docs/v1.8/tags/observability): Monitoring, tracing, and metrics for system visibility. - [One doc tagged with "ODBC"](/docs/v1.8/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/v1.8/tags/open-source): Open source licenses and community contributions. - [One doc tagged with "OpenAI"](/docs/v1.8/tags/openai): OpenAI API integration and GPT model access. - [One doc tagged with "Oracle"](/docs/v1.8/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/v1.8/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/v1.8/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/v1.8/tags/parameters): Configuration parameters and runtime options. - [One doc tagged with "Performance"](/docs/v1.8/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/v1.8/tags/persistence): Data persistence and durable storage mechanisms. - [3 docs tagged with "Reference"](/docs/v1.8/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/v1.8/tags/relational): Relational database systems and SQL operations. - [One doc tagged with "Runtime"](/docs/v1.8/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Sandbox"](/docs/v1.8/tags/sandbox): Sandbox environments and isolated testing. - [5 docs tagged with "Search"](/docs/v1.8/tags/search): Vector search, semantic search, and ranking capabilities. - [One doc tagged with "Security"](/docs/v1.8/tags/security): Security features and data protection mechanisms. - [2 docs tagged with "SpiceAI"](/docs/v1.8/tags/spiceai): SpiceAI cloud platform and managed services. - [4 docs tagged with "Spicepod"](/docs/v1.8/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [2 docs tagged with "SQL"](/docs/v1.8/tags/sql): SQL query language support and database operations. - [3 docs tagged with "Tools"](/docs/v1.8/tags/tools): Development tools and utility integrations. - [2 docs tagged with "Tracing"](/docs/v1.8/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/v1.8/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Unity Catalog"](/docs/v1.8/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [6 docs tagged with "Write"](/docs/v1.8/tags/write): Data connectors and catalogs that support write operations. - [One doc tagged with "YAML"](/docs/v1.8/tags/yaml): YAML configuration syntax and file formats. - [One doc tagged with "Zipkin"](/docs/v1.8/tags/zipkin): Zipkin distributed tracing system integration. - [Open Source Acknowledgements](/docs/v1.8/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: - [API](/docs/v1.8/api): API - [ADBC: Arrow Database Connectivity](/docs/v1.8/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/v1.8/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/v1.8/api/auth): Authentication documentation - [Generate Package](/docs/v1.8/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/v1.8/api/HTTP/get-catalogs): List Catalogs - [Get Iceberg API config](/docs/v1.8/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/v1.8/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/v1.8/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [ML Prediction](/docs/v1.8/api/HTTP/get-model-predict): Make a ML prediction using a specific model. - [List Models](/docs/v1.8/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [List Spicepods](/docs/v1.8/api/HTTP/get-spicepods): Get a list of spicepods and their details. In CSV format, it will return a summarised form. - [Check Runtime Status](/docs/v1.8/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Check Namespace exists](/docs/v1.8/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Evals](/docs/v1.8/api/HTTP/list): Return all evals available to run in the runtime. - [Send message to MCP server](/docs/v1.8/api/HTTP/mcp-event): Send message to the MCP endoint, for a given session. - [Establish an MCP SSE Connection](/docs/v1.8/api/HTTP/operation-id): Initiates a Server-Sent Events (SSE) connection using the Model Context Protocol (MCP) to interact with Spice tools. - [Update Refresh SQL](/docs/v1.8/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Run Tool](/docs/v1.8/api/HTTP/post): The request body and JSON response formats match the tool’s specification. - [Batch ML Predictions](/docs/v1.8/api/HTTP/post-batch-predict): Perform a batch of ML predictions, using multiple models, in one request. This is useful for ensembling or A/B testing different models. - [Create Chat Completion](/docs/v1.8/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/v1.8/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/v1.8/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Run Eval](/docs/v1.8/api/HTTP/post-eval): Evaluate a model against a eval spice specification - [Text-to-SQL (NSQL)](/docs/v1.8/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [Search](/docs/v1.8/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/v1.8/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/v1.8/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [runtime](/docs/v1.8/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/v1.8/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/v1.8/api/odbc): ODBC API Documentation - [API Overview](/docs/v1.8/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/v1.8/api/tls): Encryption in transit with TLS documentation - [Spice.ai OSS CLI documentation](/docs/v1.8/cli): Detailed documentation on the Spice.ai OSS CLI - [Spice.ai OSS CLI command reference](/docs/v1.8/cli/reference): Spice CLI command reference - [add](/docs/v1.8/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/v1.8/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/v1.8/cli/reference/chat): spice chat CLI documentation - [Completion](/docs/v1.8/cli/reference/completion): Generate the autocompletion script for spice for the specified shell. - [connect](/docs/v1.8/cli/reference/connect): Connect to an app on the Spice.ai Cloud Platform. - [dataset](/docs/v1.8/cli/reference/dataset): Configure a Spice dataset. - [datasets](/docs/v1.8/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [init](/docs/v1.8/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/v1.8/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/v1.8/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/v1.8/cli/reference/models): Lists models loaded by the Spice runtime - [pods](/docs/v1.8/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [refresh](/docs/v1.8/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/v1.8/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/v1.8/cli/reference/search): Performs embeddings-based searches across search configured datasets. Note: Search requires the ai feature to be installed. - [sql](/docs/v1.8/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/v1.8/cli/reference/status): Spice runtime status - [trace](/docs/v1.8/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/v1.8/cli/reference/upgrade): Upgrades the Spice CLI & Runtime to the latest release - [version](/docs/v1.8/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/v1.8/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels - [Clients and Tools](/docs/v1.8/clients): Client and tools - [DBeaver](/docs/v1.8/clients/dbeaver): Configure DBeaver to query Spice via JDBC - [JetBrains DataGrip](/docs/v1.8/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Microsoft Power BI Connector](/docs/v1.8/clients/powerbi): Use Microsoft Power BI to access, visualize and analyze Spice datasets. - [Apache Superset](/docs/v1.8/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/v1.8/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. - [Runtime Components](/docs/v1.8/components): Runtime components' - [Catalog Connectors](/docs/v1.8/components/catalogs) - [Databricks Catalog Connector](/docs/v1.8/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [Glue Catalog Connector](/docs/v1.8/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/v1.8/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Spice.ai Catalog Connector](/docs/v1.8/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/v1.8/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Data Accelerators](/docs/v1.8/components/data-accelerators) - [In-Memory Arrow Data Accelerator](/docs/v1.8/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [DuckDB Data Accelerator](/docs/v1.8/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [PostgreSQL Data Accelerator](/docs/v1.8/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [SQLite Data Accelerator](/docs/v1.8/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [Data Connectors](/docs/v1.8/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/v1.8/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ClickHouse Data Connector](/docs/v1.8/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Databricks Data Connector](/docs/v1.8/components/data-connectors/databricks): Databricks Data Connector Documentation - [Debezium Data Connector](/docs/v1.8/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/v1.8/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Dremio Data Connector](/docs/v1.8/components/data-connectors/dremio): Dremio Data Connector Documentation - [DuckDB Data Connector](/docs/v1.8/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DynamoDB Data Connector](/docs/v1.8/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [File Data Connector](/docs/v1.8/components/data-connectors/file): File Data Connector Documentation - [Flight SQL Data Connector](/docs/v1.8/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/v1.8/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GitHub Data Connector](/docs/v1.8/components/data-connectors/github): GitHub Data Connector Documentation - [Glue Data Connector](/docs/v1.8/components/data-connectors/glue): Glue Data Connector Documentation - [GraphQL Data Connector](/docs/v1.8/components/data-connectors/graphql): GraphQL Data Connector Documentation - [HTTP(s) Data Connector](/docs/v1.8/components/data-connectors/https): HTTP(s) Data Connector Documentation - [Iceberg Data Connector](/docs/v1.8/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/v1.8/components/data-connectors/imap): IMAP Data Connector Documentation - [Kafka Data Connector](/docs/v1.8/components/data-connectors/kafka): Kafka Data Connector Documentation - [Localpod Data Connector](/docs/v1.8/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/v1.8/components/data-connectors/memory): Memory Data Connector Documentation - [MongoDB Data Connector](/docs/v1.8/components/data-connectors/mongodb): MongoDB Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/v1.8/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [MySQL Data Connector](/docs/v1.8/components/data-connectors/mysql): MySQL Data Connector Documentation - [ODBC Data Connector](/docs/v1.8/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/v1.8/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/v1.8/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [Amazon Redshift Data Connector](/docs/v1.8/components/data-connectors/redshift): Connect to Amazon Redshift using the PostgreSQL connector in Spice. - [S3 Data Connector](/docs/v1.8/components/data-connectors/s3): S3 Data Connector Documentation - [SharePoint Data Connector](/docs/v1.8/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [Snowflake Data Connector](/docs/v1.8/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/v1.8/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/v1.8/components/data-connectors/spiceai): Spice.ai Data Connector Documentation - [Embedding Models](/docs/v1.8/components/embeddings): Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. - [Azure OpenAI Embedding Models](/docs/v1.8/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/v1.8/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/v1.8/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [HuggingFace Text Embedding Models](/docs/v1.8/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Local Filesystem Embedding Models](/docs/v1.8/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [Model2Vec Embedding Models](/docs/v1.8/components/embeddings/model2vec): Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning. - [OpenAI (or Compatible) Embedding Models](/docs/v1.8/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [Model Providers](/docs/v1.8/components/models): Overview of supported model providers for ML and LLMs in Spice. - [Anthropic Models](/docs/v1.8/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/v1.8/components/models/azure): Instructions for using Azure OpenAI models - [Amazon Bedrock Models](/docs/v1.8/components/models/bedrock): How to use Amazon Bedrock models with Spice. - [Databricks Model Provider](/docs/v1.8/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/v1.8/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [HuggingFace](/docs/v1.8/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [OpenAI (or Compatible) Language Models](/docs/v1.8/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [Perplexity Models](/docs/v1.8/components/models/perplexity): Instructions for using language models hosted on Perplexity with Spice. - [Spice Cloud Platform](/docs/v1.8/components/models/spiceai): Instructions for using models hosted on the Spice Cloud Platform with Spice. - [xAI Models](/docs/v1.8/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/v1.8/components/secret-stores) - [AWS Secrets Manager Secret Store](/docs/v1.8/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Environment Secret Store](/docs/v1.8/components/secret-stores/env): Environment Variables Secret Store Documentation - [Keyring Secret Store](/docs/v1.8/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/v1.8/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/v1.8/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/v1.8/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Server-Sent Events (SSE) protocol. - [Web Search Tool](/docs/v1.8/components/tools/websearch): The Web Search Tool enables Spice models to search the web for information. The tool is available through the websearch tool, and backed by different search engines. - [Vector Engines](/docs/v1.8/components/vectors) - [Amazon S3 Vectors Engine](/docs/v1.8/components/vectors/s3_vectors): Amazon S3 Vectors Engine Documentation - [Views](/docs/v1.8/components/views): Documentation for defining Views in Spice - [Workers Overview](/docs/v1.8/components/workers): Detailed documentation for workers in the Spice runtime. - [Deployment](/docs/v1.8/deployment): Deploy Spice.ai in your environment - [Deployment Architectures](/docs/v1.8/deployment/architectures): Spice.ai Open Source Deployment architectures - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/v1.8/deployment/architectures/cluster): Deploying Spice as a cluster - [Cloud Hosted](/docs/v1.8/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Microservice Deployment (Single or Multiple Replicas)](/docs/v1.8/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/v1.8/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/v1.8/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/v1.8/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/v1.8/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [Spice Cloud Platform Deployment](/docs/v1.8/deployment/cloud): Guide to deploying data and AI applications using the managed Spice Cloud Platform - [Docker - Kubernetes](/docs/v1.8/deployment/docker): Running Spice.ai as Docker container - [Docker Sandbox Guide - v1.3.0](/docs/v1.8/deployment/docker/sandbox): Migrating to v1.3.0 - [Helm - Kubernetes](/docs/v1.8/deployment/kubernetes): Deploy Spice.ai in Kubernetes using Helm. - [Frequently Asked Questions](/docs/v1.8/faq): Get answers to common questions about Spice.ai, including its features, differences from other tools, and use cases - [Features](/docs/v1.8/features): Features - [Caching](/docs/v1.8/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/v1.8/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Data Acceleration](/docs/v1.8/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/v1.8/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/v1.8/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Indexes](/docs/v1.8/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/v1.8/features/data-acceleration/partitioning): Partitioning for accelerated datasets - [Snapshots](/docs/v1.8/features/data-acceleration/snapshots): Bootstrap file-mode accelerations from managed snapshots to eliminate cold starts. - [Data Ingestion](/docs/v1.8/features/data-ingestion): Learn how to ingest data in Spice. - [Embedding Datasets](/docs/v1.8/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Large Language Models](/docs/v1.8/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models](/docs/v1.8/features/large-language-models/evals): Learn how Spice evaluates, tracks, compares, and improves language model performance for specific tasks - [Model Context Protocol (MCP)](/docs/v1.8/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/v1.8/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/v1.8/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/v1.8/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/v1.8/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/v1.8/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/v1.8/features/machine-learning-models): Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform. - [Observability & Monitoring](/docs/v1.8/features/observability): Learn how to use Spice telemetry. - [Component Metrics](/docs/v1.8/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/v1.8/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Search Functionality](/docs/v1.8/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/v1.8/features/search/full-text): Learn how Spice can perform full text search - [Vector-Based Search](/docs/v1.8/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/v1.8/features/semantic-model): Learn how to define and use semantic data models with Spice. - [Web Search](/docs/v1.8/features/web-search): Learn how Spice can perform web search - [Getting started with Spice.ai OSS](/docs/v1.8/getting-started): Get started with Spice in 5 minutes - [Community Data](/docs/v1.8/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/v1.8/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/v1.8/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. - [Spice.ai OSS Installation](/docs/v1.8/installation): Instructions for installing Spice.ai OSS - [Intelligent Applications](/docs/v1.8/intelligent-applications): Building intelligent data and AI-driven applications with Spice.ai - [Monitoring](/docs/v1.8/monitoring): Monitoring Spice.ai deployments - [Datadog](/docs/v1.8/monitoring/datadog): Monitoring Spice with Datadog - [Grafana & Prometheus](/docs/v1.8/monitoring/grafana): Monitoring Spice instances with Grafana & Prometheus - [Zipkin Integration](/docs/v1.8/monitoring/zipkin): Learn how to integrate Spice with Zipkin tracing. - [Spice.ai OSS Reference Docs](/docs/v1.8/reference): Reference documentation on the Spice API, CLI and Pod manifest syntax. - [Cron Schedules](/docs/v1.8/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/v1.8/reference/datatypes) - [Accelerator Data Types](/docs/v1.8/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/v1.8/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Duration](/docs/v1.8/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/v1.8/reference/file_format): Spice currently supports CSV, JSON, and Parquet data file-formats for data connectors that can read files from a file system or cloud object storage (i.e. s3//, file://, etc.). Support for Iceberg and other file-formats are on the roadmap. - [Managing Memory Usage](/docs/v1.8/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice.ai Open Source deployments. - [Models Grade Report](/docs/v1.8/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [YAML syntax for Spicepod manifests](/docs/v1.8/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/v1.8/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/v1.8/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/v1.8/reference/spicepod/embeddings): Embeddings YAML reference - [evals](/docs/v1.8/reference/spicepod/evals): Evaluations YAML reference - [Reserved Keywords](/docs/v1.8/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/v1.8/reference/spicepod/models): Models YAML reference - [Runtime](/docs/v1.8/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/v1.8/reference/spicepod/tools): Tools YAML reference - [Workers](/docs/v1.8/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/v1.8/reference/sql): This section provides a comprehensive reference for SQL support in Spice.ai, including syntax, data types, operators, functions, and system features. The reference is organized by topic for ease of navigation. - [Aggregate Functions](/docs/v1.8/reference/sql/aggregate_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. Note, when using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [AI Functions](/docs/v1.8/reference/sql/ai): AI functions in Spice provide direct integration with large language models (LLMs) and embedding models within SQL queries. These functions process text through configured model providers and return generated responses or vector embeddings. - [DML (Data Manipulation Language)](/docs/v1.8/reference/sql/dml): Data Manipulation Language (DML) statements for inserting and modifying data in Spice. - [Explain](/docs/v1.8/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/v1.8/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [JSON Functions and Operators](/docs/v1.8/reference/sql/json): Reference for JSON functions and operators in Spice SQL - [Operators](/docs/v1.8/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/v1.8/reference/sql/prepared_statements): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Scalar Functions](/docs/v1.8/reference/sql/scalar_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. Note, when using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [Search in SQL](/docs/v1.8/reference/sql/search): Reference for search functions and filtering in Spice SQL. - [SELECT](/docs/v1.8/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/v1.8/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/v1.8/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/v1.8/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. - [SDKs](/docs/v1.8/sdks): Connect to spice, using official Spice SDKs - [Dotnet SDK for Spice.ai](/docs/v1.8/sdks/dotnet): Connect to Spice using Spice Dotnet SDK - [Go SDK](/docs/v1.8/sdks/golang): Connect to spice using spice go SDK - [Java SDK](/docs/v1.8/sdks/java): Connect to Spice using Spice Java SDK - [JavaScript SDK](/docs/v1.8/sdks/javascript): Connect to spice using Spice.js SDK - [Python SDK](/docs/v1.8/sdks/python): Connect to spice using spice python SDK - [Rust SDK](/docs/v1.8/sdks/rust): Connect to spice using spice rust SDK - [Troubleshooting Spice](/docs/v1.8/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Spice.ai Use Cases](/docs/v1.8/use-cases): Use cases for Spice.ai' - [AI Applications and Agents](/docs/v1.8/use-cases/ai): AI Applications and Agents - [Agentic AI Applications and Agents](/docs/v1.8/use-cases/ai/agentic-apps): Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. - [Edge-Enabled AI Applications and Agents](/docs/v1.8/use-cases/ai/edge-ai): Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases. - [Federated MCP Client for Distributed Tool Ecosystems](/docs/v1.8/use-cases/ai/federated-mcp-server): Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, enhancing threat analysis. - [Object-Store Based SQL Query, Search, and LLM Inference Engine](/docs/v1.8/use-cases/ai/object-store-ai-engine): Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. - [Real-Time Decision-Making for Intelligent Applications](/docs/v1.8/use-cases/ai/real-time-decision-making): Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. - [Tool-Augmented AI with Model Context Protocol Server](/docs/v1.8/use-cases/ai/tool-calling-ai): Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. - [Data Federation, Acceleration, and SQL Query](/docs/v1.8/use-cases/data): Data Federation, Acceleration, and SQL Query - [Application Resilience and Performance Optimization](/docs/v1.8/use-cases/data/application-resilence-and-acceleration): Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance. - [Data Mesh for Unified Data Access](/docs/v1.8/use-cases/data/data-mesh): Spice.ai enables unified data access across disparate sources for health-tech applications, fostering a data mesh architecture. - [Database CDN for Enhanced Performance](/docs/v1.8/use-cases/data/database-cdn): Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience. - [ETL-free Workflows and Data Migrations](/docs/v1.8/use-cases/data/etl-free-workflows): Spice.ai enables data migrations and workflows without ETL by federating legacy and modern systems for seamless transitions. - [Object-Store Data Engine](/docs/v1.8/use-cases/data/object-store-data-engine): Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses. - [Reverse-ETL for Operational Workflows](/docs/v1.8/use-cases/data/reverse-etl): Spice.ai serves enriched data from warehouses to operational systems for real-time actions, eliminating complex ETL pipelines. - [Retrieval-Augmented-Generation (RAG)](/docs/v1.8/use-cases/rag): Retrieval-Augmented-Generation (RAG) - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/v1.8/use-cases/rag/applications): Use Spice for Retrieval-Augmented-Generation (RAG) - [Retrieval-Augmented Generation for AI-Powered Reporting](/docs/v1.8/use-cases/rag/reporting): Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. - [Search & Retrieval](/docs/v1.8/use-cases/search): Search & Retrieval - [Simplifying Real-Time Data Collection and Search](/docs/v1.8/use-cases/search/data-collection-and-search): Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic. - [Enterprise Search and Retrieval](/docs/v1.8/use-cases/search/enterprise-search): Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities. - [Object-Store Native Search Engine](/docs/v1.8/use-cases/search/object-store-search-engine): Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search. ### v1.9 Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Spice.ai Open Source](/docs/v1.9): Spice is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. - [Tags](/docs/v1.9/tags) - [One doc tagged with "Acknowledgements"](/docs/v1.9/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [One doc tagged with "ADBC"](/docs/v1.9/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/v1.9/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Arrow"](/docs/v1.9/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/v1.9/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/v1.9/tags/auth): Authentication and authorization mechanisms. - [One doc tagged with "Authentication"](/docs/v1.9/tags/authentication): User authentication methods and security protocols. - [2 docs tagged with "Azure"](/docs/v1.9/tags/azure): Microsoft Azure cloud services and integrations. - [One doc tagged with "Blob Storage"](/docs/v1.9/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Caching"](/docs/v1.9/tags/caching): Data caching strategies and performance optimization. - [5 docs tagged with "Catalogs"](/docs/v1.9/tags/catalogs): Data catalog connectors and metadata management. - [One doc tagged with "Cayenne"](/docs/v1.9/tags/cayenne): Cayenne (Vortex) data accelerator built on Vortex columnar format. - [2 docs tagged with "CLI"](/docs/v1.9/tags/cli): Spice command-line interface commands and usage. - [4 docs tagged with "Component Metrics"](/docs/v1.9/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/v1.9/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/v1.9/tags/configuration): System configuration files and runtime settings. - [One doc tagged with "Data Accelerators"](/docs/v1.9/tags/data-accelerators): Data acceleration engines for high-performance query execution. - [19 docs tagged with "Data Connectors"](/docs/v1.9/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/v1.9/tags/data-lake): Data lake architectures and storage solutions. - [2 docs tagged with "Databricks"](/docs/v1.9/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/v1.9/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debezium"](/docs/v1.9/tags/debezium): Debezium change data capture integration. - [One doc tagged with "Debugging"](/docs/v1.9/tags/debugging): Debugging techniques and troubleshooting tools. - [2 docs tagged with "Delta Lake"](/docs/v1.9/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/v1.9/tags/dependencies): Third-party dependencies and software requirements. - [3 docs tagged with "Deployment"](/docs/v1.9/tags/deployment): Production deployment using Docker and Kubernetes. - [2 docs tagged with "Docker"](/docs/v1.9/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with "DynamoDB"](/docs/v1.9/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [3 docs tagged with "Embeddings"](/docs/v1.9/tags/embeddings): Vector embeddings and semantic similarity operations. - [One doc tagged with "Evaluation"](/docs/v1.9/tags/evaluation): Model evaluation metrics and assessment frameworks. - [2 docs tagged with "Features"](/docs/v1.9/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/v1.9/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "Getting Started"](/docs/v1.9/tags/getting-started): Installation guides and quickstart tutorials. - [One doc tagged with "GitHub"](/docs/v1.9/tags/github): GitHub repository data connector and integration. - [2 docs tagged with "Glue"](/docs/v1.9/tags/glue): AWS Glue ETL service and data catalog integration. - [2 docs tagged with "Iceberg"](/docs/v1.9/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/v1.9/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/v1.9/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "JDBC"](/docs/v1.9/tags/jdbc): Java Database Connectivity driver configuration. - [One doc tagged with "Kafka"](/docs/v1.9/tags/kafka): Apache Kafka event streaming platform integration. - [2 docs tagged with "Kubernetes"](/docs/v1.9/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "Logging"](/docs/v1.9/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/v1.9/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/v1.9/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/v1.9/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/v1.9/tags/memory): In-memory data connectors and volatile storage. - [13 docs tagged with "Models"](/docs/v1.9/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MongoDB"](/docs/v1.9/tags/mongodb): MongoDB NoSQL document database integration. - [One doc tagged with "MySQL"](/docs/v1.9/tags/mysql): MySQL relational database integration. - [One doc tagged with "NoSQL"](/docs/v1.9/tags/nosql): NoSQL database systems and document stores. - [One doc tagged with "Observability"](/docs/v1.9/tags/observability): Monitoring, tracing, and metrics for system visibility. - [One doc tagged with "ODBC"](/docs/v1.9/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/v1.9/tags/open-source): Open source licenses and community contributions. - [One doc tagged with "OpenAI"](/docs/v1.9/tags/openai): OpenAI API integration and GPT model access. - [One doc tagged with "Oracle"](/docs/v1.9/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/v1.9/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/v1.9/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/v1.9/tags/parameters): Configuration parameters and runtime options. - [2 docs tagged with "Performance"](/docs/v1.9/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/v1.9/tags/persistence): Data persistence and durable storage mechanisms. - [4 docs tagged with "Reference"](/docs/v1.9/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/v1.9/tags/relational): Relational database systems and SQL operations. - [One doc tagged with "Runtime"](/docs/v1.9/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Sandbox"](/docs/v1.9/tags/sandbox): Sandbox environments and isolated testing. - [5 docs tagged with "Search"](/docs/v1.9/tags/search): Vector search, semantic search, and ranking capabilities. - [One doc tagged with "Security"](/docs/v1.9/tags/security): Security features and data protection mechanisms. - [2 docs tagged with "SpiceAI"](/docs/v1.9/tags/spiceai): SpiceAI cloud platform and managed services. - [5 docs tagged with "Spicepod"](/docs/v1.9/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [2 docs tagged with "SQL"](/docs/v1.9/tags/sql): SQL query language support and database operations. - [3 docs tagged with "Tools"](/docs/v1.9/tags/tools): Development tools and utility integrations. - [2 docs tagged with "Tracing"](/docs/v1.9/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/v1.9/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Unity Catalog"](/docs/v1.9/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [One doc tagged with "Views"](/docs/v1.9/tags/views): Virtual views and data transformation layers. - [One doc tagged with "Vortex"](/docs/v1.9/tags/vortex): Vortex columnar file format and storage engine. - [6 docs tagged with "Write"](/docs/v1.9/tags/write): Data connectors and catalogs that support write operations. - [One doc tagged with "YAML"](/docs/v1.9/tags/yaml): YAML configuration syntax and file formats. - [One doc tagged with "Zipkin"](/docs/v1.9/tags/zipkin): Zipkin distributed tracing system integration. - [Open Source Acknowledgements](/docs/v1.9/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: - [API](/docs/v1.9/api): API - [ADBC: Arrow Database Connectivity](/docs/v1.9/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/v1.9/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/v1.9/api/auth): Authentication documentation - [Generate Package](/docs/v1.9/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/v1.9/api/HTTP/get-catalogs): List Catalogs - [Get Iceberg API config](/docs/v1.9/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/v1.9/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/v1.9/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [ML Prediction](/docs/v1.9/api/HTTP/get-model-predict): Make a ML prediction using a specific model. - [List Models](/docs/v1.9/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [List Spicepods](/docs/v1.9/api/HTTP/get-spicepods): Get a list of spicepods and their details. In CSV format, it will return a summarised form. - [Check Runtime Status](/docs/v1.9/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Check Namespace exists](/docs/v1.9/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Evals](/docs/v1.9/api/HTTP/list): Return all evals available to run in the runtime. - [Send message to MCP server](/docs/v1.9/api/HTTP/mcp-event): Send message to the MCP endoint, for a given session. - [Establish an MCP SSE Connection](/docs/v1.9/api/HTTP/operation-id): Initiates a Server-Sent Events (SSE) connection using the Model Context Protocol (MCP) to interact with Spice tools. - [Update Refresh SQL](/docs/v1.9/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Run Tool](/docs/v1.9/api/HTTP/post): The request body and JSON response formats match the tool’s specification. - [Batch ML Predictions](/docs/v1.9/api/HTTP/post-batch-predict): Perform a batch of ML predictions, using multiple models, in one request. This is useful for ensembling or A/B testing different models. - [Create Chat Completion](/docs/v1.9/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/v1.9/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/v1.9/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Run Eval](/docs/v1.9/api/HTTP/post-eval): Evaluate a model against a eval spice specification - [Text-to-SQL (NSQL)](/docs/v1.9/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [Search](/docs/v1.9/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/v1.9/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/v1.9/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [runtime](/docs/v1.9/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/v1.9/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/v1.9/api/odbc): ODBC API Documentation - [API Overview](/docs/v1.9/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/v1.9/api/tls): Encryption in transit with TLS documentation - [Spice.ai OSS CLI documentation](/docs/v1.9/cli): Detailed documentation on the Spice.ai OSS CLI - [Spice.ai OSS CLI command reference](/docs/v1.9/cli/reference): Spice CLI command reference - [add](/docs/v1.9/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/v1.9/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/v1.9/cli/reference/chat): spice chat CLI documentation - [Completion](/docs/v1.9/cli/reference/completion): Generate the autocompletion script for spice for the specified shell. - [connect](/docs/v1.9/cli/reference/connect): Connect to an app on the Spice.ai Cloud Platform. - [dataset](/docs/v1.9/cli/reference/dataset): Configure a Spice dataset. - [datasets](/docs/v1.9/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [init](/docs/v1.9/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/v1.9/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/v1.9/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/v1.9/cli/reference/models): Lists models loaded by the Spice runtime - [pods](/docs/v1.9/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [refresh](/docs/v1.9/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/v1.9/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/v1.9/cli/reference/search): Performs embeddings-based searches across search configured datasets. Note: Search requires the ai feature to be installed. - [sql](/docs/v1.9/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/v1.9/cli/reference/status): Spice runtime status - [trace](/docs/v1.9/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/v1.9/cli/reference/upgrade): Upgrades the Spice CLI & Runtime to the latest release - [version](/docs/v1.9/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/v1.9/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels - [Clients and Tools](/docs/v1.9/clients): Client and tools - [DBeaver](/docs/v1.9/clients/dbeaver): Configure DBeaver to query Spice via JDBC - [JetBrains DataGrip](/docs/v1.9/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Microsoft Power BI Connector](/docs/v1.9/clients/powerbi): Use Microsoft Power BI to access, visualize and analyze Spice datasets. - [Apache Superset](/docs/v1.9/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/v1.9/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. - [Runtime Components](/docs/v1.9/components): Runtime components' - [Catalog Connectors](/docs/v1.9/components/catalogs) - [Databricks Catalog Connector](/docs/v1.9/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [Glue Catalog Connector](/docs/v1.9/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/v1.9/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Spice.ai Catalog Connector](/docs/v1.9/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/v1.9/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Data Accelerators](/docs/v1.9/components/data-accelerators) - [In-Memory Arrow Data Accelerator](/docs/v1.9/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [Cayenne Data Accelerator](/docs/v1.9/components/data-accelerators/cayenne): Cayenne Data Accelerator (Vortex) Documentation - [DuckDB Data Accelerator](/docs/v1.9/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [PostgreSQL Data Accelerator](/docs/v1.9/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [SQLite Data Accelerator](/docs/v1.9/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [Data Connectors](/docs/v1.9/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/v1.9/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ClickHouse Data Connector](/docs/v1.9/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Databricks Data Connector](/docs/v1.9/components/data-connectors/databricks): Databricks Data Connector Documentation - [Debezium Data Connector](/docs/v1.9/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/v1.9/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Dremio Data Connector](/docs/v1.9/components/data-connectors/dremio): Dremio Data Connector Documentation - [DuckDB Data Connector](/docs/v1.9/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DynamoDB Data Connector](/docs/v1.9/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [File Data Connector](/docs/v1.9/components/data-connectors/file): File Data Connector Documentation - [Flight SQL Data Connector](/docs/v1.9/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/v1.9/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GitHub Data Connector](/docs/v1.9/components/data-connectors/github): GitHub Data Connector Documentation - [Glue Data Connector](/docs/v1.9/components/data-connectors/glue): Glue Data Connector Documentation - [GraphQL Data Connector](/docs/v1.9/components/data-connectors/graphql): GraphQL Data Connector Documentation - [HTTP(s) Data Connector](/docs/v1.9/components/data-connectors/https): HTTP(s) Data Connector Documentation - [Iceberg Data Connector](/docs/v1.9/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/v1.9/components/data-connectors/imap): IMAP Data Connector Documentation - [Kafka Data Connector](/docs/v1.9/components/data-connectors/kafka): Kafka Data Connector Documentation - [Localpod Data Connector](/docs/v1.9/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/v1.9/components/data-connectors/memory): Memory Data Connector Documentation - [MongoDB Data Connector](/docs/v1.9/components/data-connectors/mongodb): MongoDB Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/v1.9/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [MySQL Data Connector](/docs/v1.9/components/data-connectors/mysql): MySQL Data Connector Documentation - [ODBC Data Connector](/docs/v1.9/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/v1.9/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/v1.9/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [Amazon Redshift Data Connector](/docs/v1.9/components/data-connectors/redshift): Connect to Amazon Redshift using the PostgreSQL connector in Spice. - [S3 Data Connector](/docs/v1.9/components/data-connectors/s3): S3 Data Connector Documentation - [SharePoint Data Connector](/docs/v1.9/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [Snowflake Data Connector](/docs/v1.9/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/v1.9/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/v1.9/components/data-connectors/spiceai): Spice.ai Data Connector Documentation - [Embedding Models](/docs/v1.9/components/embeddings): Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. - [Azure OpenAI Embedding Models](/docs/v1.9/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/v1.9/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/v1.9/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [HuggingFace Text Embedding Models](/docs/v1.9/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Local Filesystem Embedding Models](/docs/v1.9/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [Model2Vec Embedding Models](/docs/v1.9/components/embeddings/model2vec): Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning. - [OpenAI (or Compatible) Embedding Models](/docs/v1.9/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [Model Providers](/docs/v1.9/components/models): Overview of supported model providers for ML and LLMs in Spice. - [Anthropic Models](/docs/v1.9/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/v1.9/components/models/azure): Instructions for using Azure OpenAI models - [Amazon Bedrock Models](/docs/v1.9/components/models/bedrock): How to use Amazon Bedrock models with Spice. - [Databricks Model Provider](/docs/v1.9/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/v1.9/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [HuggingFace](/docs/v1.9/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [OpenAI (or Compatible) Language Models](/docs/v1.9/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [Perplexity Models](/docs/v1.9/components/models/perplexity): Instructions for using language models hosted on Perplexity with Spice. - [Spice Cloud Platform](/docs/v1.9/components/models/spiceai): Instructions for using models hosted on the Spice Cloud Platform with Spice. - [xAI Models](/docs/v1.9/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/v1.9/components/secret-stores) - [AWS Secrets Manager Secret Store](/docs/v1.9/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Environment Secret Store](/docs/v1.9/components/secret-stores/env): Environment Variables Secret Store Documentation - [Keyring Secret Store](/docs/v1.9/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/v1.9/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/v1.9/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/v1.9/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Server-Sent Events (SSE) protocol. - [Web Search Tool](/docs/v1.9/components/tools/websearch): The Web Search Tool enables Spice models to search the web for information. The tool is available through the websearch tool, and backed by different search engines. - [Vector Engines](/docs/v1.9/components/vectors) - [Amazon S3 Vectors Engine](/docs/v1.9/components/vectors/s3_vectors): Amazon S3 Vectors Engine Documentation - [Views](/docs/v1.9/components/views): Documentation for defining Views in Spice - [Workers Overview](/docs/v1.9/components/workers): Detailed documentation for workers in the Spice runtime. - [Deployment](/docs/v1.9/deployment): Deploy Spice.ai in your environment - [Deployment Architectures](/docs/v1.9/deployment/architectures): Spice.ai Open Source Deployment architectures - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/v1.9/deployment/architectures/cluster): Deploying Spice as a cluster - [Cloud Hosted](/docs/v1.9/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Microservice Deployment (Single or Multiple Replicas)](/docs/v1.9/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/v1.9/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/v1.9/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/v1.9/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/v1.9/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [Spice Cloud Platform Deployment](/docs/v1.9/deployment/cloud): Guide to deploying data and AI applications using the managed Spice Cloud Platform - [Docker - Kubernetes](/docs/v1.9/deployment/docker): Running Spice.ai as Docker container - [Docker Sandbox Guide - v1.3.0](/docs/v1.9/deployment/docker/sandbox): Migrating to v1.3.0 - [Helm - Kubernetes](/docs/v1.9/deployment/kubernetes): Deploy Spice.ai in Kubernetes using Helm. - [Frequently Asked Questions](/docs/v1.9/faq): Get answers to common questions about Spice.ai, including its features, differences from other tools, and use cases - [Features](/docs/v1.9/features): Features - [Caching](/docs/v1.9/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/v1.9/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Data Acceleration](/docs/v1.9/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/v1.9/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/v1.9/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Indexes](/docs/v1.9/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/v1.9/features/data-acceleration/partitioning): Partitioning for accelerated datasets - [Snapshots](/docs/v1.9/features/data-acceleration/snapshots): Bootstrap file-mode accelerations from managed snapshots to eliminate cold starts. - [Data Ingestion](/docs/v1.9/features/data-ingestion): Learn how to ingest data in Spice. - [Distributed Query](/docs/v1.9/features/distributed-query): Learn how to run Spice in distributed mode for larger scale queries. - [Embedding Datasets](/docs/v1.9/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Large Language Models](/docs/v1.9/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models](/docs/v1.9/features/large-language-models/evals): Learn how Spice evaluates, tracks, compares, and improves language model performance for specific tasks - [Model Context Protocol (MCP)](/docs/v1.9/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/v1.9/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/v1.9/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/v1.9/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/v1.9/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/v1.9/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/v1.9/features/machine-learning-models): Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform. - [Observability & Monitoring](/docs/v1.9/features/observability): Learn how to use Spice telemetry. - [Component Metrics](/docs/v1.9/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/v1.9/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Search Functionality](/docs/v1.9/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/v1.9/features/search/full-text): Learn how Spice can perform full text search - [Vector-Based Search](/docs/v1.9/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/v1.9/features/semantic-model): Learn how to define and use semantic data models with Spice. - [Web Search](/docs/v1.9/features/web-search): Learn how Spice can perform web search - [Getting started with Spice.ai OSS](/docs/v1.9/getting-started): Get started with Spice in 5 minutes - [Community Data](/docs/v1.9/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/v1.9/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/v1.9/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. - [Spice.ai OSS Installation](/docs/v1.9/installation): Instructions for installing Spice.ai OSS - [Intelligent Applications](/docs/v1.9/intelligent-applications): Building intelligent data and AI-driven applications with Spice.ai - [Monitoring](/docs/v1.9/monitoring): Monitoring Spice.ai deployments - [Datadog](/docs/v1.9/monitoring/datadog): Monitoring Spice with Datadog - [Grafana & Prometheus](/docs/v1.9/monitoring/grafana): Monitoring Spice instances with Grafana & Prometheus - [Zipkin Integration](/docs/v1.9/monitoring/zipkin): Learn how to integrate Spice with Zipkin tracing. - [Spice.ai OSS Reference Docs](/docs/v1.9/reference): Reference documentation on the Spice API, CLI and Pod manifest syntax. - [Cron Schedules](/docs/v1.9/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/v1.9/reference/datatypes) - [Accelerator Data Types](/docs/v1.9/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/v1.9/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Duration](/docs/v1.9/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/v1.9/reference/file_format): Spice currently supports CSV, JSON, and Parquet data file-formats for data connectors that can read files from a file system or cloud object storage (i.e. s3//, file://, etc.). Support for Iceberg and other file-formats are on the roadmap. - [Managing Memory Usage](/docs/v1.9/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice.ai Open Source deployments. - [Models Grade Report](/docs/v1.9/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [YAML syntax for Spicepod manifests](/docs/v1.9/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/v1.9/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/v1.9/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/v1.9/reference/spicepod/embeddings): Embeddings YAML reference - [evals](/docs/v1.9/reference/spicepod/evals): Evaluations YAML reference - [Reserved Keywords](/docs/v1.9/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/v1.9/reference/spicepod/models): Models YAML reference - [Runtime](/docs/v1.9/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/v1.9/reference/spicepod/tools): Tools YAML reference - [Views](/docs/v1.9/reference/spicepod/views): Views YAML reference - [Workers](/docs/v1.9/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/v1.9/reference/sql): This section provides a comprehensive reference for SQL support in Spice.ai, including syntax, data types, operators, functions, and system features. The reference is organized by topic for ease of navigation. - [Aggregate Functions](/docs/v1.9/reference/sql/aggregate_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. Note, when using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [AI Functions](/docs/v1.9/reference/sql/ai): AI functions in Spice provide direct integration with large language models (LLMs) and embedding models within SQL queries. These functions process text through configured model providers and return generated responses or vector embeddings. - [DML (Data Manipulation Language)](/docs/v1.9/reference/sql/dml): Data Manipulation Language (DML) statements for inserting and modifying data in Spice. - [Explain](/docs/v1.9/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/v1.9/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [JSON Functions and Operators](/docs/v1.9/reference/sql/json): Reference for JSON functions and operators in Spice SQL - [Operators](/docs/v1.9/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/v1.9/reference/sql/prepared_statements): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Scalar Functions](/docs/v1.9/reference/sql/scalar_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. Note, when using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [Search in SQL](/docs/v1.9/reference/sql/search): Reference for search functions and filtering in Spice SQL. - [SELECT](/docs/v1.9/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/v1.9/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/v1.9/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/v1.9/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. - [SDKs](/docs/v1.9/sdks): Connect to spice, using official Spice SDKs - [Dotnet SDK for Spice.ai](/docs/v1.9/sdks/dotnet): Connect to Spice using Spice Dotnet SDK - [Go SDK](/docs/v1.9/sdks/golang): Connect to spice using spice go SDK - [Java SDK](/docs/v1.9/sdks/java): Connect to Spice using Spice Java SDK - [JavaScript SDK](/docs/v1.9/sdks/javascript): Connect to spice using Spice.js SDK - [Python SDK](/docs/v1.9/sdks/python): Connect to spice using spice python SDK - [Rust SDK](/docs/v1.9/sdks/rust): Connect to spice using spice rust SDK - [Troubleshooting Spice](/docs/v1.9/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Spice.ai Use Cases](/docs/v1.9/use-cases): Use cases for Spice.ai' - [AI Applications and Agents](/docs/v1.9/use-cases/ai): AI Applications and Agents - [Agentic AI Applications and Agents](/docs/v1.9/use-cases/ai/agentic-apps): Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. - [Edge-Enabled AI Applications and Agents](/docs/v1.9/use-cases/ai/edge-ai): Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases. - [Federated MCP Client for Distributed Tool Ecosystems](/docs/v1.9/use-cases/ai/federated-mcp-server): Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, enhancing threat analysis. - [Object-Store Based SQL Query, Search, and LLM Inference Engine](/docs/v1.9/use-cases/ai/object-store-ai-engine): Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. - [Real-Time Decision-Making for Intelligent Applications](/docs/v1.9/use-cases/ai/real-time-decision-making): Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. - [Tool-Augmented AI with Model Context Protocol Server](/docs/v1.9/use-cases/ai/tool-calling-ai): Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. - [Data Federation, Acceleration, and SQL Query](/docs/v1.9/use-cases/data): Data Federation, Acceleration, and SQL Query - [Application Resilience and Performance Optimization](/docs/v1.9/use-cases/data/application-resilence-and-acceleration): Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance. - [Data Mesh for Unified Data Access](/docs/v1.9/use-cases/data/data-mesh): Spice.ai enables unified data access across disparate sources for health-tech applications, fostering a data mesh architecture. - [Database CDN for Enhanced Performance](/docs/v1.9/use-cases/data/database-cdn): Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience. - [ETL-free Workflows and Data Migrations](/docs/v1.9/use-cases/data/etl-free-workflows): Spice.ai enables data migrations and workflows without ETL by federating legacy and modern systems for seamless transitions. - [Object-Store Data Engine](/docs/v1.9/use-cases/data/object-store-data-engine): Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses. - [Reverse-ETL for Operational Workflows](/docs/v1.9/use-cases/data/reverse-etl): Spice.ai serves enriched data from warehouses to operational systems for real-time actions, eliminating complex ETL pipelines. - [Retrieval-Augmented-Generation (RAG)](/docs/v1.9/use-cases/rag): Retrieval-Augmented-Generation (RAG) - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/v1.9/use-cases/rag/applications): Use Spice for Retrieval-Augmented-Generation (RAG) - [Retrieval-Augmented Generation for AI-Powered Reporting](/docs/v1.9/use-cases/rag/reporting): Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. - [Search & Retrieval](/docs/v1.9/use-cases/search): Search & Retrieval - [Simplifying Real-Time Data Collection and Search](/docs/v1.9/use-cases/search/data-collection-and-search): Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic. - [Enterprise Search and Retrieval](/docs/v1.9/use-cases/search/enterprise-search): Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities. - [Object-Store Native Search Engine](/docs/v1.9/use-cases/search/object-store-search-engine): Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search. ### v2.0 Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven applications and AI agents. Learn about data federation, acceleration, RAG, and building intelligent apps. - [Spice.ai Open Source](/docs/v2.0): Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven applications and AI agents. Learn about data federation, acceleration, RAG, and building intelligent apps. - [Tags](/docs/v2.0/tags) - [One doc tagged with "Acknowledgements"](/docs/v2.0/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [3 docs tagged with "ADBC"](/docs/v2.0/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/v2.0/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Argo CD"](/docs/v2.0/tags/argocd): Argo CD declarative GitOps continuous delivery for Kubernetes. - [2 docs tagged with "Arrow"](/docs/v2.0/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/v2.0/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/v2.0/tags/auth): Authentication and authorization mechanisms. - [2 docs tagged with "Authentication"](/docs/v2.0/tags/authentication): User authentication methods and security protocols. - [5 docs tagged with "Azure"](/docs/v2.0/tags/azure): Microsoft Azure cloud services and integrations. - [2 docs tagged with "Blob Storage"](/docs/v2.0/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Caching"](/docs/v2.0/tags/caching): Data caching strategies and performance optimization. - [13 docs tagged with "Catalogs"](/docs/v2.0/tags/catalogs): Data catalog connectors and metadata management. - [2 docs tagged with "Cayenne"](/docs/v2.0/tags/cayenne): Cayenne (Vortex) data accelerator built on Vortex columnar format. - [2 docs tagged with "CLI"](/docs/v2.0/tags/cli): Spice command-line interface commands and usage. - [5 docs tagged with "Component Metrics"](/docs/v2.0/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/v2.0/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/v2.0/tags/configuration): System configuration files and runtime settings. - [2 docs tagged with "Cosmos DB"](/docs/v2.0/tags/cosmosdb): Azure Cosmos DB (NoSQL / Core SQL) data connector. - [One doc tagged with "C#"](/docs/v2.0/tags/csharp): C# programming language topics. - [7 docs tagged with "Data Accelerators"](/docs/v2.0/tags/data-accelerators): Data acceleration engines for high-performance query execution. - [48 docs tagged with "Data Connectors"](/docs/v2.0/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/v2.0/tags/data-lake): Data lake architectures and storage solutions. - [3 docs tagged with "Databricks"](/docs/v2.0/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/v2.0/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debezium"](/docs/v2.0/tags/debezium): Debezium change data capture integration. - [One doc tagged with "Debugging"](/docs/v2.0/tags/debugging): Debugging techniques and troubleshooting tools. - [3 docs tagged with "Delta Lake"](/docs/v2.0/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/v2.0/tags/dependencies): Third-party dependencies and software requirements. - [8 docs tagged with "Deployment"](/docs/v2.0/tags/deployment): Production deployment using Docker and Kubernetes. - [2 docs tagged with "Docker"](/docs/v2.0/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with ".NET"](/docs/v2.0/tags/dotnet): .NET framework topics. - [One doc tagged with "Dremio"](/docs/v2.0/tags/dremio): Dremio data lake engine integration. - [2 docs tagged with "DuckDB"](/docs/v2.0/tags/duckdb): DuckDB embedded analytical database integration. - [2 docs tagged with "DuckLake"](/docs/v2.0/tags/ducklake): DuckLake catalog integration. - [2 docs tagged with "DynamoDB"](/docs/v2.0/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [2 docs tagged with "Elasticsearch"](/docs/v2.0/tags/elasticsearch): Elasticsearch data connector and vector engine integration. - [7 docs tagged with "Embeddings"](/docs/v2.0/tags/embeddings): Vector embeddings and semantic similarity operations. - [6 docs tagged with "Features"](/docs/v2.0/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/v2.0/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "File"](/docs/v2.0/tags/file): File-based data connectors and local file access. - [One doc tagged with "Flux"](/docs/v2.0/tags/flux): Flux CD GitOps toolkit for Kubernetes. - [One doc tagged with "Flux CD"](/docs/v2.0/tags/fluxcd): Flux CD GitOps toolkit for Kubernetes. - [2 docs tagged with "Functions"](/docs/v2.0/tags/functions): User-defined SQL scalar functions and remote function endpoints. - [One doc tagged with "Getting Started"](/docs/v2.0/tags/getting-started): Installation guides and quickstart tutorials. - [3 docs tagged with "GitHub"](/docs/v2.0/tags/github): GitHub repository data connector and integration. - [3 docs tagged with "GitOps"](/docs/v2.0/tags/gitops): GitOps continuous delivery patterns for Kubernetes. - [2 docs tagged with "Glue"](/docs/v2.0/tags/glue): AWS Glue ETL service and data catalog integration. - [One doc tagged with "Go"](/docs/v2.0/tags/go): Go programming language topics. - [One doc tagged with "Golang"](/docs/v2.0/tags/golang): Go (Golang) programming language topics. - [One doc tagged with "GraphQL"](/docs/v2.0/tags/graphql): GraphQL API data connectors and query language support. - [3 docs tagged with "Helm"](/docs/v2.0/tags/helm): Helm package manager for deploying Spice.ai on Kubernetes. - [One doc tagged with "HTTPS"](/docs/v2.0/tags/https): HTTPS data connectors and secure web API access. - [2 docs tagged with "Hugging Face"](/docs/v2.0/tags/huggingface): Hugging Face model hub and transformer model integration. - [2 docs tagged with "Iceberg"](/docs/v2.0/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/v2.0/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/v2.0/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "Java"](/docs/v2.0/tags/java): Java programming language topics. - [One doc tagged with "JavaScript"](/docs/v2.0/tags/javascript): JavaScript programming language topics. - [One doc tagged with "JDBC"](/docs/v2.0/tags/jdbc): Java Database Connectivity driver configuration. - [One doc tagged with "Kafka"](/docs/v2.0/tags/kafka): Apache Kafka event streaming platform integration. - [6 docs tagged with "Kubernetes"](/docs/v2.0/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "libSQL"](/docs/v2.0/tags/libsql): libSQL database topics. - [One doc tagged with "Local"](/docs/v2.0/tags/local): Local model execution and on-premise services. - [One doc tagged with "Logging"](/docs/v2.0/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/v2.0/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/v2.0/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/v2.0/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/v2.0/tags/memory): In-memory data connectors and volatile storage. - [16 docs tagged with "Models"](/docs/v2.0/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MongoDB"](/docs/v2.0/tags/mongodb): MongoDB NoSQL document database integration. - [One doc tagged with "MSSQL"](/docs/v2.0/tags/mssql): Microsoft SQL Server database integration. - [3 docs tagged with "MySQL"](/docs/v2.0/tags/mysql): MySQL relational database integration. - [One doc tagged with "Node.js"](/docs/v2.0/tags/nodejs): Node.js runtime environment topics. - [3 docs tagged with "NoSQL"](/docs/v2.0/tags/nosql): NoSQL database systems and document stores. - [29 docs tagged with "Observability"](/docs/v2.0/tags/observability): Monitoring, tracing, and metrics for system visibility. - [One doc tagged with "ODBC"](/docs/v2.0/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/v2.0/tags/open-source): Open source licenses and community contributions. - [3 docs tagged with "OpenAI"](/docs/v2.0/tags/openai): OpenAI API integration and GPT model access. - [2 docs tagged with "Oracle"](/docs/v2.0/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/v2.0/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/v2.0/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/v2.0/tags/parameters): Configuration parameters and runtime options. - [One doc tagged with "Performance"](/docs/v2.0/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/v2.0/tags/persistence): Data persistence and durable storage mechanisms. - [4 docs tagged with "Postgres"](/docs/v2.0/tags/postgres): PostgreSQL database integration and configuration. - [One doc tagged with "Python"](/docs/v2.0/tags/python): Python programming language topics. - [3 docs tagged with "Query"](/docs/v2.0/tags/query): SQL query execution, parameterized queries, and prepared statements. - [4 docs tagged with "Reference"](/docs/v2.0/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/v2.0/tags/relational): Relational database systems and SQL operations. - [3 docs tagged with "Runtime"](/docs/v2.0/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Rust"](/docs/v2.0/tags/rust): Rust programming language topics. - [2 docs tagged with "S3"](/docs/v2.0/tags/s3): Amazon S3 object storage integration. - [One doc tagged with "S3 Express"](/docs/v2.0/tags/s3-express): Amazon S3 Express One Zone storage. - [One doc tagged with "Sandbox"](/docs/v2.0/tags/sandbox): Sandbox environments and isolated testing. - [One doc tagged with "ScyllaDB"](/docs/v2.0/tags/scylladb): ScyllaDB database integration. - [6 docs tagged with "SDK"](/docs/v2.0/tags/sdk): Software Development Kit topics and usage. - [10 docs tagged with "Search"](/docs/v2.0/tags/search): Vector search, semantic search, and ranking capabilities. - [2 docs tagged with "Security"](/docs/v2.0/tags/security): Security features and data protection mechanisms. - [One doc tagged with "Snowflake"](/docs/v2.0/tags/snowflake): Snowflake cloud data warehouse integration. - [7 docs tagged with "SpiceAI"](/docs/v2.0/tags/spiceai): SpiceAI cloud platform and managed services. - [5 docs tagged with "Spicepod"](/docs/v2.0/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [6 docs tagged with "SQL"](/docs/v2.0/tags/sql): SQL query language support and database operations. - [5 docs tagged with "Tools"](/docs/v2.0/tags/tools): Development tools and utility integrations. - [2 docs tagged with "Tracing"](/docs/v2.0/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/v2.0/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Turso"](/docs/v2.0/tags/turso): Turso data accelerator integration. - [2 docs tagged with "UDF"](/docs/v2.0/tags/udf): User-defined functions registered with the SQL engine. - [2 docs tagged with "Unity Catalog"](/docs/v2.0/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [One doc tagged with "Views"](/docs/v2.0/tags/views): Virtual views and data transformation layers. - [One doc tagged with "Vortex"](/docs/v2.0/tags/vortex): Vortex columnar file format and storage engine. - [10 docs tagged with "Write"](/docs/v2.0/tags/write): Data connectors and catalogs that support write operations. - [One doc tagged with "YAML"](/docs/v2.0/tags/yaml): YAML configuration syntax and file formats. - [One doc tagged with "Zipkin"](/docs/v2.0/tags/zipkin): Zipkin distributed tracing system integration. - [Open Source Acknowledgements](/docs/v2.0/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: - [Spice.ai API Reference](/docs/v2.0/api): Spice.ai API reference including HTTP REST API, Arrow Flight SQL, JDBC, ODBC, ADBC connectors, authentication, and TLS configuration. - [ADBC: Arrow Database Connectivity](/docs/v2.0/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/v2.0/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/v2.0/api/auth): Authentication documentation - [Generate Package](/docs/v2.0/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/v2.0/api/HTTP/get-catalogs): Returns a list of all registered catalogs (data sources). Catalogs provide metadata about schemas and tables available from external data sources. - [Get Iceberg API config](/docs/v2.0/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/v2.0/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/v2.0/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [ML Prediction](/docs/v2.0/api/HTTP/get-model-predict): Make a ML prediction using a specific model. - [List Models](/docs/v2.0/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [Check if a namespace exists.](/docs/v2.0/api/HTTP/get-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Spicepods](/docs/v2.0/api/HTTP/get-spicepods): Get a list of spicepods and their summary details. - [Check Runtime Status](/docs/v2.0/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Get a table.](/docs/v2.0/api/HTTP/get-table): This endpoint returns the table if it exists, otherwise it returns a 404 Not Found response. - [List Workers](/docs/v2.0/api/HTTP/get-workers): Returns a list of all registered workers in the runtime. Workers are configurable processing units that can perform tasks like load balancing between models or implementing fallback strategies. - [Check Namespace exists](/docs/v2.0/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [Check if a table exists.](/docs/v2.0/api/HTTP/head-table): This endpoint returns a 200 OK response if the table exists, otherwise it returns a 404 Not Found response. - [list_tables](/docs/v2.0/api/HTTP/list-tables): list_tables - [List Tools](/docs/v2.0/api/HTTP/list-tools): Returns a list of all available tools in the Spice runtime. Tools provide reusable functionality that can be invoked programmatically or by AI agents. - [Send a Model Context Protocol message](/docs/v2.0/api/HTTP/mcp-message): Send a JSON-RPC message to the Spice MCP server using the MCP Streamable HTTP transport. The response is either a single JSON-RPC response (`application/json`) or an SSE stream (`text/event-stream`), selected via the `Accept` header. Session continuity is carried via the `Mcp-Session-Id` header. - [Open an MCP server-to-client SSE stream](/docs/v2.0/api/HTTP/mcp-stream): Open a long-lived server-to-client SSE stream for the current MCP session as defined by the Streamable HTTP transport. The `Mcp-Session-Id` header must identify an existing session created via `POST /v1/mcp`. - [Terminate an MCP Streamable HTTP session](/docs/v2.0/api/HTTP/mcp-terminate-session): Terminate the MCP session identified by the `Mcp-Session-Id` header. Subsequent requests bearing the same session id will receive `404 Not Found`. - [Update Refresh SQL](/docs/v2.0/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Batch ML Predictions](/docs/v2.0/api/HTTP/post-batch-predict): Perform a batch of ML predictions, using multiple models, in one request. This is useful for ensembling or A/B testing different models. - [Create Chat Completion](/docs/v2.0/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/v2.0/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/v2.0/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Text-to-SQL (NSQL)](/docs/v2.0/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [post_responses](/docs/v2.0/api/HTTP/post-responses): post_responses - [Search](/docs/v2.0/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/v2.0/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/v2.0/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [Run Tool](/docs/v2.0/api/HTTP/run-tool): Execute a specific tool by name. The request body schema and response format are defined by each individual tool's specification. Use `GET /v1/tools` to discover available tools and their parameter schemas. - [runtime](/docs/v2.0/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/v2.0/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/v2.0/api/odbc): ODBC API Documentation - [API Overview](/docs/v2.0/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/v2.0/api/tls): Encryption in transit with TLS documentation - [Spice.ai CLI Reference](/docs/v2.0/cli): Complete CLI reference for Spice.ai including commands to create, manage Spicepods, run queries, and interact with the Spice runtime. - [Spice.ai OSS CLI command reference](/docs/v2.0/cli/reference): Spice CLI command reference - [add](/docs/v2.0/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/v2.0/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/v2.0/cli/reference/chat): spice chat CLI documentation - [completions](/docs/v2.0/cli/reference/completions): Generate shell completions for the Spice CLI. - [connect](/docs/v2.0/cli/reference/connect): Connect to an app on the Spice.ai Cloud Platform. - [dataset](/docs/v2.0/cli/reference/dataset): Add or configure dataset entries in spicepod.yaml. - [datasets](/docs/v2.0/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [feedback](/docs/v2.0/cli/reference/feedback): Open the Spice.ai community Slack in the default browser to share feedback. - [init](/docs/v2.0/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/v2.0/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/v2.0/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/v2.0/cli/reference/models): Lists models loaded by the Spice runtime - [nsql](/docs/v2.0/cli/reference/nsql): spice nsql CLI documentation - [pods](/docs/v2.0/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [query](/docs/v2.0/cli/reference/query): Submit an async query or start an interactive async query REPL against the Spice runtime's distributed query engine. - [refresh](/docs/v2.0/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/v2.0/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/v2.0/cli/reference/search): Performs embeddings-based searches across search-configured datasets. - [spiced](/docs/v2.0/cli/reference/spiced): Command-line reference for the spiced runtime binary — flags, defaults, and differences from spice run. - [sql](/docs/v2.0/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/v2.0/cli/reference/status): Spice runtime status - [trace](/docs/v2.0/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/v2.0/cli/reference/upgrade): Upgrades the Spice CLI and runtime to the latest or specified version - [validate](/docs/v2.0/cli/reference/validate): Validate a spicepod.yaml without starting the runtime. - [version](/docs/v2.0/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/v2.0/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels - [Clients and Tools](/docs/v2.0/clients): Client and tools for connecting to Spice - [DBeaver](/docs/v2.0/clients/dbeaver): Configure DBeaver to query Spice via JDBC - [JetBrains DataGrip](/docs/v2.0/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Microsoft Power BI Connector](/docs/v2.0/clients/powerbi): Use Microsoft Power BI to access, visualize and analyze Spice datasets. - [Apache Superset](/docs/v2.0/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/v2.0/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. - [How Spice Compares](/docs/v2.0/comparison): Compare Spice.ai with data platforms (Databricks, Snowflake), query engines (Trino, Dremio, ClickHouse), vector databases (Turbopuffer, LanceDB), search engines (Elasticsearch), and AI frameworks (LangChain, LlamaIndex, Ollama). - [Spice.ai Runtime Components](/docs/v2.0/components): Configure Spice.ai runtime components including data connectors, data accelerators, catalog connectors, model providers, embedding models, and secret stores. - [Catalog Connectors](/docs/v2.0/components/catalogs): Connect to external catalog providers like Unity Catalog, Databricks, Iceberg, AWS Glue, Snowflake, ADBC, PostgreSQL, MySQL, MSSQL, Oracle, and more for federated SQL query in Spice. - [ADBC Catalog Connector](/docs/v2.0/components/catalogs/adbc): Connect to databases via ADBC for automatic schema and table discovery. - [Databricks Catalog Connector](/docs/v2.0/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [DuckLake Catalog Connector](/docs/v2.0/components/catalogs/ducklake): Connect to a DuckLake catalog for federated SQL query. - [Glue Catalog Connector](/docs/v2.0/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/v2.0/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Microsoft SQL Server Catalog Connector](/docs/v2.0/components/catalogs/mssql): Connect to a Microsoft SQL Server database as a catalog provider for federated SQL query. - [MySQL Catalog Connector](/docs/v2.0/components/catalogs/mysql): Connect to a MySQL database as a catalog provider for federated SQL query. - [Oracle Catalog Connector](/docs/v2.0/components/catalogs/oracle): Connect to an Oracle database as a catalog provider for federated SQL query. - [PostgreSQL Catalog Connector](/docs/v2.0/components/catalogs/postgres): Connect to a PostgreSQL database as a catalog provider for federated SQL query. - [Snowflake Catalog Connector](/docs/v2.0/components/catalogs/snowflake): Connect to a Snowflake database as a catalog provider for federated SQL query. - [Spice.ai Catalog Connector](/docs/v2.0/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/v2.0/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Unity Catalog Catalog Connector Deployment Guide](/docs/v2.0/components/catalogs/unity-catalog/deployment): Operating guide for the Unity Catalog catalog connector in production: workspace authentication, table-type filtering, effective-permissions flow, and observability. - [Data Accelerators](/docs/v2.0/components/data-accelerators): Data acceleration engines for local materialization and query acceleration in Spice - [In-Memory Arrow Data Accelerator](/docs/v2.0/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [Arrow Data Accelerator Deployment Guide](/docs/v2.0/components/data-accelerators/arrow/deployment): Operating guide for the Arrow (in-memory) data accelerator in production: memory sizing, indexes, and observability. - [Spice Cayenne Data Accelerator](/docs/v2.0/components/data-accelerators/cayenne): Spice Cayenne Data Accelerator (Vortex) Documentation - [Cayenne Data Accelerator Deployment Guide](/docs/v2.0/components/data-accelerators/cayenne/deployment): Operating guide for Spice Cayenne in production: footer and segment caches, S3 Express, metastore durability, and observability. - [DuckDB Data Accelerator](/docs/v2.0/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [DuckDB Data Accelerator Deployment Guide](/docs/v2.0/components/data-accelerators/duckdb/deployment): Operating guide for the DuckDB data accelerator in production: memory vs file mode, checkpointing, spill, pool sizing, and observability. - [PostgreSQL Data Accelerator](/docs/v2.0/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [PostgreSQL Data Accelerator Deployment Guide](/docs/v2.0/components/data-accelerators/postgres/deployment): Operating guide for the PostgreSQL data accelerator in production: authentication, connection pooling, and observability. - [SQLite Data Accelerator](/docs/v2.0/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [SQLite Data Accelerator Deployment Guide](/docs/v2.0/components/data-accelerators/sqlite/deployment): Operating guide for the SQLite data accelerator in production: file mode, busy timeout, pool, and observability. - [Turso Data Accelerator](/docs/v2.0/components/data-accelerators/turso): Turso (libSQL) Data Accelerator Documentation - [Data Connectors](/docs/v2.0/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/v2.0/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ADBC Data Connector](/docs/v2.0/components/data-connectors/adbc): ADBC Data Connector Documentation - [ClickHouse Data Connector](/docs/v2.0/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Azure Cosmos DB Data Connector](/docs/v2.0/components/data-connectors/cosmosdb): Query Azure Cosmos DB (NoSQL / Core SQL API) containers as SQL tables in Spice. Read-only scan with schema inferred from a sample of documents. - [Azure Cosmos DB Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/cosmosdb/deployment): Operating guide for the Azure Cosmos DB data connector in production: authentication, RU sizing, resilience, metrics, and observability. - [Databricks Data Connector](/docs/v2.0/components/data-connectors/databricks): Databricks Data Connector Documentation - [Databricks Deployment Guide](/docs/v2.0/components/data-connectors/databricks/deployment): Operating guide for the Databricks connector in production: resilience controls, Unity Catalog behavior, metrics, and observability. - [Debezium Data Connector](/docs/v2.0/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/v2.0/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Delta Lake Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/delta-lake/deployment): Operating guide for the Delta Lake data connector in production: object store auth, metadata caching, metrics, and observability. - [Dremio Data Connector](/docs/v2.0/components/data-connectors/dremio): Dremio Data Connector Documentation - [Dremio Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/dremio/deployment): Operating guide for the Dremio data connector in production: authentication, Flight SQL transport, and observability. - [DuckDB Data Connector](/docs/v2.0/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DuckDB Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/duckdb/deployment): Operating guide for the DuckDB data connector in production: file vs memory mode, durability, and observability. - [DuckLake Data Connector](/docs/v2.0/components/data-connectors/ducklake): DuckLake Data Connector Documentation - [DynamoDB Data Connector](/docs/v2.0/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [DynamoDB Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/dynamodb/deployment): Operating guide for the DynamoDB data connector in production: IAM, streams, checkpointing, lag behavior, and observability. - [Elasticsearch Data Connector](/docs/v2.0/components/data-connectors/elasticsearch): Query Elasticsearch indexes as SQL tables in Spice, including kNN vector search, full-text search, and hybrid search. - [Elasticsearch Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/elasticsearch/deployment): Operating guide for the Elasticsearch data connector in production: authentication, TLS, resilience, and operational tuning. - [File Data Connector](/docs/v2.0/components/data-connectors/file): File Data Connector Documentation - [File Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/file/deployment): Operating guide for the File data connector in production: permissions, formats, performance, and observability. - [Flight SQL Data Connector](/docs/v2.0/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/v2.0/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GCS Data Connector](/docs/v2.0/components/data-connectors/gcs): GCS (Google Cloud Storage) Data Connector Documentation - [GitHub Data Connector](/docs/v2.0/components/data-connectors/github): GitHub Data Connector Documentation - [GitHub Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/github/deployment): Operating guide for the GitHub data connector in production: PATs, rate limits, pagination, and observability. - [Glue Data Connector](/docs/v2.0/components/data-connectors/glue): Connect to and query tables in an AWS Glue Data Catalog - [GraphQL Data Connector](/docs/v2.0/components/data-connectors/graphql): GraphQL Data Connector Documentation - [GraphQL Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/graphql/deployment): Operating guide for the GraphQL data connector in production: authentication, pagination, rate limits, and observability. - [HTTP(s) Data Connector](/docs/v2.0/components/data-connectors/https): HTTP(s) Data Connector Documentation - [HTTP(s) Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/https/deployment): Operating guide for the HTTP(s) data connector in production: authentication, rate control, retries, and observability. - [Iceberg Data Connector](/docs/v2.0/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/v2.0/components/data-connectors/imap): IMAP Data Connector Documentation - [Kafka Data Connector](/docs/v2.0/components/data-connectors/kafka): Kafka Data Connector Documentation - [Localpod Data Connector](/docs/v2.0/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/v2.0/components/data-connectors/memory): Memory Data Connector Documentation - [MongoDB Data Connector](/docs/v2.0/components/data-connectors/mongodb): MongoDB Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/v2.0/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [MySQL Data Connector](/docs/v2.0/components/data-connectors/mysql): MySQL Data Connector Documentation - [MySQL Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/mysql/deployment): Operating guide for the MySQL data connector in production: authentication, connection pooling, TLS, metrics, and observability. - [NFS Data Connector](/docs/v2.0/components/data-connectors/nfs): NFS Data Connector Documentation - [ODBC Data Connector](/docs/v2.0/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/v2.0/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/v2.0/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [PostgreSQL Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/postgres/deployment): Operating guide for the PostgreSQL data connector in production: authentication, connection pooling, TLS, metrics, and observability. - [Amazon Redshift Data Connector](/docs/v2.0/components/data-connectors/redshift): Connect to Amazon Redshift using the PostgreSQL connector in Spice. - [S3 Data Connector](/docs/v2.0/components/data-connectors/s3): S3 Data Connector Documentation - [S3 Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/s3/deployment): Operating guide for the S3 data connector in production: IAM, credential chains, file formats, metrics, and observability. - [ScyllaDB Data Connector](/docs/v2.0/components/data-connectors/scylladb): ScyllaDB Data Connector Documentation - [SharePoint Data Connector](/docs/v2.0/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [SMB Data Connector](/docs/v2.0/components/data-connectors/smb): SMB Data Connector Documentation - [Snowflake Data Connector](/docs/v2.0/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/v2.0/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/v2.0/components/data-connectors/spiceai): Federated SQL across Spice runtimes — Spice Cloud Platform datasets and self-hosted Spice instances (cluster-sidecar pattern). - [Spice.ai Data Connector Deployment Guide](/docs/v2.0/components/data-connectors/spiceai/deployment): Operating guide for the Spice.ai connector in production: API keys, Flight endpoints, message sizing, sidecar topology, and observability. - [Embedding Models](/docs/v2.0/components/embeddings): Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. - [Azure OpenAI Embedding Models](/docs/v2.0/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/v2.0/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/v2.0/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [Google AI Embedding Models](/docs/v2.0/components/embeddings/google): To use a hosted Google AI embedding model, specify the google path in the from field of your configuration. - [HuggingFace Text Embedding Models](/docs/v2.0/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Hugging Face Embedding Deployment Guide](/docs/v2.0/components/embeddings/huggingface/deployment): Operating guide for Hugging Face embeddings in production: tokens, download cache, pooling, device selection, and observability. - [Local Filesystem Embedding Models](/docs/v2.0/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [Local Embedding Deployment Guide](/docs/v2.0/components/embeddings/local/deployment): Operating guide for filesystem-loaded embedding models in production: formats, pooling, device selection, and observability. - [Model2Vec Embedding Models](/docs/v2.0/components/embeddings/model2vec): Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning. - [OpenAI (or Compatible) Embedding Models](/docs/v2.0/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [OpenAI Embedding Deployment Guide](/docs/v2.0/components/embeddings/openai/deployment): Operating guide for the OpenAI embedding provider in production: API keys, usage tiers, batching, retries, metrics, and observability. - [Model Providers](/docs/v2.0/components/models): Overview of supported model providers for ML and LLMs in Spice. - [Anthropic Models](/docs/v2.0/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/v2.0/components/models/azure): Instructions for using Azure OpenAI models - [Amazon Bedrock Models](/docs/v2.0/components/models/bedrock): How to use Amazon Bedrock models with Spice. - [Databricks Model Provider](/docs/v2.0/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/v2.0/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [Filesystem Model Deployment Guide](/docs/v2.0/components/models/filesystem/deployment): Operating guide for filesystem-loaded models in production: formats, device selection, memory footprint, and observability. - [Google AI Models](/docs/v2.0/components/models/google): Instructions for using language models hosted on Google AI with Spice. - [HuggingFace](/docs/v2.0/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [Hugging Face Model Deployment Guide](/docs/v2.0/components/models/huggingface/deployment): Operating guide for the Hugging Face model in production: tokens, download cache, device selection, local inference footprint, and observability. - [OpenAI (or Compatible) Language Models](/docs/v2.0/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [OpenAI Model Deployment Guide](/docs/v2.0/components/models/openai/deployment): Operating guide for the OpenAI model in production: API keys, usage tiers, rate limiting, Responses API, metrics, and observability. - [Perplexity Models (Deprecated)](/docs/v2.0/components/models/perplexity): Perplexity model support is no longer supported in Spice. - [Spice Cloud Platform](/docs/v2.0/components/models/spiceai): Instructions for using models hosted on the Spice Cloud Platform with Spice. - [xAI Models](/docs/v2.0/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/v2.0/components/secret-stores): Configure secret stores to manage sensitive data like passwords, tokens, and API keys. - [AWS Secrets Manager Secret Store](/docs/v2.0/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Azure Key Vault Secret Store](/docs/v2.0/components/secret-stores/azure-keyvault): Azure Key Vault Secret Store Documentation - [Environment Secret Store](/docs/v2.0/components/secret-stores/env): Environment Variables Secret Store Documentation - [HashiCorp Vault Secret Store](/docs/v2.0/components/secret-stores/hashicorp-vault): HashiCorp Vault Secret Store Documentation - [Keyring Secret Store](/docs/v2.0/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/v2.0/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/v2.0/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/v2.0/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Streamable HTTP transport. - [Web Search Tool (Deprecated)](/docs/v2.0/components/tools/websearch): The websearch tool is no longer supported in Spice. - [Vector Engines](/docs/v2.0/components/vectors): Configure vector engines for efficient embedding storage and similarity search in Spice. - [DuckDB Vector Engine](/docs/v2.0/components/vectors/duckdb): Use DuckDB as a vector engine in Spice for HNSW-based vector search via the DuckDB VSS extension. - [Elasticsearch Vector Engine](/docs/v2.0/components/vectors/elasticsearch): Use Elasticsearch as a vector engine in Spice for kNN vector search, full-text search, and hybrid search. - [Amazon S3 Vectors Engine](/docs/v2.0/components/vectors/s3_vectors): Amazon S3 Vectors Engine Documentation - [Spice.ai Deployment Guide](/docs/v2.0/deployment): Deploy Spice.ai in your environment using Docker, Kubernetes, AWS, Azure, or the Spice Cloud Platform. Learn about sidecar, microservice, tiered, and cluster deployment architectures. - [Deployment Architectures](/docs/v2.0/deployment/architectures): Explore Spice deployment architectures including sidecar, microservice, tiered, sharded, and cluster configurations. - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/v2.0/deployment/architectures/cluster): Deploying Spice as a cluster - [Cluster-Sidecar Deployment](/docs/v2.0/deployment/architectures/cluster-sidecar): Deploy Spice with application-local sidecars for localhost query, search, and inference, backed by a centralized cluster for ingest, acceleration, and distributed query. - [Cloud Hosted](/docs/v2.0/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Microservice Deployment (Single or Multiple Replicas)](/docs/v2.0/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/v2.0/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/v2.0/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/v2.0/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/v2.0/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [AWS Integrations](/docs/v2.0/deployment/aws/integrations): Complete guide to Spice.ai integrations with Amazon Web Services, including data connectors, AI models, vector stores, and secret management. - [Azure Deployment Options](/docs/v2.0/deployment/azure): Guide to deploying Spice.ai applications on Microsoft Azure - [Azure Integrations](/docs/v2.0/deployment/azure/integrations): Spice.ai integrations with Microsoft Azure, including data connectors, AI models, embeddings, and authentication. - [CI/CD Deployment](/docs/v2.0/deployment/ci-cd): Deploy Spice.ai applications using continuous integration and delivery pipelines, including Helm, Kubernetes GitOps with Argo CD or Flux, GitHub Actions, and the Spice Cloud deploy action. - [Spice Cloud Platform Deployment](/docs/v2.0/deployment/cloud): Guide to deploying data and AI applications using the managed Spice Cloud Platform - [Docker](/docs/v2.0/deployment/docker): Run Spice.ai as a Docker container. - [Docker Sandbox Guide - v1.3.0](/docs/v2.0/deployment/docker/sandbox): Migrating to v1.3.0 - [Google Cloud Deployment Options](/docs/v2.0/deployment/gcp): Guide to deploying Spice.ai applications on Google Cloud Platform (GCP). - [GCP Integrations](/docs/v2.0/deployment/gcp/integrations): Spice.ai integrations with Google Cloud Platform, including data connectors, AI models, embeddings, and authentication. - [Kubernetes Deployment](/docs/v2.0/deployment/kubernetes): Deploy Spice.ai on Kubernetes using Helm, Argo CD, or Flux. - [Kubernetes - Argo CD](/docs/v2.0/deployment/kubernetes/argocd): Deploy Spice.ai on self-hosted Kubernetes using Argo CD and the Spice Helm chart. - [Kubernetes - Flux](/docs/v2.0/deployment/kubernetes/flux): Deploy Spice.ai on Kubernetes using Flux CD and the Spice Helm chart. - [Kubernetes - Helm](/docs/v2.0/deployment/kubernetes/helm): Deploy Spice.ai in Kubernetes using Helm. - [Read/Write Separation](/docs/v2.0/deployment/read-write-separation): Separate write/ingest workloads (cluster) from read workloads (application sidecars, agents) using shared snapshots and live query delegation. - [Spice.ai FAQ](/docs/v2.0/faq): Answers to frequently asked questions about Spice.ai including features, use cases, differences from Trino/Presto/Dremio, federated queries, caching, and AI capabilities. - [Spice.ai Features](/docs/v2.0/features): Explore Spice.ai features including data federation, data acceleration, caching, search, LLM integration, embeddings, observability, and more for building data-driven AI applications. - [Caching](/docs/v2.0/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/v2.0/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Debezium (CDC over Kafka)](/docs/v2.0/features/cdc/debezium): Consume Debezium change events from Kafka into a Spice-accelerated dataset for sources without a native Spice CDC path. - [DynamoDB Streams (Native CDC)](/docs/v2.0/features/cdc/dynamodb-streams): Stream INSERT, UPDATE, and DELETE events from Amazon DynamoDB directly into a Spice-accelerated dataset using DynamoDB Streams. - [MongoDB Change Streams (Native CDC)](/docs/v2.0/features/cdc/mongodb-streams): Stream insert, update, replace, and delete events from MongoDB directly into a Spice-accelerated dataset using MongoDB Change Streams. - [PostgreSQL Logical Replication (Native CDC)](/docs/v2.0/features/cdc/postgres-replication): Stream INSERT, UPDATE, and DELETE events from PostgreSQL directly into a Spice-accelerated dataset using native logical replication. - [Data Acceleration](/docs/v2.0/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/v2.0/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/v2.0/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Hash Index for Arrow Acceleration](/docs/v2.0/features/data-acceleration/hash-index): Learn how to use hash indexes for O(1) point lookups on Arrow-accelerated datasets. - [Indexes](/docs/v2.0/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/v2.0/features/data-acceleration/partitioning): Partition accelerated datasets to make filtered queries faster by reading only the relevant partitions. - [Refresh Modes](/docs/v2.0/features/data-acceleration/refresh-modes): Refresh modes for accelerated datasets in Spice. - [Append Refresh Mode](/docs/v2.0/features/data-acceleration/refresh-modes/append): Incrementally append new rows to an accelerated dataset. - [Caching Refresh Mode](/docs/v2.0/features/data-acceleration/refresh-modes/caching): Learn how to use caching refresh mode for HTTP-based datasets - [Changes Refresh Mode](/docs/v2.0/features/data-acceleration/refresh-modes/changes): Apply incremental inserts, updates, and deletes via Change Data Capture. - [Full Refresh Mode](/docs/v2.0/features/data-acceleration/refresh-modes/full): Replace the entire accelerated dataset on each refresh. - [Snapshot Refresh Mode](/docs/v2.0/features/data-acceleration/refresh-modes/snapshot): Reload acceleration data exclusively from the snapshot store. - [Snapshots](/docs/v2.0/features/data-acceleration/snapshots): Bootstrap file-mode accelerations from managed snapshots to eliminate cold starts. - [Data Ingestion](/docs/v2.0/features/data-ingestion): Learn how to ingest data in Spice. - [Distributed Query](/docs/v2.0/features/distributed-query): Learn how to run Spice in distributed mode for larger scale queries, including the async queries API. - [Embedding Datasets](/docs/v2.0/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Functions](/docs/v2.0/features/functions): Define custom scalar and table SQL functions inline (SQL tier) or by calling remote HTTP services (Remote tier), automatically exposed as SQL functions and LLM tools. - [Large Language Models](/docs/v2.0/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models (Deprecated)](/docs/v2.0/features/large-language-models/evals): Language model evals are no longer supported in Spice. - [Model Context Protocol (MCP)](/docs/v2.0/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/v2.0/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/v2.0/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/v2.0/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/v2.0/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/v2.0/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/v2.0/features/machine-learning-models): Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform. - [Observability & Monitoring](/docs/v2.0/features/observability): Monitor Spice with Prometheus metrics, OpenTelemetry, and distributed tracing. - [Component Metrics](/docs/v2.0/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/v2.0/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Parameterized Queries](/docs/v2.0/features/query-federation/parameterized-queries): Learn how to use prepared statements and parameterized queries in Spice for improved security and performance. - [URL Tables](/docs/v2.0/features/query-federation/url-tables): Query object store files directly using URLs without pre-registering datasets - [Search Functionality](/docs/v2.0/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/v2.0/features/search/full-text): Learn how Spice can perform full text search - [Multi-Vector Search](/docs/v2.0/features/search/multi-vector): Embed list-of-strings columns as a column of vectors and use ColBERT-style late-interaction search in Spice. - [Reranking](/docs/v2.0/features/search/rerank): Rerank search results using dedicated reranker models or LLM-as-reranker for improved relevance. - [Vector-Based Search](/docs/v2.0/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/v2.0/features/semantic-model): Attach descriptions and metadata to datasets, views, and columns in Spice so LLMs, SQL functions, and humans share the same understanding of your data. - [Tool Registry](/docs/v2.0/features/tool-registry): Reduce per-turn token cost and improve LLM tool selection accuracy by replacing individual tool definitions with searchable tool_search and tool_invoke meta-tools backed by hybrid full-text, keyword, schema, and vector search. - [Views](/docs/v2.0/features/views): Documentation for defining Views in Spice - [Web Search](/docs/v2.0/features/web-search): Learn how Spice can perform web search - [Workers](/docs/v2.0/features/workers): Configure workers in the Spice runtime to coordinate interactions between LLMs and tools, with load-balancing, round-robin, and fallback strategies. - [Getting Started with Spice.ai OSS](/docs/v2.0/getting-started): Get started with Spice.ai in 5 minutes. Install the CLI, connect to datasets, run SQL queries, and use AI models with OpenAI-compatible APIs. - [Community Data](/docs/v2.0/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/v2.0/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/v2.0/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. - [Install Spice.ai OSS](/docs/v2.0/installation): Install Spice.ai OSS on macOS, Linux, Windows, or WSL using the install script, Homebrew, PowerShell, or direct download from GitHub releases. - [Building Intelligent AI Applications with Spice.ai](/docs/v2.0/intelligent-applications): Learn how to build intelligent, data-driven AI applications and agents with Spice.ai. Explore patterns for RAG, LLM integration, and real-time AI inference. - [Monitoring](/docs/v2.0/monitoring): Monitor Spice.ai deployments with Datadog, Grafana, New Relic, Prometheus, and Zipkin integrations. - [Datadog](/docs/v2.0/monitoring/datadog): Monitoring Spice with Datadog - [Grafana & Prometheus](/docs/v2.0/monitoring/grafana): Monitoring Spice instances with Grafana & Prometheus - [New Relic](/docs/v2.0/monitoring/new-relic): Monitoring Spice with New Relic - [Spice Cloud Platform](/docs/v2.0/monitoring/spice-cloud): Connect a self-hosted Spice runtime to the Spice Cloud Platform to centralize task history and runtime observability across deployments. - [Zipkin Integration](/docs/v2.0/monitoring/zipkin): Learn how to integrate Spice with Zipkin tracing. - [Spice.ai OSS Reference Docs](/docs/v2.0/reference): Reference documentation for Spice.ai including API reference, CLI commands, Spicepod configuration syntax, SQL reference, and data type specifications. - [Cron Schedules](/docs/v2.0/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/v2.0/reference/datatypes): Spice uses Apache Arrow data types internally, providing consistent type handling across different data sources and accelerators. This section documents how Arrow types map to specific accelerators and object store formats. - [Accelerator Data Types](/docs/v2.0/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/v2.0/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Spice Runtime Distributions](/docs/v2.0/reference/distributions): Distribution variants of the Spice runtime for different use cases and deployment scenarios, including data-only, GPU-accelerated, NAS, and allocator variants. - [Duration](/docs/v2.0/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/v2.0/reference/file_format): File-based data connectors — including s3//, file//, sftp://, and others — support multiple structured and document file formats. This page details the format-specific parameters available for each. - [Managing Memory Usage](/docs/v2.0/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice deployments. - [Models Grade Report](/docs/v2.0/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [Performance Tuning](/docs/v2.0/reference/performance-tuning): Comprehensive guide to optimizing query performance, acceleration, and resource utilization in Spice deployments. - [YAML syntax for Spicepod manifests](/docs/v2.0/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/v2.0/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/v2.0/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/v2.0/reference/spicepod/embeddings): Embeddings YAML reference - [Evals (Deprecated)](/docs/v2.0/reference/spicepod/evals): The evals Spicepod component is no longer supported in Spice. - [Functions (User-Defined Functions)](/docs/v2.0/reference/spicepod/functions): User-defined functions YAML reference - [Reserved Keywords](/docs/v2.0/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/v2.0/reference/spicepod/models): Models YAML reference - [Runtime](/docs/v2.0/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/v2.0/reference/spicepod/tools): Tools YAML reference - [Views](/docs/v2.0/reference/spicepod/views): Views YAML reference - [Workers](/docs/v2.0/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/v2.0/reference/sql): Complete SQL reference for Spice.ai including SELECT syntax, subqueries, DML statements, aggregate functions, AI functions, JSON operators, and search capabilities. - [Aggregate Functions](/docs/v2.0/reference/sql/aggregate_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [AI Functions](/docs/v2.0/reference/sql/ai): AI functions in Spice provide direct integration with large language models (LLMs) and embedding models within SQL queries. These functions process text through configured model providers and return generated responses or vector embeddings. - [DML (Data Manipulation Language)](/docs/v2.0/reference/sql/dml): Data Manipulation Language (DML) statements for inserting and modifying data in Spice. - [Explain](/docs/v2.0/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/v2.0/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [JSON Functions and Operators](/docs/v2.0/reference/sql/json): Reference for JSON functions and operators in Spice SQL - [Operators](/docs/v2.0/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/v2.0/reference/sql/prepared_statements): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Scalar Functions](/docs/v2.0/reference/sql/scalar_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [Search in SQL](/docs/v2.0/reference/sql/search): Reference for search functions and filtering in Spice SQL. - [SELECT](/docs/v2.0/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/v2.0/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/v2.0/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/v2.0/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. - [SDKs](/docs/v2.0/sdks): Connect to Spice using official SDKs - [Dotnet SDK](/docs/v2.0/sdks/dotnet): Connect to Spice using the Dotnet SDK - [Go SDK](/docs/v2.0/sdks/golang): Connect to Spice using the Go SDK - [Java SDK](/docs/v2.0/sdks/java): Connect to Spice using the Java SDK - [JavaScript SDK](/docs/v2.0/sdks/javascript): Connect to Spice using the JavaScript SDK - [Python SDK](/docs/v2.0/sdks/python): Connect to Spice using the Python SDK - [Rust SDK](/docs/v2.0/sdks/rust): Connect to Spice using the Rust SDK - [Troubleshooting Spice](/docs/v2.0/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Spice.ai Use Cases](/docs/v2.0/use-cases): Discover how to use Spice.ai for data federation, reverse-ETL, database CDN, enterprise search, RAG, and building AI-powered applications and agents. - [AI Applications and Agents](/docs/v2.0/use-cases/ai): AI Applications and Agents - [Agentic AI Applications and Agents](/docs/v2.0/use-cases/ai/agentic-apps): Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. - [Edge-Enabled AI Applications and Agents](/docs/v2.0/use-cases/ai/edge-ai): Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases. - [Federated MCP Client for Distributed Tool Ecosystems](/docs/v2.0/use-cases/ai/federated-mcp-server): Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, improving threat analysis. - [Multi-Tenant AI Agents](/docs/v2.0/use-cases/ai/multi-tenant-agents): Deploy AI agents across many SaaS tenants with strict isolation and no per-tenant ETL pipelines. - [Object-Store Based SQL Query, Search, and LLM Inference Engine](/docs/v2.0/use-cases/ai/object-store-ai-engine): Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. - [Real-Time Decision-Making for Intelligent Applications](/docs/v2.0/use-cases/ai/real-time-decision-making): Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. - [Tool-Augmented AI with Model Context Protocol Server](/docs/v2.0/use-cases/ai/tool-calling-ai): Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. - [Caching](/docs/v2.0/use-cases/caching): Caching use cases with Spice.ai, including write-through, read-through, SQL, S3, and HTTP caching. - [HTTP Cache](/docs/v2.0/use-cases/caching/http-cache): Use Spice.ai to cache HTTP API responses locally, reducing API call frequency and providing fast, SQL-queryable access to external API data. - [Read-Through Cache](/docs/v2.0/use-cases/caching/read-through-cache): Use Spice.ai as a read-through cache with the SQL results cache for federated data sources and HTTP APIs. - [S3 Cache](/docs/v2.0/use-cases/caching/s3-cache): Use Spice.ai to cache S3 and object store data locally for fast, repeatable SQL queries without re-reading from remote storage. - [SQL and Database Cache](/docs/v2.0/use-cases/caching/sql-database-cache): Use Spice.ai to cache SQL database tables and query results locally for low-latency access and reduced load on upstream databases. - [Write-Through Cache](/docs/v2.0/use-cases/caching/write-through-cache): Use Spice.ai as a write-through cache that writes data to both a local accelerator and the upstream data source. - [Data Federation, Acceleration, and SQL Query](/docs/v2.0/use-cases/data): Data Federation, Acceleration, and SQL Query - [Application Resilience and Performance Optimization](/docs/v2.0/use-cases/data/application-resilence-and-acceleration): Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance. - [Data Mesh for Unified Data Access](/docs/v2.0/use-cases/data/data-mesh): Give domain teams decentralized, real-time data ownership and access across disparate sources through a unified SQL interface. - [Database CDN for Enhanced Performance](/docs/v2.0/use-cases/data/database-cdn): Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience. - [ETL-free Workflows and Data Migrations](/docs/v2.0/use-cases/data/etl-free-workflows): Federate legacy and modern data systems without ETL for faster migrations, lower overhead, and zero application downtime. - [Object-Store Data Engine](/docs/v2.0/use-cases/data/object-store-data-engine): Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses. - [Reverse-ETL for Operational Workflows](/docs/v2.0/use-cases/data/reverse-etl): Serve enriched data from warehouses and data lakes to operational systems and applications, eliminating complex pipelines. - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/v2.0/use-cases/rag): Use Spice for Retrieval-Augmented-Generation (RAG) - [RAG for Contextual Applications](/docs/v2.0/use-cases/rag/applications): Build context-rich AI applications using Spice for Retrieval-Augmented Generation (RAG). - [Retrieval-Augmented Generation for AI-Powered Reporting](/docs/v2.0/use-cases/rag/reporting): Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. - [Search & Retrieval](/docs/v2.0/use-cases/search): Search & Retrieval - [Simplifying Real-Time Data Collection and Search](/docs/v2.0/use-cases/search/data-collection-and-search): Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic. - [Enterprise Search and Retrieval](/docs/v2.0/use-cases/search/enterprise-search): Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities. - [Object-Store Native Search Engine](/docs/v2.0/use-cases/search/object-store-search-engine): Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search. ### tags - [Tags](/docs/tags) - [One doc tagged with "Acknowledgements"](/docs/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [3 docs tagged with "ADBC"](/docs/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Argo CD"](/docs/tags/argocd): Argo CD declarative GitOps continuous delivery for Kubernetes. - [2 docs tagged with "Arrow"](/docs/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/tags/auth): Authentication and authorization mechanisms. - [2 docs tagged with "Authentication"](/docs/tags/authentication): User authentication methods and security protocols. - [5 docs tagged with "Azure"](/docs/tags/azure): Microsoft Azure cloud services and integrations. - [2 docs tagged with "Blob Storage"](/docs/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Caching"](/docs/tags/caching): Data caching strategies and performance optimization. - [13 docs tagged with "Catalogs"](/docs/tags/catalogs): Data catalog connectors and metadata management. - [2 docs tagged with "Cayenne"](/docs/tags/cayenne): Cayenne (Vortex) data accelerator built on Vortex columnar format. - [2 docs tagged with "CLI"](/docs/tags/cli): Spice command-line interface commands and usage. - [5 docs tagged with "Component Metrics"](/docs/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/tags/configuration): System configuration files and runtime settings. - [2 docs tagged with "Cosmos DB"](/docs/tags/cosmosdb): Azure Cosmos DB (NoSQL / Core SQL) data connector. - [One doc tagged with "C#"](/docs/tags/csharp): C# programming language topics. - [7 docs tagged with "Data Accelerators"](/docs/tags/data-accelerators): Data acceleration engines for high-performance query execution. - [48 docs tagged with "Data Connectors"](/docs/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/tags/data-lake): Data lake architectures and storage solutions. - [3 docs tagged with "Databricks"](/docs/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debezium"](/docs/tags/debezium): Debezium change data capture integration. - [One doc tagged with "Debugging"](/docs/tags/debugging): Debugging techniques and troubleshooting tools. - [3 docs tagged with "Delta Lake"](/docs/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/tags/dependencies): Third-party dependencies and software requirements. - [8 docs tagged with "Deployment"](/docs/tags/deployment): Production deployment using Docker and Kubernetes. - [2 docs tagged with "Docker"](/docs/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with ".NET"](/docs/tags/dotnet): .NET framework topics. - [One doc tagged with "Dremio"](/docs/tags/dremio): Dremio data lake engine integration. - [2 docs tagged with "DuckDB"](/docs/tags/duckdb): DuckDB embedded analytical database integration. - [2 docs tagged with "DuckLake"](/docs/tags/ducklake): DuckLake catalog integration. - [2 docs tagged with "DynamoDB"](/docs/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [2 docs tagged with "Elasticsearch"](/docs/tags/elasticsearch): Elasticsearch data connector and vector engine integration. - [7 docs tagged with "Embeddings"](/docs/tags/embeddings): Vector embeddings and semantic similarity operations. - [6 docs tagged with "Features"](/docs/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "File"](/docs/tags/file): File-based data connectors and local file access. - [One doc tagged with "Flux"](/docs/tags/flux): Flux CD GitOps toolkit for Kubernetes. - [One doc tagged with "Flux CD"](/docs/tags/fluxcd): Flux CD GitOps toolkit for Kubernetes. - [2 docs tagged with "Functions"](/docs/tags/functions): User-defined SQL scalar functions and remote function endpoints. - [One doc tagged with "Getting Started"](/docs/tags/getting-started): Installation guides and quickstart tutorials. - [3 docs tagged with "GitHub"](/docs/tags/github): GitHub repository data connector and integration. - [3 docs tagged with "GitOps"](/docs/tags/gitops): GitOps continuous delivery patterns for Kubernetes. - [2 docs tagged with "Glue"](/docs/tags/glue): AWS Glue ETL service and data catalog integration. - [One doc tagged with "Go"](/docs/tags/go): Go programming language topics. - [One doc tagged with "Golang"](/docs/tags/golang): Go (Golang) programming language topics. - [One doc tagged with "GraphQL"](/docs/tags/graphql): GraphQL API data connectors and query language support. - [3 docs tagged with "Helm"](/docs/tags/helm): Helm package manager for deploying Spice.ai on Kubernetes. - [One doc tagged with "HTTPS"](/docs/tags/https): HTTPS data connectors and secure web API access. - [2 docs tagged with "Hugging Face"](/docs/tags/huggingface): Hugging Face model hub and transformer model integration. - [2 docs tagged with "Iceberg"](/docs/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "Java"](/docs/tags/java): Java programming language topics. - [One doc tagged with "JavaScript"](/docs/tags/javascript): JavaScript programming language topics. - [One doc tagged with "JDBC"](/docs/tags/jdbc): Java Database Connectivity driver configuration. - [One doc tagged with "Kafka"](/docs/tags/kafka): Apache Kafka event streaming platform integration. - [6 docs tagged with "Kubernetes"](/docs/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "libSQL"](/docs/tags/libsql): libSQL database topics. - [One doc tagged with "Local"](/docs/tags/local): Local model execution and on-premise services. - [One doc tagged with "Logging"](/docs/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/tags/memory): In-memory data connectors and volatile storage. - [16 docs tagged with "Models"](/docs/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MongoDB"](/docs/tags/mongodb): MongoDB NoSQL document database integration. - [One doc tagged with "MSSQL"](/docs/tags/mssql): Microsoft SQL Server database integration. - [3 docs tagged with "MySQL"](/docs/tags/mysql): MySQL relational database integration. - [One doc tagged with "Node.js"](/docs/tags/nodejs): Node.js runtime environment topics. - [3 docs tagged with "NoSQL"](/docs/tags/nosql): NoSQL database systems and document stores. - [29 docs tagged with "Observability"](/docs/tags/observability): Monitoring, tracing, and metrics for system visibility. - [One doc tagged with "ODBC"](/docs/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/tags/open-source): Open source licenses and community contributions. - [3 docs tagged with "OpenAI"](/docs/tags/openai): OpenAI API integration and GPT model access. - [2 docs tagged with "Oracle"](/docs/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/tags/parameters): Configuration parameters and runtime options. - [One doc tagged with "Performance"](/docs/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/tags/persistence): Data persistence and durable storage mechanisms. - [4 docs tagged with "Postgres"](/docs/tags/postgres): PostgreSQL database integration and configuration. - [One doc tagged with "Python"](/docs/tags/python): Python programming language topics. - [3 docs tagged with "Query"](/docs/tags/query): SQL query execution, parameterized queries, and prepared statements. - [4 docs tagged with "Reference"](/docs/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/tags/relational): Relational database systems and SQL operations. - [3 docs tagged with "Runtime"](/docs/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Rust"](/docs/tags/rust): Rust programming language topics. - [2 docs tagged with "S3"](/docs/tags/s3): Amazon S3 object storage integration. - [One doc tagged with "S3 Express"](/docs/tags/s3-express): Amazon S3 Express One Zone storage. - [One doc tagged with "Sandbox"](/docs/tags/sandbox): Sandbox environments and isolated testing. - [One doc tagged with "ScyllaDB"](/docs/tags/scylladb): ScyllaDB database integration. - [6 docs tagged with "SDK"](/docs/tags/sdk): Software Development Kit topics and usage. - [10 docs tagged with "Search"](/docs/tags/search): Vector search, semantic search, and ranking capabilities. - [2 docs tagged with "Security"](/docs/tags/security): Security features and data protection mechanisms. - [One doc tagged with "Snowflake"](/docs/tags/snowflake): Snowflake cloud data warehouse integration. - [7 docs tagged with "SpiceAI"](/docs/tags/spiceai): SpiceAI cloud platform and managed services. - [5 docs tagged with "Spicepod"](/docs/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [6 docs tagged with "SQL"](/docs/tags/sql): SQL query language support and database operations. - [5 docs tagged with "Tools"](/docs/tags/tools): Development tools and utility integrations. - [2 docs tagged with "Tracing"](/docs/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Turso"](/docs/tags/turso): Turso data accelerator integration. - [2 docs tagged with "UDF"](/docs/tags/udf): User-defined functions registered with the SQL engine. - [2 docs tagged with "Unity Catalog"](/docs/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [One doc tagged with "Views"](/docs/tags/views): Virtual views and data transformation layers. - [One doc tagged with "Vortex"](/docs/tags/vortex): Vortex columnar file format and storage engine. - [10 docs tagged with "Write"](/docs/tags/write): Data connectors and catalogs that support write operations. - [One doc tagged with "YAML"](/docs/tags/yaml): YAML configuration syntax and file formats. - [One doc tagged with "Zipkin"](/docs/tags/zipkin): Zipkin distributed tracing system integration. ### acknowledgements Spice AI acknowledges the following open source projects for making this project possible: - [Open Source Acknowledgements](/docs/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: ### api Spice.ai API reference including HTTP REST API, Arrow Flight SQL, JDBC, ODBC, ADBC connectors, authentication, and TLS configuration. - [Spice.ai API Reference](/docs/api): Spice.ai API reference including HTTP REST API, Arrow Flight SQL, JDBC, ODBC, ADBC connectors, authentication, and TLS configuration. - [ADBC: Arrow Database Connectivity](/docs/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/api/auth): Authentication documentation - [Generate Package](/docs/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/api/HTTP/get-catalogs): Returns a list of all registered catalogs (data sources). Catalogs provide metadata about schemas and tables available from external data sources. - [Get Iceberg API config](/docs/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [ML Prediction](/docs/api/HTTP/get-model-predict): Make a ML prediction using a specific model. - [List Models](/docs/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [Check if a namespace exists.](/docs/api/HTTP/get-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Spicepods](/docs/api/HTTP/get-spicepods): Get a list of spicepods and their summary details. - [Check Runtime Status](/docs/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Get a table.](/docs/api/HTTP/get-table): This endpoint returns the table if it exists, otherwise it returns a 404 Not Found response. - [List Workers](/docs/api/HTTP/get-workers): Returns a list of all registered workers in the runtime. Workers are configurable processing units that can perform tasks like load balancing between models or implementing fallback strategies. - [Check Namespace exists](/docs/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [Check if a table exists.](/docs/api/HTTP/head-table): This endpoint returns a 200 OK response if the table exists, otherwise it returns a 404 Not Found response. - [list_tables](/docs/api/HTTP/list-tables): list_tables - [List Tools](/docs/api/HTTP/list-tools): Returns a list of all available tools in the Spice runtime. Tools provide reusable functionality that can be invoked programmatically or by AI agents. - [Send a Model Context Protocol message](/docs/api/HTTP/mcp-message): Send a JSON-RPC message to the Spice MCP server using the MCP Streamable HTTP transport. The response is either a single JSON-RPC response (`application/json`) or an SSE stream (`text/event-stream`), selected via the `Accept` header. Session continuity is carried via the `Mcp-Session-Id` header. - [Open an MCP server-to-client SSE stream](/docs/api/HTTP/mcp-stream): Open a long-lived server-to-client SSE stream for the current MCP session as defined by the Streamable HTTP transport. The `Mcp-Session-Id` header must identify an existing session created via `POST /v1/mcp`. - [Terminate an MCP Streamable HTTP session](/docs/api/HTTP/mcp-terminate-session): Terminate the MCP session identified by the `Mcp-Session-Id` header. Subsequent requests bearing the same session id will receive `404 Not Found`. - [Update Refresh SQL](/docs/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Batch ML Predictions](/docs/api/HTTP/post-batch-predict): Perform a batch of ML predictions, using multiple models, in one request. This is useful for ensembling or A/B testing different models. - [Create Chat Completion](/docs/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Text-to-SQL (NSQL)](/docs/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [post_responses](/docs/api/HTTP/post-responses): post_responses - [Search](/docs/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [Run Tool](/docs/api/HTTP/run-tool): Execute a specific tool by name. The request body schema and response format are defined by each individual tool's specification. Use `GET /v1/tools` to discover available tools and their parameter schemas. - [runtime](/docs/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/api/odbc): ODBC API Documentation - [API Overview](/docs/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/api/tls): Encryption in transit with TLS documentation ### cli Complete CLI reference for Spice.ai including commands to create, manage Spicepods, run queries, and interact with the Spice runtime. - [Spice.ai CLI Reference](/docs/cli): Complete CLI reference for Spice.ai including commands to create, manage Spicepods, run queries, and interact with the Spice runtime. - [Spice.ai OSS CLI command reference](/docs/cli/reference): Spice CLI command reference - [add](/docs/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/cli/reference/chat): spice chat CLI documentation - [completions](/docs/cli/reference/completions): Generate shell completions for the Spice CLI. - [connect](/docs/cli/reference/connect): Connect to an app on the Spice.ai Cloud Platform. - [dataset](/docs/cli/reference/dataset): Add or configure dataset entries in spicepod.yaml. - [datasets](/docs/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [feedback](/docs/cli/reference/feedback): Open the Spice.ai community Slack in the default browser to share feedback. - [init](/docs/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/cli/reference/models): Lists models loaded by the Spice runtime - [nsql](/docs/cli/reference/nsql): spice nsql CLI documentation - [pods](/docs/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [query](/docs/cli/reference/query): Submit an async query or start an interactive async query REPL against the Spice runtime's distributed query engine. - [refresh](/docs/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/cli/reference/search): Performs embeddings-based searches across search-configured datasets. - [spiced](/docs/cli/reference/spiced): Command-line reference for the spiced runtime binary — flags, defaults, and differences from spice run. - [sql](/docs/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/cli/reference/status): Spice runtime status - [trace](/docs/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/cli/reference/upgrade): Upgrades the Spice CLI and runtime to the latest or specified version - [validate](/docs/cli/reference/validate): Validate a spicepod.yaml without starting the runtime. - [version](/docs/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels ### clients Client and tools for connecting to Spice - [Clients and Tools](/docs/clients): Client and tools for connecting to Spice - [DBeaver](/docs/clients/dbeaver): Configure DBeaver to query Spice via JDBC - [JetBrains DataGrip](/docs/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Microsoft Power BI Connector](/docs/clients/powerbi): Use Microsoft Power BI to access, visualize and analyze Spice datasets. - [Apache Superset](/docs/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. ### comparison Compare Spice.ai with data platforms (Databricks, Snowflake), query engines (Trino, Dremio, ClickHouse), vector databases (Turbopuffer, LanceDB), search engines (Elasticsearch), and AI frameworks (LangChain, LlamaIndex, Ollama). - [How Spice Compares](/docs/comparison): Compare Spice.ai with data platforms (Databricks, Snowflake), query engines (Trino, Dremio, ClickHouse), vector databases (Turbopuffer, LanceDB), search engines (Elasticsearch), and AI frameworks (LangChain, LlamaIndex, Ollama). ### components Configure Spice.ai runtime components including data connectors, data accelerators, catalog connectors, model providers, embedding models, and secret stores. - [Spice.ai Runtime Components](/docs/components): Configure Spice.ai runtime components including data connectors, data accelerators, catalog connectors, model providers, embedding models, and secret stores. - [Catalog Connectors](/docs/components/catalogs): Connect to external catalog providers like Unity Catalog, Databricks, Iceberg, AWS Glue, Snowflake, ADBC, PostgreSQL, MySQL, MSSQL, Oracle, and more for federated SQL query in Spice. - [ADBC Catalog Connector](/docs/components/catalogs/adbc): Connect to databases via ADBC for automatic schema and table discovery. - [Databricks Catalog Connector](/docs/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [DuckLake Catalog Connector](/docs/components/catalogs/ducklake): Connect to a DuckLake catalog for federated SQL query. - [Glue Catalog Connector](/docs/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Microsoft SQL Server Catalog Connector](/docs/components/catalogs/mssql): Connect to a Microsoft SQL Server database as a catalog provider for federated SQL query. - [MySQL Catalog Connector](/docs/components/catalogs/mysql): Connect to a MySQL database as a catalog provider for federated SQL query. - [Oracle Catalog Connector](/docs/components/catalogs/oracle): Connect to an Oracle database as a catalog provider for federated SQL query. - [PostgreSQL Catalog Connector](/docs/components/catalogs/postgres): Connect to a PostgreSQL database as a catalog provider for federated SQL query. - [Snowflake Catalog Connector](/docs/components/catalogs/snowflake): Connect to a Snowflake database as a catalog provider for federated SQL query. - [Spice.ai Catalog Connector](/docs/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Unity Catalog Catalog Connector Deployment Guide](/docs/components/catalogs/unity-catalog/deployment): Operating guide for the Unity Catalog catalog connector in production: workspace authentication, table-type filtering, effective-permissions flow, and observability. - [Data Accelerators](/docs/components/data-accelerators): Data acceleration engines for local materialization and query acceleration in Spice - [In-Memory Arrow Data Accelerator](/docs/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [Arrow Data Accelerator Deployment Guide](/docs/components/data-accelerators/arrow/deployment): Operating guide for the Arrow (in-memory) data accelerator in production: memory sizing, indexes, and observability. - [Spice Cayenne Data Accelerator](/docs/components/data-accelerators/cayenne): Spice Cayenne Data Accelerator (Vortex) Documentation - [Cayenne Data Accelerator Deployment Guide](/docs/components/data-accelerators/cayenne/deployment): Operating guide for Spice Cayenne in production: footer and segment caches, S3 Express, metastore durability, and observability. - [DuckDB Data Accelerator](/docs/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [DuckDB Data Accelerator Deployment Guide](/docs/components/data-accelerators/duckdb/deployment): Operating guide for the DuckDB data accelerator in production: memory vs file mode, checkpointing, spill, pool sizing, and observability. - [PostgreSQL Data Accelerator](/docs/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [PostgreSQL Data Accelerator Deployment Guide](/docs/components/data-accelerators/postgres/deployment): Operating guide for the PostgreSQL data accelerator in production: authentication, connection pooling, and observability. - [SQLite Data Accelerator](/docs/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [SQLite Data Accelerator Deployment Guide](/docs/components/data-accelerators/sqlite/deployment): Operating guide for the SQLite data accelerator in production: file mode, busy timeout, pool, and observability. - [Turso Data Accelerator](/docs/components/data-accelerators/turso): Turso (libSQL) Data Accelerator Documentation - [Data Connectors](/docs/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ADBC Data Connector](/docs/components/data-connectors/adbc): ADBC Data Connector Documentation - [ClickHouse Data Connector](/docs/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Azure Cosmos DB Data Connector](/docs/components/data-connectors/cosmosdb): Query Azure Cosmos DB (NoSQL / Core SQL API) containers as SQL tables in Spice. Read-only scan with schema inferred from a sample of documents. - [Azure Cosmos DB Data Connector Deployment Guide](/docs/components/data-connectors/cosmosdb/deployment): Operating guide for the Azure Cosmos DB data connector in production: authentication, RU sizing, resilience, metrics, and observability. - [Databricks Data Connector](/docs/components/data-connectors/databricks): Databricks Data Connector Documentation - [Databricks Deployment Guide](/docs/components/data-connectors/databricks/deployment): Operating guide for the Databricks connector in production: resilience controls, Unity Catalog behavior, metrics, and observability. - [Debezium Data Connector](/docs/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Delta Lake Data Connector Deployment Guide](/docs/components/data-connectors/delta-lake/deployment): Operating guide for the Delta Lake data connector in production: object store auth, metadata caching, metrics, and observability. - [Dremio Data Connector](/docs/components/data-connectors/dremio): Dremio Data Connector Documentation - [Dremio Data Connector Deployment Guide](/docs/components/data-connectors/dremio/deployment): Operating guide for the Dremio data connector in production: authentication, Flight SQL transport, and observability. - [DuckDB Data Connector](/docs/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DuckDB Data Connector Deployment Guide](/docs/components/data-connectors/duckdb/deployment): Operating guide for the DuckDB data connector in production: file vs memory mode, durability, and observability. - [DuckLake Data Connector](/docs/components/data-connectors/ducklake): DuckLake Data Connector Documentation - [DynamoDB Data Connector](/docs/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [DynamoDB Data Connector Deployment Guide](/docs/components/data-connectors/dynamodb/deployment): Operating guide for the DynamoDB data connector in production: IAM, streams, checkpointing, lag behavior, and observability. - [Elasticsearch Data Connector](/docs/components/data-connectors/elasticsearch): Query Elasticsearch indexes as SQL tables in Spice, including kNN vector search, full-text search, and hybrid search. - [Elasticsearch Data Connector Deployment Guide](/docs/components/data-connectors/elasticsearch/deployment): Operating guide for the Elasticsearch data connector in production: authentication, TLS, resilience, and operational tuning. - [File Data Connector](/docs/components/data-connectors/file): File Data Connector Documentation - [File Data Connector Deployment Guide](/docs/components/data-connectors/file/deployment): Operating guide for the File data connector in production: permissions, formats, performance, and observability. - [Flight SQL Data Connector](/docs/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GCS Data Connector](/docs/components/data-connectors/gcs): GCS (Google Cloud Storage) Data Connector Documentation - [GitHub Data Connector](/docs/components/data-connectors/github): GitHub Data Connector Documentation - [GitHub Data Connector Deployment Guide](/docs/components/data-connectors/github/deployment): Operating guide for the GitHub data connector in production: PATs, rate limits, pagination, and observability. - [Glue Data Connector](/docs/components/data-connectors/glue): Connect to and query tables in an AWS Glue Data Catalog - [GraphQL Data Connector](/docs/components/data-connectors/graphql): GraphQL Data Connector Documentation - [GraphQL Data Connector Deployment Guide](/docs/components/data-connectors/graphql/deployment): Operating guide for the GraphQL data connector in production: authentication, pagination, rate limits, and observability. - [HTTP(s) Data Connector](/docs/components/data-connectors/https): HTTP(s) Data Connector Documentation - [HTTP(s) Data Connector Deployment Guide](/docs/components/data-connectors/https/deployment): Operating guide for the HTTP(s) data connector in production: authentication, rate control, retries, and observability. - [Iceberg Data Connector](/docs/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/components/data-connectors/imap): IMAP Data Connector Documentation - [Kafka Data Connector](/docs/components/data-connectors/kafka): Kafka Data Connector Documentation - [Localpod Data Connector](/docs/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/components/data-connectors/memory): Memory Data Connector Documentation - [MongoDB Data Connector](/docs/components/data-connectors/mongodb): MongoDB Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [MySQL Data Connector](/docs/components/data-connectors/mysql): MySQL Data Connector Documentation - [MySQL Data Connector Deployment Guide](/docs/components/data-connectors/mysql/deployment): Operating guide for the MySQL data connector in production: authentication, connection pooling, TLS, metrics, and observability. - [NFS Data Connector](/docs/components/data-connectors/nfs): NFS Data Connector Documentation - [ODBC Data Connector](/docs/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [PostgreSQL Data Connector Deployment Guide](/docs/components/data-connectors/postgres/deployment): Operating guide for the PostgreSQL data connector in production: authentication, connection pooling, TLS, metrics, and observability. - [Amazon Redshift Data Connector](/docs/components/data-connectors/redshift): Connect to Amazon Redshift using the PostgreSQL connector in Spice. - [S3 Data Connector](/docs/components/data-connectors/s3): S3 Data Connector Documentation - [S3 Data Connector Deployment Guide](/docs/components/data-connectors/s3/deployment): Operating guide for the S3 data connector in production: IAM, credential chains, file formats, metrics, and observability. - [ScyllaDB Data Connector](/docs/components/data-connectors/scylladb): ScyllaDB Data Connector Documentation - [SharePoint Data Connector](/docs/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [SMB Data Connector](/docs/components/data-connectors/smb): SMB Data Connector Documentation - [Snowflake Data Connector](/docs/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/components/data-connectors/spiceai): Federated SQL across Spice runtimes — Spice Cloud Platform datasets and self-hosted Spice instances (cluster-sidecar pattern). - [Spice.ai Data Connector Deployment Guide](/docs/components/data-connectors/spiceai/deployment): Operating guide for the Spice.ai connector in production: API keys, Flight endpoints, message sizing, sidecar topology, and observability. - [Embedding Models](/docs/components/embeddings): Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. - [Azure OpenAI Embedding Models](/docs/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [Google AI Embedding Models](/docs/components/embeddings/google): To use a hosted Google AI embedding model, specify the google path in the from field of your configuration. - [HuggingFace Text Embedding Models](/docs/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Hugging Face Embedding Deployment Guide](/docs/components/embeddings/huggingface/deployment): Operating guide for Hugging Face embeddings in production: tokens, download cache, pooling, device selection, and observability. - [Local Filesystem Embedding Models](/docs/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [Local Embedding Deployment Guide](/docs/components/embeddings/local/deployment): Operating guide for filesystem-loaded embedding models in production: formats, pooling, device selection, and observability. - [Model2Vec Embedding Models](/docs/components/embeddings/model2vec): Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning. - [OpenAI (or Compatible) Embedding Models](/docs/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [OpenAI Embedding Deployment Guide](/docs/components/embeddings/openai/deployment): Operating guide for the OpenAI embedding provider in production: API keys, usage tiers, batching, retries, metrics, and observability. - [Model Providers](/docs/components/models): Overview of supported model providers for ML and LLMs in Spice. - [Anthropic Models](/docs/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/components/models/azure): Instructions for using Azure OpenAI models - [Amazon Bedrock Models](/docs/components/models/bedrock): How to use Amazon Bedrock models with Spice. - [Databricks Model Provider](/docs/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [Filesystem Model Deployment Guide](/docs/components/models/filesystem/deployment): Operating guide for filesystem-loaded models in production: formats, device selection, memory footprint, and observability. - [Google AI Models](/docs/components/models/google): Instructions for using language models hosted on Google AI with Spice. - [HuggingFace](/docs/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [Hugging Face Model Deployment Guide](/docs/components/models/huggingface/deployment): Operating guide for the Hugging Face model in production: tokens, download cache, device selection, local inference footprint, and observability. - [OpenAI (or Compatible) Language Models](/docs/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [OpenAI Model Deployment Guide](/docs/components/models/openai/deployment): Operating guide for the OpenAI model in production: API keys, usage tiers, rate limiting, Responses API, metrics, and observability. - [Perplexity Models (Deprecated)](/docs/components/models/perplexity): Perplexity model support is no longer supported in Spice. - [Spice Cloud Platform](/docs/components/models/spiceai): Instructions for using models hosted on the Spice Cloud Platform with Spice. - [xAI Models](/docs/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/components/secret-stores): Configure secret stores to manage sensitive data like passwords, tokens, and API keys. - [AWS Secrets Manager Secret Store](/docs/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Azure Key Vault Secret Store](/docs/components/secret-stores/azure-keyvault): Azure Key Vault Secret Store Documentation - [Environment Secret Store](/docs/components/secret-stores/env): Environment Variables Secret Store Documentation - [HashiCorp Vault Secret Store](/docs/components/secret-stores/hashicorp-vault): HashiCorp Vault Secret Store Documentation - [Keyring Secret Store](/docs/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Streamable HTTP transport. - [Web Search Tool (Deprecated)](/docs/components/tools/websearch): The websearch tool is no longer supported in Spice. - [Vector Engines](/docs/components/vectors): Configure vector engines for efficient embedding storage and similarity search in Spice. - [DuckDB Vector Engine](/docs/components/vectors/duckdb): Use DuckDB as a vector engine in Spice for HNSW-based vector search via the DuckDB VSS extension. - [Elasticsearch Vector Engine](/docs/components/vectors/elasticsearch): Use Elasticsearch as a vector engine in Spice for kNN vector search, full-text search, and hybrid search. - [Amazon S3 Vectors Engine](/docs/components/vectors/s3_vectors): Amazon S3 Vectors Engine Documentation ### deployment Deploy Spice.ai in your environment using Docker, Kubernetes, AWS, Azure, or the Spice Cloud Platform. Learn about sidecar, microservice, tiered, and cluster deployment architectures. - [Spice.ai Deployment Guide](/docs/deployment): Deploy Spice.ai in your environment using Docker, Kubernetes, AWS, Azure, or the Spice Cloud Platform. Learn about sidecar, microservice, tiered, and cluster deployment architectures. - [Deployment Architectures](/docs/deployment/architectures): Explore Spice deployment architectures including sidecar, microservice, tiered, sharded, and cluster configurations. - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/deployment/architectures/cluster): Deploying Spice as a cluster - [Cluster-Sidecar Deployment](/docs/deployment/architectures/cluster-sidecar): Deploy Spice with application-local sidecars for localhost query, search, and inference, backed by a centralized cluster for ingest, acceleration, and distributed query. - [Cloud Hosted](/docs/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Microservice Deployment (Single or Multiple Replicas)](/docs/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [AWS Integrations](/docs/deployment/aws/integrations): Complete guide to Spice.ai integrations with Amazon Web Services, including data connectors, AI models, vector stores, and secret management. - [Azure Deployment Options](/docs/deployment/azure): Guide to deploying Spice.ai applications on Microsoft Azure - [Azure Integrations](/docs/deployment/azure/integrations): Spice.ai integrations with Microsoft Azure, including data connectors, AI models, embeddings, and authentication. - [CI/CD Deployment](/docs/deployment/ci-cd): Deploy Spice.ai applications using continuous integration and delivery pipelines, including Helm, Kubernetes GitOps with Argo CD or Flux, GitHub Actions, and the Spice Cloud deploy action. - [Spice Cloud Platform Deployment](/docs/deployment/cloud): Guide to deploying data and AI applications using the managed Spice Cloud Platform - [Docker](/docs/deployment/docker): Run Spice.ai as a Docker container. - [Docker Sandbox Guide - v1.3.0](/docs/deployment/docker/sandbox): Migrating to v1.3.0 - [Google Cloud Deployment Options](/docs/deployment/gcp): Guide to deploying Spice.ai applications on Google Cloud Platform (GCP). - [GCP Integrations](/docs/deployment/gcp/integrations): Spice.ai integrations with Google Cloud Platform, including data connectors, AI models, embeddings, and authentication. - [Kubernetes Deployment](/docs/deployment/kubernetes): Deploy Spice.ai on Kubernetes using Helm, Argo CD, or Flux. - [Kubernetes - Argo CD](/docs/deployment/kubernetes/argocd): Deploy Spice.ai on self-hosted Kubernetes using Argo CD and the Spice Helm chart. - [Kubernetes - Flux](/docs/deployment/kubernetes/flux): Deploy Spice.ai on Kubernetes using Flux CD and the Spice Helm chart. - [Kubernetes - Helm](/docs/deployment/kubernetes/helm): Deploy Spice.ai in Kubernetes using Helm. - [Read/Write Separation](/docs/deployment/read-write-separation): Separate write/ingest workloads (cluster) from read workloads (application sidecars, agents) using shared snapshots and live query delegation. ### faq Answers to frequently asked questions about Spice.ai including features, use cases, differences from Trino/Presto/Dremio, federated queries, caching, and AI capabilities. - [Spice.ai FAQ](/docs/faq): Answers to frequently asked questions about Spice.ai including features, use cases, differences from Trino/Presto/Dremio, federated queries, caching, and AI capabilities. ### features Explore Spice.ai features including data federation, data acceleration, caching, search, LLM integration, embeddings, observability, and more for building data-driven AI applications. - [Spice.ai Features](/docs/features): Explore Spice.ai features including data federation, data acceleration, caching, search, LLM integration, embeddings, observability, and more for building data-driven AI applications. - [Caching](/docs/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Debezium (CDC over Kafka)](/docs/features/cdc/debezium): Consume Debezium change events from Kafka into a Spice-accelerated dataset for sources without a native Spice CDC path. - [DynamoDB Streams (Native CDC)](/docs/features/cdc/dynamodb-streams): Stream INSERT, UPDATE, and DELETE events from Amazon DynamoDB directly into a Spice-accelerated dataset using DynamoDB Streams. - [MongoDB Change Streams (Native CDC)](/docs/features/cdc/mongodb-streams): Stream insert, update, replace, and delete events from MongoDB directly into a Spice-accelerated dataset using MongoDB Change Streams. - [PostgreSQL Logical Replication (Native CDC)](/docs/features/cdc/postgres-replication): Stream INSERT, UPDATE, and DELETE events from PostgreSQL directly into a Spice-accelerated dataset using native logical replication. - [Data Acceleration](/docs/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Hash Index for Arrow Acceleration](/docs/features/data-acceleration/hash-index): Learn how to use hash indexes for O(1) point lookups on Arrow-accelerated datasets. - [Indexes](/docs/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/features/data-acceleration/partitioning): Partition accelerated datasets to make filtered queries faster by reading only the relevant partitions. - [Refresh Modes](/docs/features/data-acceleration/refresh-modes): Refresh modes for accelerated datasets in Spice. - [Append Refresh Mode](/docs/features/data-acceleration/refresh-modes/append): Incrementally append new rows to an accelerated dataset. - [Caching Refresh Mode](/docs/features/data-acceleration/refresh-modes/caching): Learn how to use caching refresh mode for HTTP-based datasets - [Changes Refresh Mode](/docs/features/data-acceleration/refresh-modes/changes): Apply incremental inserts, updates, and deletes via Change Data Capture. - [Full Refresh Mode](/docs/features/data-acceleration/refresh-modes/full): Replace the entire accelerated dataset on each refresh. - [Snapshot Refresh Mode](/docs/features/data-acceleration/refresh-modes/snapshot): Reload acceleration data exclusively from the snapshot store. - [Snapshots](/docs/features/data-acceleration/snapshots): Bootstrap file-mode accelerations from managed snapshots to eliminate cold starts. - [Data Ingestion](/docs/features/data-ingestion): Learn how to ingest data in Spice. - [Distributed Query](/docs/features/distributed-query): Learn how to run Spice in distributed mode for larger scale queries, including the async queries API. - [Embedding Datasets](/docs/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Functions](/docs/features/functions): Define custom scalar and table SQL functions inline (SQL tier) or by calling remote HTTP services (Remote tier), automatically exposed as SQL functions and LLM tools. - [Large Language Models](/docs/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models (Deprecated)](/docs/features/large-language-models/evals): Language model evals are no longer supported in Spice. - [Model Context Protocol (MCP)](/docs/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/features/machine-learning-models): Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform. - [Observability & Monitoring](/docs/features/observability): Monitor Spice with Prometheus metrics, OpenTelemetry, and distributed tracing. - [Component Metrics](/docs/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Parameterized Queries](/docs/features/query-federation/parameterized-queries): Learn how to use prepared statements and parameterized queries in Spice for improved security and performance. - [URL Tables](/docs/features/query-federation/url-tables): Query object store files directly using URLs without pre-registering datasets - [Search Functionality](/docs/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/features/search/full-text): Learn how Spice can perform full text search - [Multi-Vector Search](/docs/features/search/multi-vector): Embed list-of-strings columns as a column of vectors and use ColBERT-style late-interaction search in Spice. - [Reranking](/docs/features/search/rerank): Rerank search results using dedicated reranker models or LLM-as-reranker for improved relevance. - [Vector-Based Search](/docs/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/features/semantic-model): Attach descriptions and metadata to datasets, views, and columns in Spice so LLMs, SQL functions, and humans share the same understanding of your data. - [Tool Registry](/docs/features/tool-registry): Reduce per-turn token cost and improve LLM tool selection accuracy by replacing individual tool definitions with searchable tool_search and tool_invoke meta-tools backed by hybrid full-text, keyword, schema, and vector search. - [Views](/docs/features/views): Documentation for defining Views in Spice - [Web Search](/docs/features/web-search): Learn how Spice can perform web search - [Workers](/docs/features/workers): Configure workers in the Spice runtime to coordinate interactions between LLMs and tools, with load-balancing, round-robin, and fallback strategies. ### getting-started Get started with Spice.ai in 5 minutes. Install the CLI, connect to datasets, run SQL queries, and use AI models with OpenAI-compatible APIs. - [Getting Started with Spice.ai OSS](/docs/getting-started): Get started with Spice.ai in 5 minutes. Install the CLI, connect to datasets, run SQL queries, and use AI models with OpenAI-compatible APIs. - [Community Data](/docs/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. ### installation Install Spice.ai OSS on macOS, Linux, Windows, or WSL using the install script, Homebrew, PowerShell, or direct download from GitHub releases. - [Install Spice.ai OSS](/docs/installation): Install Spice.ai OSS on macOS, Linux, Windows, or WSL using the install script, Homebrew, PowerShell, or direct download from GitHub releases. ### intelligent-applications Learn how to build intelligent, data-driven AI applications and agents with Spice.ai. Explore patterns for RAG, LLM integration, and real-time AI inference. - [Building Intelligent AI Applications with Spice.ai](/docs/intelligent-applications): Learn how to build intelligent, data-driven AI applications and agents with Spice.ai. Explore patterns for RAG, LLM integration, and real-time AI inference. ### monitoring Monitor Spice.ai deployments with Datadog, Grafana, New Relic, Prometheus, and Zipkin integrations. - [Monitoring](/docs/monitoring): Monitor Spice.ai deployments with Datadog, Grafana, New Relic, Prometheus, and Zipkin integrations. - [Datadog](/docs/monitoring/datadog): Monitoring Spice with Datadog - [Grafana & Prometheus](/docs/monitoring/grafana): Monitoring Spice instances with Grafana & Prometheus - [New Relic](/docs/monitoring/new-relic): Monitoring Spice with New Relic - [Spice Cloud Platform](/docs/monitoring/spice-cloud): Connect a self-hosted Spice runtime to the Spice Cloud Platform to centralize task history and runtime observability across deployments. - [Zipkin Integration](/docs/monitoring/zipkin): Learn how to integrate Spice with Zipkin tracing. ### reference Reference documentation for Spice.ai including API reference, CLI commands, Spicepod configuration syntax, SQL reference, and data type specifications. - [Spice.ai OSS Reference Docs](/docs/reference): Reference documentation for Spice.ai including API reference, CLI commands, Spicepod configuration syntax, SQL reference, and data type specifications. - [Cron Schedules](/docs/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/reference/datatypes): Spice uses Apache Arrow data types internally, providing consistent type handling across different data sources and accelerators. This section documents how Arrow types map to specific accelerators and object store formats. - [Accelerator Data Types](/docs/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Spice Runtime Distributions](/docs/reference/distributions): Distribution variants of the Spice runtime for different use cases and deployment scenarios, including data-only, GPU-accelerated, NAS, and allocator variants. - [Duration](/docs/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/reference/file_format): File-based data connectors — including s3//, file//, sftp://, and others — support multiple structured and document file formats. This page details the format-specific parameters available for each. - [Managing Memory Usage](/docs/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice deployments. - [Models Grade Report](/docs/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [Performance Tuning](/docs/reference/performance-tuning): Comprehensive guide to optimizing query performance, acceleration, and resource utilization in Spice deployments. - [YAML syntax for Spicepod manifests](/docs/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/reference/spicepod/embeddings): Embeddings YAML reference - [Evals (Deprecated)](/docs/reference/spicepod/evals): The evals Spicepod component is no longer supported in Spice. - [Functions (User-Defined Functions)](/docs/reference/spicepod/functions): User-defined functions YAML reference - [Reserved Keywords](/docs/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/reference/spicepod/models): Models YAML reference - [Runtime](/docs/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/reference/spicepod/tools): Tools YAML reference - [Views](/docs/reference/spicepod/views): Views YAML reference - [Workers](/docs/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/reference/sql): Complete SQL reference for Spice.ai including SELECT syntax, subqueries, DML statements, aggregate functions, AI functions, JSON operators, and search capabilities. - [Aggregate Functions](/docs/reference/sql/aggregate_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [AI Functions](/docs/reference/sql/ai): AI functions in Spice provide direct integration with large language models (LLMs) and embedding models within SQL queries. These functions process text through configured model providers and return generated responses or vector embeddings. - [DML (Data Manipulation Language)](/docs/reference/sql/dml): Data Manipulation Language (DML) statements for inserting and modifying data in Spice. - [Explain](/docs/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [JSON Functions and Operators](/docs/reference/sql/json): Reference for JSON functions and operators in Spice SQL - [Operators](/docs/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/reference/sql/prepared_statements): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Scalar Functions](/docs/reference/sql/scalar_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [Search in SQL](/docs/reference/sql/search): Reference for search functions and filtering in Spice SQL. - [SELECT](/docs/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. ### sdks Connect to Spice using official SDKs - [SDKs](/docs/sdks): Connect to Spice using official SDKs - [Dotnet SDK](/docs/sdks/dotnet): Connect to Spice using the Dotnet SDK - [Go SDK](/docs/sdks/golang): Connect to Spice using the Go SDK - [Java SDK](/docs/sdks/java): Connect to Spice using the Java SDK - [JavaScript SDK](/docs/sdks/javascript): Connect to Spice using the JavaScript SDK - [Python SDK](/docs/sdks/python): Connect to Spice using the Python SDK - [Rust SDK](/docs/sdks/rust): Connect to Spice using the Rust SDK ### troubleshooting Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Troubleshooting Spice](/docs/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. ### use-cases Discover how to use Spice.ai for data federation, reverse-ETL, database CDN, enterprise search, RAG, and building AI-powered applications and agents. - [Spice.ai Use Cases](/docs/use-cases): Discover how to use Spice.ai for data federation, reverse-ETL, database CDN, enterprise search, RAG, and building AI-powered applications and agents. - [AI Applications and Agents](/docs/use-cases/ai): AI Applications and Agents - [Agentic AI Applications and Agents](/docs/use-cases/ai/agentic-apps): Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. - [Edge-Enabled AI Applications and Agents](/docs/use-cases/ai/edge-ai): Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases. - [Federated MCP Client for Distributed Tool Ecosystems](/docs/use-cases/ai/federated-mcp-server): Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, improving threat analysis. - [Multi-Tenant AI Agents](/docs/use-cases/ai/multi-tenant-agents): Deploy AI agents across many SaaS tenants with strict isolation and no per-tenant ETL pipelines. - [Object-Store Based SQL Query, Search, and LLM Inference Engine](/docs/use-cases/ai/object-store-ai-engine): Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. - [Real-Time Decision-Making for Intelligent Applications](/docs/use-cases/ai/real-time-decision-making): Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. - [Tool-Augmented AI with Model Context Protocol Server](/docs/use-cases/ai/tool-calling-ai): Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. - [Caching](/docs/use-cases/caching): Caching use cases with Spice.ai, including write-through, read-through, SQL, S3, and HTTP caching. - [HTTP Cache](/docs/use-cases/caching/http-cache): Use Spice.ai to cache HTTP API responses locally, reducing API call frequency and providing fast, SQL-queryable access to external API data. - [Read-Through Cache](/docs/use-cases/caching/read-through-cache): Use Spice.ai as a read-through cache with the SQL results cache for federated data sources and HTTP APIs. - [S3 Cache](/docs/use-cases/caching/s3-cache): Use Spice.ai to cache S3 and object store data locally for fast, repeatable SQL queries without re-reading from remote storage. - [SQL and Database Cache](/docs/use-cases/caching/sql-database-cache): Use Spice.ai to cache SQL database tables and query results locally for low-latency access and reduced load on upstream databases. - [Write-Through Cache](/docs/use-cases/caching/write-through-cache): Use Spice.ai as a write-through cache that writes data to both a local accelerator and the upstream data source. - [Data Federation, Acceleration, and SQL Query](/docs/use-cases/data): Data Federation, Acceleration, and SQL Query - [Application Resilience and Performance Optimization](/docs/use-cases/data/application-resilence-and-acceleration): Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance. - [Data Mesh for Unified Data Access](/docs/use-cases/data/data-mesh): Give domain teams decentralized, real-time data ownership and access across disparate sources through a unified SQL interface. - [Database CDN for Enhanced Performance](/docs/use-cases/data/database-cdn): Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience. - [ETL-free Workflows and Data Migrations](/docs/use-cases/data/etl-free-workflows): Federate legacy and modern data systems without ETL for faster migrations, lower overhead, and zero application downtime. - [Object-Store Data Engine](/docs/use-cases/data/object-store-data-engine): Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses. - [Reverse-ETL for Operational Workflows](/docs/use-cases/data/reverse-etl): Serve enriched data from warehouses and data lakes to operational systems and applications, eliminating complex pipelines. - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/use-cases/rag): Use Spice for Retrieval-Augmented-Generation (RAG) - [RAG for Contextual Applications](/docs/use-cases/rag/applications): Build context-rich AI applications using Spice for Retrieval-Augmented Generation (RAG). - [Retrieval-Augmented Generation for AI-Powered Reporting](/docs/use-cases/rag/reporting): Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. - [Search & Retrieval](/docs/use-cases/search): Search & Retrieval - [Simplifying Real-Time Data Collection and Search](/docs/use-cases/search/data-collection-and-search): Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic. - [Enterprise Search and Retrieval](/docs/use-cases/search/enterprise-search): Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities. - [Object-Store Native Search Engine](/docs/use-cases/search/object-store-search-engine): Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search. ## Optional - [GitHub Repository](https://github.com/spiceai/spiceai): Spice.ai OSS source code and issue tracker - [Cookbook](https://github.com/spiceai/cookbook): Ready-to-use recipes and examples for Spice.ai - [Spice Cloud Platform](https://spice.ai): Managed cloud platform for Spice.ai --- # Full Documentation Content # 🧑‍🍳 Spice.ai OSS Cookbook 79 guides and samples to help you build data-grounded AI apps and agents with Spice.ai Open-Source. Find ready-to-use examples for data acceleration, AI agents, LLM memory, and more. [ Contribute to the Cookbook on GitHub!](https://github.com/spiceai/cookbook/blob/trunk/README.md) ## Featured Recipes ### Federated SQL Query Join S3, PostgreSQL, and Dremio data in one SQL query. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) ### Run Llama3 Locally Use Llama models from HuggingFace with Spice. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md)[ Video](https://youtu.be/I2i6uZKBbd4) ### Data Acceleration with DuckDB Speed up queries using DuckDB. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md)[ Video](https://youtu.be/hFvVz5NGpaw) ### LLM Memory Persistent memory for language models [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/llm-memory/README.md)[ Video](https://youtu.be/uc8TCAPu1IM) ## Sample Applications and Guides Example apps and guides for real-world Spice.ai usage and best practices. ### Command Query Responsibility Segregation (CQRS) Sample application implementing the CQRS pattern with Spice. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/cqrs/README.md) ## Core Features Discover core capabilities like data federation, acceleration, search, and LLM inference to enhance your applications. ### Federated SQL Query Query data from S3, PostgreSQL, and Dremio in a single query. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) ### OpenAI SDK Use the OpenAI SDK to connect to models hosted on Spice. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/openai_sdk/README.md) ### OpenAI Responses API Use the OpenAI Responses API with Spice [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/openai-responses-api/README.md) ### DuckDB Data Accelerator Accelerate data locally using DuckDB. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md)[ Video](https://youtu.be/hFvVz5NGpaw) ## Models, AI, and Agents Integrate with popular AI models, LLMs, and build intelligent agents using Spice.ai. ### Azure OpenAI Models Connect and use Azure OpenAI models with Spice. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/azure_openai/README.md) ### Running Llama3 Locally Use the Llama family of models locally from HuggingFace using Spice. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md)[ Video](https://youtu.be/I2i6uZKBbd4) ### OpenAI SDK Use the OpenAI SDK to connect to models hosted on Spice. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/openai_sdk/README.md) ### OpenAI Responses API Use the OpenAI Responses API with Spice [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/openai-responses-api/README.md) ### LLM Memory Persistent memory for language models. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/llm-memory/README.md)[ Video](https://youtu.be/uc8TCAPu1IM) ### Text to SQL (NSQL) Ask natural language (NLP) questions of your datasets using the built-in text-to-SQL tool. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/text-to-sql/README.md) ### AI SQL Function Invoke LLMs directly within SQL queries using the AI SQL function. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/ai/README.md) ### Nvidia NIM Deploy Nvidia NIM infrastructure, on Kubernetes, with GPUs connected to Spice. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/nvidia-nim/ec2/README.md) ### Searching GitHub Files Search GitHub files with embeddings and vector similarity search. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md)[ Video](https://youtu.be/5y26MveEJ8c) ### Hybrid-Search with RRF Combine multiple search methods using Reciprocal Rank Fusion (RRF) for improved search results. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/search/README.md) ### xAI Models Use xAI models such as Grok. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/models/xai/README.md)[ Video](https://youtu.be/-7RkAsqQLdk) ### DeepSeek Model Use DeepSeek model through Spice. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/deepseek/README.md) ### Model-Context-Protocol (MCP) Use Spice to connect to or host MCP servers. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/mcp/README.md) ### Amazon S3 Vectors Use Amazon S3 Vectors to store embeddings and perform efficient vector search. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/vectors/s3/README.md)[ Video](https://www.youtube.com/watch?v=QPbqPf5W36g) ## Data Acceleration, Materialization, and Federation Optimize query performance with local acceleration, data materialization, and federation techniques. ### DuckDB Data Accelerator Accelerate data locally using DuckDB. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md)[ Video](https://youtu.be/hFvVz5NGpaw) ### PostgreSQL Data Accelerator Accelerate data locally using PostgreSQL. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/postgres/accelerator/README.md) ### SQLite Data Accelerator Accelerate data locally using SQLite. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/sqlite/accelerator/README.md) ### Apache Arrow Data Accelerator Accelerate data using Apache Arrow. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/arrow/README.md) ### Hashed Partitioning with DuckDB Use hashed partitioning for performance with DuckDB. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/hashed_partitioning/README.md) ### Accelerated Views Use view materialization for improved performance. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/views/README.md) ### Indexes on Accelerated Data Create and manage indexes on accelerated data. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/acceleration/indexes/README.md) ## Search & Embeddings Implement advanced search capabilities and leverage embeddings for vector similarity search. ### Searching GitHub Files Search GitHub files with embeddings and vector similarity search. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md)[ Video](https://youtu.be/5y26MveEJ8c) ### Hybrid-Search with RRF Combine multiple search methods using Reciprocal Rank Fusion (RRF) for improved search results. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/search/README.md) ### Amazon S3 Vectors Use Amazon S3 Vectors to store embeddings and perform efficient vector search. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/vectors/s3/README.md)[ Video](https://www.youtube.com/watch?v=QPbqPf5W36g) ## Data Connectors Connect to various data sources and systems to query, analyze, and manage your data efficiently. ### PostgreSQL Connector Connect to and query PostgreSQL databases. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/postgres/connector/README.md) ### AWS RDS PostgreSQL Connect to AWS RDS PostgreSQL instances. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/postgres/rds/README.md) ### Supabase PostgreSQL Connect to Supabase PostgreSQL databases. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/postgres/supabase/README.md) ### MySQL Connector Connect to and query MySQL databases. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/mysql/connector/README.md) ### AWS RDS Aurora MySQL Connect to AWS RDS Aurora with MySQL compatibility. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/mysql/rds-aurora/README.md) ### PlanetScale MySQL Connect to PlanetScale MySQL databases. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/mysql/planetscale/README.md) ### Clickhouse Connector Connect to and query Clickhouse databases. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/clickhouse/README.md) ### Databricks Connector Connect to and query Databricks instances using Delta Lake or Spark Connect. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/databricks/README.md) ### Delta Lake Connector Query data from Delta Lake tables. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/delta-lake/README.md) ### Debezium CDC from Postgres Stream changes from PostgreSQL using Debezium CDC. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/cdc-debezium/README.md) ### Debezium CDC with SASL/SCRAM Stream MySQL changes using Debezium with SASL/SCRAM authentication. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/cdc-debezium/sasl-scram/README.md) ### Dremio Connector Connect to and query Dremio. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/dremio/README.md) ### DuckDB Connector Query DuckDB databases with sample TPCH data. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/connector/README.md) ### File Connector Query data from local files. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/file/README.md) ### FTP Connector Query data from FTP servers. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/ftp/README.md) ### GitHub Connector Connect to and query GitHub data. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/github/README.md)[ Video](https://youtu.be/mxwt0HEF1VQ) ### GraphQL Connector Query data from GraphQL endpoints. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/graphql/README.md) ### MSSQL Connector Connect to Microsoft SQL Server databases. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/mssql/README.md) ### ODBC Connector Connect to databases using ODBC. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/odbc/README.md) ### Oracle Connector Connect to and query Oracle databases. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/oracle/README.md) ### Glue Connector Connect to AWS Glue. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/glue/README.md) ### S3 Connector Query data from S3 compatible storage. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/s3/README.md) ### SharePoint Connector Connect to SharePoint and OneDrive for Business. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/sharepoint/README.md) ### Snowflake Connector Connect to and query Snowflake databases. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/snowflake/README.md) ### Spice.ai Cloud Connector Connect to the Spice.ai Cloud Platform. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/spiceai/README.md) ### Apache Spark Connector Connect to and query Apache Spark. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/spark/README.md) ### IMAP Emails Federated SQL query of mail across IMAP email servers. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/imap/README.md) ### MongoDB Connector Connect to and query MongoDB databases. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/mongodb/connector/README.md) ### Live Orders Analytics with Apache Kafka Data Connector Combine real-time data streaming from Kafka with other datasets using Spice. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/kafka/README.md) ## Catalog Connectors Connect to data catalogs to discover, manage, and utilize your data assets effectively. ### Spice.ai Cloud Platform Catalog Connect to the Spice.ai Cloud Platform catalog. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/catalogs/spiceai/README.md) ### Databricks Unity Catalog Connect to Databricks Unity catalog. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/catalogs/databricks/README.md) ### Unity Catalog Connect to Unity catalog. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/catalogs/unity_catalog/README.md) ### Iceberg Catalog Connector Connect to Iceberg catalog with support for reading and writing Iceberg tables. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/catalogs/iceberg/README.md) ## Visualization Visualize data with BI and analytics tools. ### Sales BI with Apache Superset Visualize data in Spice with Apache Superset. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/sales-bi/README.md) ### Grafana Datasource Add Spice as a Grafana datasource. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/grafana-datasource/README.md) ## API Clients Use API clients for data access and integration. ### Python ADBC Client Query Spice using ADBC and Parameterized Queries with Python. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/clients/adbc/README.md) ### Java JDBC Client Query Spice.ai using the Java JDBC client. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/clients/java/README.md) ### Scala JDBC Client Query Spice.ai using the Scala JDBC client. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/clients/scala/README.md) ## Deployment Deploy Spice.ai in different environments. ### Deploying to Kubernetes Deploy Spice.ai on Kubernetes. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/kubernetes/README.md) ### Running in Docker Run Spice.ai in Docker containers. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/docker/README.md) ## Performance and Benchmarking Measure and optimize performance with benchmarks and best practices for your Spice.ai deployment. ### TPC-H Benchmarking Benchmark performance using TPC-H. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/tpc-h/README.md) ### Results Caching Cache query results for improved performance. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/caching/accelerator/README.md) ### Indexes on Accelerated Data Create and manage indexes on accelerated data. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/acceleration/indexes/README.md) ## Configuration Fine-tune your Spice.ai deployment with advanced configuration options for optimal performance. ### Data Retention Policy Configure data retention policies. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/retention/README.md) ### Refresh Data Window Configure data refresh windows. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/refresh-data-window/README.md) ### Advanced Data Refresh Advanced configuration for data refresh. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/acceleration/data-refresh/README.md) ### Data Quality with Constraints Add data quality constraints. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/acceleration/constraints/README.md) ### Cron Dataset Schedules Schedule dataset refreshes using cron syntax. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/acceleration/cron/README.md) ## SDKs Use SDKs for different programming languages. ### OpenAI SDK Use the OpenAI SDK to connect to models hosted on Spice. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/openai_sdk/README.md) ### Rust SDK Query Spice.ai using the Rust SDK. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/client-sdk/spice-rs-sdk-sample/README.md) ### Python SDK Query Spice.ai using the Python SDK. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/client-sdk/spicepy-sdk-sample/README.md) ### Go SDK Query Spice.ai using the Go SDK. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/client-sdk/gospice-sdk-sample/README.md) ### Spice.js JavaScript (Node.js) SDK Query Spice.ai using the JavaScript (Node.js) SDK with examples. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/client-sdk/spice.js-sdk-sample/README.md) ### Java SDK Query Spice.ai using the Java SDK. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/client-sdk/spice-java-sdk-sample/README.md) ## Security Secure your Spice.ai deployment and data access with robust security practices and configurations. ### Intelligent Security Copilot Analyze real-time data access patterns with Spice.ai. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/guides/security-analyzer/README.md) ### TLS Encryption Enable encryption in transit using TLS. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/tls/README.md) ### API Key Authentication Secure access with API key authentication. [ Recipe](https://github.com/spiceai/cookbook/blob/trunk/api_key/README.md) --- ## .[​](#. "Direct link to .") * [.NET1](/docs/next/tags/dotnet ".NET framework topics.") *** --- ## [Open Source Acknowledgements](/docs/next/acknowledgements) Spice AI acknowledges the following open source projects for making this project possible: --- ## [ADBC Catalog Connector](/docs/next/components/catalogs/adbc) Connect to databases via ADBC for automatic schema and table discovery. --- ## [ADBC: Arrow Database Connectivity](/docs/next/api/adbc) ADBC API Documentation --- ## [Kubernetes - Argo CD](/docs/next/deployment/kubernetes/argocd) Deploy Spice.ai on self-hosted Kubernetes using Argo CD and the Spice Helm chart. --- ## [ADBC: Arrow Database Connectivity](/docs/next/api/adbc) ADBC API Documentation --- ## [Arrow Flight SQL API](/docs/next/api/arrow-flight-sql) Query Spice using JDBC/ODBC/ADBC --- ## [Authentication](/docs/next/api/auth) Authentication documentation --- ## [login](/docs/next/cli/reference/login) Login to the Spice.ai Platform, or other services with sub-commands. --- ## [Azure BlobFS Data Connector](/docs/next/components/data-connectors/abfs) Azure BlobFS Data Connector Documentation --- ## [Azure BlobFS Data Connector](/docs/next/components/data-connectors/abfs) Azure BlobFS Data Connector Documentation --- ## [Caching](/docs/next/features/caching) Learn how to use Spice in-memory caching --- ## [ADBC Catalog Connector](/docs/next/components/catalogs/adbc) Connect to databases via ADBC for automatic schema and table discovery. --- ## [Cayenne Data Accelerator Deployment Guide](/docs/next/components/data-accelerators/cayenne/deployment) Operating guide for Spice Cayenne in production: footer and segment caches, S3 Express, metastore durability, and observability. --- ## [Configuring Trace Levels](/docs/next/cli/tracing) Configuring Spice.ai OSS trace output verbosity levels --- ## [Component Metrics](/docs/next/features/observability/component_metrics) Learn how to enable optional component metrics. --- ## [Embedding Models](/docs/next/components/embeddings) Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. --- ## [Language Model Overrides](/docs/next/features/large-language-models/parameter_overrides) Learn how to override default LLM hyperparameters in Spice. --- ## [Azure Cosmos DB Data Connector](/docs/next/components/data-connectors/cosmosdb) Query Azure Cosmos DB (NoSQL / Core SQL API) containers as SQL tables in Spice. Read-only scan with schema inferred from a sample of documents. --- ## [Dotnet SDK](/docs/next/sdks/dotnet) Connect to Spice using the Dotnet SDK --- ## [Arrow Data Accelerator Deployment Guide](/docs/next/components/data-accelerators/arrow/deployment) Operating guide for the Arrow (in-memory) data accelerator in production: memory sizing, indexes, and observability. --- ## [ADBC Catalog Connector](/docs/next/components/catalogs/adbc) Connect to databases via ADBC for automatic schema and table discovery. --- ## [Delta Lake Data Connector](/docs/next/components/data-connectors/delta-lake) Delta Lake Data Connector Documentation --- ## [Databricks Catalog Connector](/docs/next/components/catalogs/databricks) Connect to a Databricks Unity Catalog provider. --- ## [Datasets](/docs/next/reference/spicepod/datasets) Datasets YAML reference --- ## [Debezium Data Connector](/docs/next/components/data-connectors/debezium) Debezium Data Connector Documentation --- ## [Troubleshooting Spice](/docs/next/troubleshooting) Review and debug runtime tasks, logs, and diagnostic steps in Spice. --- ## [Databricks Data Connector](/docs/next/components/data-connectors/databricks) Databricks Data Connector Documentation --- ## [Open Source Acknowledgements](/docs/next/acknowledgements) Spice AI acknowledges the following open source projects for making this project possible: --- ## [CI/CD Deployment](/docs/next/deployment/ci-cd) Deploy Spice.ai applications using continuous integration and delivery pipelines, including Helm, Kubernetes GitOps with Argo CD or Flux, GitHub Actions, and the Spice Cloud deploy action. --- ## [Docker](/docs/next/deployment/docker) Run Spice.ai as a Docker container. --- ## [Dotnet SDK](/docs/next/sdks/dotnet) Connect to Spice using the Dotnet SDK --- ## [Dremio Data Connector Deployment Guide](/docs/next/components/data-connectors/dremio/deployment) Operating guide for the Dremio data connector in production: authentication, Flight SQL transport, and observability. --- ## [DuckDB Data Accelerator Deployment Guide](/docs/next/components/data-accelerators/duckdb/deployment) Operating guide for the DuckDB data accelerator in production: memory vs file mode, checkpointing, spill, pool sizing, and observability. --- ## [DuckLake Catalog Connector](/docs/next/components/catalogs/ducklake) Connect to a DuckLake catalog for federated SQL query. --- ## [DynamoDB Data Connector](/docs/next/components/data-connectors/dynamodb) DynamoDB Data Connector Documentation --- ## [Elasticsearch Data Connector](/docs/next/components/data-connectors/elasticsearch) Query Elasticsearch indexes as SQL tables in Spice, including kNN vector search, full-text search, and hybrid search. --- ## [Azure OpenAI Embedding Models](/docs/next/components/embeddings/azure) To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: --- ## [Data Ingestion](/docs/next/features/data-ingestion) Learn how to ingest data in Spice. --- ## [Data Connectors](/docs/next/components/data-connectors) Learn how to use Data Connector to query external data. --- ## [File Data Connector Deployment Guide](/docs/next/components/data-connectors/file/deployment) Operating guide for the File data connector in production: permissions, formats, performance, and observability. --- ## [Kubernetes - Flux](/docs/next/deployment/kubernetes/flux) Deploy Spice.ai on Kubernetes using Flux CD and the Spice Helm chart. --- ## [Kubernetes - Flux](/docs/next/deployment/kubernetes/flux) Deploy Spice.ai on Kubernetes using Flux CD and the Spice Helm chart. --- ## [Functions](/docs/next/features/functions) Define custom scalar and table SQL functions inline (SQL tier) or by calling remote HTTP services (Remote tier), automatically exposed as SQL functions and LLM tools. --- ## [Spicepods](/docs/next/getting-started/spicepods) An introduction to Spicepods --- ## [CI/CD Deployment](/docs/next/deployment/ci-cd) Deploy Spice.ai applications using continuous integration and delivery pipelines, including Helm, Kubernetes GitOps with Argo CD or Flux, GitHub Actions, and the Spice Cloud deploy action. --- ## [CI/CD Deployment](/docs/next/deployment/ci-cd) Deploy Spice.ai applications using continuous integration and delivery pipelines, including Helm, Kubernetes GitOps with Argo CD or Flux, GitHub Actions, and the Spice Cloud deploy action. --- ## [Glue Catalog Connector](/docs/next/components/catalogs/glue) Connect to an AWS Glue Data Catalog. --- ## [Go SDK](/docs/next/sdks/golang) Connect to Spice using the Go SDK --- ## [Go SDK](/docs/next/sdks/golang) Connect to Spice using the Go SDK --- ## [GraphQL Data Connector Deployment Guide](/docs/next/components/data-connectors/graphql/deployment) Operating guide for the GraphQL data connector in production: authentication, pagination, rate limits, and observability. --- ## [CI/CD Deployment](/docs/next/deployment/ci-cd) Deploy Spice.ai applications using continuous integration and delivery pipelines, including Helm, Kubernetes GitOps with Argo CD or Flux, GitHub Actions, and the Spice Cloud deploy action. --- ## [HTTP(s) Data Connector Deployment Guide](/docs/next/components/data-connectors/https/deployment) Operating guide for the HTTP(s) data connector in production: authentication, rate control, retries, and observability. --- ## [Hugging Face Embedding Deployment Guide](/docs/next/components/embeddings/huggingface/deployment) Operating guide for Hugging Face embeddings in production: tokens, download cache, pooling, device selection, and observability. --- ## [Iceberg Catalog Connector](/docs/next/components/catalogs/iceberg) Connect to an Iceberg catalog provider. --- ## [Memory Data Connector](/docs/next/components/data-connectors/memory) Memory Data Connector Documentation --- ## [GitHub Data Connector](/docs/next/components/data-connectors/github) GitHub Data Connector Documentation --- ## [Java SDK](/docs/next/sdks/java) Connect to Spice using the Java SDK --- ## [JavaScript SDK](/docs/next/sdks/javascript) Connect to Spice using the JavaScript SDK --- ## [Arrow Flight SQL API](/docs/next/api/arrow-flight-sql) Query Spice using JDBC/ODBC/ADBC --- ## [Kafka Data Connector](/docs/next/components/data-connectors/kafka) Kafka Data Connector Documentation --- ## [CI/CD Deployment](/docs/next/deployment/ci-cd) Deploy Spice.ai applications using continuous integration and delivery pipelines, including Helm, Kubernetes GitOps with Argo CD or Flux, GitHub Actions, and the Spice Cloud deploy action. --- ## [Turso Data Accelerator](/docs/next/components/data-accelerators/turso) Turso (libSQL) Data Accelerator Documentation --- ## [Local Embedding Deployment Guide](/docs/next/components/embeddings/local/deployment) Operating guide for filesystem-loaded embedding models in production: formats, pooling, device selection, and observability. --- ## [Configuring Trace Levels](/docs/next/cli/tracing) Configuring Spice.ai OSS trace output verbosity levels --- ## [login](/docs/next/cli/reference/login) Login to the Spice.ai Platform, or other services with sub-commands. --- ## [YAML syntax for Spicepod manifests](/docs/next/reference/spicepod) Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) --- ## [Model Context Protocol (MCP)](/docs/next/features/large-language-models/mcp) Learn how to use the Model Context Protocol (MCP) with Spice. --- ## [Language Model Memory](/docs/next/features/large-language-models/memory) Learn how to provide LLMs with memory --- ## [Embedding Models](/docs/next/components/embeddings) Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. --- ## [MongoDB Data Connector](/docs/next/components/data-connectors/mongodb) MongoDB Data Connector Documentation --- ## [Microsoft SQL Server Catalog Connector](/docs/next/components/catalogs/mssql) Connect to a Microsoft SQL Server database as a catalog provider for federated SQL query. --- ## [MySQL Catalog Connector](/docs/next/components/catalogs/mysql) Connect to a MySQL database as a catalog provider for federated SQL query. --- ## [JavaScript SDK](/docs/next/sdks/javascript) Connect to Spice using the JavaScript SDK --- ## [Azure Cosmos DB Data Connector](/docs/next/components/data-connectors/cosmosdb) Query Azure Cosmos DB (NoSQL / Core SQL API) containers as SQL tables in Spice. Read-only scan with schema inferred from a sample of documents. --- ## [Arrow Data Accelerator Deployment Guide](/docs/next/components/data-accelerators/arrow/deployment) Operating guide for the Arrow (in-memory) data accelerator in production: memory sizing, indexes, and observability. --- ## [Arrow Flight SQL API](/docs/next/api/arrow-flight-sql) Query Spice using JDBC/ODBC/ADBC --- ## [Open Source Acknowledgements](/docs/next/acknowledgements) Spice AI acknowledges the following open source projects for making this project possible: --- ## [Azure OpenAI Embedding Models](/docs/next/components/embeddings/azure) To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: --- ## [Oracle Catalog Connector](/docs/next/components/catalogs/oracle) Connect to an Oracle database as a catalog provider for federated SQL query. --- ## [Language Model Overrides](/docs/next/features/large-language-models/parameter_overrides) Learn how to override default LLM hyperparameters in Spice. --- ## [Catalog Connectors](/docs/next/components/catalogs) Connect to external catalog providers like Unity Catalog, Databricks, Iceberg, AWS Glue, Snowflake, ADBC, PostgreSQL, MySQL, MSSQL, Oracle, and more for federated SQL query in Spice. --- ## [Language Model Overrides](/docs/next/features/large-language-models/parameter_overrides) Learn how to override default LLM hyperparameters in Spice. --- ## [Spice Cayenne Data Accelerator](/docs/next/components/data-accelerators/cayenne) Spice Cayenne Data Accelerator (Vortex) Documentation --- ## [Language Model Memory](/docs/next/features/large-language-models/memory) Learn how to provide LLMs with memory --- ## [PostgreSQL Catalog Connector](/docs/next/components/catalogs/postgres) Connect to a PostgreSQL database as a catalog provider for federated SQL query. --- ## [Python SDK](/docs/next/sdks/python) Connect to Spice using the Python SDK --- ## [Parameterized Queries](/docs/next/features/query-federation/parameterized-queries) Learn how to use prepared statements and parameterized queries in Spice for improved security and performance. --- ## [Datasets](/docs/next/reference/spicepod/datasets) Datasets YAML reference --- ## [MySQL Data Connector](/docs/next/components/data-connectors/mysql) MySQL Data Connector Documentation --- ## [Language Models Tools](/docs/next/features/large-language-models/tools) Learn how LLMs interact with the Spice runtime. --- ## [Rust SDK](/docs/next/sdks/rust) Connect to Spice using the Rust SDK --- ## [S3 Data Connector Deployment Guide](/docs/next/components/data-connectors/s3/deployment) Operating guide for the S3 data connector in production: IAM, credential chains, file formats, metrics, and observability. --- ## [Spice Cayenne Data Accelerator](/docs/next/components/data-accelerators/cayenne) Spice Cayenne Data Accelerator (Vortex) Documentation --- ## [Docker Sandbox Guide - v1.3.0](/docs/next/deployment/docker/sandbox) Migrating to v1.3.0 --- ## [ScyllaDB Data Connector](/docs/next/components/data-connectors/scylladb) ScyllaDB Data Connector Documentation --- ## [Dotnet SDK](/docs/next/sdks/dotnet) Connect to Spice using the Dotnet SDK --- ## [Elasticsearch Data Connector](/docs/next/components/data-connectors/elasticsearch) Query Elasticsearch indexes as SQL tables in Spice, including kNN vector search, full-text search, and hybrid search. --- ## [Authentication](/docs/next/api/auth) Authentication documentation --- ## [Snowflake Catalog Connector](/docs/next/components/catalogs/snowflake) Connect to a Snowflake database as a catalog provider for federated SQL query. --- ## [Docker](/docs/next/deployment/docker) Run Spice.ai as a Docker container. --- ## [Datasets](/docs/next/reference/spicepod/datasets) Datasets YAML reference --- ## [Arrow Flight SQL API](/docs/next/api/arrow-flight-sql) Query Spice using JDBC/ODBC/ADBC --- ## [Functions](/docs/next/features/functions) Define custom scalar and table SQL functions inline (SQL tier) or by calling remote HTTP services (Remote tier), automatically exposed as SQL functions and LLM tools. --- ## [Configuring Trace Levels](/docs/next/cli/tracing) Configuring Spice.ai OSS trace output verbosity levels --- ## [Troubleshooting Spice](/docs/next/troubleshooting) Review and debug runtime tasks, logs, and diagnostic steps in Spice. --- ## [Turso Data Accelerator](/docs/next/components/data-accelerators/turso) Turso (libSQL) Data Accelerator Documentation --- ## [Functions](/docs/next/features/functions) Define custom scalar and table SQL functions inline (SQL tier) or by calling remote HTTP services (Remote tier), automatically exposed as SQL functions and LLM tools. --- ## [Unity Catalog Catalog Connector](/docs/next/components/catalogs/unity-catalog) Connect to a Unity Catalog provider. --- ## [Views](/docs/next/reference/spicepod/views) Views YAML reference --- ## [Spice Cayenne Data Accelerator](/docs/next/components/data-accelerators/cayenne) Spice Cayenne Data Accelerator (Vortex) Documentation --- ## [Data Ingestion](/docs/next/features/data-ingestion) Learn how to ingest data in Spice. --- ## [YAML syntax for Spicepod manifests](/docs/next/reference/spicepod) Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) --- ## [Zipkin Integration](/docs/next/monitoring/zipkin) Learn how to integrate Spice with Zipkin tracing. --- # Spice.ai Open Source **Spice** is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. ![Spice.ai Open Source Data Query \& AI-Inference Compute Engine](/assets/images/spice.ai-compute-engine-new-c9b4c58d2c615acbe2537023202c09c9.png) Spice provides four industry-standard APIs in a lightweight, portable runtime (single \~140 MB binary): 1. **SQL Query & Search APIs**: Supports HTTP, Arrow Flight, Arrow Flight SQL, ODBC, JDBC, ADBC, and `vector_search` and `text_search` UDTFs. 2. **OpenAI-Compatible APIs**: Provides HTTP APIs for OpenAI SDK compatibility, local model serving (CUDA/Metal accelerated), and hosted model gateway. 3. **Iceberg Catalog REST APIs**: Offers a unified API for Iceberg Catalog. 4. **MCP HTTP+SSE APIs**: Enables integration with external tools via Model Context Protocol (MCP) using HTTP and Server-Sent Events (SSE). Goal 🎯 Developers can focus on building data apps and AI agents confidently, knowing they are grounded in data. Spice is primarily used for: * **Data Federation**: SQL query across any database, data warehouse, or data lake. [Learn More](https://spiceai.org/docs/features/query-federation). * **Data Materialization and Acceleration**: Materialize, accelerate, and cache database queries. [Read the MaterializedView interview - Building a CDN for Databases](https://materializedview.io/p/building-a-cdn-for-databases-spice-ai). * **Enterprise Search**: Keyword, vector, and full-text search with Tantivy-powered BM25 and vector similarity search for structured and unstructured data. [Learn More](https://github.com/spiceai/cookbook/blob/trunk/vectors/s3/README.md). * **AI Apps and Agents**: An AI-database powering retrieval-augmented generation (RAG) and intelligent agents. [Learn More](https://spiceai.org/docs/use-cases/rag). Watch 🎥 * [CMU Databases - Accelerating Data and AI with Spice.ai Open-Source](https://www.youtube.com/watch?v=tyM-ec1lKfU) * [Query Data using Spice, OpenAI, and MCP](https://www.youtube.com/watch?v=TFAu4qxjTPk\&list=PLesJrUXEx3U-dQul0PqLV3TGTdUmr3B6e\&index=8) * [Search with Amazon S3 Vectors](https://www.youtube.com/watch?v=QPbqPf5W36g) Spice is built on industry-leading technologies including [Apache DataFusion](https://datafusion.apache.org), Apache Arrow, Arrow Flight, SQLite, and DuckDB. If you want to build with DataFusion or DuckDB, Spice provides a simple, flexible, and production-ready engine. ![How Spice works](https://github.com/spiceai/spiceai/assets/80174/7d93ae32-d6d8-437b-88d3-d64fe089e4b7) Announcement 📣 Read the [Spice.ai 1.0-stable announcement](https://spice.ai/blog/announcing-1.0-stable). ## Why Spice?[​](#why-spice "Direct link to Why Spice?") Spice simplifies building data-driven AI applications and agents by combining SQL query, search, and LLM inference in a single runtime. Instead of stitching together separate databases, caching layers, and AI services, developers deploy Spice alongside their applications to: * **Query any data source with SQL**: Join data across PostgreSQL, Snowflake, S3, and other sources without building ETL pipelines. * **Accelerate queries locally**: Materialize datasets in-memory or on disk for sub-second query performance, with automatic refresh from source. * **Ground AI in real data**: Connect LLMs to datasets through built-in tools so AI responses are based on actual data, not hallucinations. * **Search across data**: Run vector, keyword, and hybrid search through SQL functions like `vector_search()` and `text_search()`. ![Spice.ai](https://github.com/spiceai/spiceai/assets/80174/29e4421d-8942-4f2a-8397-e9d4fdeda36b) ### How is Spice different?[​](#how-is-spice-different "Direct link to How is Spice different?") 1. **AI-Native Runtime**: Combines data query and AI inference in a single engine for data-grounded, accurate AI. 2. **Application-Focused**: Designed for distributed deployment at the application or agent level, often as a 1:1 or 1 :N mapping, unlike centralized databases serving multiple apps. Multiple Spice instances can be deployed, even one per tenant or customer. 3. **Dual-Engine Acceleration**: Supports **OLAP** (Arrow/DuckDB) and **OLTP** (SQLite/PostgreSQL) engines at the dataset level for flexible performance across analytical and transactional workloads. 4. **Disaggregated Storage**: Separates compute from storage, co-locating materialized working datasets with applications, dashboards, or ML pipelines while accessing source data in its original storage. 5. **Edge to Cloud Native**: Deploy as a standalone instance, Kubernetes sidecar, microservice, or cluster across edge, on-prem, and public clouds. Chain multiple Spice instances for tier-optimized, distributed deployments. For detailed comparisons with other data query engines and AI frameworks, see [How Spice Compares](/docs/next/comparison). ## Example Use-Cases[​](#example-use-cases "Direct link to Example Use-Cases") ### Data-Grounded Agentic AI Applications[​](#data-grounded-agentic-ai-applications "Direct link to Data-Grounded Agentic AI Applications") * **OpenAI-Compatible API**: Connect to hosted models (OpenAI, Anthropic, xAI) or deploy locally (Llama, NVIDIA NIM). [AI Gateway Recipe](https://github.com/spiceai/cookbook/blob/trunk/openai_sdk/README.md) * **Federated Data Access**: Query using SQL and NSQL (text-to-SQL) across databases, data warehouses, and data lakes with advanced query push-down. [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) * **Search and RAG**: Perform keyword, vector, and full-text search with Tantivy-powered BM25 and vector similarity search (VSS) integrated into SQL queries using `vector_search` and `text_search`. Supports multi-column vector search with reciprocal rank fusion. [Amazon S3 Vectors Recipe](https://github.com/spiceai/cookbook/blob/trunk/vectors/s3/README.md) * **LLM Memory and Observability**: Store and retrieve history and context for AI agents with visibility into data flows, model performance, and traces. [LLM Memory Recipe](https://github.com/spiceai/cookbook/blob/trunk/llm-memory/README.md) | [Observability Documentation](https://spiceai.org/docs/features/observability) ### Database CDN and Query Mesh[​](#database-cdn-and-query-mesh "Direct link to Database CDN and Query Mesh") * **Data Acceleration**: Co-locate materialized datasets in Arrow, SQLite, or DuckDB for sub-second queries. [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md) * **Resiliency and Local Dataset Replication**: Maintain availability with local replicas of critical datasets. [Local Dataset Replication Recipe](https://github.com/spiceai/cookbook/blob/trunk/localpod/README.md) * **Responsive Dashboards**: Enable fast, real-time analytics for frontends and BI tools. [Sales BI Dashboard Demo](https://github.com/spiceai/cookbook/blob/trunk/sales-bi/README.md) * **Simplified Legacy Migration**: Unify legacy systems with modern infrastructure via federated SQL querying. [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) ### Retrieval-Augmented Generation (RAG)[​](#retrieval-augmented-generation-rag "Direct link to Retrieval-Augmented Generation (RAG)") * **Unified Search with Vector Similarity**: Perform efficient vector similarity search across structured and unstructured data, with native support for Amazon S3 Vectors for petabyte-scale storage and querying. Supports distance metrics like cosine similarity, Euclidean distance, or dot product. [Amazon S3 Vectors Recipe](https://github.com/spiceai/cookbook/blob/trunk/vectors/s3/README.md) * **Semantic Knowledge Layer**: Define a semantic context model to enrich data for AI. [Semantic Model Documentation](/docs/next/features/semantic-model) * **Text-to-SQL**: Convert natural language queries into SQL using built-in NSQL and sampling tools. [Text-to-SQL Recipe](https://github.com/spiceai/cookbook/blob/trunk/text-to-sql/README.md) ## FAQ[​](#faq "Direct link to FAQ") * **Is Spice a cache?** No, but its data acceleration acts as an active cache, materialization, or prefetcher. Unlike traditional caches that fetch on miss, Spice prefetches and materializes filtered data on intervals, triggers, or via CDC. It also supports [results caching](https://spiceai.org/docs/features/caching). * **Is Spice a CDN for databases?** Yes, Spice enables shipping working datasets to where they're accessed most, like data-intensive applications or AI contexts, similar to a CDN. [Docs FAQ](/docs/next/faq) ### Watch a 30-second BI dashboard acceleration demo[​](#watch-a-30-second-bi-dashboard-acceleration-demo "Direct link to Watch a 30-second BI dashboard acceleration demo") See more demos on [YouTube](https://www.youtube.com/playlist?list=PLesJrUXEx3U9anekJvbjyyTm7r9A26ugK). ### Intelligent Applications and Agents[​](#intelligent-applications-and-agents "Direct link to Intelligent Applications and Agents") Spice enables developers to build data-grounded AI applications and agents by co-locating data and ML models with applications. Read more about the vision for [intelligent AI-driven applications](/docs/next/use-cases). ### Connect with Us[​](#connect-with-us "Direct link to Connect with Us") * Build with Spice and share feedback at , [Slack](https://spice.ai/slack), [X](https://twitter.com/spice_ai), or [LinkedIn](https://www.linkedin.com/company/74148478). * [File an issue](https://github.com/spiceai/spiceai/issues/new) for bugs or issues. * Join our team ([We're hiring!](https://spice.ai/careers)). * Contribute code or documentation ([CONTRIBUTING.md](https://github.com/spiceai/spiceai/blob/trunk/CONTRIBUTING)). * ⭐️ Star the [Spice.ai repo](https://github.com/spiceai/spiceai) to show support! --- # Open Source Acknowledgements Spice AI acknowledges the following open source projects for making this project possible: ## Rust Crates[​](#rust-crates "Direct link to Rust Crates") * adbc\_core 0.23.0, Apache-2.0
* adbc\_driver\_manager 0.23.0, Apache-2.0
* aegis 0.9.8, MIT
* aes 0.8.4, Apache-2.0 OR MIT
* ahash 0.7.8, Apache-2.0 OR MIT
* ahash 0.8.12, Apache-2.0 OR MIT
* anyhow 1.0.102, Apache-2.0 OR MIT
* arc-swap 1.8.0, Apache-2.0 OR MIT
* arrow 57.2.0, Apache-2.0
* arrow-array 57.2.0, Apache-2.0
* arrow-buffer 57.2.0, Apache-2.0
* arrow-cast 57.2.0, Apache-2.0
* arrow-csv 57.2.0, Apache-2.0
* arrow-flight 57.2.0, Apache-2.0
* arrow-ipc 57.2.0, Apache-2.0
* arrow-json 57.2.0, Apache-2.0
* arrow-odbc 21.0.0, MIT
* arrow-row 57.2.0, Apache-2.0
* arrow-schema 57.2.0, Apache-2.0
* assert\_cmd 2.1.2, Apache-2.0 OR MIT
* async-channel 1.9.0, Apache-2.0 OR MIT
* async-channel 2.5.0, Apache-2.0 OR MIT
* async-compression 0.4.41, Apache-2.0 OR MIT
* async-graphql 7.2.1, Apache-2.0 OR MIT
* async-graphql-axum 7.2.1, Apache-2.0 OR MIT
* async-openai 0.32.0, MIT
* async-stream 0.3.6, MIT
* async-trait 0.1.89, Apache-2.0 OR MIT
* aws-config 1.8.16, Apache-2.0
* aws-credential-types 1.2.14, Apache-2.0
* aws-runtime 1.7.3, Apache-2.0
* aws-sdk-bedrockruntime 1.130.0, Apache-2.0
* aws-sdk-cognitoidentity 1.99.0, Apache-2.0
* aws-sdk-cognitoidentityprovider 1.116.0, Apache-2.0
* aws-sdk-dynamodb 1.111.0, Apache-2.0
* aws-sdk-dynamodbstreams 1.99.0, Apache-2.0
* aws-sdk-glue 1.145.1, Apache-2.0
* aws-sdk-s3 1.132.0, Apache-2.0
* aws-sdk-s3vectors 1.24.0, Apache-2.0
* aws-sdk-secretsmanager 1.104.0, Apache-2.0
* aws-sdk-sts 1.103.0, Apache-2.0
* aws-smithy-async 1.2.14, Apache-2.0
* aws-smithy-runtime 1.11.1, Apache-2.0
* aws-smithy-runtime-api 1.12.0, Apache-2.0
* aws-smithy-types 1.4.7, Apache-2.0
* axum 0.8.8, MIT
* axum-extra 0.12.5, MIT
* azure\_core 0.21.0, MIT
* azure\_core 0.31.0, MIT
* azure\_data\_cosmos 0.30.0, MIT
* azure\_identity 0.31.0, MIT
* azure\_security\_keyvault\_secrets 0.10.0, MIT
* azure\_storage 0.21.0, MIT
* azure\_storage\_blobs 0.21.0, MIT
* backoff 0.4.0, Apache-2.0 OR MIT
* ballista-core 52.0.0, Apache-2.0
* ballista-executor 52.0.0, Apache-2.0
* ballista-scheduler 52.0.0, Apache-2.0
* base64 0.13.1, Apache-2.0 OR MIT
* base64 0.21.7, Apache-2.0 OR MIT
* base64 0.22.1, Apache-2.0 OR MIT
* bb8 0.9.1, MIT
* bb8-oracle 0.3.0, Apache-2.0 OR MIT
* bcder 0.7.6, BSD-3-Clause
* bigdecimal 0.4.10, Apache-2.0 OR MIT
* bindgen 0.69.5, BSD-3-Clause
* bindgen 0.71.1, BSD-3-Clause
* bindgen 0.72.1, BSD-3-Clause
* blake3 1.8.3, Apache-2.0 OR Apache-2.0 WITH LLVM-exception OR CC0-1.0
* bollard 0.18.1, Apache-2.0
* brotli 8.0.2, BSD-3-Clause AND MIT
* byte-unit 5.2.0, MIT
* bytes 1.11.1, MIT
* calamine 0.28.0, MIT
* charset 0.1.5, Apache-2.0 OR MIT
* chbench-driver 2.0.0, ../../LICENSE
* chrono 0.4.44, Apache-2.0 OR MIT
* chrono-tz 0.8.6, Apache-2.0 OR MIT
* chrono-tz 0.9.0, Apache-2.0 OR MIT
* chrono-tz 0.10.4, Apache-2.0 OR MIT
* clap 4.5.60, Apache-2.0 OR MIT
* clap\_complete 4.6.0, Apache-2.0 OR MIT
* clickhouse-rs 1.1.0-alpha.1, MIT
* cmac 0.7.2, Apache-2.0 OR MIT
* comfy-table 7.1.2, MIT
* connector-clickhouse 2.0.0, ../../../LICENSE
* connector-databricks 2.0.0, ../../../LICENSE
* connector-delta-lake 2.0.0, ../../../LICENSE
* connector-dremio 2.0.0, ../../../LICENSE
* connector-duckdb 2.0.0, ../../../LICENSE
* connector-elasticsearch 2.0.0, ../../../LICENSE
* connector-flightsql 2.0.0, ../../../LICENSE
* connector-ftp 2.0.0, ../../../LICENSE
* connector-graphql 2.0.0, ../../../LICENSE
* connector-imap 2.0.0, ../../../LICENSE
* connector-mongodb 2.0.0, ../../../LICENSE
* connector-mssql 2.0.0, ../../../LICENSE
* connector-mysql 2.0.0, ../../../LICENSE
* connector-nfs 1.11.0-unstable, Apache-2.0
* connector-odbc 2.0.0, ../../../LICENSE
* connector-oracle 2.0.0, ../../../LICENSE
* connector-postgres 2.0.0, ../../../LICENSE
* connector-scylladb 2.0.0, ../../../LICENSE
* connector-sftp 2.0.0, ../../../LICENSE
* connector-sharepoint 2.0.0, ../../../LICENSE
* connector-smb 2.0.0, ../../../LICENSE
* connector-snowflake 2.0.0, ../../../LICENSE
* connector-spark 2.0.0, ../../../LICENSE
* criterion 0.5.1, Apache-2.0 OR MIT
* criterion 0.6.0, Apache-2.0 OR MIT
* criterion 0.7.0, Apache-2.0 OR MIT
* croner 3.0.1, MIT
* csv 1.4.0, MIT OR Unlicense
* ctor 0.8.0, Apache-2.0 OR MIT
* ctrlc 3.5.1, Apache-2.0 OR MIT
* cudarc 0.17.8, Apache-2.0 OR MIT
* cudarc 0.19.4, Apache-2.0 OR MIT
* dashmap 6.1.0, MIT
* datafusion 52.5.0, Apache-2.0
* datafusion-catalog 52.5.0, Apache-2.0
* datafusion-common 52.5.0, Apache-2.0
* datafusion-common-runtime 52.5.0, Apache-2.0
* datafusion-datasource 52.5.0, Apache-2.0
* datafusion-execution 52.5.0, Apache-2.0
* datafusion-expr 52.5.0, Apache-2.0
* datafusion-federation 0.4.2, Apache-2.0
* datafusion-flightsql 0.1.0, Apache-2.0
* datafusion-functions 52.5.0, Apache-2.0
* datafusion-functions-json 0.52.0, Apache-2.0
* datafusion-physical-expr 52.5.0, Apache-2.0
* datafusion-physical-expr-adapter 52.5.0, Apache-2.0
* datafusion-physical-expr-common 52.5.0, Apache-2.0
* datafusion-physical-plan 52.5.0, Apache-2.0
* datafusion-proto 52.5.0, Apache-2.0
* datafusion-proto-common 52.5.0, Apache-2.0
* datafusion-pruning 52.5.0, Apache-2.0
* datafusion-spark 52.5.0, Apache-2.0
* datafusion-substrait 52.5.0, Apache-2.0
* datafusion-table-providers 0.1.0, Apache-2.0
* delta\_kernel 0.18.2, Apache-2.0
* dialoguer 0.12.0, MIT
* dirs 6.0.0, Apache-2.0 OR MIT
* docx-rs 0.4.17, MIT
* dotenvy 0.15.7, MIT
* duckdb 1.10503.1, MIT
* duration-parse 2.0.0, Apache-2.0
* dyn-clone 1.0.20, Apache-2.0 OR MIT
* either 1.15.0, Apache-2.0 OR MIT
* elasticsearch 2.0.0, ../../LICENSE
* flatbuffers 25.12.19, Apache-2.0
* flate2 1.1.9, Apache-2.0 OR MIT
* fundu 2.0.1, MIT
* futures 0.3.32, Apache-2.0 OR MIT
* geodatafusion 0.3.0, Apache-2.0 OR MIT
* gethostname 1.1.0, Apache-2.0
* git2 0.20.4, Apache-2.0 OR MIT
* globset 0.4.18, MIT OR Unlicense
* governor 0.10.4, MIT
* graph-rs-sdk 2.0.1, MIT
* graphql-parser 0.4.1, Apache-2.0 OR MIT
* headers 0.4.1, MIT
* headers-accept 0.3.0, MIT
* hex 0.4.3, Apache-2.0 OR MIT
* hf-hub 0.4.3, Apache-2.0
* hickory-resolver 0.25.2, Apache-2.0 OR MIT
* hmac 0.12.1, Apache-2.0 OR MIT
* hmac 0.13.0, Apache-2.0 OR MIT
* hostname 0.3.1, MIT
* hostname 0.4.2, MIT
* http 0.2.12, Apache-2.0 OR MIT
* http 1.4.0, Apache-2.0 OR MIT
* http-body 0.4.6, MIT
* http-body 1.0.1, MIT
* http-body-util 0.1.3, MIT
* httpdate 1.0.3, Apache-2.0 OR MIT
* humantime 2.3.0, Apache-2.0 OR MIT
* hyper 0.14.32, MIT
* hyper 1.8.1, MIT
* hyper-util 0.1.20, MIT
* iceberg 0.9.1, Apache-2.0
* iceberg-catalog-glue 0.9.1, Apache-2.0
* iceberg-catalog-rest 0.9.1, Apache-2.0
* iceberg-datafusion 0.9.1, Apache-2.0
* iceberg-storage-opendal 0.9.1, Apache-2.0
* iceberg\_test\_utils 0.9.1, Apache-2.0
* im 15.1.0, MPL-2.0+
* imap 3.0.0-alpha.14, Apache-2.0 OR MIT
* indexmap 1.9.3, Apache-2.0 OR MIT
* indexmap 2.14.0, Apache-2.0 OR MIT
* indicatif 0.17.11, MIT
* indicatif 0.18.4, MIT
* insta 1.46.3, Apache-2.0
* itertools 0.10.5, Apache-2.0 OR MIT
* itertools 0.11.0, Apache-2.0 OR MIT
* itertools 0.12.1, Apache-2.0 OR MIT
* itertools 0.13.0, Apache-2.0 OR MIT
* itertools 0.14.0, Apache-2.0 OR MIT
* jsonpath-rust 0.7.5, MIT
* jsonwebtoken 9.3.1, MIT
* jsonwebtoken 10.3.0, MIT
* keyring 3.6.3, Apache-2.0 OR MIT
* libc 0.2.182, Apache-2.0 OR MIT
* libnfs 0.1.0, Apache-2.0
* linkme 0.3.35, Apache-2.0 OR MIT
* logos 0.16.1, Apache-2.0 OR MIT
* mailparse 0.16.1, 0BSD
* md-5 0.10.6, Apache-2.0 OR MIT
* md-5 0.11.0, Apache-2.0 OR MIT
* md4 0.10.2, Apache-2.0 OR MIT
* mediatype 0.21.0, MIT
* mimalloc 0.1.48, MIT
* mistralrs 0.8.1, MIT
* mistralrs-core 0.8.1, MIT
* model2vec-rs 0.1.3, LICENSE
* moka 0.12.13, (Apache-2.0 OR MIT) AND Apache-2.0
* mongodb 3.5.2, Apache-2.0
* mysql\_async 0.36.2, Apache-2.0 OR MIT
* native-tls 0.2.14, Apache-2.0 OR MIT
* ndarray 0.15.6, Apache-2.0 OR MIT
* ndarray 0.16.1, Apache-2.0 OR MIT
* nix 0.29.0, MIT
* nix 0.30.1, MIT
* nix 0.31.2, MIT
* notify 8.2.0, CC0-1.0
* num\_cpus 1.17.0, Apache-2.0 OR MIT
* object\_store 0.12.4, Apache-2.0 OR MIT
* object\_store\_occ 0.1.0, Apache-2.0
* octocrab 0.49.5, Apache-2.0 OR MIT
* odbc-api 20.2.0, MIT
* open 5.3.3, MIT
* opendal 0.55.0, Apache-2.0
* opentelemetry 0.31.0, Apache-2.0
* opentelemetry-http 0.31.0, Apache-2.0
* opentelemetry-otlp 0.31.1, Apache-2.0
* opentelemetry-prometheus 0.31.0, Apache-2.0
* opentelemetry-proto 0.31.0, Apache-2.0
* opentelemetry-zipkin 0.31.0, Apache-2.0
* opentelemetry\_sdk 0.31.0, Apache-2.0
* oracle 0.6.3, Apache-2.0 OR UPL-1.0
* parking\_lot 0.11.2, Apache-2.0 OR MIT
* parking\_lot 0.12.5, Apache-2.0 OR MIT
* parquet 57.2.0, Apache-2.0
* paste 1.0.15, Apache-2.0 OR MIT
* pdf-extract 0.8.0, MIT
* pem 3.0.6, MIT
* percent-encoding 2.3.2, Apache-2.0 OR MIT
* pgwire-replication 0.3.1, Apache-2.0 OR MIT
* pin-project 1.1.10, Apache-2.0 OR MIT
* pingora-lru 0.7.0, Apache-2.0
* pkcs8 0.9.0, Apache-2.0 OR MIT
* pkcs8 0.10.2, Apache-2.0 OR MIT
* postcard 1.1.3, Apache-2.0 OR MIT
* postgres-native-tls 0.5.2, Apache-2.0 OR MIT
* predicates 3.1.4, Apache-2.0 OR MIT
* prometheus 0.14.0, Apache-2.0
* prost 0.11.9, Apache-2.0
* prost 0.14.3, Apache-2.0
* quick-xml 0.31.0, MIT
* quick-xml 0.37.5, MIT
* quick-xml 0.38.4, MIT
* r2d2 0.8.10, Apache-2.0 OR MIT
* rand 0.7.3, Apache-2.0 OR MIT
* rand 0.8.5, Apache-2.0 OR MIT
* rand 0.9.4, Apache-2.0 OR MIT
* rand 0.10.1, Apache-2.0 OR MIT
* rayon 1.11.0, Apache-2.0 OR MIT
* rcgen 0.14.7, Apache-2.0 OR MIT
* rdkafka 0.39.0, MIT
* regex 1.12.3, Apache-2.0 OR MIT
* reqwest 0.12.24, Apache-2.0 OR MIT
* reqwest 0.13.2, Apache-2.0 OR MIT
* reqwest-eventsource 0.6.0, Apache-2.0 OR MIT
* rmcp 1.5.0, Apache-2.0
* roaring 0.11.3, Apache-2.0 OR MIT
* rstest 0.25.0, Apache-2.0 OR MIT
* rstest 0.26.1, Apache-2.0 OR MIT
* rusqlite 0.37.0, MIT
* rustls 0.21.12, Apache-2.0 OR ISC OR MIT
* rustls 0.23.36, Apache-2.0 OR ISC OR MIT
* rustls-native-certs 0.6.3, Apache-2.0 OR ISC OR MIT
* rustls-native-certs 0.8.3, Apache-2.0 OR ISC OR MIT
* rustls-pemfile 1.0.4, Apache-2.0 OR ISC OR MIT
* rustls-pemfile 2.2.0, Apache-2.0 OR ISC OR MIT
* rustyline 17.0.2, MIT
* schemars 0.8.22, MIT
* schemars 0.9.0, MIT
* schemars 1.2.1, MIT
* scopeguard 1.2.0, Apache-2.0 OR MIT
* scylla 1.4.1, Apache-2.0 OR MIT
* secrecy 0.10.3, Apache-2.0 OR MIT
* semver 1.0.27, Apache-2.0 OR MIT
* serde 1.0.228, Apache-2.0 OR MIT
* serde-value 0.7.0, MIT
* serde\_json 1.0.149, Apache-2.0 OR MIT
* sha2 0.10.9, Apache-2.0 OR MIT
* sha2 0.11.0, Apache-2.0 OR MIT
* simsimd 6.5.12, Apache-2.0
* snafu 0.8.9, Apache-2.0 OR MIT
* snmalloc-rs 0.3.8, MIT
* snowflake-api 0.14.0, Apache-2.0
* spark-connect-rs 0.0.1-beta.4, Apache-2.0
* spiceai 3.2.0, Apache-2.0
* spicepod-validator 2.0.0, ../../LICENSE
* ssh2 0.9.5, Apache-2.0 OR MIT
* subtle 2.6.1, BSD-3-Clause
* suppaftp 6.3.0, Apache-2.0 OR MIT
* sysinfo 0.36.1, MIT
* sysinfo 0.38.4, MIT
* tantivy 0.26.0, MIT
* tar 0.4.45, Apache-2.0 OR MIT
* tempfile 3.26.0, Apache-2.0 OR MIT
* tera 1.20.1, MIT
* test-log 0.2.19, Apache-2.0 OR MIT
* text-embeddings-backend 1.9.3,
* text-embeddings-backend-candle 1.9.3,
* text-embeddings-backend-core 1.9.3,
* text-embeddings-core 1.9.3,
* text-splitter 0.18.1, MIT
* tiberius 0.12.3, Apache-2.0 OR MIT
* tiktoken-rs 0.6.0, MIT
* tiktoken-rs 0.9.1, MIT
* tikv-jemallocator 0.6.1, Apache-2.0 OR MIT
* time 0.3.47, Apache-2.0 OR MIT
* tokenizers 0.21.4, Apache-2.0
* tokenizers 0.22.2, Apache-2.0
* tokio 1.49.0, MIT
* tokio-postgres 0.7.16, Apache-2.0 OR MIT
* tokio-rusqlite 0.7.0, MIT
* tokio-rustls 0.24.1, Apache-2.0 OR MIT
* tokio-rustls 0.26.4, Apache-2.0 OR MIT
* tokio-stream 0.1.18, MIT
* tokio-util 0.7.18, MIT
* tonic 0.14.5, MIT
* tonic-health 0.14.2, MIT
* tonic-prost 0.14.2, MIT
* tonic-prost-build 0.14.5, MIT
* tower 0.4.13, MIT
* tower 0.5.3, MIT
* tower-http 0.6.8, MIT
* tracing 0.1.44, MIT
* tracing-futures 0.2.5, MIT
* tracing-log 0.2.0, MIT
* tracing-opentelemetry 0.32.1, MIT
* tracing-subscriber 0.3.22, MIT
* tracing-test 0.2.5, MIT
* tract-onnx 0.22.0, Apache-2.0 OR MIT
* turso 0.6.0, MIT
* twox-hash 2.1.2, MIT
* url 2.5.8, Apache-2.0 OR MIT
* urlencoding 2.1.3, MIT
* utoipa 5.4.0, Apache-2.0 OR MIT
* utoipa-swagger-ui 9.0.2, Apache-2.0 OR MIT
* uuid 0.8.2, Apache-2.0 OR MIT
* uuid 1.21.0, Apache-2.0 OR MIT
* vortex 0.1.0, Apache-2.0
* vortex-scan 0.1.0, Apache-2.0
* vortex-session 0.1.0, Apache-2.0
* vortex-session 0.1.0, Apache-2.0
* vortex-utils 0.1.0, Apache-2.0
* vortex-utils 0.1.0, Apache-2.0
* walkdir 2.5.0, MIT OR Unlicense
* wasmtime 44.0.1, Apache-2.0 WITH LLVM-exception
* wat 1.248.0, Apache-2.0 OR Apache-2.0 WITH LLVM-exception OR MIT
* winver 1.0.0, MIT
* wiremock 0.6.5, Apache-2.0 OR MIT
* x509-certificate 0.25.0, MPL-2.0
* yaml-rust2 0.11.0, Apache-2.0 OR MIT
* zip 3.0.0, MIT
* zip 4.6.1, MIT
* zip 6.0.0, MIT
* zip 7.2.0, MIT
* zip 8.5.1, MIT
* zstd 0.13.3, MIT
--- # Spice.ai API Reference ## [📄️Overview](/docs/next/api/overview) [Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools.](/docs/next/api/overview) ## [📄️ADBC](/docs/next/api/adbc) [ADBC API Documentation](/docs/next/api/adbc) ## [📄️JDBC](/docs/next/api/jdbc) [JDBC API Documentation](/docs/next/api/jdbc) ## [📄️TLS](/docs/next/api/tls) [Encryption in transit with TLS documentation](/docs/next/api/tls) ## [📄️Authentication](/docs/next/api/auth) [Authentication documentation](/docs/next/api/auth) ## [📄️Arrow Flight SQL](/docs/next/api/arrow-flight-sql) [Query Spice using JDBC/ODBC/ADBC](/docs/next/api/arrow-flight-sql) ## [📄️ODBC](/docs/next/api/odbc) [ODBC API Documentation](/docs/next/api/odbc) ## [🗃HTTP](/docs/next/api/HTTP/runtime) [29 items](/docs/next/api/HTTP/runtime) --- # ADBC: Arrow Database Connectivity [ADBC](https://arrow.apache.org/adbc) is a set of APIs and libraries for Arrow-native access to databases. Spice supports ADBC clients using the [FlightSQL driver](https://arrow.apache.org/adbc/current/driver/flight_sql.html). ## Quickstart[​](#quickstart "Direct link to Quickstart") Get started with ADBC using Python. ### Installation[​](#installation "Direct link to Installation") Start a Python environment. ``` python ``` Install the ADBC driver manager, FlightSQL driver, and PyArrow. ``` pip install adbc_driver_manager adbc_driver_flightsql pyarrow ``` ### Create a connection to Spice over ADBC[​](#create-a-connection-to-spice-over-adbc "Direct link to Create a connection to Spice over ADBC") ``` >>> import adbc_driver_flightsql.dbapi >>> conn = adbc_driver_flightsql.dbapi.connect('grpc://localhost:50051') ``` ### Create a cursor[​](#create-a-cursor "Direct link to Create a cursor") ``` >>> cursor = conn.cursor() ``` ### Executing a query[​](#executing-a-query "Direct link to Executing a query") DBAPI interface: ``` >>> cursor.execute("SELECT 1, 2.0, 'Hello, world!'") >>> cursor.fetchone() (1, 2.0, 'Hello, world!') >>> cursor.execute("SHOW TABLES") >>> cursor.fetchall() [('spice', 'public', 'messages', 'BASE TABLE'), ('spice', 'runtime', 'task_history', 'BASE TABLE'), ('spice', 'information_schema', 'tables', 'VIEW'), ('spice', 'information_schema', 'views', 'VIEW'), ('spice', 'information_schema', 'columns', 'VIEW'), ('spice', 'information_schema', 'df_settings', 'VIEW'), ('spice', 'information_schema', 'schemata', 'VIEW')] ``` Arrow: ``` >>> cursor.execute("SELECT 1, 2.0, 'Hello, world!'") >>> cursor.fetch_arrow_table() pyarrow.Table 1: int64 2.0: double 'Hello, world!': string ---- 1: [[1]] 2.0: [[2]] 'Hello, world!': [["Hello, world!"]] ``` ## Parameterized Queries[​](#parameterized-queries "Direct link to Parameterized Queries") Spice supports parameterized queries when using ADBC clients. Parameterized queries help prevent SQL injection and improve code clarity by separating query logic from data values. The following example demonstrates how to use parameterized queries with the Python ADBC FlightSQL driver: ``` from adbc_driver_flightsql import DatabaseOptions from adbc_driver_flightsql.dbapi import connect with connect( "grpc://127.0.0.1:50051", ) as conn: with conn.cursor() as cur: cur.execute("SELECT $1 + 1 AS the_answer", parameters=(41,)) table = cur.fetch_arrow_table() print(table) cur.execute("SELECT 1 AS one") table = cur.fetch_arrow_table() print(table) conn.close() ``` --- # Arrow Flight SQL API [Arrow Flight SQL](https://arrow.apache.org/docs/format/FlightSql.html) is a protocol for interacting with SQL databases using the Arrow in-memory format and the Flight RPC framework. Spice implements the Flight SQL protocol, enabling querying of the datasets configured in Spice via tools that support connecting via one of the Arrow Flight SQL drivers, such as [DBeaver](https://dbeaver.io), [Tableau](https://www.tableau.com/), or [Power BI](https://www.microsoft.com/en-us/power-platform/products/power-bi). ![arrow flight and spice](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/0a8bc474-03c3-4c1c-8003-d250cd52b300/public) ## Authentication[​](#authentication "Direct link to Authentication") API Key authentication is supported for the Arrow Flight SQL endpoint. For more details, see [API Key Authentication](/docs/next/api/auth). --- # Authentication Spice supports adding optional authentication to its API endpoints via configurable API keys. Use the `auth` section as a child to `runtime` to provide the API keys. Multiple API keys can be specified, and any of the keys can be used to authenticate requests. ``` runtime: auth: api-key: enabled: true keys: - ${ secrets:api_key } # Use the secret replacement syntax to load the API key from a secret store - 1234567890 # Or specify the API key directly ``` To learn more about secrets, see [Secret Stores](/docs/next/components/secret-stores). info The API key authentication is applied on startup and changes will not take effect until the runtime is restarted. ## HTTP[​](#http "Direct link to HTTP") For HTTP routes, the API key can be supplied in either the `X-API-Key` header or the `Authorization: Bearer ${ api_key }` header. The `Bearer` scheme is matched case-insensitively. If both headers are present, `X-API-Key` takes precedence. ``` > curl -i "http://localhost:8090/v1/sql" -H "X-API-Key: 1234567890" -d 'SELECT 1' HTTP/1.1 200 OK content-type: text/plain; charset=utf-8 x-cache: Miss from spiceai content-length: 16 date: Fri, 08 Nov 2024 07:14:24 GMT [{"Int64(1)":1}] ``` Or using `Authorization: Bearer`: ``` > curl -i "http://localhost:8090/v1/sql" -H "Authorization: Bearer 1234567890" -d 'SELECT 1' ``` The `/health` and `/v1/ready` endpoints are not protected and can be accessed without an API key. ## Flight SQL[​](#flight-sql "Direct link to Flight SQL") For the Flight SQL endpoint, the API key is expected to be included in the `Authorization` header as a Bearer token, i.e. `Authorization: Bearer ${ api_key }`. ## Spice CLI[​](#spice-cli "Direct link to Spice CLI") When API key authentication is enabled, the Spice CLI can connect to the runtime by specifying the `--api-key` argument. ``` spice sql --api-key 1234567890 spice status --api-key 1234567890 spice refresh taxi_trips --api-key 1234567890 # etc. ``` --- # Generate Package ``` POST /v1/packages/generate ``` This endpoint generates a zip package from a specified GitHub source. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 500 Package generated successfully Invalid request parameters Internal server error --- # List Catalogs ``` GET /v1/catalogs ``` Returns a list of all registered catalogs (data sources). Catalogs provide metadata about schemas and tables available from external data sources. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of catalogs Internal server error occurred while processing catalogs --- # Get Iceberg API config ``` GET /v1/iceberg/config ``` This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. ## Responses[​](#responses "Direct link to Responses") * 200 API configuration retrieved successfully --- # List Datasets ``` GET /v1/datasets ``` This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, and additional filters can be applied using query parameters. Use `status=true` query parameter to include the current status of each dataset in the response. Possible status values: `initializing`, `ready`, `disabled`, `error`, `refreshing`, `shuttingdown`. When `status=true` and a dataset is in `Error`, the response also includes: * `error`: structured code object with `category`, `type`, and stable `code` * `error_message`: user-visible details ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of datasets. When `status=true` is specified, each dataset includes `status` and error metadata (`error`, `error_message`) when applicable. Internal server error occurred while processing datasets --- # List Iceberg namespaces ``` GET /v1/iceberg/namespaces ``` This endpoint retrieves namespaces available in the Iceberg catalog. If a `parent` namespace is provided, it will list the child namespaces under the specified parent. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 404 * 500 Namespaces retrieved successfully Bad request Namespace not found Internal server error --- # List Models ``` GET /v1/models ``` List all models, both machine learning and language models, available in the runtime. When `status=true` and a model is in `Error`, the response also includes: * `error`: structured code object with `category`, `type`, and stable `code` * `error_message`: user-visible details ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of models in JSON format Internal server error occurred while processing models --- # Check if a namespace exists. ``` GET /v1/iceberg/namespaces/:namespace ``` This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 404 Namespace exists Invalid namespace format Namespace does not exist --- # List Spicepods ``` GET /v1/spicepods ``` Get a list of spicepods and their summary details. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of spicepods Internal server error --- # Check Runtime Status ``` GET /v1/status ``` Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of connection statuses Error converting to CSV --- # Get a table. ``` GET /v1/iceberg/namespaces/:namespace/tables/:table ``` This endpoint returns the table if it exists, otherwise it returns a 404 Not Found response. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 500 Table exists Table does not exist An internal server error occurred while getting the table --- # List Workers ``` GET /v1/workers ``` Returns a list of all registered workers in the runtime. Workers are configurable processing units that can perform tasks like load balancing between models or implementing fallback strategies. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of workers in JSON format Internal server error occurred while processing workers --- # Check Namespace exists ``` HEAD /v1/iceberg/namespaces/:namespace ``` This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 404 Namespace exists Invalid namespace format Namespace does not exist --- # Check if a table exists. ``` HEAD /v1/iceberg/namespaces/:namespace/tables/:table ``` This endpoint returns a 200 OK response if the table exists, otherwise it returns a 404 Not Found response. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 Table exists Table does not exist --- # list\_tables ``` GET /v1/iceberg/namespaces/:namespace/tables ``` list\_tables ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 404 Tables retrieved successfully Invalid namespace format Namespace does not exist --- # List Tools ``` GET /v1/tools ``` Returns a list of all available tools in the Spice runtime. Tools provide reusable functionality that can be invoked programmatically or by AI agents. ## Responses[​](#responses "Direct link to Responses") * 200 All tools available in the Spice runtime --- # Send a Model Context Protocol message ``` POST /v1/mcp ``` Send a JSON-RPC message to the Spice MCP server using the MCP Streamable HTTP transport. The response is either a single JSON-RPC response (`application/json`) or an SSE stream (`text/event-stream`), selected via the `Accept` header. Session continuity is carried via the `Mcp-Session-Id` header. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 202 * 400 * 401 * 403 * 404 * 413 JSON-RPC response. Returned as `application/json` for a single response or `text/event-stream` when the server streams additional messages. Message accepted (for JSON-RPC notifications / responses that do not require a reply). Malformed JSON-RPC payload. Unauthorized. The `/v1/mcp` endpoint requires `runtime.auth` to be configured. Configure an API key provider in your Spicepod and retry with credentials. Forbidden. The `Host` header value is not in the `runtime.mcp.allowed_hosts` list. Configure `runtime.mcp.allowed_hosts` or set it to `["*"]` to allow all hosts. Unknown or expired `Mcp-Session-Id`. Payload too large. Maximum allowed size is 32 MiB. --- # Open an MCP server-to-client SSE stream ``` GET /v1/mcp ``` Open a long-lived server-to-client SSE stream for the current MCP session as defined by the Streamable HTTP transport. The `Mcp-Session-Id` header must identify an existing session created via `POST /v1/mcp`. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 401 * 403 * 404 SSE stream (`text/event-stream`) of server-originated MCP messages. Unauthorized. The `/v1/mcp` endpoint requires `runtime.auth` to be configured. Configure an API key provider in your Spicepod and retry with credentials. Forbidden. The `Host` header value is not in the `runtime.mcp.allowed_hosts` list. Unknown or expired `Mcp-Session-Id`. --- # Terminate an MCP Streamable HTTP session ``` DELETE /v1/mcp ``` Terminate the MCP session identified by the `Mcp-Session-Id` header. Subsequent requests bearing the same session id will receive `404 Not Found`. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 204 * 401 * 403 * 404 Session terminated. Unauthorized. The `/v1/mcp` endpoint requires `runtime.auth` to be configured. Configure an API key provider in your Spicepod and retry with credentials. Forbidden. The `Host` header value is not in the `runtime.mcp.allowed_hosts` list. Unknown or already-terminated `Mcp-Session-Id`. --- # Update Refresh SQL ``` PATCH /v1/datasets/:name/acceleration ``` Update the refresh SQL for a dataset's acceleration. This endpoint allows for updating the `refresh_sql` parameter for a dataset's acceleration at runtime. The change is **temporary** and will revert to the `spicepod.yml` definition at the next runtime restart. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 500 The refresh SQL was updated successfully. The specified dataset was not found An internal server error occurred while updating the refresh SQL --- # Create Chat Completion ``` POST /v1/chat/completions ``` Creates a model response for the given chat conversation. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 500 Chat completion generated successfully The specified model was not found An internal server error occurred while processing the chat completion --- # Refresh Dataset ``` POST /v1/datasets/:name/acceleration/refresh ``` Trigger an on-demand refresh for an accelerated dataset. This endpoint triggers an on-demand refresh for an accelerated dataset. The refresh only applies to `full` and `append` refresh modes (not `changes` mode). ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 201 * 400 * 404 * 500 Dataset refresh triggered successfully Acceleration not enabled for the dataset Dataset not found Internal server error occurred while processing refresh --- # Create Embeddings ``` POST /v1/embeddings ``` Creates an embedding vector representing the input text. Get a vector representation of a given input that can be easily consumed by machine learning models and algorithms. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 500 Embedding created successfully Model not found Internal server error --- # Text-to-SQL (NSQL) ``` POST /v1/nsql ``` Generate and optionally execute a natural-language text-to-SQL (NSQL) query. This endpoint generates a SQL query using a natural language query (NSQL) and optionally executes it. The SQL query is generated by the specified model and executed if the `Accept` header is not set to `application/sql`. When `stream` is true, the response is streamed as Server-Sent Events (SSE). ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 500 SQL query executed successfully Invalid request parameters Internal server error --- # post\_responses ``` POST /v1/responses ``` post\_responses ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 500 Response generated successfully The specified model was not found An internal server error occurred while processing the response --- # Search ``` POST /v1/search ``` Perform a vector similarity search (VSS) operation on a dataset. The search operation will return the most relevant matches based on cosine similarity with the input `text`. The datasets queries should have an embedding column, and the appropriate embedding model loaded. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 500 Search completed successfully Invalid request parameters Internal server error --- # SQL Query ``` POST /v1/sql ``` Execute a SQL query and return the results. This endpoint allows users to execute SQL queries directly from an HTTP request. The SQL query is sent as plain text in the request body. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 500 SQL query executed successfully Invalid SQL query or malformed input Internal server error --- # Check Readiness ``` GET /v1/ready ``` Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message "ready". If not, it returns a 503 status with the message "not ready". The behavior for when an accelerated dataset is considered ready is configurable via the `ready_state` parameter. See [Data refresh](https://spiceai.org/docs/components/data-accelerators/data-refresh#ready-state) for more details. In distributed (scheduler) mode the readiness response can additionally be gated on executor availability via the `min_ready_executors` and `min_ready_executors_percent` query parameters (both optional). Both gates must pass when supplied. Pass `verbose=true` to get a multi-line diagnostic body explaining each gate. ### Readiness Probe[​](#readiness-probe "Direct link to Readiness Probe") In production deployments, the /v1/ready endpoint can be used as a readiness probe for a Spice deployment to ensure traffic is routed to the Spice runtime only after all datasets have finished loading. Example Kubernetes readiness probe: ``` readinessProbe: httpGet: path: /v1/ready port: 8090 ``` Example with executor gating (scheduler role): ``` readinessProbe: httpGet: path: /v1/ready?min_ready_executors=3&min_ready_executors_percent=80 port: 8090 ``` ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 503 Service is ready Invalid query parameter or executor gate requested outside scheduler role Service is not ready --- # Run Tool ``` POST /v1/tools/:name ``` Execute a specific tool by name. The request body schema and response format are defined by each individual tool's specification. Use `GET /v1/tools` to discover available tools and their parameter schemas. ## Request[​](#request "Direct link to request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 500 Tool Specific response, in JSON format Tool not found An error occurred while calling the tool --- Version: 2.0.0-unstable # runtime The spiced runtime ### License --- # JDBC: Java Database Connectivity [JDBC](https://docs.oracle.com/javase/tutorial/jdbc/basics/index.html) (Java Database Connectivity) is a standard API for connecting to and interacting with databases. Spice supports JDBC clients through a JDBC driver implementation based on the [Flight SQL](https://arrow.apache.org/docs/format/FlightSql.html) protocol. This enables any JDBC-compatible application to connect to Spice, execute queries, and retrieve data. ## Download and install the Flight SQL JDBC driver[​](#download-and-install-the-flight-sql-jdbc-driver "Direct link to Download and install the Flight SQL JDBC driver") ### Download the Flight SQL JDBC driver[​](#download-the-flight-sql-jdbc-driver "Direct link to Download the Flight SQL JDBC driver") * Find the appropriate [Flight SQL JDBC driver](https://central.sonatype.com/artifact/org.apache.arrow/flight-sql-jdbc-driver/versions) version. * Click **Browse** next to the version you want to download * Click the `flight-sql-jdbc-driver-XX.XX.XX.jar` file (with only the `.jar` file extension) from the list of files to download the driver jar file ### Add the driver to your application[​](#add-the-driver-to-your-application "Direct link to Add the driver to your application") Follow the instructions specific to your application for adding a custom JDBC driver. Examples: **Tableau**: * Windows: `C:\Program Files\Tableau\Drivers` * Mac: `~/Library/Tableau/Drivers` * Linux: `/opt/tableau/tableau_driver/jdbc` * Start or restart Tableau [Full instruction](/docs/next/clients/tableau) **JetBrains DataGrip**: * In Database Explorer menu, select "+" and choose "Driver" * Follow the steps to add the JDBC `.jar` file [Full instruction](/docs/next/clients/jetbrains-datagrip) **DBeaver**: * In the DBeaver application menu bar, open the "Database" menu and choose: "Driver Manager" * Click the "New" button and follow instructions to add JDBC `.jar` file. [Full instruction](/docs/next/clients/dbeaver) ## Configure JDBC connection[​](#configure-jdbc-connection "Direct link to Configure JDBC connection") 1. Use the following configuration settings: * **URL**: `jdbc:arrow-flight-sql://{host}:{port}` * **Dialect**: `PostgreSQL` For example: ![](/img/tableau/tableau-jdbc-conn.png) 1. **Ensure Spice is running** 2. Click **Connect** info Spice has [TLS support](/docs/next/api/tls). For testing or non-production use cases for Spice without TLS, the following JDBC connection URL will bypass TLS `jdbc:arrow-flight-sql://{host}:{port}?useEncryption=false&disableCertificateVerification=true`. ### Authentication[​](#authentication "Direct link to Authentication") If [API Key authentication](/docs/next/api/auth) is enabled, the API key can be provided in the JDBC connection URL as a query parameter: `jdbc:arrow-flight-sql://{host}:{port}?user=&password=` Replace `` with the API key value. The `user` and `password` parameters are required by the JDBC driver, but only the `password` parameter is used for the API key. ## Execute Test Query[​](#execute-test-query "Direct link to Execute Test Query") In the configured application, run a sample query, such as `SELECT * FROM taxi_trips;` ![Query Results](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/0e9f3c0f-2e03-47f9-8d5e-65e078d7e900/public "Query Results") ## Parameterized Queries[​](#parameterized-queries "Direct link to Parameterized Queries") Spice supports parameterized queries with JDBC. Parameterized queries help prevent SQL injection and improve code clarity by separating query logic from data values. --- # ODBC: Open Database Connectivity [ODBC](https://learn.microsoft.com/en-us/sql/odbc/microsoft-open-database-connectivity-odbc) (Open Database Connectivity) is a low-level, high-performance interface that is designed specifically for relational data stores as a standard way to connect to, and interact with a database. Spice supports ODBC clients through an ODBC driver implementation based on the [Flight SQL](https://arrow.apache.org/docs/format/FlightSql.html) protocol. This enables ODBC-compatible applications to connect to Spice, execute queries, and retrieve data. Limitations 1. ODBC support is currently in alpha, and not all functionality is supported 2. The Arrow Flight SQL ODBC driver is not available for 32-bit Windows versions 3. The Arrow Flight SQL ODBC driver is not supported on the Apple ARM architecture ## Install and configure the Flight SQL ODBC driver[​](#install-and-configure-the-flight-sql-odbc-driver "Direct link to Install and configure the Flight SQL ODBC driver") ### Download and install the Flight SQL ODBC driver[​](#download-and-install-the-flight-sql-odbc-driver "Direct link to Download and install the Flight SQL ODBC driver") * Download and install the driver from the [ODBC driver download page](https://www.dremio.com/drivers/odbc/) * [Windows instructions](https://docs.dremio.com/current/sonar/client-applications/drivers/arrow-flight-sql-odbc-driver/#downloading-and-installing-on-windows) * [Linux](https://docs.dremio.com/current/sonar/client-applications/drivers/arrow-flight-sql-odbc-driver/#downloading-and-installing-on-linux) * [macOS instructions](https://docs.dremio.com/current/sonar/client-applications/drivers/arrow-flight-sql-odbc-driver/#downloading-and-installing-on-macos) ### Configure Flight SQL ODBC driver[​](#configure-flight-sql-odbc-driver "Direct link to Configure Flight SQL ODBC driver") * Windows * Linux * macOS - Open **Start Menu** -> **Windows Administrative Tools** -> click **ODBC Data Sources (64-bit)** - In the **ODBC Data Source Administrator (64-bit)** dialog, click **System DSN** ![ODBC Data Source Administrator](/img/odbc/spice-odbc-windows-config.png) * Select **Arrow Flight SQL ODBC DSN** and click **Configure** * Specify Spice.ai OSS runtime `HOST`, `PORT`, in the `UseEncryption` field, specify one of these values: * `true`, if [Spice is configured for encrypted communication (TLS)](https://docs.spiceai.org/api/tls) * `false`, otherwise * For descriptions of all the parameters, see [ODBC Connection Parameters](#odbc-connection-parameters). - Ensure that `unixODBC` is installed. To verify whether `unixODBC` is installed, execute the following commands: ``` which odbcinst which isql ``` * Copy the content of the `odbc.ini` and `odbcinst.ini` from the `/opt/arrow-flight-sql-odbc-driver/conf` and paste into your system `/etc/odbc.ini` and `/etc/odbcinst.ini` files * Edit `odbc.ini`: specify Spice.ai OSS runtime `HOST`, `PORT`, in the `UseEncryption` field, specify one of these values: * `true`, if [Spice is configured for encrypted communication (TLS)](https://docs.spiceai.org/api/tls) * `false`, otherwise * For descriptions of all the parameters, see [ODBC Connection Parameters](#odbc-connection-parameters) * Run this command to verify `unixODBC` configuration ``` odbcinst -j ``` * Ensure that [ODBC Manager](http://www.odbcmanager.net/) is installed. * Launch **ODBC Manager** -> **System DSN** page, select **Arrow Flight SQL ODBC DSN** and click **Configure**. * Specify Spice.ai OSS runtime `HOST`, `PORT`, in the `UseEncryption` field, specify one of these values: * `true`, if [Spice is configured for encrypted communication (TLS)](https://docs.spiceai.org/api/tls) * `false`, otherwise ![ODBC Data Source Administrator](/img/odbc/spice-odbc-macos-config.png) * For descriptions of all the parameters, see [ODBC Connection Parameters](#odbc-connection-parameters). ### ODBC Connection Parameters[​](#odbc-connection-parameters "Direct link to ODBC Connection Parameters") | Name | Type | Description | | ------------------------------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | host | string | The IP address or hostname for the Spice runtime. | | port | integer | The Spice runtime Arrow Flight endpoint port number | | useEncryption | integer | Configures the driver to use an SSL-encrypted connection. Accepted values: `true` (default) - The client communicates with the Spice runtime only using SSL encryption and `false` - SSL encryption is disabled. | | disableCertificateVerification | integer | Specifies whether the driver should verify the host certificate against the trust store. Default is `false` | | useSystemTrustStore | integer | Controls whether to use a CA certificate from the system's trust store, or from a specified .pem file. If `true` - The driver verifies the connection using a certificate in the system trust store. IF `false` - The driver verifies the connection using the .pem file specified by the `trustedCerts` parameter. `true` on Windows and macOS, `false` on Linux by default | | trustedCerts | string | The full path of the .pem file containing certificates trusted by a CA, for the purpose of verifying the server. If this option is not set, then the driver defaults to using the trusted CA certificates .pem file installed by the driver. | note The ODBC driver for Arrow Flight SQL does not support password-protected `.pem/.crt` files or multiple `.crt` certificates in a single `.pem/.crt` file. ## Execute Test Query[​](#execute-test-query "Direct link to Execute Test Query") Ensure Spice runtime has started: ``` spiced --flight 0.0.0.0:50051 ``` ``` 2024-10-06T20:06:50.084015Z INFO runtime::flight: Spice Runtime Flight listening on 0.0.0.0:50051 2024-10-06T20:06:50.086948Z INFO runtime::http: Spice Runtime HTTP listening on 127.0.0.1:8090 2024-10-06T20:06:50.297512Z INFO runtime: Initialized results cache; max size: 128.00 MiB, item ttl: 1s 2024-10-06T20:06:50.308775Z INFO runtime: Tool [search] ready to use 2024-10-06T20:06:50.308803Z INFO runtime: Tool [table_schema] ready to use 2024-10-06T20:06:50.308814Z INFO runtime: Tool [sql] ready to use 2024-10-06T20:06:50.308829Z INFO runtime: Tool [list_datasets] ready to use 2024-10-06T20:06:50.921420Z INFO runtime: Dataset taxi_trips registered (s3://spiceai-demo-datasets/taxi_trips/2024/), results cache enabled. ``` Configure the client app to use `Arrow Flight SQL ODBC Driver`. ![Example ODBC Client Configuration](/img/odbc/spice-odbc-example-config.png) Run a sample query, such as ``` SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10; ``` ![Example Query Results](/img/odbc/spice-odbc-example-query.png) ## Parameterized Queries[​](#parameterized-queries "Direct link to Parameterized Queries") Spice supports parameterized queries with ODBC. Parameterized queries help prevent SQL injection and improve code clarity by separating query logic from data values. --- # API Overview Spice provides high-performance, industry-standard APIs: ### SQL Query APIs[​](#sql-query-apis "Direct link to SQL Query APIs") * **Arrow Flight** / **Arrow Flight SQL**: High-performance SQL query. * **ODBC**, **JDBC**, **ADBC**: Standard SQL interfaces for database clients and analytics tools. ### OpenAI-Compatible APIs[​](#openai-compatible-apis "Direct link to OpenAI-Compatible APIs") * **HTTP APIs**: Compatible with the OpenAI SDK and AI SDK. Supports local model serving (CUDA/Metal accelerated) and gateway to hosted models. ### Iceberg Catalog REST APIs[​](#iceberg-catalog-rest-apis "Direct link to Iceberg Catalog REST APIs") * **HTTP APIs**: Unified API for consuming Apache Iceberg catalogs in data lake architectures. ### MCP API[​](#mcp-api "Direct link to MCP API") * **HTTP APIs**: The Model Context Protocol (MCP) helps integrate external tools and services into the Spice runtime. MCP tools can be accessed via HTTP APIs for tool integration and orchestration. For details, see the [MCP documentation](/docs/next/features/large-language-models/mcp). --- # TLS: Transport Layer Security [Transport Layer Security](https://en.wikipedia.org/wiki/Transport_Layer_Security) (TLS) is a cryptographic protocol that secures communication over a network. TLS is the successor to deprecated Secure-Sockets-Layer (SSL). Learn how to configure Spice to use TLS for encryption in transit. ## Pre-requisites[​](#pre-requisites "Direct link to Pre-requisites") A valid TLS certificate and private key in [PEM](https://en.wikipedia.org/wiki/Privacy-Enhanced_Mail) format are required. To generate certificates for testing, follow the [TLS Cookbook](https://github.com/spiceai/cookbook/tree/trunk/tls). ## Enable TLS via command line arguments[​](#enable-tls-via-command-line-arguments "Direct link to Enable TLS via command line arguments") Use `--tls-enabled true` to enable TLS from the command line. The arguments `--tls-certificate-file` and `--tls-key-file` specify the paths to the certificate and private key files. ``` # Provide the TLS certicate and key PEM files to the Spice runtime spiced --tls-enabled true --tls-certificate-file /path/to/cert.pem --tls-key-file /path/to/key.pem ``` Alternatively, to pass PEM-encoded certificate and private key strings directly, use the `--tls-certificate` and `--tls-key` arguments. ``` # Provide the TLS certicate and key using PEM-encoded strings to the Spice runtime export TLS_CERT=$(cat /path/to/cert.pem) export TLS_KEY=$(cat /path/to/key.pem) spiced --tls-enabled true --tls-certificate "$TLS_CERT" --tls-key "$TLS_KEY" ``` When using the Spice CLI, arguments, including the TLS arguments, are passed to `spice run` automatically. ``` # Run Spice using the CLI and provide the TLS certicate and key as PEM files spice run -- --tls-enabled true --tls-certificate-file /path/to/cert.pem --tls-key-file /path/to/key.pem ``` Note that `--` is used to separate the `spice run` arguments from the Spice runtime arguments. ## Enable TLS via spicepod.yaml[​](#enable-tls-via-spicepodyaml "Direct link to Enable TLS via spicepod.yaml") Use the `tls` section as a child to `runtime` to provide the certificate and key files/strings. ``` runtime: tls: enabled: true # Using filesystem paths certificate_file: /path/to/cert.pem key_file: /path/to/key.pem ``` ``` runtime: tls: enabled: true # Specify the certificate and key directly certificate: | -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- key: | -----BEGIN PRIVATE KEY----- ... -----END PRIVATE KEY----- ``` ``` runtime: tls: enabled: true # Provide the certificate and key using secrets certificate: ${secrets:tls_cert} key: ${secrets:tls_key} ``` To learn more about secrets, see [Secret Stores](/docs/next/components/secret-stores). ## Certificate Hot-Reload[​](#certificate-hot-reload "Direct link to Certificate Hot-Reload") When TLS is configured using **file paths** (`certificate_file` / `key_file` or `--tls-certificate-file` / `--tls-key-file`), the runtime automatically watches the certificate and key files for changes and reloads them without restarting. This is useful when certificates are rotated by external tools such as SPIRE, cert-manager, or kubelet. * In-flight TLS connections are unaffected — only new handshakes use the rotated certificate. * If a rotated file contains invalid PEM data, the runtime logs the error and continues serving with the previous certificate. * File changes are detected via polling (every 2 seconds). Atomic file renames are handled correctly. When TLS is configured using **inline values** (`certificate` / `key`, including `${secrets:…}` references), certificates are loaded once at startup and are not automatically reloaded. The `runtime_tls_reload_total` OTel counter tracks reload attempts: | Label | Values | | -------- | ------------------------------- | | `scope` | `public`, `cluster` | | `result` | `ok`, `io_error`, `parse_error` | info When using inline certificates or secrets (`certificate` / `key`), changes are not applied at runtime and will only take effect on restart. ## Output[​](#output "Direct link to Output") When TLS is enabled, the runtime output will print the TLS certificate details. ``` INFO runtime: All endpoints secured with TLS using certificate: CN=spiced.localhost, OU=IT, O=Widgets, Inc., L=Seattle, S=Washington, C=US ``` ## Using the Spice CLI[​](#using-the-spice-cli "Direct link to Using the Spice CLI") When TLS is enabled in the runtime, the Spice CLI can be configured to connect to the runtime using TLS by specifying the `--tls-root-certificate-file` argument, providing the path to the root certificate file. ``` spice sql --tls-root-certificate-file /path/to/root.pem ``` ## Mutual TLS (mTLS)[​](#mutual-tls-mtls "Direct link to Mutual TLS (mTLS)") Enterprise Feature mTLS (client certificate authentication) is included in the Enterprise distribution of Spice.ai. [Learn more](https://docs.spice.ai/docs/enterprise). mTLS extends standard TLS by requiring the client to also present a certificate during the TLS handshake. This provides cryptographic authentication of both the server and the client. ### Enable mTLS via spicepod.yaml[​](#enable-mtls-via-spicepodyaml "Direct link to Enable mTLS via spicepod.yaml") Set `client_auth_mode` to `request` or `required` and provide a CA bundle to verify client certificates: ``` runtime: tls: enabled: true certificate_file: /path/to/server.crt key_file: /path/to/server.key client_auth_mode: required client_auth_ca_file: /path/to/client-ca.pem ``` ### Enable mTLS via command line[​](#enable-mtls-via-command-line "Direct link to Enable mTLS via command line") ``` spiced --tls-enabled true \ --tls-certificate-file /path/to/server.crt \ --tls-key-file /path/to/server.key \ --tls-client-auth-mode required \ --tls-client-auth-ca-file /path/to/client-ca.pem ``` ### Client auth modes[​](#client-auth-modes "Direct link to Client auth modes") | Mode | Behavior | | ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `none` *(default)* | Standard one-way TLS. No client certificate is requested. | | `request` | The server requests a client certificate but accepts connections without one. Presented certificates are verified against the CA. Useful for migration or audit-only deployments. | | `required` | A valid client certificate is required. The Flight listener rejects no-cert connections at the TLS handshake. The HTTP listener admits no-cert connections so `/health` and `/v1/ready` remain accessible for Kubernetes probes, but all other HTTP endpoints return 401 without a verified client certificate. | ### Connecting with a client certificate[​](#connecting-with-a-client-certificate "Direct link to Connecting with a client certificate") Use cURL with a client certificate: ``` curl --cacert ca.pem --cert client.crt --key client.key \ https://localhost:8090/v1/sql -d 'SELECT 1' ``` Use the Spice CLI with a client certificate: ``` spice sql --tls-root-certificate-file ./ca.pem \ --client-tls-certificate-file ./client.crt \ --client-tls-key-file ./client.key ``` ### Probe and metrics access[​](#probe-and-metrics-access "Direct link to Probe and metrics access") Kubernetes liveness/readiness probes (`/health`, `/v1/ready`) and the metrics endpoint (`/metrics`) are always accessible without a client certificate, even under `client_auth_mode: required`. ### Client CA hot-reload[​](#client-ca-hot-reload "Direct link to Client CA hot-reload") When `client_auth_ca_file` is used, the CA bundle is watched for changes and reloaded atomically alongside the server certificate and key. When `client_auth_ca` (inline) is used, the CA is loaded once at startup. For a complete walkthrough, see the [mTLS Cookbook recipe](https://github.com/spiceai/cookbook/tree/trunk/mtls). --- # Spice.ai CLI Reference The Spice CLI is a set of commands to create and manage Spicepods and interact with the Spice runtime. ## Install[​](#install "Direct link to Install") The Spice CLI can be installed by: * macOS, Linux, and WSL * Windows - Running `curl https://install.spiceai.org | /bin/bash` - Using `brew`: `brew install spiceai/spiceai/spice` - Downloading the binary from [GitHub Releases](https://github.com/spiceai/spiceai/releases) * Shell command ``` curl -L "https://install.spiceai.org/Install.ps1" -o Install.ps1 && PowerShell -ExecutionPolicy Bypass -File ./Install.ps1 ``` * Downloading the binary from [GitHub Releases](https://github.com/spiceai/spiceai/releases) The `spice` program will be added to the PATH automatically for **bash**, **fish**, and **zsh** shells. After installing the Spice CLI for the first time, verify the installation by running `spice version`. Expected output: ``` CLI version: 1.x.x Runtime version: (not installed) ``` The runtime is downloaded and installed automatically upon first run of `spice run`. ## Getting started[​](#getting-started "Direct link to Getting started") For getting started with Spice using the Spice CLI, see the [Getting Started Guide](/docs/next/getting-started). Use `spice help` for all commands and `spice [command] --help` for more information about a command. A typical command-line workflow might be as follows: ``` # Start the runtime spice run ``` Run new shell in the same folder: ``` # Init new app spice init spice_app # Add the Quickstart Spicepod spice add spiceai/quickstart ``` Common commands are: | Command | Description | | --------------- | ---------------------------------------------------------------- | | `spice run` | Start the Spice runtime, installing if necessary | | `spice sql` | Open an interactive SQL REPL connected to the running runtime | | `spice chat` | Open an interactive chat REPL for AI model conversations | | `spice init` | Initialize a new Spice app with a `spicepod.yaml` | | `spice add` | Add a Spicepod dependency to the project | | `spice login` | Log in to Spice.ai Cloud Platform or other data sources | | `spice trace` | Inspect traces for AI chat, SQL queries, and other runtime tasks | | `spice version` | Display CLI and runtime versions | | `spice upgrade` | Upgrade the Spice CLI to the latest release | | `spice help` | Display help for any command | See [Spice CLI command reference](/docs/next/reference) for the full list of available commands. ## Direct Command Shortcuts[​](#direct-command-shortcuts "Direct link to Direct Command Shortcuts") For quick scripting and shell workflows, the CLI accepts top-level shortcuts that skip the subcommand name: | Shortcut | Equivalent | Description | | ---------------------- | -------------------------------- | ------------------------------------ | | `spice -sql ` | `spice sql --query ` | Run a single SQL statement and exit. | | `spice -p ` | `spice chat ` (one-shot) | Send a single chat prompt and exit. | | `spice -chat ` | `spice chat ` (one-shot) | Alias for `-p`. | The shortcut consumes the first non-flag argument as the value; any other flags (including global flags such as `--cloud`, `--api-key`, `--endpoint`, and `--machine`) are forwarded to the underlying command. Use `--` to terminate flag parsing when the prompt or query begins with a `-`. ``` # One-shot SQL against the local runtime spice -sql "select count(*) from taxi_trips" # One-shot chat prompt spice -p "summarize the orders table" # Same against Spice.ai Cloud spice --cloud --api-key $SPICE_KEY -sql "select 1" ``` ## Machine-readable Mode[​](#machine-readable-mode "Direct link to Machine-readable Mode") Pass `--machine` (alias `--programmatic`) as a global flag to make the CLI output JSON wherever supported and emit structured JSON for errors. This is intended for LLM agents, CI scripts, and other automation that needs to parse `spice` output reliably. ``` # Default human output spice version # Same command, JSON output for automation spice --machine version # Works with subcommands spice --machine status spice --machine query list spice --machine cloud secrets list # Combine with direct shortcuts spice --machine -sql "select 1" ``` `--machine` is global and may appear before or after the subcommand (`spice status --machine` is also valid). When set, mutating commands that support JSON output also default to JSON, and clap parsing errors are emitted as JSON to stderr. ## Updating[​](#updating "Direct link to Updating") To update to latest CLI, run the upgrade command. ``` spice upgrade ``` note Upgrade command is supported from CLI v0.3.1. For version < 0.3.1 users have to re-run the [install](#install) script. ## Uninstall[​](#uninstall "Direct link to Uninstall") The Spice CLI is installed by default to `$HOME/.spice/bin/spice` and a line added to the shell config, such as `.zshrc` It can be uninstalled by deleting the `spice` binary and removing the PATH addition from the rc file. --- # Spice.ai OSS CLI command reference ## spice[​](#spice "Direct link to spice") ### Usage[​](#usage "Direct link to Usage") ``` spice [command] [--help] ``` ### Full Command Reference[​](#full-command-reference "Direct link to Full Command Reference") | Command | Description | | --------------------------------------------------- | --------------------------------------------------------------------------- | | acceleration | Manage dataset acceleration features | | [add](/docs/next/cli/reference/add) | Add Spicepod - adds a Spicepod to the project | | catalog | Add or configure catalog entries in `spicepod.yaml` | | [catalogs](/docs/next/cli/reference/catalogs) | List [catalogs](/docs/next/components/catalogs) loaded by the Spice runtime | | [chat](/docs/next/cli/reference/chat) | Chat with an LLM | | [completions](/docs/next/cli/reference/completions) | Generate shell completions for the Spice CLI | | cloud | Manage Spice Cloud resources | | cluster | Cluster operations for the Spice runtime | | [connect](/docs/next/cli/reference/connect) | Enroll this host with Spice Cloud (Cloud Connect) | | [dataset](/docs/next/cli/reference/dataset) | Add or configure dataset entries in `spicepod.yaml` | | [datasets](/docs/next/cli/reference/datasets) | Lists datasets loaded by the Spice runtime | | embedding | Add or configure embedding entries in `spicepod.yaml` | | extension | Add or configure entries under `extensions:` in `spicepod.yaml` | | [feedback](/docs/next/cli/reference/feedback) | Open the Spice.ai community Slack to share feedback | | function | Add or configure function entries in `spicepod.yaml` | | help | Help about any command | | [init](/docs/next/cli/reference/init) | Initialize Spice app - creates a new spicepod.yaml | | [install](/docs/next/cli/reference/install) | Install or reinstall the Spice.ai runtime | | [login](/docs/next/cli/reference/login) | Login to Spice.ai or configure credentials for data sources | | management | Configure the `management:` section of `spicepod.yaml` | | metadata | Add or configure key/value entries under `metadata:` in `spicepod.yaml` | | model | Add or configure model entries in `spicepod.yaml` | | [models](/docs/next/cli/reference/models) | Lists models loaded by the Spice runtime | | [nsql](/docs/next/cli/reference/nsql) | Text-to-SQL REPL - translate natural language to SQL | | [pods](/docs/next/cli/reference/pods) | Lists Spicepods loaded by the Spice runtime | | [query](/docs/next/cli/reference/query) | Submit an async query or start an interactive async query REPL | | [refresh](/docs/next/cli/reference/refresh) | Refresh a dataset loaded by the Spice runtime | | reranker | Add or configure reranker entries in `spicepod.yaml` | | [run](/docs/next/cli/reference/run) | Run Spice - starts the Spice runtime, installing if necessary | | runtime | Configure the `runtime:` section of `spicepod.yaml` | | [search](/docs/next/cli/reference/search) | Search datasets with embeddings | | secret | Add or configure secret entries in `spicepod.yaml` | | snapshots | Configure the `snapshots:` section of `spicepod.yaml` | | [sql](/docs/next/cli/reference/sql) | Start an interactive SQL query session against the Spice runtime | | [spiced](/docs/next/cli/reference/spiced) | Spice runtime binary — direct invocation reference | | [status](/docs/next/cli/reference/status) | Spice runtime status | | tool | Add or configure tool entries in `spicepod.yaml` | | [trace](/docs/next/cli/reference/trace) | Return traces for operations that occurred in Spice | | [upgrade](/docs/next/cli/reference/upgrade) | Upgrades the Spice CLI and runtime to the latest or specified version | | [validate](/docs/next/cli/reference/validate) | Validate a spicepod.yaml without starting the runtime | | [version](/docs/next/cli/reference/version) | Spice CLI version | | view | Add or configure view entries in `spicepod.yaml` | | worker | Add or configure worker entries in `spicepod.yaml` | | workers | Lists workers loaded by the Spice runtime | ### Spicepod manifest editors[​](#spicepod-manifest-editors "Direct link to Spicepod manifest editors") `catalog`, `dataset`, `embedding`, `function`, `model`, `reranker`, `secret`, `tool`, `view`, and `worker` all edit the matching list section of `spicepod.yaml` and share the same two subcommands and body flags: ``` spice add [body flags] # add a new entry; fails if it already exists spice configure [body flags] # add or update an entry in place ``` `runtime`, `management`, and `snapshots` edit singleton sections, so they expose only `configure`. `extension` edits the `extensions:` map (`add` and `configure`), and `metadata` edits the `metadata:` map (`add`, `configure`, and `set `). See [`spice dataset`](/docs/next/cli/reference/dataset) for the shared body-flag reference. The other list editors accept the same flags, plus the type-specific ones: `--sql` / `--sql-ref` for views, `--cron` for workers, and `--body` / `--body-ref` for functions. Run `spice --help` for the per-command help. ### Command Flags[​](#command-flags "Direct link to Command Flags") The following flags are global — they are accepted by every command: | Flag | Description | | ------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------- | | `-h`, `--help` | Print the help message. | | `-v`, `--verbose` | Increase log verbosity. `-v` for debug, `-vv` for trace. | | `--machine` | Machine-readable mode for LLMs and automation: prefer JSON output where supported, and always emit structured JSON errors. Alias: `--programmatic`. | | `--api-key ` | API key used to authenticate with the runtime or the Spice.ai Cloud Platform. Also read from the `SPICE_API_KEY` environment variable. | | `--cloud` | Target the Spice.ai Cloud Platform instead of a local runtime. Requires `--api-key`. | | `--cloud-region ` | Spice.ai Cloud Platform runtime endpoint region, used with `--cloud`. Defaults to `us-east-1`. | | `--http-endpoint ` | HTTP endpoint of the Spice runtime to talk to. Defaults to `http://127.0.0.1:8090`. | | `--tls-root-certificate-file ` | Path to a PEM root certificate used to verify the runtime's TLS server certificate. | --- # add Add a Spicepod to the project. ### Usage[​](#usage "Direct link to Usage") ``` spice add [spicepod path] [flags] ``` * `spicepod path`: Either a [spicerack.org](https://spicerack.org) slug (optionally pinned to a version with `@`), or a path to a Spicepod directory on the local filesystem. | Form | Example | Resolves to | | --------------------------- | ---------------------------- | -------------------------------------------------------------- | | Spicerack slug | `spiceai/quickstart` | The latest published version of the Spicepod on spicerack.org. | | Spicerack slug with version | `spiceai/quickstart@v1.0` | The `v1.0` version of that Spicepod. | | Local directory | `../shared/pods/analytics` | A Spicepod directory copied from the local filesystem. | | Local `file://` URL | `file:///srv/pods/analytics` | The same, written as a URL. | A path is treated as local when it starts with `/`, `../`, or `file://`, or when it already exists on disk; anything else is fetched from spicerack.org. #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message ### Examples[​](#examples "Direct link to Examples") Adding a Spicepod from Spicerack (like `spiceai/quickstart`): ``` > spice add spiceai/quickstart ``` Pinning a published Spicepod to a version: ``` > spice add spiceai/quickstart@v1.0 ``` Adding a Spicepod from a local directory: ``` > spice add ../shared/pods/analytics ``` A local Spicepod is copied into `spicepods/[pod-name]`, where `pod-name` is the source directory's name lowercased, and the root `spicepod.yaml` records the dependency as a path relative to the app directory rather than as a slug. **Directory Structure**: The command makes two main modifications to the directory structure: 1. It creates the `spicepods` directory in the project root if it does not exist. 2. It adds the Spicepod defined by the Spicerack Slug in the relative path in the `spicepods` directory. For this example, the command would create the directories `spicepods/spiceai` and `spicepods/spiceai/quickstart`, instantiating a Spicepod under the latter. More generally, the Spicepod is placed under `spicepods/[slug]`, where `slug` is the Spicerack slug associated with that Spicepod. After running the command, the directory structure looks like this: ``` ├── spicepods/ │ ├── spiceai/ │ ├── quickstart/ │ ├── spicepod.yaml ├── spicepod.yaml └── ... ``` Any other Spicepods added using `spice add` are placed in the `spicepods` directory. `spice add` also creates the appropriate Spicepod for the given Spicerack slug. For this example with `spiceai/quickstart`, the command creates the following the Spicepod under `spicepods/spiceai/quickstart`: ``` # File: ./spicepods/spiceai/quickstart/spicepod.yaml version: v1 kind: Spicepod name: quickstart datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips description: taxi trips in s3 params: file_format: parquet acceleration: enabled: true ``` The `add` command also includes the above Spicepod as a dependency in the root `spicepod.yaml`, creating this file if it does not exist: ``` # File: ./spicepod.yaml version: v1 kind: Spicepod name: Spice AI quickstart dependencies: - spiceai/quickstart ``` --- # catalogs List [catalogs](/docs/next/components/catalogs) currently loaded by the Spice runtime. ### Requirements[​](#requirements "Direct link to Requirements") * Spice runtime must be running ### Usage[​](#usage "Direct link to Usage") ``` > spice catalogs [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `-o`, `--output ` Output format: `table` (default) or `json`. ### Example[​](#example "Direct link to Example") ``` > spice catalogs FROM NAME spiceai spiceai databricks us_census_data ``` --- # chat Start an interactive or one-shot chat with a [model](/docs/next/components/models) registered in the Spice runtime. ## Requirements[​](#requirements "Direct link to Requirements") * Spice runtime must be running * At least one model defined in `spicepod.yaml` and the model is ready ## Usage[​](#usage "Direct link to Usage") ### Interative Chat: Invoke the command without arguments to open a REPL[​](#interative-chat-invoke-the-command-without-arguments-to-open-a-repl "Direct link to Interative Chat: Invoke the command without arguments to open a REPL") ``` spice chat [flags] ``` ### One-shot Chat: Pass a single message as the argument to send a one-shot chat request and print the response[​](#one-shot-chat-pass-a-single-message-as-the-argument-to-send-a-one-shot-chat-request-and-print-the-response "Direct link to One-shot Chat: Pass a single message as the argument to send a one-shot chat request and print the response") ``` spice chat [flags] [] ``` ## Flags[​](#flags "Direct link to Flags") * `--model`, `-m` Target model for the chat request. When omitted, the CLI uses the single ready model or prompts for a choice if several models are ready. * `--temperature ` Model temperature used for chat request. * `--endpoint ` Specifies the remote Spice instance HTTP endpoint (e.g., `http://localhost:8090`). * `--headers ` Custom HTTP headers in format `Key:Value` (can be specified multiple times). * `--output `, `-o` Output format: `table` (default) or `json`. ## Examples[​](#examples "Direct link to Examples") When exactly one model is **ready**, `spice chat` opens a REPL that uses that model automatically: ``` > spice chat Using model: openai chat> hello Hello! How can I assist you today? Time: 0.57s (first token 0.53s). Tokens: 18. Prompt: 8. Completion: 10 (325.04/s). ``` #### Remote and Cloud Examples[​](#remote-and-cloud-examples "Direct link to Remote and Cloud Examples") ``` # Chat with Spice Cloud spice chat --cloud --api-key --model # Chat with a remote spiced instance over HTTP spice chat --endpoint http://my-remote-host:8090 --model # Chat with a remote spiced instance over Arrow Flight SQL (gRPC) spice chat --endpoint grpc://my-remote-host:50051 --model ``` When multiple models are **ready**, the command prompts for a selection before starting the REPL: ``` > spice chat Use the arrow keys to navigate: ↓ ↑ → ← ? Select model: ▸ openai llama Using model: openai chat> hello Hello! How can I assist you today? Time: 0.55s (first token 0.43s). Tokens: 18. Prompt: 8. Completion: 10 (80.09/s). ``` Passing `--model` skips the prompt and directs the request to the specified model. The flag works both in REPL mode and in one‑shot mode: ``` # REPL spice chat --model openai chat> hello Hello! How can I assist you today? Time: 0.61s (first token 0.58s). Tokens: 18. Prompt: 8. Completion: 10 (285.90/s). ``` Single prompt: ``` # One‑shot spice chat --model openai "hello" Hello! How can I assist you today? Time: 1.10s (first token 0.80s). Tokens: 18. Prompt: 8. Completion: 10 (33.74/s). ``` --- # completions Generate shell completions for the Spice CLI. By default, `spice completions` auto-detects the user's shell and writes the completion script to the standard completion directory for that shell. If the shell cannot be detected, specify it as an argument. ### Usage[​](#usage "Direct link to Usage") ``` spice completions [shell] [flags] ``` `shell` - the shell to generate completions for. Supported values: `bash`, `zsh`, `fish`, `elvish`, `powershell`. If omitted, the shell is detected from the `$SHELL` environment variable. #### Flags[​](#flags "Direct link to Flags") * `--stdout` Print completions to stdout instead of writing to a file (useful for piping) * `-h`, `--help` Print this help message ### Default completion directories[​](#default-completion-directories "Direct link to Default completion directories") When writing to a file, `spice completions` uses the following directories by default: | Shell | Directory | | ---------- | ----------------------------------------------------------------------------------- | | bash | `~/.local/share/bash-completion/completions/spice` (Homebrew on macOS if available) | | zsh | `~/.local/share/zsh/site-functions/_spice` (Homebrew on macOS if available) | | fish | `~/.config/fish/completions/spice.fish` | | elvish | `~/.config/elvish/lib/spice.elv` | | powershell | `~/.config/powershell/completions/spice.ps1` | Parent directories are created automatically if they do not exist. ### Examples[​](#examples "Direct link to Examples") #### Auto-detect shell and install completions[​](#auto-detect-shell-and-install-completions "Direct link to Auto-detect shell and install completions") ``` spice completions ``` #### Generate completions for a specific shell[​](#generate-completions-for-a-specific-shell "Direct link to Generate completions for a specific shell") ``` spice completions zsh ``` #### Print completions to stdout[​](#print-completions-to-stdout "Direct link to Print completions to stdout") ``` spice completions bash --stdout ``` #### Pipe completions to a custom file[​](#pipe-completions-to-a-custom-file "Direct link to Pipe completions to a custom file") ``` spice completions zsh --stdout > ~/.zfunc/_spice ``` --- # connect Enroll this host with Spice Cloud for remote management (Cloud Connect). `spice connect ` **enrolls and exits**: the runtime identity is issued and persisted locally, the instance is registered with Spice Cloud, and — unless `--install` is passed — nothing is left running. Start the managed runtime separately with [`spiced --cloud-connect`](/docs/next/cli/reference/spiced#cloud-connect-flags), or pass [`--install`](#flags) to have enrollment install and start it as a persistent system service. If the runtime is not installed, enrollment installs it first so that next step works immediately. ### Usage[​](#usage "Direct link to Usage") ``` spice connect [TARGET] [flags] spice connect status spice connect remove ``` * `TARGET`: a Spice Cloud adoption code (`SPICE-ADOPT-XXXXX-XXXXX-XXXXX-XXXXX`, each segment five uppercase letters or digits), obtained from the Spice Cloud portal. Required on a host that has not run [`login`](/docs/next/cli/reference/login). With no argument, `spice connect` resolves in this order: 1. **Already enrolled** (an identity exists for this directory) — reports the enrollment state, the same as `spice connect status`. The directory is never re-enrolled, which would create a second registry row for one host. 2. **An interrupted enroll is staged** — resumes it with the staged code. 3. **Otherwise** — mints a single-use adoption code from the `spice login` credential on this host and redeems it in the same command, so there is nothing to copy from the portal. The minted code is never displayed and never written to disk. #### Subcommands[​](#subcommands "Direct link to Subcommands") * `status` — Show the current enrollment state. * `remove` — Release this instance: report the release to Spice Cloud, uninstall the service when one was installed, and clear the local Cloud Connect identity from disk. Asks for confirmation first — pass `--yes` to skip the prompt, which is required when stdin is not a terminal. Needs root when a service was installed. A running `spiced` keeps its in-memory identity until it is restarted or the cloud sends a Remove command (a dropped stream just reconnects with the same identity), so restart `spiced` to stop remote management immediately. #### Flags[​](#flags "Direct link to Flags") * `--endpoint ` The Spice Cloud enroll endpoint the adoption code is presented to. Default: `https://api.spice.ai`. The gateway (stream) address is issued by the enroll response and is not configured here. * `--dir ` Root this instance's state at `/.spice`, so several instances on one host enroll independently. Defaults to the current directory. Applies to enrollment, `status`, and `remove`. * `--org ` Which Spice Cloud org to enroll into, when the `spice login` credential on this host belongs to several. The org is resolved from the token, so naming an org the login does not belong to is an error rather than a silent enroll into a different one. Ignored when an explicit adoption code is given — a code already carries its own org scope. * `--region ` Where this instance runs, e.g. `us-west-2` or `on-prem-syd`. A customer-declared label, not a probed fact: Spice Cloud displays it on the registry row and resolves this instance's gateway from it, falling back to the deployment's home stamp for a label it cannot rank. Any label of 2–64 lowercase letters, digits, and hyphens, starting and ending with a letter or digit, is accepted — it need not be a real cloud region. Omitting it on a re-enroll leaves an existing region untouched. * `--app-name ` Attach the instance to the existing Spice Cloud app of this name at enroll time, instead of attaching it later in the portal. When no such app exists the command fails **without consuming the adoption code** — pass `--create` to create it. * `--create` With `--app-name`: create the app when it does not exist, then attach the instance to it. An absent app without this flag is an error; apps are never created silently. * `--install` Install and start [`spiced --cloud-connect`](/docs/next/cli/reference/spiced#cloud-connect-flags) as a persistent system service running from the instance directory, so the instance survives reboots and closed terminals. Requires root, and either Linux with systemd or macOS with launchd. Combinable with an adoption code, or run on its own after a prior enroll. Re-running is the idempotent in-place upgrade path: latest binary, rewritten service definition, service restarted, staged identity untouched. * `-y`, `--yes` Skip the confirmation prompt. Applies to `remove`. * `-h`, `--help` Print this help message #### Environment variables[​](#environment-variables "Direct link to Environment variables") For hosts provisioned without an interactive CLI invocation: | Variable | Equivalent to | Purpose | | ------------------------------ | ------------- | -------------------------------------------------------------------------------------------------------------------------- | | `SPICE_CONNECT_ADOPT_CODE` | `TARGET` | Adoption code to enroll with. | | `SPICE_CONNECT_ADOPT_APP_NAME` | `--app-name` | Spice Cloud app to attach the instance to at enroll. | | `SPICE_CONNECT_ADOPT_CREATE` | `--create` | Create the app named above when it does not exist. | | `SPICE_CONNECT_ADOPT_REGION` | `--region` | Where this instance runs. | | `SPICE_CLOUD_ENDPOINT` | `--endpoint` | Enroll endpoint override. | | `SPICE_CONFIG_DIR` | — | Overrides where per-instance state (identity and staged adoption code) is written, in full. Takes precedence over `--dir`. | Flags take precedence over their corresponding environment variables. The enroll endpoint resolves in full as `--endpoint`, then `SPICE_CLOUD_ENDPOINT`, then a `cloud-endpoint` file written into the config directory by a previous enroll, then the `https://api.spice.ai` default — so later `spiced` starts reach the same control plane the enroll used. Containers use this environment-variable flow together with `spiced --cloud-connect` under the container runtime's restart policy; `--install` is Linux/systemd and macOS/launchd only, and Windows enrolls and runs under the user's own supervisor. ### Examples[​](#examples "Direct link to Examples") Enroll this host, then start the managed runtime: ``` > spice connect SPICE-ADOPT-7K2PX-9XYZ2-A1B2C-D3E4F > spiced --cloud-connect ``` Enroll a host already logged in with [`login`](/docs/next/cli/reference/login), with no code to copy from the portal: ``` > spice connect ``` Enroll and install a persistent service in one step, so the instance survives reboots: ``` > sudo spice connect --install ``` Record where the instance runs, and attach it to an app: ``` > spice connect --region on-prem-syd --app-name edge-fleet ``` Enroll a second instance on the same host, rooted at its own directory: ``` > spice connect SPICE-ADOPT-7K2PX-9XYZ2-A1B2C-D3E4F --dir /opt/edge-1 ``` Attach the instance to a Spice Cloud app at enroll time, creating the app if it does not exist: ``` > spice connect SPICE-ADOPT-7K2PX-9XYZ2-A1B2C-D3E4F --app-name edge-fleet --create ``` Inspect the enrollment, then release the instance (root because an installed service is uninstalled too): ``` > spice connect status > sudo spice connect remove ``` ### Deprecated: adding a cloud-hosted Spicepod[​](#deprecated-adding-a-cloud-hosted-spicepod "Direct link to Deprecated: adding a cloud-hosted Spicepod") `spice connect /` added a Spicepod hosted on the Spice.ai Cloud Platform, using Spice.ai Cloud authentication from [`login`](/docs/next/cli/reference/login). It is **deprecated**, prints a warning, and will be removed in a future release — use [`spice add /`](/docs/next/cli/reference/add) instead: ``` > spice add spiceai/quickstart ``` --- # dataset Add or configure dataset entries in `spicepod.yaml`. ### Usage[​](#usage "Direct link to Usage") ``` spice dataset [command] [name] [flags] ``` Available `command`s: * `add`: Add a new dataset entry. Fails if a dataset with that name already exists. * `configure`: Add or update a dataset entry in place. Run with no name and no body flags to configure a dataset through interactive prompts instead. **Note**: In order to run `spice dataset configure`, there *must* be a `spicepod.yaml` file in the root of your project directory. To create this file, see [`spice init`](/docs/next/cli/reference/init). #### Flags[​](#flags "Direct link to Flags") Both `add` and `configure` accept the same body flags, shared with the other Spicepod manifest editors (`spice model`, `spice catalog`, `spice view`, and so on): * `--from ` Provider or URI for the dataset, e.g. `s3://bucket/key` or `databricks:catalog.schema.table` * `--ref ` Add a dataset reference (`ref:`) instead of an inline definition * `--description ` Human-readable description * `--param ` Add a `params:` entry. Repeatable. Values are stored as strings unless prefixed with `yaml:` (parsed as typed YAML) or `string:` (stored as a literal string after the prefix). * `--env ` Add an `env:` entry. Repeatable. Values are always stored as strings — the `yaml:` and `string:` prefixes do not apply. * `--set ` Set any schema field by dotted path, e.g. `--set acceleration.enabled=yaml:true`. Repeatable, with the same value-prefix rules as `--param`. * `--depends-on ` Append an entry to `dependsOn:`. Repeatable. * `--enable` Set `enabled: true` * `--disable` Set `enabled: false` * `--file ` Read the dataset body from a YAML or JSON file * `--stdin` Read the dataset body from stdin * `--manifest ` Edit a non-default Spicepod file * `-h`, `--help` Print this help message ### Examples[​](#examples "Direct link to Examples") Add a Parquet dataset on S3: ``` >>> spice dataset add taxi_trips --from s3://my-bucket/trips.parquet --param file_format=parquet ``` Enable acceleration on an existing dataset: ``` >>> spice dataset configure taxi_trips --set acceleration.enabled=yaml:true ``` ### Interactive example[​](#interactive-example "Direct link to Interactive example") When running `spice dataset configure` with no name and no body flags, Spice will prompt for four inputs: 1. The name of the dataset, labelled by `(1)` below. 2. The description of the dataset, labelled by `(2)` below. 3. The source of the dataset, labelled by `(3)` below. Consult [Spice's supported data connectors](/docs/next/components/data-connectors) to see possible values for this field. Note: Spice may prompt for a file format if necessary, as shown in the example below. 4. Whether or not to enable acceleration for this dataset, labelled by `(4)`. The default value for this input is `y`, enabling acceleration for this dataset. Learn more about acceleration in the [dataset acceleration reference](/docs/next/components/data-accelerators). ``` > spice dataset configure dataset name: (spiceai) taxi-trips # (1) description: Taxi Trips in S3 # (2) from: s3://spiceai-demo-datasets/taxi_trips/2024/ # (3) file_format (parquet/csv) (parquet) parquet locally accelerate (y/n)? (y) y # (4) 2025/01/10 14:07:46 INFO Saved datasets/test/dataset.yaml ``` After execution, the directory structure looks like this for the above example: ``` ├── datasets │ ├── taxi-trips │ ├── dataset.yaml ├── spicepod.yaml └── ... ``` The datasets folder includes the datasets for your project configured by using `spice dataset configure` or added manually. The `dataset.yaml` file in `./datasets/taxi-trips` is configured as defined by the inputs provided to `spice dataset configure`. For this example, the `dataset.yaml` file looks as follows: ``` from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi-trips description: Taxi trips in s3 acceleration: enabled: true ``` The command additionally updates the root `spicepod.yaml` file to include the configured dataset as a reference (`ref`). For this example, `spicepod.yaml` would include the following: ``` version: v1 kind: Spicepod name: Taxi Trips with Spice datasets: - ref: datasets/taxi-trips ``` To learn more about Spice datasets and Spicepods, visit the [Spice dataset reference](/docs/next/reference/spicepod/datasets) and [Spicepod reference](/docs/next/reference/spicepod). --- # datasets Lists datasets loaded by the Spice runtime ### Usage[​](#usage "Direct link to Usage") ``` spice datasets [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `-o`, `--output ` Output format: `table` (default) or `json`. * `-h`, `--help` help for datasets ### Examples:[​](#examples "Direct link to Examples:") ``` >>> spice datasets NAME FROM REPLICATION ACCELERATION STATUS ERROR taxi_trips spice.ai/spiceai/quickstart/datasets/taxi_trips false false Ready tpch.customer spice.ai/spiceai/tpch/datasets/tpch.customer false false Ready tpch.lineitem spice.ai/spiceai/tpch/datasets/tpch.lineitem false false Ready tpch.nation spice.ai/spiceai/tpch/datasets/tpch.nation false false Ready tpch.orders spice.ai/spiceai/tpch/datasets/tpch.orders false false Ready tpch.part spice.ai/spiceai/tpch/datasets/tpch.part false false Ready tpch.partsupp spice.ai/spiceai/tpch/datasets/tpch.partsupp false false Ready tpch.region spice.ai/spiceai/tpch/datasets/tpch.region false false Ready tpch.supplier spice.ai/spiceai/tpch/datasets/tpch.supplier false false Ready ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice datasets --tls-root-certificate-file /path/to/cert.pem NAME FROM REPLICATION ACCELERATION STATUS ERROR taxi_trips spice.ai/spiceai/quickstart/datasets/taxi_trips false false Ready tpch.customer spice.ai/spiceai/tpch/datasets/tpch.customer false false Ready tpch.lineitem spice.ai/spiceai/tpch/datasets/tpch.lineitem false false Ready tpch.nation spice.ai/spiceai/tpch/datasets/tpch.nation false false Ready tpch.orders spice.ai/spiceai/tpch/datasets/tpch.orders false false Ready tpch.part spice.ai/spiceai/tpch/datasets/tpch.part false false Ready tpch.partsupp spice.ai/spiceai/tpch/datasets/tpch.partsupp false false Ready tpch.region spice.ai/spiceai/tpch/datasets/tpch.region false false Ready tpch.supplier spice.ai/spiceai/tpch/datasets/tpch.supplier false false Ready ``` --- # feedback Open the Spice.ai community Slack in the default browser to share feedback. ### Usage[​](#usage "Direct link to Usage") ``` spice feedback ``` #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message ### Example[​](#example "Direct link to Example") ``` > spice feedback Opening Spice.ai community Slack in your default browser: https://spice.ai/slack If the browser does not open, visit the URL above manually. ``` --- # init Initialize Spice app in the current working directory. ### Usage[​](#usage "Direct link to Usage") ``` spice init [app_name] ``` * `app_name`: The name of the app. If this is not provided, Spice will prompt for the name, defaulting to the name of the current working directory. #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message ### Examples[​](#examples "Direct link to Examples") ``` > spice init taxi-trips Initialized taxi-trips/spicepod.yaml Next steps: cd taxi-trips spice dataset configure # add a dataset interactively spice run # start the runtime Docs: https://spiceai.org/docs/ ``` The command creates a `spicepod.yaml` with basic initial metadata about the app, including a `yaml-language-server` schema directive for editor support (VS Code, Neovim, IntelliJ). For this example, the `spicepod.yaml` file is initialized to the following: ``` # File: ./taxi-trips/spicepod.yaml # yaml-language-server: $schema=https://raw.githubusercontent.com/spiceai/spiceai/trunk/.schema/spicepod.schema.json version: v2 kind: Spicepod name: taxi-trips ``` If no app name is provided, Spice initializes the Spicepod in the current working directory. For example: ``` > spice init name: (taxi-trips)? Initialized spicepod.yaml Next steps: spice dataset configure # add a dataset interactively spice run # start the runtime Docs: https://spiceai.org/docs/ ``` After execution, the current working directory contains the file `spicepod.yaml` with the same configuration as the previous example: ``` # File: ./spicepod.yaml # yaml-language-server: $schema=https://raw.githubusercontent.com/spiceai/spiceai/trunk/.schema/spicepod.schema.json version: v2 kind: Spicepod name: taxi-trips ``` --- # install Download and install the latest version of the Spice runtime. ### Usage[​](#usage "Direct link to Usage") ``` spice install [flavor] [flags] ``` #### flavor[​](#flavor "Direct link to flavor") * \`\` Install the core runtime that only includes data components * `cuda` Install the runtime with CUDA GPU acceleration Metal/CUDA acceleration is auto-detected by default. #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message * `-f`, `--force` Force installation of the latest released runtime ### Examples[​](#examples "Direct link to Examples") ``` spice install ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` spice install cuda ``` --- # login Login to the Spice.ai Platform, or other services with sub-commands. ### Usage[​](#usage "Direct link to Usage") ``` spice login [command] [flags] ``` ### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message * `-k`, `--key` string API key (for spice.ai) * `-o`, `--output` string Where to store the resulting credentials. One of `env` (default; appends to a local `.env` file), `json` (prints the credentials as JSON to stdout), or `keychain` (stores them in the platform keychain, e.g. macOS Keychain). note `--output` applies to `spice login` itself (the Spice.ai login, with or without `--key`). The provider subcommands below always write to `.env`. #### Available Commands[​](#available-commands "Direct link to Available Commands") * `abfs` Login to a Azure Storage Account * `databricks` Login to a Databricks instance * `delta-lake` Configure credentials to access a Delta Lake table * `dremio` Login to a Dremio instance * `postgres` Login to a Postgres instance * `s3` Login to an s3 storage * `sharepoint` Login to a Microsoft 365 sharepoint account * `snowflake` Login to a Snowflake warehouse * `spark` Login to a Spark Connect remote #### Examples[​](#examples "Direct link to Examples") ``` spice login ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` spice login --key ``` ### Browser Login Flow[​](#browser-login-flow "Direct link to Browser Login Flow") Running `spice login` without `--key` prints an auth code, opens the Spice.ai authorization page, and then polls the token exchange once per second while the browser flow is completed. * **The wait is bounded at 5 minutes.** If the authorization is not completed in that window, the command exits with `Authentication timed out. Please try again.` — followed by the last retryable error, when there was one. It does not poll indefinitely. * **A refusal stops immediately.** An explicitly denied authorization exits with `Access denied` without waiting out the deadline, as does a rejection the endpoint cannot answer differently on a retry (for example `400`, `401`, `403`, or `410`). * **Transient failures keep polling.** Server errors (`5xx`), `408`, `429`, and network failures are retried until the deadline. So is `404`, which is the normal answer while the auth code has not been authorized yet — but an endpoint that answers `404` for the whole 5 minutes (typically a `SPICE_BASE_URL` pointing at the wrong deployment) stops at the deadline and says so. `spice cloud login subscription` polls the same endpoint under the same 5-minute deadline. ### Credentials Stay on Their Origin[​](#credentials-stay-on-their-origin "Direct link to Credentials Stay on Their Origin") The auth code, device code, and access token in these flows are sent in the **request body**, which an HTTP `307` or `308` redirect replays verbatim to the redirect target. The CLI's credential-bearing HTTP clients therefore follow redirects only **within the same origin** (scheme, host, and port), and refuse any hop that leaves it rather than forwarding the credential to another host. This applies to the `spice login` flows, the Spice Cloud client, and the CLI's own runtime and Cloud Platform requests made with `--api-key`. An off-origin redirect surfaces as the `3xx` response itself rather than as a transport error, so a misconfigured endpoint stays diagnosable. ## `spice cloud login`[​](#spice-cloud-login "Direct link to spice-cloud-login") Authenticate with the Spice Cloud Platform. Running `spice cloud login` without a subcommand opens an interactive method chooser when stdin is a TTY. Non-interactive callers must specify a method explicitly. ### Methods[​](#methods "Direct link to Methods") #### `spice cloud login subscription`[​](#spice-cloud-login-subscription "Direct link to spice-cloud-login-subscription") Browser-based OAuth login flow. Automatically opens a browser for authentication. ``` spice cloud login subscription ``` Use `--device` to print the URL and one-time code without opening a browser (useful for SSH/headless environments): ``` spice cloud login subscription --device ``` #### `spice cloud login pat`[​](#spice-cloud-login-pat "Direct link to spice-cloud-login-pat") Authenticate with a personal access token. ``` spice cloud login pat --token ``` The token can also be provided via the `SPICE_CLOUD_PAT` environment variable: ``` export SPICE_CLOUD_PAT= spice cloud login pat ``` #### `spice cloud login api`[​](#spice-cloud-login-api "Direct link to spice-cloud-login-api") Authenticate using OAuth2 client credentials for CI/automation workflows. ``` spice cloud login api --client-id --client-secret ``` Credentials can also be provided via environment variables: ``` export SPICE_CLOUD_CLIENT_ID= export SPICE_CLOUD_CLIENT_SECRET= spice cloud login api ``` ### Environment Variables[​](#environment-variables "Direct link to Environment Variables") | Variable | Used by | Description | | --------------------------- | ----------- | --------------------- | | `SPICE_CLOUD_PAT` | `login pat` | Personal access token | | `SPICE_CLOUD_CLIENT_ID` | `login api` | OAuth2 client ID | | `SPICE_CLOUD_CLIENT_SECRET` | `login api` | OAuth2 client secret | --- # models Lists models loaded by the Spice runtime ### Usage[​](#usage "Direct link to Usage") ``` spice models [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `-o`, `--output ` Output format: `table` (default) or `json`. * `-h`, `--help` help for models ### Examples[​](#examples "Direct link to Examples") ``` >>> spice models ID OWNED_BY STATUS ERROR modlz local Ready ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice models --tls-root-certificate-file /path/to/cert.pem ID OWNED_BY STATUS ERROR modlz local Ready ``` --- # nsql Text-to-SQL REPL — translate natural language queries into SQL using a model loaded by the Spice runtime. The REPL sends the input to the runtime's `/v1/nsql` endpoint and prints the generated SQL along with the executed results. ## Requirements[​](#requirements "Direct link to Requirements") * Spice runtime must be running * At least one [model](/docs/next/components/models) defined in `spicepod.yaml` and ready ## Usage[​](#usage "Direct link to Usage") ``` spice nsql [flags] spice nsql [flags] [command] ``` ### Flags[​](#flags "Direct link to Flags") * `--model`, `-m ` Target model for Text-to-SQL generation. When omitted, the CLI uses the single ready model or prompts for a choice if several models are ready. * `-h`, `--help` Print usage information. ### Subcommands[​](#subcommands "Direct link to Subcommands") * [`analyze`](#analyze) Analyze Text-to-SQL performance by comparing the generated SQL against an expected SQL query. ## Examples[​](#examples "Direct link to Examples") Start an interactive Text-to-SQL session: ``` $ spice nsql Welcome to the Spice.ai NSQL REPL! Using model: openai Enter a query in natural language. nsql> show the top 5 longest taxi trips ``` Pass `--model` to select a specific model when more than one is ready: ``` spice nsql --model openai ``` Type `exit`, `quit`, `.exit`, or `.quit` — or press `Ctrl+D` — to leave the REPL. Inputs are saved to `nsql_history.txt` for recall with the up-arrow key. ## analyze[​](#analyze "Direct link to analyze") The `spice nsql analyze` subcommand evaluates Text-to-SQL quality by comparing a generated SQL query against an expected SQL query and reporting accuracy and performance metrics. ### Usage[​](#usage-1 "Direct link to Usage") ``` spice nsql analyze --query --expected --model ``` ### Flags[​](#flags-1 "Direct link to Flags") * `--query ` Natural language query to analyze. Required. * `--expected ` Expected SQL query to compare the generated SQL against. Required. * `--model`, `-m ` Model to use for Text-to-SQL. Required. ### Metrics[​](#metrics "Direct link to Metrics") Functional metrics (generated vs. expected SQL): * `exact_match` — `1.0` if the generated SQL exactly matches the expected SQL, `0.0` otherwise. * `correct_tables` — Intersection-over-Union (IoU) of tables referenced. * `correct_projections` — IoU of projected columns/expressions. * `correct_schema` — IoU of output schema fields. Performance metrics (read from `runtime.task_history` via the request's W3C trace ID): * `input_tokens` — Total prompt tokens used by LLM calls. * `output_tokens` — Total completion tokens generated by the LLM. * `latency_ms` — End-to-end latency of the nsql request. * `sql_duration_ms` — Total time spent executing SQL queries. * `llm_duration_ms` — Total time spent in LLM inference. * `sql_query_count` — Number of SQL queries executed. * `llm_count` — Number of LLM completion calls made. ### Example[​](#example "Direct link to Example") ``` spice nsql analyze \ --model openai \ --query "how many taxi trips are there?" \ --expected "SELECT COUNT(*) FROM taxi_trips" ``` --- # pods Lists Spicepods loaded by the Spice runtime ### Usage[​](#usage "Direct link to Usage") ``` spice pods [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `-o`, `--output ` Output format: `table` (default) or `json`. * `-h`, `--help` help for pods ### Examples[​](#examples "Direct link to Examples") ``` >>> spice pods NAME VERSION DATASETS MODELS DEPENDENCIES demo v2 2 1 0 another_pod v2 3 0 1 ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice pods --tls-root-certificate-file /path/to/cert.pem NAME VERSION DATASETS MODELS DEPENDENCIES demo v2 2 1 0 another_pod v2 3 0 1 ``` --- # query Submit an async query or start an interactive async query REPL against the Spice runtime's distributed query engine. ### Usage[​](#usage "Direct link to Usage") ``` spice query [flags] [SQL] ``` When invoked with a SQL statement, submits it as an async query and waits for the result. When invoked without arguments, starts an interactive REPL. #### Flags[​](#flags "Direct link to Flags") * `--no-wait` Submit the query and return immediately without waiting for results. * `--timeout ` Maximum client-side wait time (e.g., `30s`, `5m`). The query continues running on timeout. * `-o`, `--output ` Output format: `table` (default) or `json`. * `-h`, `--help` Print this help message. #### Subcommands[​](#subcommands "Direct link to Subcommands") | Subcommand | Description | | ------------------------------------------- | ---------------------------------- | | `spice query list [--status X] [--limit N]` | List queries | | `spice query status ` | Check query status | | `spice query results ` | Fetch results of a completed query | | `spice query cancel ` | Cancel a running query | ### Examples[​](#examples "Direct link to Examples") #### Submit and Wait[​](#submit-and-wait "Direct link to Submit and Wait") ``` $ spice query "SELECT * FROM orders WHERE total > 100 LIMIT 50;" ``` The CLI auto-polls with a spinner and displays results when ready. Press `Ctrl+C` to stop waiting — the query continues running in the background. #### Submit Without Waiting[​](#submit-without-waiting "Direct link to Submit Without Waiting") ``` $ spice query "SELECT * FROM large_table;" --no-wait ``` #### Interactive REPL[​](#interactive-repl "Direct link to Interactive REPL") ``` $ spice query query> SELECT COUNT(*) > FROM large_table > WHERE status = 'active'; Submitted query: 01ABC-DEF-456-7890AB (PENDING) Press Ctrl+C to stop waiting (query continues in background) ⠹ RUNNING (2.3s)... ✓ SUCCEEDED (5.1s) +----------+ | count(*) | +----------+ | 42000 | +----------+ Time: 5.10000000 seconds. 1 rows. ``` **REPL Commands**: | Command | Description | | ---------------------- | ---------------------------------------------- | | `.list` | List all queries tracked in this REPL session | | `.status ` | Show detailed status of a query | | `.results ` | Fetch and display results of a completed query | | `.wait ` | Resume waiting for a query to complete | | `.cancel ` | Cancel a running query | | `.clear` | Clear the local tracked queries list | | `.clear history` | Clear command history | | `.help` | Show available commands | | `.exit`, `.quit`, `.q` | Exit the REPL | Query IDs can be abbreviated if they uniquely identify a query within the tracked session. note The `spice query` command requires the runtime to be running in [cluster mode](/docs/features/distributed-query) with `--role scheduler` and `scheduler.state_location` configured. --- # refresh Refreshes an accelerated dataset loaded by the Spice runtime ### Usage[​](#usage "Direct link to Usage") ``` spice refresh [dataset] [flags] ``` `dataset` - an accelerated dataset name #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `--refresh-sql` SQL used to refresh the dataset, see [Refresh SQL docs](/docs/next/features/data-acceleration/data-refresh#refresh-sql). * `--refresh-mode` Refresh mode to use, see [Refresh Modes docs](/docs/next/features/data-acceleration/data-refresh#refresh-modes). * `--refresh-jitter-max` Maximum [jitter](/docs/next/features/data-acceleration/data-refresh#refresh-jitter) applied to the refresh, as a [duration](/docs/next/reference/duration) (e.g. `1m`). * `-o`, `--output ` Output format: `table` (default) or `json`. * `-h`, `--help` Print this help message ### Examples[​](#examples "Direct link to Examples") ``` >>> spice refresh taxi_trips --refresh-sql "SELECT * FROM taxi_trips WHERE trip_amount > 10.0" ``` Refreshing dataset taxi\_trips ... Dataset refresh triggered for taxi\_trips. ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice refresh taxi_trips --refresh-mode append ``` Refreshing dataset taxi\_trips with append mode... Dataset refresh triggered for taxi\_trips. ``` ``` --- # run Run Spice - starts the Spice runtime, installing if necessary. `spice run` is a wrapper around the [`spiced`](/docs/next/cli/reference/spiced) runtime binary. It installs `spiced` on first use, applies developer-friendly defaults, and forwards arguments after `--` to the runtime. To invoke the runtime directly (for containers, systemd units, or CI), see the [`spiced` reference](/docs/next/cli/reference/spiced). ### Usage[​](#usage "Direct link to Usage") ``` spice run [flags] spice run [flags] -- [spiced flags] ``` #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message. * `--endpoint` Configure the runtime endpoint. The URL scheme determines the endpoint type: `http://` or `https://` sets the HTTP endpoint, `grpc://` or `grpc+tls://` sets the Flight endpoint. Cannot be combined with `--http-endpoint` or `--flight-endpoint` for the same endpoint type. * `--flight-endpoint` Configure runtime Flight endpoint. Defaults to `http://127.0.0.1:50051`. * `--http-endpoint` Configure runtime HTTP endpoint. Defaults to `http://127.0.0.1:8090`. * `--metrics-endpoint` Configure the runtime Prometheus metrics endpoint (disabled by default). #### Spiced Flags[​](#spiced-flags "Direct link to Spiced Flags") Flags that are passed to the `spiced` runtime directly using `--`. * `--http` Configure runtime HTTP address \[default: 127.0.0.1:8090] * `--flight` Configure runtime Flight address \[default: 127.0.0.1:50051] * `--metrics` Enable and configure the Prometheus metrics endpoint (disabled by default) * `--tls-enabled` Enable TLS * `--tls-certificate` The TLS PEM-encoded certificate * `--tls-certificate-file` Path to the TLS PEM-encoded certificate file * `--tls-key` The TLS PEM-encoded key * `--tls-key-file` Path to the TLS PEM-encoded key file * `--telemetry-enabled` Enable or disable anonymous telemetry * `--pods-watcher-enabled` Enable the pods watcher (disabled by default) * `--repl` Start a SQL REPL against the runtime's Flight endpoint * `-v`, `--verbose` Enable verbose logging (use `-vv` for more detail) * `--very-verbose` Enable very verbose logging * `--set-runtime` Override [runtime configuration](/docs/next/reference/spicepod#runtime) with a name/value pair specified as `name=value`. Multiple overrides can be specified by using the flag multiple times. * `[PATH]` Positional argument specifying the path to a Spicepod directory or file. Supports local paths and `s3://` remote URLs. Directories must contain a `spicepod.yaml` (or `spicepod.yml`); file paths can use any name as long as the file declares `kind: Spicepod` and a recognized `version`. ### Examples[​](#examples "Direct link to Examples") #### `--set-runtime`[​](#--set-runtime "Direct link to --set-runtime") The `--set-runtime` flag overrides runtime configuration values. It can be specified multiple times to set multiple values. It is used like this: `--set-runtime name=value`. The Spicepod YAML equivalent of that is: ``` runtime: name: value ``` Examples: `--set-runtime task_history.captured_output=none`: ``` runtime: task_history: captured_output: none ``` `--set-runtime results_cache.enabled=false`: ``` runtime: results_cache: enabled: false ``` `--set-runtime runtime.tls.enabled=true --set-runtime runtime.tls.certificate_file=/path/to/cert.pem --set-runtime runtime.tls.key_file=/path/to/key.pem`: ``` runtime: tls: enabled: true certificate_file: /path/to/cert.pem key_file: /path/to/key.pem ``` #### `--endpoint`[​](#--endpoint "Direct link to --endpoint") ``` # Set the HTTP bind address using scheme-based routing spice run --endpoint http://0.0.0.0:8090 # Set the Flight bind address using scheme-based routing spice run --endpoint grpc://0.0.0.0:50051 ``` #### No arguments[​](#no-arguments "Direct link to No arguments") ``` spice run ``` #### `--captured-outputs none`[​](#--captured-outputs-none "Direct link to --captured-outputs-none") ``` # Set task history captured outputs to none via --set-runtime spice run -- --set-runtime task_history.captured_output=none ``` #### `--http`[​](#--http "Direct link to --http") ``` # Expose the HTTP server on all interfaces spice run -- --http 0.0.0.0:8090 ``` #### `--flight`[​](#--flight "Direct link to --flight") ``` # Expose the HTTP & Flight servers on all interfaces with TLS spice run -- --http 0.0.0.0:8090 --flight 0.0.0.0:50051 --tls-enabled true --tls-certificate-file /path/to/cert.pem --tls-key-file /path/to/key.pem ``` --- # search Performs embeddings-based searches across search-configured datasets. ### Usage[​](#usage "Direct link to Usage") ``` spice search [query] [flags] ``` `query` - a search query #### Flags[​](#flags "Direct link to Flags") * `--limit`, `-l` Limit number of search results. Default: `10`. * `--cache-control ` Control whether the results cache is used for searches. Values: `cache` (default), `no-cache`. * `--model ` Model to use for search. * `--endpoint ` Specifies the remote Spice instance HTTP endpoint (e.g., `http://localhost:8090`). * `--headers ` Custom HTTP headers in format `Key:Value` (can be specified multiple times). * `-o`, `--output ` Output format: `table` (default) or `json`. ### Examples[​](#examples "Direct link to Examples") ``` >>> spice search --limit 2 ``` #### Remote Example[​](#remote-example "Direct link to Remote Example") ``` # Search with a remote spiced instance spice search --endpoint http://my-remote-host:8090 ``` ``` search> artificial intelligence Rank 1, Score: 20.6, Datasets [pdf] Undergraduate Texts in Mathematics Editors: F. W. Gehring P. R. Halmos · Advisory Board: C. DePrima I. Herstein J. Kiefer W. LeVeque Kai Lai Chung Elementary Probability Theory with Stochastic Processes Springer Science+Business Media, LLC ... Rank 2, Score: 17.8, Datasets [pdf] Forecasting at Scale Sean J. Taylor y Facebook, Menlo Park, California, United States sjt@fb.com and Benjamin Letham y Facebook, Menlo Park, California, United States bletham@fb.com Abstract Forecasting is a common data science... ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice search --model gpt-3 --limit 1 ``` ``` search> machine learning Rank 1, Score: 25.4, Datasets [pdf] Machine Learning Yearning by Andrew Ng Machine Learning Yearning is a technical book by Andrew Ng that provides practical advice on how to structure machine learning projects. ... ``` --- # spiced `spiced` is the Spice.ai runtime binary. It hosts the HTTP and Flight servers, loads a Spicepod, and serves queries. Most users invoke `spiced` indirectly via [`spice run`](/docs/next/cli/reference/run), which installs the runtime, applies developer-friendly defaults, and forwards CLI arguments. This page documents `spiced` itself for operators who run the binary directly (for example in containers, systemd units, or CI). ### Usage[​](#usage "Direct link to Usage") ``` spiced [flags] [SPICEPOD_PATH] ``` If `SPICEPOD_PATH` is omitted, `spiced` uses the current working directory. When the path resolves to a directory, the runtime loads `spicepod.yaml` (or `spicepod.yml`) from inside it. When the path resolves to a file, any YAML filename is accepted (e.g. `spiced ./configs/my-app.yaml`); the file must declare `kind: Spicepod` and a recognized `version` — otherwise `spiced` exits with an explicit validation error. ### Runtime flags[​](#runtime-flags "Direct link to Runtime flags") * `--http ` — HTTP server bind address. Default: `127.0.0.1:8090`. * `--flight ` — Arrow Flight server bind address. Default: `127.0.0.1:50051`. * `--metrics ` — Enable the Prometheus scrape endpoint on the given address. Disabled by default. * `--pods-watcher-enabled` — Watch the Spicepod directory for changes and hot-reload. Disabled by default. * `--telemetry-enabled ` — Override anonymous telemetry collection. When unset, the value is taken from `runtime.telemetry.enabled` in the Spicepod. * `--version` — Print the runtime version and exit. * `-v`, `--verbose` — Increase log verbosity. Repeat (`-vv`) for debug output. * `--very-verbose` — Equivalent to `-vv`. * `--set-runtime ` — Override a runtime configuration value. Can be repeated; see [examples in `spice run`](/docs/next/cli/reference/run#--set-runtime). ### TLS flags[​](#tls-flags "Direct link to TLS flags") These configure TLS for the HTTP and Flight servers. * `--tls-enabled` — Enable TLS. Default: `false`. * `--tls-certificate ` — TLS certificate as an inline PEM string. * `--tls-certificate-file ` — Path to a PEM-encoded certificate file. * `--tls-key ` — TLS private key as an inline PEM string. * `--tls-key-file ` — Path to a PEM-encoded private key file. ### Cluster flags[​](#cluster-flags "Direct link to Cluster flags") Used when running `spiced` as part of a [distributed cluster](/docs/next/features/distributed-query). Omit these for standalone operation. * `--role ` — Explicit cluster role. If omitted but `--scheduler-address` is set, the role defaults to `executor`. * `--scheduler-address ` — URL of the scheduler service. Required on executors. * `--node-bind-address ` — Bind address for the internal cluster gRPC service. Default: `0.0.0.0:50052`. * `--node-advertise-address ` — Hostname or IP this node advertises to the rest of the cluster. * `--node-mtls-ca-certificate-file ` — CA certificate used to validate peer node identities. * `--node-mtls-certificate-file ` — Certificate file used for both server TLS and client mTLS on the node port. * `--node-mtls-key-file ` — Private key file for the node certificate. * `--allow-insecure-connections` — Allow cluster communication without mTLS. Use only in development or testing environments. ### Cloud Connect flags[​](#cloud-connect-flags "Direct link to Cloud Connect flags") * `--cloud-connect` — Connect this runtime to Spice Cloud for remote management (Cloud Connect). Default: `false`. Requires an enrolled identity or a staged adoption code — see [`spice connect`](/docs/next/cli/reference/connect). When the flag is omitted the client still activates if such adoption state exists, so instances enrolled before the flag existed keep connecting across an upgrade; a `spiced` with no adoption state never connects to the cloud. ### SQL REPL flags[​](#sql-repl-flags "Direct link to SQL REPL flags") * `--repl` — Start a SQL REPL against the runtime's Flight endpoint instead of serving requests. * `--repl-flight-endpoint ` — Flight endpoint the REPL connects to. Default: `http://localhost:50051`. * `--http-endpoint ` — HTTP endpoint the REPL connects to. Default: `http://localhost:8090`. * `--tls-root-certificate-file ` — Root certificate used to verify the runtime's TLS certificate. * `--client-tls-certificate-file ` — Client certificate for mTLS authentication. Must be used with `--client-tls-key-file`. ### Differences from `spice run`[​](#differences-from-spice-run "Direct link to differences-from-spice-run") `spice run` is a thin launcher around `spiced`. When it spawns the runtime, it adds a few flags automatically that `spiced` does **not** apply when invoked directly: | Behavior | `spice run` | `spiced` | | ---------------------------- | -------------------------------------------------------- | ---------------------------------------------------- | | Runtime installation | Auto-installs `spiced` if missing | Must already be installed | | Pods watcher | Enabled (`--pods-watcher-enabled`) | Disabled by default | | Task history captured output | Forced to `truncated` via `--set-runtime` | Uses the Spicepod value (default: `full`) | | `--endpoint` scheme routing | Supported — `http://…` sets HTTP, `grpc://…` sets Flight | Not supported — use `--http` and `--flight` directly | Operators running `spiced` directly who want `spice run`-equivalent behavior should pass `--pods-watcher-enabled` and, if desired, `--set-runtime task_history.captured_output=truncated`. ### Examples[​](#examples "Direct link to Examples") #### Start the runtime with a Spicepod in the current directory[​](#start-the-runtime-with-a-spicepod-in-the-current-directory "Direct link to Start the runtime with a Spicepod in the current directory") ``` spiced ``` #### Bind the HTTP and Flight servers to all interfaces[​](#bind-the-http-and-flight-servers-to-all-interfaces "Direct link to Bind the HTTP and Flight servers to all interfaces") ``` spiced --http 0.0.0.0:8090 --flight 0.0.0.0:50051 ``` #### Enable Prometheus metrics[​](#enable-prometheus-metrics "Direct link to Enable Prometheus metrics") ``` spiced --metrics 0.0.0.0:9090 ``` #### Enable TLS[​](#enable-tls "Direct link to Enable TLS") ``` spiced --tls-enabled \ --tls-certificate-file /path/to/cert.pem \ --tls-key-file /path/to/key.pem ``` #### Run as a cluster executor[​](#run-as-a-cluster-executor "Direct link to Run as a cluster executor") ``` spiced --role executor \ --scheduler-address https://scheduler.example.internal:50052 \ --node-advertise-address executor-1.example.internal \ --node-mtls-ca-certificate-file /etc/spice/ca.pem \ --node-mtls-certificate-file /etc/spice/node.pem \ --node-mtls-key-file /etc/spice/node.key ``` #### Override runtime configuration at launch[​](#override-runtime-configuration-at-launch "Direct link to Override runtime configuration at launch") ``` spiced --set-runtime task_history.captured_output=none \ --set-runtime results_cache.enabled=false ``` --- # sql Start an interactive SQL query session against the Spice runtime ### Usage[​](#usage "Direct link to Usage") ``` spice sql [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--endpoint ` Specifies the remote Spice instance endpoint. Supports `http://`, `https://`, `grpc://`, or `grpc+tls://` schemes. If not provided, uses the local spiced runtime. * `--flight-endpoint ` (Deprecated) Specifies the remote Spice instance Flight endpoint (treated as gRPC endpoint). If not provided, uses the local spiced runtime. * `--cache-control ` Control whether the results cache is used for queries. Default: `cache`. * `--tls-root-certificate-file ` The path to the root certificate file used to verify the Spice.ai runtime server certificate. * `--client-tls-certificate-file ` The path to the client certificate file for mTLS authentication. Required when connecting to a cluster node that enforces mutual TLS. Must be used together with `--client-tls-key-file`. * `--client-tls-key-file ` The path to the client private key file for mTLS authentication. Must be used together with `--client-tls-certificate-file`. * `--headers ` Custom HTTP headers in format `Key:Value` (can be specified multiple times). * `-x`, `--expanded` Start the REPL in expanded view, rendering each column on its own line per record. Useful for wide tables. Can be toggled at runtime with the `.expanded` meta-command. * `-o`, `--output ` Output format: `table` (default) or `json`. * `-h`, `--help` Print this help message. #### REPL Meta-commands[​](#repl-meta-commands "Direct link to REPL Meta-commands") Inside the REPL, the following meta-commands are available: * `.expanded` Toggle expanded (column-per-line) display, similar to PostgreSQL's `\x`. Pass `.expanded on` or `.expanded off` to set the mode explicitly. * `help` Print the list of available commands. ### Examples[​](#examples "Direct link to Examples") ``` $ spice sql Welcome to the Spice.ai SQL REPL! Type 'help' for help. show tables; -- list available tables sql> show tables +---------------+--------------------+---------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------------+---------------+------------+ | datafusion | public | tmp_view_test | VIEW | | datafusion | information_schema | tables | VIEW | | datafusion | information_schema | views | VIEW | | datafusion | information_schema | columns | VIEW | | datafusion | information_schema | df_settings | VIEW | +---------------+--------------------+---------------+------------+ ``` ### Additional Examples[​](#additional-examples "Direct link to Additional Examples") ``` $ spice sql --tls-root-certificate-file /path/to/cert.pem Welcome to the Spice.ai SQL REPL! Type 'help' for help. ``` #### mTLS Client Authentication[​](#mtls-client-authentication "Direct link to mTLS Client Authentication") ``` $ spice sql --tls-root-certificate-file /path/to/ca.pem \ --client-tls-certificate-file /path/to/client-cert.pem \ --client-tls-key-file /path/to/client-key.pem Welcome to the Spice.ai SQL REPL! Type 'help' for help. ``` #### Expanded view for wide tables[​](#expanded-view-for-wide-tables "Direct link to Expanded view for wide tables") Use `--expanded` (or `-x`) to start the REPL with each column on its own line, which is easier to read when tables are wider than the terminal. The mode can also be toggled at runtime with `.expanded`: ``` $ spice sql --expanded Welcome to the Spice.ai SQL REPL! Type 'help' for help. sql> SELECT * FROM eth.recent_blocks LIMIT 1; -[ RECORD 1 ]----------+---------------------------------------- number | 18000000 hash | 0x9c2f1e8... timestamp | 2023-08-14T00:00:00Z gas_used | 12500000 ``` #### Remote and Cloud Examples[​](#remote-and-cloud-examples "Direct link to Remote and Cloud Examples") ``` # Connect to Spice Cloud spice sql --cloud --api-key # Connect to a remote spiced instance over HTTP spice sql --endpoint http://my-remote-host:8090 # Connect to a remote spiced instance over Arrow Flight SQL (gRPC) spice sql --endpoint grpc://my-remote-host:50051 show tables; -- list available tables sql> show tables +---------------+--------------------+---------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------------+---------------+------------+ | datafusion | public | tmp_view_test | VIEW | | datafusion | information_schema | tables | VIEW | | datafusion | information_schema | views | VIEW | | datafusion | information_schema | columns | VIEW | | datafusion | information_schema | df_settings | VIEW | +---------------+--------------------+---------------+------------+ ``` --- # status Spice runtime status ### Usage[​](#usage "Direct link to Usage") ``` spice status [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `-o`, `--output ` Output format: `table` (default) or `json`. * `-h`, `--help` help for status ### Examples[​](#examples "Direct link to Examples") ``` >>> spice status NAME ENDPOINT STATUS http 127.0.0.1:8090 Ready flight 127.0.0.1:50051 Ready metrics N/A Disabled ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice status --tls-root-certificate-file /path/to/cert.pem NAME ENDPOINT STATUS http 127.0.0.1:8090 Ready flight 127.0.0.1:50051 Ready metrics N/A Disabled ``` --- # trace Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the `runtime.task_history` table. ### Usage[​](#usage "Direct link to Usage") ``` spice trace [task] [flags] ``` `task` - The name of the task whose trace is requested. Supported tasks include: * `acceleration_refresh` * `ai` * `ai_chat` * `ai_completion` * `eval_run` * `sql_query` * `nsql` * `text_embed` * `tool_use::search` * `tool_use::list_datasets` * `tool_use::sql` * `tool_use::table_schema` * `tool_use::sample_data` * `tool_use::load_memory` * `tool_use::store_memory` * `search` * `scheduled_worker` Tools proxied from MCP servers and custom tools are recorded dynamically as `tool_use::` or `tool_use::/` (for example, `tool_use::github/search_code`). Any such `tool_use::`-prefixed task name can be traced, not just the built-in tools listed above. These tasks are from the `task` column in the Spice SQL `runtime.task_history` table. #### Flags[​](#flags "Direct link to Flags") * `--trace-id` Retrieve the trace with the given trace ID (the column `trace_id` from `runtime.task_history`). * `--id` Retrieve the trace with the given `id` label (i.e. the task has a valid `id` within the `labels` column of `runtime.task_history`). * `--api-key` Specify the API key for authentication. * `--include-output`: Include, as an additional column, the captured output to each span (i.e. the `captured_output` column from `runtime.task_history`). Note: If captured outputs are not being stored, this will return an empty row. * `--include-input`: Include, as an additional column, the input to each span (i.e. the `input` column from `runtime.task_history`). * `--truncate []` Truncate the `input` and `captured_output` columns to the given number of characters. Defaults to `80` characters when the flag is passed without a value; when omitted, the columns are not truncated. * `-o`, `--output ` Output format: `table` (default) or `json`. The latest trace for the task will be used if neither `--trace-id` nor `--id` is specified. ### Examples[​](#examples "Direct link to Examples") #### Retrieve the trace for the last text-to-SQL operation[​](#retrieve-the-trace-for-the-last-text-to-sql-operation "Direct link to Retrieve the trace for the last text-to-SQL operation") ``` spice trace nsql ``` #### Retrieve the trace for a specific task by ID[​](#retrieve-the-trace-for-a-specific-task-by-id "Direct link to Retrieve the trace for a specific task by ID") ``` spice trace ai_chat --id chatcmpl-At6ZmDE8iAYRPeuQLA0FLlWxGKNnM ``` #### Retrieve a trace by `trace-id`[​](#retrieve-a-trace-by-trace-id "Direct link to retrieve-a-trace-by-trace-id") ``` spice trace sql_query --trace-id d5c6f1eed9f27257 ``` ### Output Example[​](#output-example "Direct link to Output Example") ``` TREE STATUS DURATION TASK a97f52ccd7687e64 ✅ 673.14ms ai_chat ├── 4eebde7b04321803 ✅ 0.04ms tool_use::list_datasets └── 4c9049e1bf1c3500 ✅ 671.91ms ai_completion ``` This output represents a structured trace of executed tasks. ### Output Example (with `--include-output`)[​](#output-example-with---include-output "Direct link to output-example-with---include-output") ``` TREE STATUS DURATION TASK OUTPUT a97f52ccd7687e64 ✅ 673.14ms ai_chat The capital of New York is Albany. ├── 4eebde7b04321803 ✅ 0.04ms tool_use::list_datasets [] └── 4c9049e1bf1c3500 ✅ 671.91ms ai_completion [{"content":"The capital of New York is Albany.","refusal":null,"tool_calls":null,"role":"assistant","function_call":null,"audio":null}] ``` --- # upgrade Upgrades the Spice CLI and runtime to the latest or specified version ### Usage[​](#usage "Direct link to Usage") ``` spice upgrade [target_version] [flags] ``` `target_version` - an optional release version to install, including the leading `v` (e.g. `v1.8.3`). A version without the leading `v` is rejected. When omitted, the latest release is used. #### Flags[​](#flags "Direct link to Flags") * `-f`, `--force` Reinstall the CLI and runtime even when the target version is already installed * `-h`, `--help` help for upgrade ### Examples[​](#examples "Direct link to Examples") ``` spice upgrade ``` ### Additional Examples[​](#additional-examples "Direct link to Additional Examples") Upgrade to a specific version: ``` spice upgrade v1.8.3 ``` Reinstall the currently installed version: ``` spice upgrade --force ``` --- # validate Validate a `spicepod.yaml` without starting the runtime. Checks YAML syntax and schema, component references, duplicate component names, reserved keywords, and nested pod includes (`dependsOn`). ### Usage[​](#usage "Direct link to Usage") ``` spice validate [path] ``` * `path`: Path to a Spicepod YAML file or a directory containing one. Defaults to the current directory (`.`). Directories must contain a `spicepod.yaml` (or `spicepod.yml`); file paths accept any name as long as the file declares `kind: Spicepod` and a recognized `version`. #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message ### Examples[​](#examples "Direct link to Examples") Validate the spicepod in the current directory: ``` > spice validate OK my_app (datasets: 2, models: 1, views: 0, tools: 0, workers: 0) ``` Validate a specific directory: ``` > spice validate ./my-app OK my_app (datasets: 3, models: 0, views: 1, tools: 2, workers: 0) ``` Validate a specific file: ``` > spice validate path/to/spicepod.yaml OK taxi_trips (datasets: 1, models: 0, views: 0, tools: 0, workers: 0) ``` When validation fails, the error describes the issue: ``` > spice validate ./broken-app Invalid: spicepod validation failed: missing field `kind` ``` --- # version Outputs the current version of the Spice CLI and runtime ### Usage[​](#usage "Direct link to Usage") ``` spice version [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--cli-only` Show only the CLI version, skipping the runtime version lookup * `-o`, `--output ` Output format: `table` (default) or `json`. * `-h`, `--help` help for version ### Sample output[​](#sample-output "Direct link to Sample output") **Upgrade available**: ``` > spice version CLI version: v1.0.6 Runtime version: v1.0.6+models CLI version v1.1.0 is now available! To upgrade, run "spice upgrade". ``` Learn more about upgrading the Spice CLI and runtime using `spice upgrade` [here.](/docs/next/cli/reference/upgrade) **Latest Version**: ``` > spice version CLI version: v1.1.0 Runtime version: v1.1.0+models ``` --- # Configuring Trace Levels Trace output verbosity is determined by the following sources, listed in order of precedence: 1. Verbosity flags (`-v`/`--verbose`, `-vv`/`--very-verbose`). If these flags are provided, they override all other settings. 2. The `SPICED_LOG` environment variable. This is used only if verbosity flags are not set 3. The `runtime.output_level` YAML configuration file. This is used only if neither verbosity flags nor the environment variable are set. ### Default[​](#default "Direct link to Default") The default trace level is `INFO`, suitable for general information about the system. ``` SPICED_LOG="task_history=INFO,spiced=INFO,runtime=INFO,secrets=INFO,data_components=INFO,cache=INFO,extensions=INFO,spice_cloud=INFO,llms=INFO,reqwest_retry::middleware=off,WARN" ``` The equivalent `runtime.output_level` configuration is `info`: ``` runtime: output_level: info ``` ### Enabling Debug Mode[​](#enabling-debug-mode "Direct link to Enabling Debug Mode") Use the `-v`/`--verbose` CLI flags to enable detailed logs, useful for debugging. ``` spice run -v spiced -v ``` Alternatively you can use `runtime.output_level` yaml configuration: ``` runtime: output_level: verbose ``` This sets `SPICED_LOG` to `DEBUG` level: ``` SPICED_LOG="task_history=DEBUG,spiced=DEBUG,runtime=DEBUG,secrets=DEBUG,data_components=DEBUG,cache=DEBUG,extensions=DEBUG,spice_cloud=DEBUG,llms=DEBUG,DEBUG" spice run ``` ### Enabling Trace Mode[​](#enabling-trace-mode "Direct link to Enabling Trace Mode") Use the `-vv`/`--very-verbose` CLI flag to enable the most detailed logs, typically for in-depth troubleshooting. ``` spice run -vv spiced -vv ``` Alternatively you can use `runtime.output_level` yaml configuration: ``` runtime: output_level: very_verbose ``` This sets `SPICED_LOG` to `TRACE` level: ``` SPICED_LOG="task_history=TRACE,spiced=TRACE,runtime=TRACE,secrets=TRACE,data_components=TRACE,cache=TRACE,extensions=TRACE,spice_cloud=TRACE,llms=TRACE,TRACE" spice run ``` ### Granular Configuration[​](#granular-configuration "Direct link to Granular Configuration") For specific component trace configuration, adjust the trace levels as needed: ``` SPICED_LOG="spiced=INFO,runtime=DEBUG,data_components=WARN,cache=WARN" spice run ``` --- # Clients and Tools Spice supports a variety of clients and tools for querying data, including SQL clients, BI tools, and programmatic interfaces. Use these integrations to connect applications and dashboards to the Spice runtime. ### Connection Protocols[​](#connection-protocols "Direct link to Connection Protocols") Most SQL clients connect to Spice using one of the following protocols: | Protocol | Endpoint | Best For | | ---------------- | ------------------------ | ------------------------------------------------- | | Arrow Flight SQL | `grpc://localhost:50051` | High-performance data transfer, JDBC/ODBC drivers | | HTTP | `http://localhost:8090` | REST API calls, curl, web applications | | ADBC | Flight SQL driver | Python and R data science workflows | For programmatic access, see [SDKs](/docs/next/sdks). ## [📄️DBeaver](/docs/next/clients/dbeaver) [Configure DBeaver to query Spice via JDBC](/docs/next/clients/dbeaver) ## [📄️JetBrains DataGrip](/docs/next/clients/jetbrains-datagrip) [Configure JetBrains Datagrip to query Spice via JDBC](/docs/next/clients/jetbrains-datagrip) ## [📄️Apache Superset](/docs/next/clients/superset) [Use Apache Superset to query and visualize datasets loaded in Spice.](/docs/next/clients/superset) ## [📄️Tableau](/docs/next/clients/tableau) [Use Tableau to to access, visualise and analyse datasets loaded in Spice.](/docs/next/clients/tableau) ## [📄️Microsoft Power BI](/docs/next/clients/powerbi) [Use Microsoft Power BI to access, visualize and analyze Spice datasets.](/docs/next/clients/powerbi) --- # DBeaver 1. Start the Spice runtime with a dataset loaded. Follow the [quickstart guide](/docs/next/getting-started) to get started. 2. Download [DBeaver Community Edition](https://dbeaver.io). 3. Download the [Apache Arrow Flight SQL JDBC driver](https://search.maven.org/search?q=a:flight-sql-jdbc-driver) - choose the "jar" option. 4. Launch DBeaver 5. In the DBeaver application menu bar, open the "Database" menu and choose: "Driver Manager": ![Driver manager menu option](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/691d1f83-c1d0-4ad8-ec8d-d8f37ccc9d00/public "Driver manager menu option") 6. Click the "New" button on the right: ![Driver manager new button](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/5783d944-daae-4735-99e9-976f974bc100/public "Driver manager new button") 7. Add the JDBC jar file: 1. Click the "Libraries" tab 2. Click the: "Add File" button 3. Choose the "flight-sql-jdbc-driver-19.0.0.jar" jar file (the file downloaded in step 3 above) - and click "Open" ![Select jar file](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/19900f7a-f00f-473d-780e-4a28c2ecd800/public "Select jar file") 4. Close the Driver editor window with the blue "OK" button on the lower-right 8. Enter the driver settings: 1. Click the "Settings" tab 2. In the "Driver Name" field - enter: `Apache Arrow Flight SQL` 3. In the "URL Template" field - enter: `jdbc:arrow-flight-sql://{host}:{port}?useEncryption=false&disableCertificateVerification=true` * If [API key authentication](/docs/next/api/auth) is enabled, the URL template should be: `jdbc:arrow-flight-sql://{host}:{port}?useEncryption=false&disableCertificateVerification=true&user=&password=` - where `` is the API key value 1. In the "Driver Type" drop-down box - choose: "SQLite" 2. Select "No authentication" * This should be selected even if API key authentication is enabled in the runtime, as the API key is supplied via the URL template above. 1. The driver manager "Edit Driver" window should look like this: ![Driver Manager completed](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/20348c42-117b-4763-80d2-6e615b23ae00/public "Driver Manager completed") 2. Click the blue "OK" button on the lower-right to save the driver 3. Close the "Driver Manager" window by clicking the blue "Close" button on the lower-right. 9. Create a new Database Connection: 1. In the DBeaver application menu bar, open the "Database" menu and choose: "New Database Connection": ![New Database Connection](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/acdf7251-4238-44ee-9639-0c557518da00/public "New Database Connection") 2. In the "Connect to a database" window - type: `Flight` in the search bar 3. Choose the `Apache Arrow Flight SQL` driver - the window should look like this: ![Connect to a database window](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/61cee5fe-dc75-4ac1-e558-eea3aff4c100/public "Connect to a database window") 4. Click the blue "Next >" button on the bottom of the window 5. On the next screen, the JDBC URL should be filled out already - just supply the Host (`localhost`) and Port (`50051`) values for the Spice runtime. The window should look like this: ![Connect to a database window 2](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/2a2b2fdc-00db-49d3-5359-059b12342b00/public "Connect to a database window 2") 6. Click the "Test Connection" button - the window should look like this: ![Test Connection results](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/a3fc5f5f-a39f-47ce-7955-4b384ec1ae00/public "Test Connection results") 7. Click the blue "OK" button to close the Connection test window 8. Click the "Connection details (name, type, ...)" button on the right 9. In the "General" section, enter: `Spice Runtime` for the "Connection name". It should look like this: ![Name the Database Connection](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/f6d04fe1-92a1-4082-d4ea-e9daacaca200/public) 10. Click the blue "Finish" button to save the connection 10. Run a query: 1. Right-click on the Database Connection on the left - choose: "SQL Editor", and then: "Open SQL Console" as shown here: ![Open SQL Console](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/642a5885-9e3f-4dd7-ef43-72bfce27bb00/public "Open SQL Console") 2. In the Console window - run a query - something like: `SELECT * FROM taxi_trips;` 3. Click the triangle button to execute the SQL statement - as shown below (or use keyboard shortcut: Ctrl+Enter): ![Execute SQL](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/2134e47b-a066-47e9-1d48-06352675f400/public "Execute SQL") 4. See the query results as shown in this screenshot: ![Query Results](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/0e9f3c0f-2e03-47f9-8d5e-65e078d7e900/public "Query Results") 5. DBeaver is now configured to query the Spice runtime using SQL! 🎉 --- # JetBrains DataGrip 1. Start the Spice runtime with a dataset loaded. Follow the [quickstart guide](/docs/next/getting-started) to get started. 2. Download [JetBrains DataGrip](https://www.jetbrains.com/datagrip). 3. Download the [Apache Arrow Flight SQL JDBC driver](https://search.maven.org/search?q=a:flight-sql-jdbc-driver) - Select "Versions", tab click "Browse" on most recent version and then download `flight-sql-jdbc-driver-.jar`. 4. Launch DataGrip 5. In Database Explorer menu, select "+" and choose "Driver" ![Data Sources and Drivers menu option](/assets/images/datagrip-1-12175a5760ca0dfde19f28d8253f2263.png "Data Sources and Drivers menu option") 6. Add the JSBC jar file: 1. Click the "+" button in "Driver Files" selection 2. Click the "Custom JARs" button 3. Choose the `flight-sql-jdbc-driver-.jar` jar file (the file downloaded in step 3 above) - and click "Open" 4. Click the "Class:" selector 5. Select `org.apache.arrow.driver.jdbc.ArrowFlightJdbcDriver` ![Driver Class selector](/assets/images/datagrip-3-62e98abc87e0f4705a08f2921bf2dbbc.png "Driver Class selector") 7. Enter the driver settings: 1. In the "Name" field - enter: `Apache Arrow Flight SQL` 2. Add "URL Template" Default: `jdbc:arrow-flight-sql://{host}:{port}\?useEncryption=false&disableCertificateVerification=true` 3. Click "Ok" ![Driver creation window](/assets/images/datagrip-4-6082661b869e86a8d39e08c804546735.png "Driver creation window") 8. Create a new Database Connection: 1. In Database Explorer menu, select "+", choose "Data Source" > "Arrow Flight JDBC" 2. Set the host to `localhost` and the port to `50051` 3. In "Authentication" select "No auth" 4. Click "Test Connection" to verify ![New Data Source](/assets/images/datagrip-5-08d8b998640048850709308946762442.png "New Data Source") 10. Run a query: 1. Right-click on the connection in Database Explorer and choose "New" > "Query Console" ![Create new Query Console](/assets/images/datagrip-6-7ca8eafad45c1890065deb2048c3e625.png "Create new Query Console") 2. In the Console window - add a query - something like: `SELECT * FROM taxi_trips;` and click the triangle button to execute the SQL statement 3. See the query results: ![Query Results](/assets/images/datagrip-7-81710c69e680dd7e49aad30e951aa2f0.png "Query Results") DataGrip is now configured to query the Spice runtime using SQL! 🎉 --- # Microsoft Power BI Connector Use the instructions below to get started with the **[Spice.ai Power BI Connector](https://github.com/spiceai/powerbi-connector)**—an [ADBC](https://github.com/apache/arrow-adbc)-based connector that enables [Microsoft Power BI](https://www.microsoft.com/en-us/power-platform/products/power-bi) users to easily connect to and visualize data loaded in [Spice.ai Enterprise](https://spiceai.org/) and [Spice Cloud Platform](https://spice.ai/) instances. ## Manual Connector Installation[​](#manual-connector-installation "Direct link to Manual Connector Installation") ### Power BI Desktop[​](#power-bi-desktop "Direct link to Power BI Desktop") 1. Download the latest `spice_adbc.mez` file from the [releases page](https://github.com/spiceai/powerbi-connector/releases) 2. Copy to your Power BI `Custom Connectors` directory: `C:\Users\[USERNAME]\Documents\Microsoft Power BI Desktop\Custom Connectors` ``` Invoke-WebRequest -Uri "https://github.com/spiceai/powerbi-connector/releases/latest/download/spice_adbc.mez" -OutFile "C:\Users\[USERNAME]\Documents\Microsoft Power BI Desktop\Custom Connectors\spice_adbc.mez" ``` 3. [Enable Uncertified Connectors](https://learn.microsoft.com/en-us/power-bi/connect-data/desktop-connector-extensibility#custom-connectors) in Power BI Desktop settings and restart Power BI Desktop. ## Adding Spice as a Data Source[​](#adding-spice-as-a-data-source "Direct link to Adding Spice as a Data Source") 1. Open Power BI Desktop. 2. Click on `Get Data` → `More...`. 3. In the dialog, select `Spice.ai` connector. ![Spice.ai connector](/img/powerbi/powerbi-spice-connector.png) 4. Click `Connect`. 5. Enter the **ADBC (Arrow Flight SQL) Endpoint**: * For Spice Cloud Platform:
`grpc+tls://flight.spiceai.io:443`
*(Use the region-specific address if applicable.)* * For on-premises/self-hosted Spice.ai: * Without TLS (default): `grpc://:50051` * With TLS: `grpc+tls://:50051` ![Spice.ai Connection Dialog](/img/powerbi/powerbi-spice-connection-dlg.png) 6. Select the `Data Connectivity` mode: * **Import**: Data is loaded into Power BI, enabling extensive functionality but requiring periodic refreshes and sufficient local memory to accommodate the dataset. * **DirectQuery**: Queries are executed directly against Spice in real-time, providing fast performance even on large datasets by leveraging Spice's optimized query engine. 7. Click `OK`. 8. Select `Authentication` option: * **Anonymous**: Select for unauthenticated on-premises deployments. * **API Key**: Your Spice.ai API key for authentication (required for Spice Cloud). Follow the [guide](https://docs.spice.ai/portal/apps/api-keys) to obtain it from the Spice Cloud portal. ![Spice.ai Authentication](/img/powerbi/powerbi-spice-auth-dlg.png) 9. Click `Connect` to establish the connection. ## Working with Spice datasets[​](#working-with-spice-datasets "Direct link to Working with Spice datasets") After establishing a connection, Spice datasets appear under their respective schemas, with the default schema being `spice.public`. When writing native queries, use the `PostgreSQL` dialect, as Spice is built on this standard. ![Spice PowerBI Example](/img/powerbi/powerbi-spice-example.png) ## Supported Data Types[​](#supported-data-types "Direct link to Supported Data Types") The following Apache Arrow / DataFusion SQL types are supported. Other types will result in a `Unable to understand the type for column` error. Please [report an issue](https://github.com/spiceai/powerbi-connector/issues) if support for additional types is required. | Arrow Type | DataFusion SQL Type | Power Query M Type | | ----------------------------------------------------------- | ------------------- | ------------------ | | Boolean | BOOLEAN | Logical | | Int16 | SMALLINT | Int16 | | Int32 | INTEGER | Int32 | | Int64 | BIGINT | Int64 | | Float32 | REAL | Single | | Float64 | DOUBLE | Double | | Decimal128 / Decimal256 | DECIMAL | Decimal | | Utf8 | VARCHAR | Text | | Date32 / Date64 | DATE | Date | | Time32 / Time64 | TIME | Time | | Timestamp | TIMESTAMP | DateTime | | List / LargeList / FixedSizeList / ListView / LargeListView | ARRAY | Text | | Interval | INTERVAL | Text | | Struct | STRUCT | Text | ## Limitations[​](#limitations "Direct link to Limitations") ### LargeUtf8 Data Type Is Not Supported[​](#largeutf8-data-type-is-not-supported "Direct link to LargeUtf8 Data Type Is Not Supported") To work around this limitation, use [views](/docs/next/features/views) to manually convert `LargeUtf8` columns to `Utf8` by casting them with `::TEXT`. **Example:** ``` views: - name: taxi_zone_lookup sql: | SELECT LocationID as LocationID, Borough::TEXT as Borough, Zone::TEXT as Zone, service_zone::TEXT as service_zone FROM taxi_zone_lookup_temp; ``` ### Date Time Arithmetic Operations Are Not Supported[​](#date-time-arithmetic-operations-are-not-supported "Direct link to Date Time Arithmetic Operations Are Not Supported") Due to lack of support for the `timestampdiff` function in the [DataFusion query engine](https://datafusion.apache.org/user-guide/sql/scalar_functions.html), date and time arithmetic operations—such as subtracting or adding timestamps and intervals—are not supported and will result in an error similar to `Invalid function 'timestampdiff'.\nDid you mean 'to_timestamp'? (Internal; ExecuteQuery)`. For example: ``` (parameter) => let Sorted = Table.Sort(parameter[taxi_table], {"RecordID"}), T2 = Table.SelectColumns(Sorted, {"PULocationID","lpep_pickup_datetime"}), T3 = Table.Sort(T2, {"PULocationID"}), T4 = Table.AddColumn(T3, "Diff1", each [lpep_pickup_datetime] - #datetime(1999,1,5,0,0,0)) TA = Table.FirstN(T6, 4) in TA ``` ``` ADBC: InternalError [] [FlightSQL] [FlightSQL] Error during planning: Invalid function 'timestampdiff'.\nDid you mean 'to_timestamp'? (Internal; ExecuteQuery) ``` Please [report an issue](https://github.com/spiceai/powerbi-connector/issues) if support for date or time arithmetic operations is required. --- # Apache Superset Use [Apache Superset](https://superset.apache.org/) to query and visualize datasets loaded in Spice. > Apache Superset is a modern, enterprise-ready business intelligence web application. It is fast, lightweight, intuitive, and loaded with options that make it easy for users of all skill sets to explore and visualize their data, from simple pie charts to highly detailed deck.gl geospatial charts. > > – [Apache Superset documentation](https://superset.apache.org/docs/intro/) ## Start Apache Superset with Flight SQL & DataFusion SQL Dialect support[​](#start-apache-superset-with-flight-sql--datafusion-sql-dialect-support "Direct link to Start Apache Superset with Flight SQL & DataFusion SQL Dialect support") Superset requires a Python [DB API 2](https://peps.python.org/pep-0249/) database driver and a [SQLAlchemy](https://www.sqlalchemy.org/) dialect to be installed for each connected datastore. Spice implements a Flight SQL server that understands the DataFusion SQL Dialect. The [`flightsql-dbapi`](https://pypi.org/project/flightsql-dbapi/) library for Python provides the required DB API 2 driver and SQLAlchemy dialect. Select the appropriate tab based on whether you are experimenting with this feature or integrating it into an existing Superset instance. * Experimenting * Integrating with Existing Superset The easiest way to connect Apache Superset and Spice is to follow the [`Sales BI` Cookbook Recipe](https://github.com/spiceai/cookbook/tree/trunk/sales-bi). This recipe builds a local Docker image based on Apache Superset that is pre-configured with the `flightsql-dbapi` library needed to connect to Spice. Clone the Spice cookbook repository and navigate to the `sales-bi` directory: ``` git clone https://github.com/spiceai/cookbook.git cd cookbook/sales-bi ``` Start Apache Superset along with the Spice runtime in Docker Compose: ``` make start ``` Log into Apache Superset at with the username and password `admin/admin`. Follow the below steps to configure a database connection to Spice manually, or run `make import-dashboards` to automatically configure the connection and create a sample dashboard. ## Generic / Virtual Machine[​](#generic--virtual-machine "Direct link to Generic / Virtual Machine") Install the `flightsql-dbapi` library in your existing Apache Superset environment: ``` pip install flightsql-dbapi ``` ## Docker Container[​](#docker-container "Direct link to Docker Container") Install the library in the Dockerfile: ``` FROM apache/superset # Switching to root to install the required packages USER root # https://github.com/influxdata/flightsql-dbapi RUN pip install flightsql-dbapi # Switching back to using the `superset` user USER superset ``` Re-deploy Apache Superset with the updated Docker image. ## Temporary Docker Container Modification[​](#temporary-docker-container-modification "Direct link to Temporary Docker Container Modification") It's possible to modify a running Docker container to install the library, but the change will be lost on container restart. ``` docker exec -u root -it superset /bin/bash pip install flightsql-dbapi ``` *** ## Configure a Spice Connection[​](#configure-a-spice-connection "Direct link to Configure a Spice Connection") Once Apache Superset is up and running, and you are logged in, you can configure a connection to Spice. Hover over the `Settings` menu and select `Database Connections`. ![](/img/superset/superset-docs-connection-settings.png) Click the `+ Database` button to configure the connection. ![](/img/superset/superset-docs-new-db.png) Under `Supported Databases` select `Other`. Set the Display Name to `Spice` and the SQL Alchemy URI to `datafusion+flightsql://spiceai_host:[spiceai_port]`. Specify `?insecure=true` to skip connecting over TLS. Example: `datafusion+flightsql://spiceai-sales-bi-demo:50051?insecure=true`. Click `Test Connection` to verify the connection. ![](/img/superset/superset-docs-test-conn.png) Click `Connect` to save the connection. Start exploring the datasets loaded in Spice by creating a new dataset in Apache Superset to match one of the existing tables. --- # Tableau Use instructions below to install the **Spice.ai Tableau Connector** that enables [Tableau](https://www.tableau.com/) users to easily connect to and visualize data loaded in Spice. > Tableau is the world's leading analytics platform. Tableau is the broadest and deepest end-to-end data and analytics platform. Ensure the responsible use of data and drive better business outcomes with fully integrated data management and governance, visual analytics and data storytelling, and collaboration – all with Salesforce’s industry-leading Einstein built right in. > > – [The Tableau platform](https://www.tableau.com/) ## Step 1. Install the Arrow Flight SQL JDBC Driver[​](#step-1-install-the-arrow-flight-sql-jdbc-driver "Direct link to Step 1. Install the Arrow Flight SQL JDBC Driver") [JDBC](https://docs.oracle.com/javase/tutorial/jdbc/basics/index.html) (Java Database Connectivity) is a standard interface for connecting to and interacting with databases. The Flight SQL driver is a JDBC driver implementation based on the [Arrow Flight SQL](https://arrow.apache.org/docs/format/FlightSql.html) protocol. As Spice supports the Flight SQL protocol, the driver helps establish a connection between Tableau and Spice, enabling Tableau to execute queries and retrieve data from Spice efficiently. Download the [flight-sql-jdbc-driver.jar](https://repo1.maven.org/maven2/org/apache/arrow/flight-sql-jdbc-driver/) file to the Tableau drivers folder: * Windows * macOS * Linux **PowerShell Install Script** ``` Invoke-WebRequest -Uri "https://repo1.maven.org/maven2/org/apache/arrow/flight-sql-jdbc-driver/19.0.0/flight-sql-jdbc-driver-19.0.0.jar" -OutFile "C:\Program Files\Tableau\Drivers\flight-sql-jdbc-driver-19.0.0.jar" ``` **Install Script** ``` curl -L https://repo1.maven.org/maven2/org/apache/arrow/flight-sql-jdbc-driver/19.0.0/flight-sql-jdbc-driver-19.0.0.jar -o ~/Library/Tableau/Drivers/flight-sql-jdbc-driver-19.0.0.jar ``` **Install Script** ``` curl -L https://repo1.maven.org/maven2/org/apache/arrow/flight-sql-jdbc-driver/19.0.0/flight-sql-jdbc-driver-19.0.0.jar -o /opt/tableau/tableau_driver/jdbc/flight-sql-jdbc-driver-19.0.0.jar ``` ## Step 2. Install Spice.ai Tableau Connector[​](#step-2-install-spiceai-tableau-connector "Direct link to Step 2. Install Spice.ai Tableau Connector") ### Tableau Server[​](#tableau-server "Direct link to Tableau Server") 1. Download the latest `spiceai.taco` file from [Releases](https://github.com/spicehq/tableau-connector/releases) 2. Copy to the Tableau connectors directory * Windows * Linux **PowerShell Install Script** ``` Invoke-WebRequest -Uri "https://github.com/spicehq/tableau-connector/releases/latest/download/spiceai.taco" -OutFile "C:\Program Files\Tableau\Connectors\spiceai.taco" ``` **Install Script** ``` curl -L https://github.com/spicehq/tableau-connector/releases/latest/download/spiceai.taco -o /opt/tableau/connectors/spiceai.taco ``` 3. Restart server: `tsm restart` ### Tableau Desktop[​](#tableau-desktop "Direct link to Tableau Desktop") 1. Download the latest `spiceai.taco` file from [Releases](https://github.com/spiceai/tableau-connector/releases) 2. Copy to the Tableau connectors directory * Windows * macOS * Linux **PowerShell Install Script** ``` Invoke-WebRequest -Uri "https://github.com/spicehq/tableau-connector/releases/latest/download/spiceai.taco" -OutFile "C:\Users\[USERNAME]\Documents\My Tableau Repository\Connectors\spiceai.taco" ``` **Install Script** ``` curl -L https://github.com/spicehq/tableau-connector/releases/latest/download/spiceai.taco -o ~/Documents/My\ Tableau\ Repository/Connectors/spiceai.taco ``` **Install Script** ``` curl -L https://github.com/spicehq/tableau-connector/releases/latest/download/spiceai.taco -o /opt/tableau/connectors/spiceai.taco ``` ## Configure a Spice connection[​](#configure-a-spice-connection "Direct link to Configure a Spice connection") 1. Open **Tableau** 2. In the **Connect** column, under **To a Server**, select **Spice.ai by Spice AI, Inc**. 3. Configure a Spice connection to **Spice.ai OSS Self-Hosted** instance or to **Spice Cloud Platform**. ![Spice Tableau Connection Dialog](/img/tableau/tableau-spice-dialog.png) 4. Click **Sign In** ## Working with Spice datasets[​](#working-with-spice-datasets "Direct link to Working with Spice datasets") After establishing a connection, Spice datasets appear under their respective schemas, with the default schema being `spice.public`. When writing queries, use the `PostgreSQL` dialect, as Spice is built on this standard. ![Spice Tableau Example](/img/tableau/tableau-spice-example.png) --- # How Spice Compares Spice combines SQL query, search, and LLM inference in a single runtime. This page compares Spice to data platforms, query engines, vector databases, and AI frameworks. ## Data Query and Analytics[​](#data-query-and-analytics "Direct link to Data Query and Analytics") | Feature | **Spice** | Databricks | Snowflake | Trino / Presto | Dremio | ClickHouse | | -------------------------------- | -------------------------------------- | ------------------------------ | ------------------------- | -------------------- | --------------------- | -------------------------------------- | | **Primary Use-Case** | Data & AI apps/agents | Data lakehouse | Data warehouse | Big data analytics | Interactive analytics | Real-time analytics | | **Primary Deployment Model** | Sidecar | Cloud | Cloud | Cluster | Cluster | Cluster | | **Federated Query Support** | ✅ | Limited (Lakehouse Federation) | Limited (external tables) | ✅ | ✅ | Limited (integration engines) | | **Acceleration/Materialization** | ✅ (Arrow, SQLite, DuckDB, PostgreSQL) | ✅ (Delta Lake, MVs) | ✅ (Materialized views) | Intermediate storage | Reflections (Iceberg) | Materialized views | | **Catalog Support** | ✅ (Iceberg, Unity Catalog, AWS Glue) | ✅ (Unity Catalog) | ✅ (Snowflake Catalog) | ✅ | ✅ | Limited (Iceberg, Delta Lake, Hudi) | | **Query Result Caching** | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | **Multi-Modal Acceleration** | ✅ (OLAP + OLTP) | ❌ | ❌ | ❌ | ❌ | ❌ | | **Change Data Capture (CDC)** | ✅ (Debezium) | ✅ (Delta Live Tables) | Limited (Streams) | ❌ | ❌ | Limited (MaterializedMySQL/PostgreSQL) | ## Search and Vector Databases[​](#search-and-vector-databases "Direct link to Search and Vector Databases") | Feature | **Spice** | Elasticsearch | Turbopuffer | LanceDB | ClickHouse | | ------------------------- | ----------------------------------------- | --------------- | ------------- | ------------------ | ----------------------------- | | **Primary Use-Case** | Data & AI apps/agents | Search engine | Vector search | Embedded vector DB | Real-time analytics | | **SQL Query Support** | ✅ (Full SQL) | Limited (ES | QL) | ❌ | Limited (SQL-like API) | | **Vector Search** | ✅ | ✅ (kNN) | ✅ | ✅ | ✅ | | **Full-Text Search** | ✅ (Tantivy BM25) | ✅ (Lucene) | ✅ (BM25) | ✅ (Tantivy) | ✅ (inverted index) | | **Hybrid Search** | ✅ (Reciprocal Rank Fusion) | ✅ (RRF) | ✅ | ✅ (reranking) | Limited | | **Federated Data Access** | ✅ (40+ connectors) | ❌ | ❌ | ❌ | Limited (integration engines) | | **Data Acceleration** | ✅ (Arrow, DuckDB, SQLite, Cayenne) | ❌ | ❌ | ❌ | ❌ | | **Built-in Embeddings** | ✅ (Local + hosted models) | Limited (ELSER) | ❌ | ✅ (pluggable) | ❌ | | **LLM Integration** | ✅ (OpenAI-compatible API, tools, MCP) | ❌ | ❌ | ❌ | ❌ | | **Storage Format** | Pluggable (Arrow, Vortex, DuckDB, SQLite) | Lucene | Proprietary | Lance | MergeTree | ## AI Apps and Agents[​](#ai-apps-and-agents "Direct link to AI Apps and Agents") | Feature | **Spice** | LangChain | LlamaIndex | Ollama | | ----------------------------- | ------------------------------------ | ------------------ | ------------------ | ----------------------------- | | **Primary Use-Case** | Data & AI apps | Agentic workflows | RAG apps | LLM apps | | **Programming Language** | Any language (HTTP interface) | JavaScript, Python | Python, TypeScript | Any language (HTTP interface) | | **Unified Data + AI Runtime** | ✅ | ❌ | ❌ | ❌ | | **Federated Data Query** | ✅ | ❌ | ❌ | ❌ | | **Accelerated Data Access** | ✅ | ❌ | ❌ | ❌ | | **Tools/Functions** | ✅ (MCP HTTP+SSE) | ✅ | ✅ | ✅ | | **LLM Memory** | ✅ | ✅ | ✅ | ❌ | | **Search** | ✅ (Keyword, Vector, Full-Text) | ✅ | ✅ | Limited | | **Caching** | ✅ (Query and results caching) | Limited | ❌ | ❌ | | **Embeddings** | ✅ (Built-in & pluggable models/DBs) | ✅ | ✅ | ✅ | ✅ = Fully supported ❌ = Not supported Limited = Partial or restricted support --- # Spice.ai Runtime Components Spice runtime components are the building blocks for configuring data access, acceleration, AI models, embeddings, and secrets. Each component is defined in the `spicepod.yaml` manifest. **[Data Connectors](/docs/next/components/data-connectors)** connect to databases, data warehouses, data lakes, and file systems for federated SQL queries. Spice supports over 40 connectors including PostgreSQL, MySQL, S3, Snowflake, Databricks, and DuckDB. **[Data Accelerators](/docs/next/components/data-accelerators)** materialize datasets locally in memory or on disk for faster query performance. Choose from Arrow (in-memory), DuckDB, SQLite, PostgreSQL, or Cayenne depending on workload characteristics. **[Catalog Connectors](/docs/next/components/catalogs)** integrate with data catalogs like Apache Iceberg, Unity Catalog, AWS Glue, and DuckLake to discover and register datasets from existing catalog infrastructure. **[Models](/docs/next/components/models)** configure LLM providers for AI inference through an OpenAI-compatible API. Connect to hosted models (OpenAI, Anthropic, xAI) or serve models locally with CUDA/Metal acceleration. **[Embeddings](/docs/next/components/embeddings)** generate vector representations of text for semantic search and RAG workflows, using built-in models or external providers. **[Tools](/docs/next/components/tools)** define callable functions that LLMs can invoke during inference, including MCP (Model Context Protocol) integrations for connecting to external services. **[Vector Engines](/docs/next/components/vectors)** index dataset embedding columns and serve nearest-neighbour search backed by DuckDB, Elasticsearch, or Amazon S3 Vectors. **[Secret Stores](/docs/next/components/secret-stores)** manage credentials and sensitive configuration values using environment variables, files, or external secret managers like AWS Secrets Manager and Azure Key Vault. ## [🗃Data Connectors](/docs/next/components/data-connectors) [40 items](/docs/next/components/data-connectors) ## [🗃Data Accelerators](/docs/next/components/data-accelerators) [6 items](/docs/next/components/data-accelerators) ## [🗃Secret Stores](/docs/next/components/secret-stores) [6 items](/docs/next/components/secret-stores) ## [🗃Catalog Connectors](/docs/next/components/catalogs) [12 items](/docs/next/components/catalogs) ## [🗃Model Providers](/docs/next/components/models) [11 items](/docs/next/components/models) ## [🗃Embeddings](/docs/next/components/embeddings) [8 items](/docs/next/components/embeddings) ## [🗃LLM Tools](/docs/next/components/tools) [2 items](/docs/next/components/tools) ## [🗃Vector Engines](/docs/next/components/vectors) [3 items](/docs/next/components/vectors) --- # Catalog Connectors In Spice, datasets are organized hierarchically with catalogs, schemas, and tables. A catalog, at the top level, contains multiple schemas. Each schema, in turn, contains multiple tables where the actual data is stored. By default a catalog named `spice` is created with all of the datasets defined in the `datasets` section of the Spicepod. ![](/img/catalog-schema-table.png) Creating schemas and tables within the `spice` catalog is configured by the `name` field in the dataset configuration. A name with a period (`.`) will create a schema, i.e. a dataset defined with `name: foo.bar` would have a full path of `spice.foo.bar`. If the name does not contain a period, the dataset will be created in the `public` schema of the `spice` catalog. For example, a dataset defined with `name: foo` would have a full path of `spice.public.foo`. Attempting to create a dataset with a name that contains a catalog name will result in an error. Adding catalogs to Spice is done via Catalog Connectors. Catalog Connectors connect to external catalog providers and make their tables available for federated SQL query in Spice. The schema hierarchy of the external catalog is preserved in Spice. Accelerating a catalog as a whole is supported by the [PostgreSQL Catalog Connector](/docs/next/components/catalogs/postgres), which CDC-accelerates every discovered table from a single replication slot — see [Catalog-Level CDC Acceleration](/docs/next/components/catalogs/postgres#catalog-level-cdc-acceleration). For every other Catalog Connector, an `acceleration` block on the catalog is a configuration error rather than a silent no-op. To accelerate an individual table from one of those catalogs, define it as a [dataset](/docs/next/reference/spicepod/datasets) with its own `acceleration` block. Supported Catalog Connectors include: | Name | Description | Status | Protocol/Format | | --------------- | ----------------------- | ------ | ---------------------------- | | `unity_catalog` | Unity Catalog | Stable | Delta Lake | | `databricks` | Databricks | Beta | Spark Connect, S3/Delta Lake | | `iceberg` | Apache Iceberg | Beta | Parquet | | `spice.ai` | Spice.ai Cloud Platform | Beta | Arrow Flight | | `ducklake` | DuckLake | Beta | Parquet | | `glue` | AWS Glue | Alpha | Parquet, Iceberg | | `snowflake` | Snowflake | Alpha | Snowflake SQL | | `pg` | PostgreSQL / Redshift | Alpha | PostgreSQL Wire Protocol | | `mysql` | MySQL | Alpha | MySQL Wire Protocol | | `mssql` | Microsoft SQL Server | Alpha | TDS | | `adbc` | ADBC | Alpha | Arrow (ADBC) | | `oracle` | Oracle | Alpha | Oracle Net | ## Catalog Connector Docs[​](#catalog-connector-docs "Direct link to Catalog Connector Docs") Catalogs are configured using a Catalog Connector in the `catalogs` section of the Spicepod. See the specific Catalog Connector documentation for configuration details. ### `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. Example: ``` catalogs: - from: spice.ai name: spiceai include: - 'tpch.*' # Include only the "tpch" tables. ``` ### `exclude`[​](#exclude "Direct link to exclude") Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. Every Catalog Connector applies `exclude`. Example: ``` catalogs: - from: pg name: my_pg include: - 'public.*' # Consider every table in the "public" schema... exclude: - 'public.*_audit' # ...except the audit tables. ``` ## [📄️Databricks](/docs/next/components/catalogs/databricks) [Connect to a Databricks Unity Catalog provider.](/docs/next/components/catalogs/databricks) ## [🗃Unity Catalog](/docs/next/components/catalogs/unity-catalog) [1 item](/docs/next/components/catalogs/unity-catalog) ## [📄️Spice.ai](/docs/next/components/catalogs/spiceai) [Connect to the Spice.ai built-in catalog.](/docs/next/components/catalogs/spiceai) ## [📄️Iceberg](/docs/next/components/catalogs/iceberg) [Connect to an Iceberg catalog provider.](/docs/next/components/catalogs/iceberg) ## [📄️Glue](/docs/next/components/catalogs/glue) [Connect to an AWS Glue Data Catalog.](/docs/next/components/catalogs/glue) ## [📄️DuckLake](/docs/next/components/catalogs/ducklake) [Connect to a DuckLake catalog for federated SQL query.](/docs/next/components/catalogs/ducklake) ## [📄️Snowflake](/docs/next/components/catalogs/snowflake) [Connect to a Snowflake database as a catalog provider for federated SQL query.](/docs/next/components/catalogs/snowflake) ## [📄️PostgreSQL](/docs/next/components/catalogs/postgres) [Connect to a PostgreSQL database as a catalog provider for federated SQL query.](/docs/next/components/catalogs/postgres) ## [📄️MySQL](/docs/next/components/catalogs/mysql) [Connect to a MySQL database as a catalog provider for federated SQL query.](/docs/next/components/catalogs/mysql) ## [📄️MS SQL Server](/docs/next/components/catalogs/mssql) [Connect to a Microsoft SQL Server database as a catalog provider for federated SQL query.](/docs/next/components/catalogs/mssql) ## [📄️ADBC](/docs/next/components/catalogs/adbc) [Connect to databases via ADBC for automatic schema and table discovery.](/docs/next/components/catalogs/adbc) ## [📄️Oracle](/docs/next/components/catalogs/oracle) [Connect to an Oracle database as a catalog provider for federated SQL query.](/docs/next/components/catalogs/oracle) --- # ADBC Catalog Connector Connect to any database via [ADBC](https://arrow.apache.org/adbc/) (Arrow Database Connectivity) as a catalog provider for federated SQL query. The ADBC Catalog Connector uses the ADBC metadata API to automatically discover schemas and tables from the connected database. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") An ADBC-compatible driver must be installed on the system where Spice runs. See the [ADBC Data Connector](/docs/next/components/data-connectors/adbc#prerequisites) for driver installation instructions. ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: adbc name: my_catalog include: - 'my_schema.*' # include all tables from my_schema params: adbc_driver: bigquery adbc_uri: "bigquery:///my-gcp-project" ``` ### `from`[​](#from "Direct link to from") The `from` field is used to specify the catalog provider. For ADBC, specify `adbc`. ### `name`[​](#name "Direct link to name") The `name` field specifies the name of the catalog in Spice. Tables from the ADBC-connected database will be available under this catalog name in Spice. The schema hierarchy of the external database is preserved in Spice. ### `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. ### `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ### `params`[​](#params "Direct link to params") The following parameters are supported: | Parameter Name | Description | | -------------------------- | ---------------------------------------------------------------------------------------------------------- | | `adbc_driver` | Required. The ADBC driver name (e.g., `bigquery`, `trino`, `snowflake`, `redshift`, `duckdb`, `postgres`). | | `adbc_driver_path` | Optional. Absolute path to the ADBC driver shared library. If not specified, the driver is loaded by name. | | `adbc_uri` | Required. Database URI/connection string for the ADBC driver. | | `adbc_username` | Optional. Username for database authentication. | | `adbc_password` | Optional. Password for database authentication. | | `adbc_driver_options` | Optional. Semicolon-delimited driver-specific database options (e.g., `key1=value1;key2=value2`). | | `connection_pool_size` | Optional. Maximum number of connections in the connection pool. Default: `5`. | | `connection_pool_min_idle` | Optional. Minimum number of idle connections to keep open in the pool. Default: `1`. | ## Examples[​](#examples "Direct link to Examples") ### BigQuery[​](#bigquery "Direct link to BigQuery") ``` catalogs: - from: adbc name: bq_catalog include: - 'my_dataset.*' params: adbc_driver: bigquery adbc_uri: "bigquery:///my-gcp-project" adbc_driver_options: >- adbc.bigquery.sql.dataset_id=my_dataset ``` ### Snowflake[​](#snowflake "Direct link to Snowflake") ``` catalogs: - from: adbc name: sf_catalog include: - 'PUBLIC.*' params: adbc_driver: snowflake adbc_uri: ${secrets:snowflake_uri} adbc_username: ${secrets:snowflake_user} adbc_password: ${secrets:snowflake_pass} ``` --- # Databricks Catalog Connector Connect to a [Databricks Unity Catalog](https://www.databricks.com/product/unity-catalog) as a catalog provider for federated SQL query using [Spark Connect](https://www.databricks.com/blog/2022/07/07/introducing-spark-connect-the-power-of-apache-spark-everywhere.html), directly from [Delta Lake](https://delta.io/) tables, or using the [SQL Statement Execution API](https://docs.databricks.com/aws/en/dev-tools/sql-execution-tutorial). ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog # tables from this catalog will be available in the "uc_catalog" catalog in Spice include: - '*.my_table_name' # include only the "my_table_name" tables params: mode: delta_lake # or spark_connect or sql_warehouse databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com dataset_params: # delta_lake S3 parameters databricks_aws_region: us-west-2 databricks_aws_access_key_id: ${secrets:aws_access_key_id} databricks_aws_secret_access_key: ${secrets:aws_secret_access_key} databricks_aws_endpoint: s3.us-west-2.amazonaws.com # spark_connect parameters databricks_cluster_id: 1234-567890-abcde123 # sql_warehouse parameters databricks_sql_warehouse_id: 2b4e24cff378fb24 ``` ## `from`[​](#from "Direct link to from") The `from` field is used to specify the catalog provider. For Databricks, use `databricks:`. The `catalog_name` is the name of the catalog in the Databricks Unity Catalog you want to connect to. ## `name`[​](#name "Direct link to name") The `name` field is used to specify the name of the catalog in Spice. Tables from the Databricks catalog will be available in the schema with this name in Spice. The schema hierarchy of the external catalog is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. ## `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ## `params`[​](#params "Direct link to params") The following parameters are supported for configuring the connection to the Databricks Unity Catalog: | Parameter Name | Definition | | ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mode` | The execution mode for querying against Databricks. `spark_connect` uses Spark Connect to query against Databricks requires a Spark cluster to be available. `delta_lake` queries directly from Delta Tables and requires the object store credentials to be provided. Default is `spark_connect`. | | `databricks_endpoint` | The Databricks workspace endpoint, e.g. `dbc-a12cd3e4-56f7.cloud.databricks.com` | | `databricks_token` | The Databricks API token to authenticate with the Unity Catalog API. Use the [secret replacement syntax](/docs/next/components/secret-stores) to reference a secret, e.g. `${secrets:my_databricks_token}`. | | `databricks_use_ssl` | If true, use a TLS connection to connect to the Databricks endpoint. Default is `true`. | | `databricks_auth_mode` | Optional. Pins the authentication flow instead of inferring it from the credentials that are set. One of `auto`, `token`, `m2m`, `u2m` (case-insensitive). Defaults to `auto`. See [Pinning the authentication mode](#pinning-the-authentication-mode). | To locate the Databricks endpoint, do the following: 1. Log in to your Databricks workspace. 2. In the sidebar, click Compute. 3. In the list of available clusters, click the target cluster's name. 4. On the Configuration tab, expand Advanced options. 5. Click the JDBC/ODBC tab. 6. The endpoint is the Server Hostname. ## Authentication[​](#authentication "Direct link to Authentication") ### Personal access token[​](#personal-access-token "Direct link to Personal access token") To learn more about how to set up personal access tokens, see [Databricks PAT docs](https://docs.databricks.com/aws/en/dev-tools/auth/pat). ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog include: - '*.my_table_name' params: databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com databricks_token: ${secrets:DATABRICKS_TOKEN} # PAT ``` ### Databricks service principal[​](#databricks-service-principal "Direct link to Databricks service principal") Spice supports the Machine-to-Machine (M2M) OAuth flow with service principal credentials by utilizing the `databricks_client_id` and `databricks_client_secret` parameters. The runtime will automatically refresh the token. Ensure that you grant your service principal the "Data Reader" privilege preset for the catalog and "Can Attach" cluster permissions when using Spark Connect mode. To learn more about how to set up the service principal, see [Databricks M2M OAuth docs](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m). ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog include: - '*.my_table_name' params: databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} # service principal client id databricks_client_secret: ${secrets:DATABRICKS_CLIENT_SECRET} # service principal client secret ``` ### Pinning the authentication mode[​](#pinning-the-authentication-mode "Direct link to Pinning the authentication mode") By default (`databricks_auth_mode: auto`) the catalog connector infers the flow from whichever credentials resolve: a `databricks_token` alone selects personal access token, `databricks_client_id` alone selects U2M, and `databricks_client_id` together with `databricks_client_secret` selects M2M. Credentials do not only come from the Spicepod. `databricks_token`, `databricks_client_secret`, and the other secret parameters are auto-loaded from the [secret stores](/docs/next/components/secret-stores) and the environment (e.g. `DATABRICKS_CLIENT_SECRET`, `SPICE_DATABRICKS_CLIENT_SECRET`) when the Spicepod omits them — so an ambient client secret on the host is enough to switch a U2M catalog to machine-to-machine, which then fails with a service principal `401`. Set `databricks_auth_mode` to pin the flow. A pinned mode ignores the credentials the flow does not use: | Value | Flow | Requires | | ---------------- | --------------------------------- | -------------------------------------------------- | | `auto` (default) | Inferred from the credentials set | — | | `token` | Personal access token | `databricks_token` | | `m2m` | Service principal (M2M) OAuth | `databricks_client_id`, `databricks_client_secret` | | `u2m` | User-to-machine OAuth | `databricks_client_id` | Values are matched case-insensitively, so `M2M` and `U2M` are also accepted. A required parameter that is missing for the pinned mode fails at load with an error naming it; an unrecognized value is rejected. ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog params: databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com databricks_auth_mode: m2m # ignore any ambient databricks_token databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} databricks_client_secret: ${secrets:DATABRICKS_CLIENT_SECRET} ``` ## `dataset_params`[​](#dataset_params "Direct link to dataset_params") The `dataset_params` field is used to configure the dataset-specific parameters for the catalog. The following parameters are supported: ### Spark Connect parameters[​](#spark-connect-parameters "Direct link to Spark Connect parameters") | Dataset Parameter Name | Definition | | ----------------------- | ---------------------------------------------------------------------------------------------- | | `databricks_cluster_id` | The ID of the compute cluster in Databricks to use for the query. e.g. `1234-567890-abcde123`. | To locate the cluster ID, do the following: 1. Log in to your Databricks workspace. 2. In the sidebar, click Compute. 3. In the list of available clusters, click the target cluster's name. 4. On the Configuration tab, expand Advanced options. 5. Click the JDBC/ODBC tab. 6. The cluster ID is the prefix of the Server Hostname. ### Delta Lake object store parameters[​](#delta-lake-object-store-parameters "Direct link to Delta Lake object store parameters") Configure the connection to the object store when using `mode: delta_lake`. Use the [secret replacement syntax](/docs/next/components/secret-stores) to reference a secret, e.g. `${secrets:aws_access_key_id}`. ### SQL Warehouse parameters[​](#sql-warehouse-parameters "Direct link to SQL Warehouse parameters") * `databricks_sql_warehouse_id`: The ID of the SQL Warehouse in Databricks to use for the query. e.g. `2b4e24cff378fb24`. To locate your SQL Warehouse ID, do the following: 1. Log in to your Databricks workspace. 2. In the sidebar, click SQL -> SQL Warehouses. 3. In the list of available warehouses, click the target warehouse's name. 4. Next to the **Name** field, the ID follows the name in parentheses. For example: `My Serverless Warehouse (ID: 2b4e24cff378fb24)` #### SQL Warehouse connection tuning[​](#sql-warehouse-connection-tuning "Direct link to SQL Warehouse connection tuning") The following parameters control resilience and concurrency for the SQL Statement Execution API when using `mode: sql_warehouse`: | Parameter Name | Description | Default | | ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------- | | `connect_timeout` | Timeout for establishing TCP/TLS connections to the Databricks API. Accepts durations like `10s`. | `10s` | | `client_timeout` | Per-HTTP-call timeout (statement submit, status poll, chunk fetch). Set to the longest expected single call, not total query duration. Accepts durations like `30s` or `2m`. | `30s` | | `max_concurrent_requests` | Maximum number of concurrent HTTP requests to the SQL Warehouse API. | `8` | | `http_max_retries` | Maximum number of HTTP-level retries for transient failures (429, 5xx). | `3` | | `backoff_method` | Backoff strategy for transient HTTP retries. Options: `fibonacci`, `exponential`. | `fibonacci` | | `statement_max_retries` | Maximum number of poll retries when waiting for async statement completion. | `14` | | `disable_on_permanent_error` | When `true`, non-retryable errors (401, 403, 404) permanently disable the connector to prevent a thundering herd of failed requests. | `true` | ### Delta Lake parameters[​](#delta-lake-parameters "Direct link to Delta Lake parameters") * `client_timeout`: HTTP client request timeout. In `delta_lake` mode, applies to the object store client. In `sql_warehouse` mode, applies per-HTTP-call. Accepts durations like `30s` or `5m`. Default: `30s`. #### AWS S3[​](#aws-s3 "Direct link to AWS S3") | Dataset Parameter Name | Definition | | ---------------------------------- | --------------------------------------------------------------------------------------------------- | | `databricks_aws_region` | The AWS region for the S3 object store. E.g. `us-west-2`. | | `databricks_aws_access_key_id` | The access key ID for the S3 object store. | | `databricks_aws_secret_access_key` | The secret access key for the S3 object store. | | `databricks_aws_session_token` | Optional. The AWS session token for the S3 object store. Required with temporary (STS) credentials. | | `databricks_aws_endpoint` | The endpoint for the S3 object store. E.g. `s3.us-west-2.amazonaws.com`. | Example: ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog include: - '*.my_table_name' params: mode: delta_lake databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com dataset_params: databricks_aws_region: us-west-2 databricks_aws_access_key_id: ${secrets:aws_access_key_id} databricks_aws_secret_access_key: ${secrets:aws_secret_access_key} databricks_aws_endpoint: s3.us-west-2.amazonaws.com ``` #### Azure Blob[​](#azure-blob "Direct link to Azure Blob") Note One of the following auth values must be provided for Azure Blob: * `databricks_azure_storage_account_key`, * `databricks_azure_storage_client_id` and `databricks_azure_storage_client_secret`, or * `databricks_azure_storage_sas_key`. | Dataset Parameter Name | Definition | | ---------------------------------------- | ---------------------------------------------------------------------- | | `databricks_azure_storage_account_name` | The Azure Storage account name. | | `databricks_azure_storage_account_key` | The Azure Storage master key for accessing the storage account. | | `databricks_azure_storage_client_id` | The service principal client id for accessing the storage account. | | `databricks_azure_storage_client_secret` | The service principal client secret for accessing the storage account. | | `databricks_azure_storage_sas_key` | The shared access signature key for accessing the storage account. | | `databricks_azure_storage_endpoint` | The endpoint for the Azure Blob storage account. | Example: ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog include: - '*.my_table_name' params: mode: delta_lake databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com dataset_params: databricks_azure_storage_account_name: myaccount databricks_azure_storage_account_key: ${secrets:azure_storage_account_key} databricks_azure_storage_endpoint: myaccount.blob.core.windows.net ``` #### Google Storage (GCS)[​](#google-storage-gcs "Direct link to Google Storage (GCS)") | Dataset Parameter Name | Definition | | ----------------------------------- | ------------------------------------------------------------ | | `databricks_google_service_account` | Filesystem path to the Google service account JSON key file. | Example: ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog include: - '*.my_table_name' params: mode: delta_lake databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com dataset_params: databricks_google_service_account: /path/to/service-account.json ``` ## Limitations[​](#limitations "Direct link to Limitations") * Databricks catalog connector (mode: delta\_lake) does not support reading Delta tables with the `V2Checkpoint` feature enabled. To use the Databricks catalog connector (mode: delta\_lake) with such tables, drop the `V2Checkpoint` feature by executing the following command: ``` ALTER TABLE DROP FEATURE v2Checkpoint [TRUNCATE HISTORY]; ``` For more details on dropping Delta table features, refer to the official documentation: [Drop Delta table features](https://docs.databricks.com/en/delta/drop-feature.html#:~:text=Databricks%20provides%20limited%20support%20for,data%20files%20backing%20the%20table.) * The Databricks Catalog Connector (`mode: spark_connect`) does not yet support streaming query results from Spark. Memory Considerations When using the Databricks (mode: delta\_lake) Catalog connector without acceleration, data is loaded into memory during query execution. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. --- # DuckLake Catalog Connector Connect to a [DuckLake](https://ducklake.select/) catalog for federated SQL query. DuckLake is an open lakehouse format that stores metadata in a SQLite-compatible database (or PostgreSQL) and data in Parquet files, providing lakehouse-style operations without a separate metadata service. For connecting to individual DuckLake tables, see the [DuckLake Data Connector documentation](/docs/next/components/data-connectors/ducklake). ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: ducklake:s3://my-bucket/path/metadata.ducklake name: my_lakehouse # access: read_write # Optional. Enable write operations. params: ducklake_name: ducklake # Optional. Name to attach the catalog as in DuckDB. Defaults to 'ducklake'. ducklake_open: /path/to/local.duckdb # Optional. Path to a DuckDB file for persistent storage. ``` ## `from`[​](#from "Direct link to from") The `from` field specifies the DuckLake catalog connection. Use `ducklake:`, where `connection_string` is the location of the DuckLake metadata. Supported connection string formats: | Backend | Example | | ---------- | ---------------------------------------------------------------------------- | | Local file | `ducklake:/path/to/metadata.ducklake` | | AWS S3 | `ducklake:s3://bucket/path/metadata.ducklake` | | PostgreSQL | `ducklake:postgres:dbname=mydb host=localhost user=postgres password=secret` | The connection string can also be provided via the `ducklake_connection_string` parameter. ## `name`[​](#name "Direct link to name") The `name` field specifies the name of the catalog in Spice. Tables from the DuckLake catalog will be available using this name in Spice. The schema hierarchy of the DuckLake catalog is preserved. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. ``` catalogs: - from: ducklake:s3://my-bucket/metadata.ducklake name: my_lakehouse include: - 'main.*' # Include all tables in the "main" schema ``` ## `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ## `access`[​](#access "Direct link to access") The `access` field controls what operations are allowed on the catalog: | Access Mode | Description | | ------------------- | ------------------------------------------------------------------------------------ | | `read` (default) | Query tables only. DuckDB opens in read-only mode. | | `read_write` | Query and write data (INSERT). DuckDB opens in read-write mode. | | `read_write_create` | Full access including CREATE/DROP SCHEMA and TABLE. DuckDB opens in read-write mode. | ## `params`[​](#params "Direct link to params") | Parameter Name | Description | | -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `ducklake_connection_string` | The DuckLake metadata location (e.g., `s3://bucket/path/metadata.ducklake`). If omitted, the value from `from: ducklake:` is used. | | `ducklake_name` | The name to attach the DuckLake catalog as in DuckDB. Default: `ducklake`. | | `ducklake_open` | Path to an existing DuckDB file for persistent storage. If not provided, an in-memory DuckDB instance is used. | | `ducklake_aws_region` | Optional. The AWS region for S3 storage. Default: `us-east-1` when explicit credentials are provided. | | `ducklake_aws_access_key_id` | Optional. The AWS access key ID for S3 storage. Must be set together with `ducklake_aws_secret_access_key`. | | `ducklake_aws_secret_access_key` | Optional. The AWS secret access key for S3 storage. Must be set together with `ducklake_aws_access_key_id`. | | `ducklake_aws_session_token` | Optional. The AWS session token for S3 storage. Required with temporary (STS) credentials. Ignored, with a warning, unless `ducklake_aws_access_key_id` is also set. | | `ducklake_aws_endpoint` | Optional. Custom S3-compatible endpoint URL (e.g., for MinIO). | | `ducklake_aws_allow_http` | Optional. Set to `true` to allow HTTP (non-TLS) connections to S3. Default: `false`. | | `ducklake_automatic_migration` | Optional. Set to `true` to automatically migrate an older DuckLake catalog schema to the version required by the DuckLake extension on attach. Default: `false`. Migration rewrites catalog metadata and **cannot be undone**. | ## Authentication[​](#authentication "Direct link to Authentication") ### AWS S3[​](#aws-s3 "Direct link to AWS S3") When no explicit S3 credentials are configured, DuckDB falls back to its built-in credential chain provider: 1. Environment variables (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN`) 2. Shared credentials file (`~/.aws/credentials`) 3. IAM instance profiles (on EC2/ECS) To provide explicit S3 credentials, use the `ducklake_aws_*` parameters: ``` catalogs: - from: ducklake:s3://my-bucket/metadata.ducklake name: my_lakehouse params: ducklake_aws_region: us-west-2 ducklake_aws_access_key_id: ${secrets:AWS_ACCESS_KEY_ID} ducklake_aws_secret_access_key: ${secrets:AWS_SECRET_ACCESS_KEY} ``` For S3-compatible storage (e.g., MinIO), use `ducklake_aws_endpoint`: ``` catalogs: - from: ducklake:s3://my-bucket/metadata.ducklake name: my_lakehouse params: ducklake_aws_endpoint: http://minio:9000 ducklake_aws_access_key_id: ${secrets:MINIO_ACCESS_KEY} ducklake_aws_secret_access_key: ${secrets:MINIO_SECRET_KEY} ducklake_aws_allow_http: true ``` ## Examples[​](#examples "Direct link to Examples") ### Local DuckLake catalog[​](#local-ducklake-catalog "Direct link to Local DuckLake catalog") ``` catalogs: - from: ducklake:/path/to/metadata.ducklake name: my_lakehouse ``` ### S3-backed DuckLake catalog[​](#s3-backed-ducklake-catalog "Direct link to S3-backed DuckLake catalog") ``` catalogs: - from: ducklake:s3://my-bucket/lakehouse/metadata.ducklake name: cloud_lakehouse ``` ### PostgreSQL metadata backend[​](#postgresql-metadata-backend "Direct link to PostgreSQL metadata backend") Use PostgreSQL as the metadata storage for multi-user access (note: `from` field does not support secrets replacement): ``` catalogs: - from: ducklake name: my_lakehouse params: ducklake_connection_string: "postgres:dbname=ducklake_catalog host=localhost user=postgres password=${secrets:PASSWORD}" ``` ### Read-write with DDL support[​](#read-write-with-ddl-support "Direct link to Read-write with DDL support") ``` catalogs: - from: ducklake:s3://my-bucket/metadata.ducklake name: my_lakehouse access: read_write_create ``` ``` -- Create a new schema CREATE SCHEMA my_lakehouse.analytics; -- Create a table CREATE TABLE my_lakehouse.analytics.events ( id BIGINT, event_type VARCHAR, timestamp TIMESTAMP ); -- Insert data INSERT INTO my_lakehouse.analytics.events VALUES (1, 'click', '2026-03-01T10:00:00'); -- Drop a table DROP TABLE my_lakehouse.analytics.events; -- Drop an empty schema DROP SCHEMA my_lakehouse.analytics; ``` ## Write Support[​](#write-support "Direct link to Write Support") This catalog supports writing data to DuckLake tables using SQL [`INSERT INTO`](/docs/next/reference/sql/dml#insert) statements when `access` is set to `read_write` or `read_write_create`. ``` catalogs: - from: ducklake:s3://my-bucket/metadata.ducklake name: my_lakehouse access: read_write ``` ``` INSERT INTO my_lakehouse.main.customers (id, name, email) VALUES (1, 'Acme Corp', 'info@acme.com'); ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). Limitations * Spice uses DuckDB 1.5.3, which supports DuckLake 1.0. Older DuckLake catalogs require a metadata migration before use — set `ducklake_automatic_migration: true` to perform it on attach (this rewrites catalog metadata and cannot be undone). See [DuckLake migration guide](https://ducklake.select/docs/stable/duckdb/guides/troubleshooting#connecting-to-an-older-ducklake). * The DuckLake DuckDB extension is downloaded at runtime on first use, requiring network connectivity. * The `information_schema` and `pg_catalog` system schemas are automatically filtered out during discovery. * Catalog refresh is non-incremental — a full re-query of `information_schema` is performed on each refresh cycle. * If a table fails to load during catalog refresh, it is skipped with a warning and does not fail the entire catalog. --- # Glue Catalog Connector Connect to an [AWS Glue Data Catalog](https://docs.aws.amazon.com/glue/latest/dg/start-data-catalog.html) as a catalog provider for federated SQL query. ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: glue name: my_glue_catalog # tables from this catalog will be available in the "my_glue_catalog" catalog in Spice include: - '*.my_table_name' # include only the "my_table_name" tables params: glue_region: us-east-1 # Region of the AWS Glue Data Catalog. glue_key: ${secrets:aws_access_key_id} # Optional. Access key ID for the AWS Glue Data Catalog. glue_secret: ${secrets:aws_secret_access_key} # Optional. Secret access key for the AWS Glue Data Catalog. ``` ### `from`[​](#from "Direct link to from") The `from` field is used to specify the catalog provider. For Glue, use either `glue` (which targets the default Glue catalog for the AWS account and region) or `glue:` to target a specific Glue catalog. The catalog id appears after the first `:` in the `from` value. For [Amazon S3 Tables](https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tables.html), use the format `glue::s3tablescatalog/`. The catalog id is otherwise the AWS account id when overriding the default catalog explicitly. Examples: ``` catalogs: # Default Glue catalog for the configured AWS account and region. - from: glue name: glue_default # Specific S3 Tables-backed Glue catalog. - from: glue:123456789012:s3tablescatalog/my_table_bucket name: glue_s3_tables ``` ### `name`[​](#name "Direct link to name") The `name` field is used to specify the name of the catalog in Spice. Tables from the AWS Glue Data Catalog will be available in the schema with this name in Spice. The schema hierarchy of the external catalog is preserved in Spice. ### `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. ### `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ### `params`[​](#params "Direct link to params") The following parameters are supported for configuring the connection to the Glue Data Catalog: | Parameter Name | Definition | | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `glue_region` | The AWS region for the Glue Data Catalog. E.g. `us-west-2`. | | `glue_key` | Access key (e.g. AWS\_ACCESS\_KEY\_ID for AWS). If not provided, credentials will be loaded from environment variables or IAM roles. | | `glue_secret` | Secret key (e.g. AWS\_SECRET\_ACCESS\_KEY for AWS). If not provided, credentials will be loaded from environment variables or IAM roles. | | `glue_session_token` | Session token (e.g. AWS\_SESSION\_TOKEN for AWS) for temporary credentials | | `glue_iam_role_source` | Optional. IAM role credential source. `auto` (default) uses the default AWS credential chain, `metadata` uses only instance/container metadata (IMDS, ECS, EKS/IRSA), `env` uses only environment variables. | The following parameters control how the embedded S3 reader fetches Parquet/CSV data files referenced by Glue table metadata. They are inherited from the [S3 data connector](/docs/next/components/data-connectors/s3) and do not apply to Iceberg-format tables, whose object I/O is handled by the Iceberg client. | Parameter Name | Definition | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `glue_endpoint` | Optional. Custom S3-compatible endpoint URL used when reading Parquet/CSV data files (e.g. `https://s3.us-east-1.amazonaws.com`, `http://minio.local:9000`). Leave unset for AWS S3. | | `glue_url_style` | Optional. S3 URL addressing style for Parquet/CSV data files. One of `vhost` or `path`. Auto-detected from the endpoint when unset. | | `glue_versioning` | Optional. Enables S3 object versioning support for Parquet/CSV data files when set to `enabled`. Defaults to `enabled`. | | `client_timeout` | Optional. Timeout for the underlying S3 client used to fetch Parquet/CSV data files. E.g. `30s`. | | `allow_http` | Optional. Set to `true` to allow insecure HTTP for the S3 endpoint used to read Parquet/CSV data files. Defaults to `false`. Required when `glue_endpoint` uses an `http://` scheme. | | To target a non-default Glue catalog (for example, an S3 Tables catalog), specify the catalog id in the `from` field as `glue:` — see [`from`](#from) above. | | ## Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. These credentials will be used to connect to the S3 bucket as well as the Glue catalog. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3/Glue permissions (e.g., `s3:ListBucket`, `glue:GetTable`) to access the tables. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ### Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access Iceberg tables in S3/Glue: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::company-bucketname-datasets" }, { "Effect": "Allow", "Action": ["s3:GetObject", "s3:PutObject"], "Resource": "arn:aws:s3:::company-bucketname-datasets/*" }, { "Effect": "Allow", "Action": [ "glue:GetCatalog", "glue:GetDatabases", "glue:GetDatabase", "glue:GetTable", "glue:GetTables" ], "Resource": "*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------- | -------------------------------------------------------------- | | `s3:ListBucket` | Required. Allows scanning all objects from the bucket | | `s3:GetObject` | Required. Allows fetching objects | | `s3:PutObject` | Required for write operations. Allows writing objects | | `glue:GetCatalog` | Required. Retrieve metadata about the specified catalog. | | `glue:GetDatabases` | Required. List the databases available in the current catalog. | | `glue:GetDatabase` | Required. Retrieve metadata about the specified database. | | `glue:GetTable` | Required. Retrieve metadata about the specified table. | | `glue:GetTables` | Required. List the tables available in the current database. | ## Write Support[​](#write-support "Direct link to Write Support") This catalog supports writing data to Glue-managed Iceberg tables using SQL [`INSERT INTO`](/docs/next/reference/sql/dml#insert) statements. Writes are currently append-only — inserted data is added as new data files and registered through a new Iceberg table snapshot. Schema validation ensures inserted data matches the target table schema. To enable writes for all tables in the catalog, set `access: read_write` on the catalog: ``` catalogs: - from: glue name: my_glue_catalog access: read_write params: glue_region: us-east-1 ``` ``` -- Insert with values INSERT INTO my_glue_catalog.my_database.my_table (id, name, amount) VALUES (1, 'Alice', 100.0); -- Insert from another table INSERT INTO my_glue_catalog.my_database.my_table SELECT * FROM staging_table; ``` Inserting into partitioned Iceberg tables is supported. `UPDATE` and `DELETE` operations are not currently supported. Write operations require `s3:PutObject` permission on the target S3 bucket in addition to the read permissions listed above. For more details, see [Data Ingestion](/docs/next/features/data-ingestion). ## Limitations[​](#limitations "Direct link to Limitations") warning * This catalog connector is limited to tables that use the S3 data source. Kinesis and Kafka data sources are not currently supported. * This catalog connector is currently limited to Iceberg tables, tables with parquet or CSV data format only. ## Cookbook[​](#cookbook "Direct link to Cookbook") There is a [cookbook recipe](https://github.com/spiceai/cookbook/tree/trunk/catalogs/glue) to configure an AWS Glue Data Connector in Spice. --- # Iceberg Catalog Connector The Iceberg Catalog Connector helps connect Spice to an [Apache Iceberg](https://iceberg.apache.org/) catalog, making Iceberg tables and schemas available for federated SQL queries. Iceberg table format versions V1, V2, and V3 are supported. Every Iceberg table must be registered in a catalog, which manages table metadata and access. Using a catalog connector is the recommended approach for working with multiple Iceberg datasets, as it helps organize tables and schemas efficiently and mirrors the structure of the source catalog provider. For connecting to a single Iceberg table, see the [Iceberg Data Connector documentation](/docs/next/components/data-connectors/iceberg). For AWS Glue-based catalogs, see the [AWS Glue Catalog Connector documentation](/docs/next/components/catalogs/glue). Iceberg catalogs can be of several types: * **Iceberg REST Catalog**: The most common and recommended approach. REST Catalogs expose Iceberg table metadata over HTTP(S) endpoints and are compatible with most managed Iceberg services and cloud providers. * **AWS Glue Catalog**: Integrates with AWS Glue as a catalog provider, supporting Iceberg tables stored in S3. This is the preferred method for AWS environments. * **Hadoop-style Catalogs**: Use file-based storage (e.g., `file://`, `s3://`, `s3a://`) to manage table metadata. This approach is typically used for local development or legacy deployments. Hadoop-style Catalogs For production and cloud environments, REST and AWS Glue catalogs are recommended. Hadoop-style catalogs are supported but less common and not recommended for most new deployments. ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_catalog name: ice # tables from this catalog will be available in the "ice" catalog in Spice include: - '*.my_table_name' # include only the "my_table_name" tables params: iceberg_token: ${secrets:iceberg_token} # Optional. Bearer token value to use for Authorization header. iceberg_oauth2_credential: ${secrets:client_id}:${secrets:client_secret} # Optional. Credential to use for OAuth2 client credential flow when initializing the catalog. Separated by a colon as :. iceberg_oauth2_scope: catalog # Optional. Scope to use for OAuth2 client credential flow when initializing the catalog (default: catalog). iceberg_oauth2_server_url: https://iceberg-catalog-host.com/oauth2/token # Optional. URL of the OAuth2 server tokens endpoint for the client credential flow. iceberg_s3_endpoint: http://localhost:9000 # Optional. S3-compatible endpoint where the Iceberg tables are stored. iceberg_s3_region: us-west-2 # Optional. Region of the S3-compatible endpoint. iceberg_s3_access_key_id: ${secrets:aws_access_key_id} # Optional. Access key ID for the S3-compatible endpoint. iceberg_s3_secret_access_key: ${secrets:aws_secret_access_key} # Optional. Secret access key for the S3-compatible endpoint. iceberg_s3_session_token: ${secrets:aws_session_token} # Optional. Session token for the S3-compatible endpoint. iceberg_s3_role_arn: arn:aws:iam::123456789012:role/my-role # Optional. ARN of the IAM role to assume when accessing the S3-compatible endpoint. iceberg_s3_role_session_name: my-session # Optional. Session name to use when assuming the IAM role. # AWS Glue Catalog (see also the [AWS Glue Catalog Connector documentation](./glue)) - from: iceberg:https://glue.us-east-1.amazonaws.com/iceberg/v1/catalogs/123456789012/namespaces name: glue params: iceberg_sigv4_enabled: true ``` ## `from`[​](#from "Direct link to from") The `from` field specifies the catalog provider. For Iceberg, use `iceberg:`, where `namespace_path` is the URL to the Iceberg namespace in the catalog provider. The format is `http[s]:///v1/{prefix}/namespaces/`. For AWS Glue catalogs, the URL format is `https://glue..amazonaws.com/iceberg/v1/catalogs//namespaces`, where `` is the AWS account ID. While possible to connect to Iceberg tables hosted by Glue using this generic connector, it is recommended to instead use the [AWS Glue Catalog Connector](/docs/next/components/catalogs/glue) for connecting to Iceberg tables managed by Glue for a better experience. The selected namespace must have sub-namespaces where the tables are stored. Example: With this Iceberg catalog structure: ``` . ├── blockchain │ └── eth │ ├── blocks │ └── transactions ├── spice │ ├── tpch │ │ ├── orders │ │ └── customers │ ├── info │ └── extra │ └── tpch_orders_metadata └── unity └── very └── nested └── namespace └── foobar ``` A valid `from` value would be `iceberg:https://iceberg-catalog-host.com/v1/namespaces/spice`, and would load the following tables: * `.tpch.orders` * `.tpch.customers` * `.extra.tpch_orders_metadata` For loading a multi-part namespace, separate the namespace parts with the `%1F` character. For example, `/v1/namespaces/unity%1Fvery%1Fnested` would load the `foobar` table from the `unity/very/nested/namespace` namespace as `.namespace.foobar`. To connect to a single Iceberg table directly, see the [Iceberg Data Connector documentation](/docs/next/components/data-connectors/iceberg). ## `name`[​](#name "Direct link to name") The `name` field is used to specify the name of the catalog in Spice. Tables from the Iceberg catalog will be available in the schema with this name in Spice. The schema hierarchy of the external catalog is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. ## `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ## `params`[​](#params "Direct link to params") The following parameters are supported for configuring the connection to the Iceberg catalog, file, or S3 storage: | Parameter Name | Description | | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `iceberg_token` | Bearer token value to use for Authorization header. | | `iceberg_oauth2_credential` | Credential to use for OAuth2 client credential flow when initializing the catalog. Separated by a colon as `:`. | | `iceberg_oauth2_scope` | The scope to use for OAuth2 token endpoint (default: `catalog`). | | `iceberg_oauth2_server_url` | URL of the OAuth2 server tokens endpoint. | | `iceberg_sigv4_enabled` | Enable SigV4 authentication for the catalog (for connecting to AWS Glue). | | `iceberg_signing_region` | The region to use when signing the request for SigV4. Defaults to the region in the catalog URL if available. | | `iceberg_signing_name` | The name to use when signing the request for SigV4. Default: `glue`. | | `iceberg_warehouse` | Name of the Iceberg warehouse. Used with Glue-like catalog services (e.g., Lakekeeper). | | `iceberg_s3_endpoint` | Configure an alternative endpoint for the S3 service. This can be any S3-compatible object storage service (e.g., Minio, R2). | | `iceberg_s3_access_key_id` | The AWS access key ID to use for S3 storage. If not provided, credentials will be loaded from environment variables or IAM roles. | | `iceberg_s3_secret_access_key` | The AWS secret access key to use for S3 storage. If not provided, credentials will be loaded from environment variables or IAM roles. | | `iceberg_s3_session_token` | Configure the static session token used for S3 storage. | | `iceberg_s3_region` | The AWS S3 region to use. | | `iceberg_s3_role_session_name` | An optional identifier for the assumed role session for auditing purposes. | | `iceberg_s3_role_arn` | The Amazon Resource Name (ARN) of the role to assume. If provided instead of iceberg\_s3\_access\_key\_id and iceberg\_s3\_secret\_access\_key, temporary credentials will be fetched by assuming this role. | | `iceberg_s3_iam_role_source` | Optional. IAM role credential source. `auto` (default) uses the default AWS credential chain, `metadata` uses only instance/container metadata (IMDS, ECS, EKS/IRSA), `env` uses only environment variables. | | `iceberg_s3_connect_timeout` | Connection timeout in seconds for the S3-compatible endpoint. Default: `60`. **Note:** This parameter is currently accepted but has no effect — it is not consumed by any code path. | | `iceberg_s3_path_style_access` | Controls S3 addressing style. `true` (default) uses path-style (`endpoint/bucket`), required for object stores such as MinIO; `false` uses virtual-hosted-style (`bucket.endpoint`). | | `iceberg_gcs_project_id` | The Google Cloud project ID for GCS storage. | | `iceberg_gcs_credentials` | Base64-encoded Google Cloud service account credentials JSON for GCS storage. | | `iceberg_gcs_token` | OAuth2 token to use for GCS authentication. | | `iceberg_gcs_service_path` | Custom endpoint URL for GCS (for emulators or custom endpoints). | | `iceberg_gcs_no_auth` | Set to `true` to allow anonymous access to GCS (for public buckets). | The Iceberg Catalog Connector supports both REST and Hadoop-style Catalogs. In both cases, the warehouse path (for example, s3://bucket/warehouse/) specifies the object store location where tables are physically stored. With a Hadoop-style Catalog, the metadata is resolved directly from the filesystem using the Hadoop convention (by reading a `version-hint.txt`), rather than through a catalog service. The warehouse path itself does not change between using a REST Catalog and a Hadoop-style Catalog — only how the metadata is discovered and managed differs. The warehouse path is discovered automatically from the catalog service, but must be explicitly specified when using Hadoop-style Iceberg tables. Hadoop-style catalogs are most commonly used for local development or legacy deployments. Example using Hadoop Catalog with a local warehouse: ``` catalogs: - from: iceberg:file:///tmp/hadoop_warehouse/ name: local_hadoop ``` Example using Hadoop Catalog with S3: ``` catalogs: - from: iceberg:s3a://my-bucket/hadoop_warehouse/ name: s3_hadoop ``` ### AWS Authentication[​](#aws-authentication "Direct link to AWS Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. These credentials will be used to connect to the S3 bucket as well as the Glue catalog (if configured). 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3/Glue permissions (e.g., `s3:ListBucket`, `s3:GetObject`) to access the tables. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access Iceberg tables in S3/Glue: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::company-bucketname-datasets" }, { "Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::company-bucketname-datasets/*" }, { "Effect": "Allow", "Action": [ "glue:GetCatalog", "glue:GetDatabases", "glue:GetDatabase", "glue:GetTable", "glue:GetTables" ], "Resource": "*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------- | -------------------------------------------------------------- | | `s3:ListBucket` | Required. Allows scanning all objects from the bucket | | `s3:GetObject` | Required. Allows fetching objects | | `glue:GetCatalog` | Required. Retrieve metadata about the specified catalog. | | `glue:GetDatabases` | Required. List the databases available in the current catalog. | | `glue:GetDatabase` | Required. Retrieve metadata about the specified database. | | `glue:GetTable` | Required. Retrieve metadata about the specified table. | | `glue:GetTables` | Required. List the tables available in the current database. | ## Write Support[​](#write-support "Direct link to Write Support") This catalog supports writing data to Iceberg tables using SQL [`INSERT INTO`](/docs/next/reference/sql/dml#insert) statements. Writes are currently append-only — inserted data is added as new data files and registered through a new Iceberg table snapshot. Schema validation ensures inserted data matches the target table schema. To enable writes for all tables in the catalog, set `access: read_write` on the catalog: ``` catalogs: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_catalog name: ice access: read_write params: iceberg_token: ${secrets:iceberg_token} ``` ``` -- Insert with values INSERT INTO ice.sales.transactions (id, order_date, amount, status) VALUES (1001, '2025-01-15', 299.99, 'completed'); -- Insert from another table INSERT INTO ice.sales.transactions SELECT * FROM staging_transactions; ``` Inserting into partitioned Iceberg tables is supported. `DELETE FROM` is supported via equality delete files (Iceberg v2+ tables only). `UPDATE` operations are not currently supported. Write operations require `s3:PutObject` permission on the target S3 bucket in addition to the read permissions listed above. For more details, see [Data Ingestion](/docs/next/features/data-ingestion). ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Iceberg as a catalog connector in Spice. [Iceberg Catalog Connector](https://github.com/spiceai/cookbook/tree/trunk/catalogs/iceberg#readme) --- # Microsoft SQL Server Catalog Connector Connect to a [Microsoft SQL Server](https://www.microsoft.com/en-us/sql-server) database as a catalog provider for federated SQL query. The MSSQL Catalog Connector automatically discovers schemas and tables within a SQL Server database and makes them available for querying in Spice. For connecting to individual SQL Server tables, see the [Microsoft SQL Server Data Connector documentation](/docs/next/components/data-connectors/mssql). ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: mssql name: my_mssql include: - 'dbo.*' # include all tables from the dbo schema params: mssql_connection_string: "Server=localhost,1433;Database=my_database;User Id=${secrets:MSSQL_USER};Password=${secrets:MSSQL_PASS};" ``` ## `from`[​](#from "Direct link to from") The `from` field specifies the catalog provider. For Microsoft SQL Server, use `mssql`. ## `name`[​](#name "Direct link to name") The `name` field specifies the name of the catalog in Spice. Tables from the SQL Server database will be available under this catalog name. The schema hierarchy of the SQL Server database is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` from any schema. Multiple `include` patterns are OR'ed together. ## `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ## `params`[​](#params "Direct link to params") Connection can be configured using a connection string or individual parameters. ### Connection string[​](#connection-string "Direct link to Connection string") | Parameter Name | Description | | ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mssql_connection_string` | An [ADO.NET connection string](https://learn.microsoft.com/en-us/sql/connect/ado-net/connection-string-syntax). E.g. `Server=host,port;Database=mydb;User Id=sa;Password=pass;`. | ### Individual parameters[​](#individual-parameters "Direct link to Individual parameters") | Parameter Name | Description | | -------------------------------- | -------------------------------------------------------------------- | | `mssql_host` | The SQL Server host address. | | `mssql_port` | The SQL Server port number. | | `mssql_database` | The SQL Server database name. | | `mssql_username` | The SQL Server username for authentication. | | `mssql_password` | The SQL Server password for authentication. | | `mssql_encrypt` | Encryption mode: `true` (default), `false`, `require`, or `disable`. | | `mssql_trust_server_certificate` | Whether to trust the server certificate: `true` or `false`. | ## Authentication[​](#authentication "Direct link to Authentication") ### Connection string[​](#connection-string-1 "Direct link to Connection string") ``` catalogs: - from: mssql name: my_mssql params: mssql_connection_string: "Server=localhost,1433;Database=my_database;User Id=${secrets:MSSQL_USER};Password=${secrets:MSSQL_PASS};" ``` ### Individual parameters[​](#individual-parameters-1 "Direct link to Individual parameters") ``` catalogs: - from: mssql name: my_mssql params: mssql_host: localhost mssql_port: '1433' mssql_database: my_database mssql_username: ${secrets:MSSQL_USER} mssql_password: ${secrets:MSSQL_PASS} mssql_encrypt: 'true' mssql_trust_server_certificate: 'true' ``` Limitations The MSSQL Catalog Connector supports SQL Server authentication (SQL Login and Password) only. Windows Authentication and Azure Active Directory authentication are not currently supported. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). --- # MySQL Catalog Connector Connect to a [MySQL](https://www.mysql.com/) database as a catalog provider for federated SQL query. The MySQL Catalog Connector automatically discovers schemas and tables within a MySQL server and makes them available for querying in Spice. For connecting to individual MySQL tables, see the [MySQL Data Connector documentation](/docs/next/components/data-connectors/mysql). ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: mysql name: my_mysql include: - 'my_database.*' # include all tables from my_database params: mysql_connection_string: mysql://${secrets:MYSQL_USER}:${secrets:MYSQL_PASS}@localhost:3306/my_database ``` ## `from`[​](#from "Direct link to from") The `from` field specifies the catalog provider. For MySQL, use `mysql`. ## `name`[​](#name "Direct link to name") The `name` field specifies the name of the catalog in Spice. Tables from the MySQL server will be available under this catalog name. The schema hierarchy of the MySQL server is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` from any schema. Multiple `include` patterns are OR'ed together. ## `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ## `params`[​](#params "Direct link to params") Connection can be configured using a connection string or individual parameters. ### Connection string[​](#connection-string "Direct link to Connection string") | Parameter Name | Description | | ------------------------- | ------------------------------------------------------------------------- | | `mysql_connection_string` | A MySQL connection string. E.g. `mysql://user:password@host:port/dbname`. | ### Individual parameters[​](#individual-parameters "Direct link to Individual parameters") | Parameter Name | Description | | ------------------- | ---------------------------------------------------------------------------------------------- | | `mysql_host` | The MySQL host address. Default: `localhost`. | | `mysql_tcp_port` | The MySQL port number. Default: `3306`. | | `mysql_db` | The MySQL database name. | | `mysql_user` | The MySQL username for authentication. | | `mysql_pass` | The MySQL password for authentication. | | `mysql_sslmode` | The SSL mode for the connection (`disabled`, `required`, or `preferred`). Default: `required`. | | `mysql_sslrootcert` | Path to the SSL root certificate file. | ## Authentication[​](#authentication "Direct link to Authentication") ### Connection string[​](#connection-string-1 "Direct link to Connection string") ``` catalogs: - from: mysql name: my_mysql params: mysql_connection_string: mysql://${secrets:MYSQL_USER}:${secrets:MYSQL_PASS}@localhost:3306/my_database ``` ### Individual parameters[​](#individual-parameters-1 "Direct link to Individual parameters") ``` catalogs: - from: mysql name: my_mysql params: mysql_host: localhost mysql_tcp_port: '3306' mysql_db: my_database mysql_user: ${secrets:MYSQL_USER} mysql_pass: ${secrets:MYSQL_PASS} mysql_sslmode: required ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). --- # Oracle Catalog Connector Connect to an [Oracle](https://www.oracle.com/database/) database as a catalog provider for federated SQL query. The Oracle Catalog Connector automatically discovers schemas and tables within an Oracle database and makes them available for querying in Spice. Supports on-premises instances, Oracle Cloud User-Managed Databases, and Oracle Cloud Autonomous Databases (ADB). For connecting to individual Oracle tables, see the [Oracle Data Connector documentation](/docs/next/components/data-connectors/oracle). ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: oracle name: my_oracle include: - 'HR.*' # include all tables from the HR schema params: oracle_host: localhost oracle_port: '1521' oracle_username: ${secrets:ORACLE_USER} oracle_password: ${secrets:ORACLE_PASS} oracle_service_name: XEPDB1 ``` ## `from`[​](#from "Direct link to from") The `from` field specifies the catalog provider. For Oracle, use `oracle`. ## `name`[​](#name "Direct link to name") The `name` field specifies the name of the catalog in Spice. Tables from the Oracle database will be available under this catalog name. The schema hierarchy of the Oracle database is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` from any schema. Multiple `include` patterns are OR'ed together. ## `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ## `params`[​](#params "Direct link to params") Connection can be configured using a connection string or individual parameters. In both cases, `oracle_username` and `oracle_password` are required. ### Connection string[​](#connection-string "Direct link to Connection string") | Parameter Name | Description | | -------------------------- | ----------------------------------------------------------- | | `oracle_connection_string` | An Oracle connection string. E.g. `//myhost:1521/ORCLPDB1`. | | `oracle_username` | The Oracle username for authentication. | | `oracle_password` | The Oracle password for authentication. | ### Individual parameters[​](#individual-parameters "Direct link to Individual parameters") | Parameter Name | Description | | --------------------- | ------------------------------------------- | | `oracle_host` | The Oracle host address. | | `oracle_port` | The Oracle port number. Default: `1521`. | | `oracle_service_name` | The Oracle service name. Default: `XEPDB1`. | | `oracle_username` | The Oracle username for authentication. | | `oracle_password` | The Oracle password for authentication. | ### Wallet parameters (for Oracle Cloud ADB)[​](#wallet-parameters-for-oracle-cloud-adb "Direct link to Wallet parameters (for Oracle Cloud ADB)") | Parameter Name | Description | | ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `oracle_wallet` | Path to the Oracle wallet directory used for mTLS connections. Also used as the destination directory when `oracle_wallet_sso_cert` is set. Default: `.oracle`. | | `oracle_wallet_sso_cert` | Base64-encoded `cwallet.sso` (wallet auto-login certificate) content. The runtime writes the decoded certificate to `oracle_wallet`. | ## Authentication[​](#authentication "Direct link to Authentication") ### Basic authentication[​](#basic-authentication "Direct link to Basic authentication") ``` catalogs: - from: oracle name: my_oracle params: oracle_host: localhost oracle_port: '1521' oracle_username: ${secrets:ORACLE_USER} oracle_password: ${secrets:ORACLE_PASS} oracle_service_name: ORCLPDB1 ``` ### Connection string[​](#connection-string-1 "Direct link to Connection string") ``` catalogs: - from: oracle name: my_oracle params: oracle_connection_string: //myhost:1521/ORCLPDB1 oracle_username: ${secrets:ORACLE_USER} oracle_password: ${secrets:ORACLE_PASS} ``` ### Oracle Cloud Autonomous Database (mTLS)[​](#oracle-cloud-autonomous-database-mtls "Direct link to Oracle Cloud Autonomous Database (mTLS)") ``` catalogs: - from: oracle name: my_oracle params: oracle_connection_string: "(description=(retry_count=20)(retry_delay=3)(address=(protocol=tcps)(port=1522)(host=adb.us-ashburn-1.oraclecloud.com))(connect_data=(service_name=abc_high.adb.oraclecloud.com))(security=(ssl_server_dn_match=yes)))" oracle_username: ${secrets:ORACLE_USER} oracle_password: ${secrets:ORACLE_PASS} oracle_wallet: /path/to/wallet ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). --- # PostgreSQL Catalog Connector Connect to a [PostgreSQL](https://www.postgresql.org/) database as a catalog provider for federated SQL query. The PostgreSQL Catalog Connector automatically discovers schemas and tables within a PostgreSQL database and makes them available for querying in Spice. This connector also works with PostgreSQL-compatible databases such as [Amazon Redshift](https://aws.amazon.com/redshift/). For connecting to individual PostgreSQL tables, see the [PostgreSQL Data Connector documentation](/docs/next/components/data-connectors/postgres). ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: pg name: my_pg include: - 'public.*' # include all tables from the public schema params: pg_connection_string: postgresql://${secrets:PG_USER}:${secrets:PG_PASS}@localhost:5432/my_database ``` ## `from`[​](#from "Direct link to from") The `from` field specifies the catalog provider. For PostgreSQL, use `pg`. ## `name`[​](#name "Direct link to name") The `name` field specifies the name of the catalog in Spice. Tables from the PostgreSQL database will be available under this catalog name. The schema hierarchy of the PostgreSQL database is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` from any schema. Multiple `include` patterns are OR'ed together. ## `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against `schema.table` using the same glob syntax as `include`, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ``` catalogs: - from: pg name: my_pg include: - 'public.*' # Consider every table in the "public" schema... exclude: - 'public.*_audit' # ...except the audit tables. params: pg_connection_string: postgresql://${secrets:PG_USER}:${secrets:PG_PASS}@localhost:5432/my_database ``` Excluded tables are filtered out before their schema is read, so narrow patterns also reduce the metadata load each refresh places on the source — see [Catalog Refresh](#catalog-refresh). A common use is to keep tables that cannot be CDC-accelerated out of an accelerated catalog's scope, which also suppresses their skip warnings — see [Table eligibility](#table-eligibility). ## `params`[​](#params "Direct link to params") Connection can be configured using a connection string or individual parameters. ### Connection string[​](#connection-string "Direct link to Connection string") | Parameter Name | Description | | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `pg_connection_string` | A [PostgreSQL connection string](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNSTRING). E.g. `postgresql://user:password@host:port/dbname`. | ### Individual parameters[​](#individual-parameters "Direct link to Individual parameters") | Parameter Name | Description | | ---------------- | ---------------------------------------------------------------------- | | `pg_host` | The PostgreSQL host address. | | `pg_port` | The PostgreSQL port number. | | `pg_db` | The PostgreSQL database name. | | `pg_user` | The PostgreSQL username for authentication. | | `pg_pass` | The PostgreSQL password for authentication. | | `pg_sslmode` | The SSL mode for the connection (e.g. `require`, `prefer`, `disable`). | | `pg_sslrootcert` | Path to the SSL root certificate file, or inline PEM content. | ## `dataset_params`[​](#dataset_params "Direct link to dataset_params") Optional. Parameters applied to every table discovered through the catalog. | Parameter Name | Description | | ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | `unsupported_type_action` | Action to take when a discovered table contains a column of a type that cannot be mapped. One of `string` (default), `error`, `warn`, `ignore`. | The supported `unsupported_type_action` values are: * `string` — Default. Attempt to convert the unsupported type to a string (e.g. PostgreSQL `JSONB`). This matches the default of the [PostgreSQL Data Connector](/docs/next/components/data-connectors/postgres). * `error` — Fail catalog registration when an unsupported type is encountered. * `warn` — Log a warning and drop the column containing the unsupported type. * `ignore` — Drop the column containing the unsupported type without logging. An invalid value returns a configuration error rather than being silently ignored. ``` catalogs: - from: pg name: my_pg dataset_params: unsupported_type_action: warn # string (default) | error | warn | ignore ``` ## Authentication[​](#authentication "Direct link to Authentication") ### Connection string[​](#connection-string-1 "Direct link to Connection string") ``` catalogs: - from: pg name: my_pg params: pg_connection_string: postgresql://${secrets:PG_USER}:${secrets:PG_PASS}@localhost:5432/my_database ``` ### Individual parameters[​](#individual-parameters-1 "Direct link to Individual parameters") ``` catalogs: - from: pg name: my_pg params: pg_host: localhost pg_port: '5432' pg_db: my_database pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} pg_sslmode: require ``` ### Amazon Redshift[​](#amazon-redshift "Direct link to Amazon Redshift") The PostgreSQL Catalog Connector can also be used with Amazon Redshift: ``` catalogs: - from: pg name: my_redshift params: pg_connection_string: postgresql://${secrets:REDSHIFT_USER}:${secrets:REDSHIFT_PASS}@my-cluster.abc123.us-east-1.redshift.amazonaws.com:5439/my_database?sslmode=require ``` ## Discovered Relations[​](#discovered-relations "Direct link to Discovered Relations") The connector discovers the following PostgreSQL relation types from each included schema: * Base tables * Standard views * Materialized views * Foreign tables Relations are read directly from `pg_catalog.pg_class`. Only relations the connecting role holds `SELECT` privilege on (checked via `has_table_privilege`) are registered, so the catalog does not surface relations that cannot be read. For declaratively-partitioned tables (and legacy table inheritance), only the partitioned parent is registered — child partitions are not registered as separate tables. Querying the parent returns rows from every partition. ## Foreign Key Discovery[​](#foreign-key-discovery "Direct link to Foreign Key Discovery") The PostgreSQL Catalog Connector automatically discovers foreign key relationships by querying `information_schema.referential_constraints` and `key_column_usage` during catalog refresh. Discovered FK metadata is attached to each table's Arrow schema and surfaces through: * The `table_schema` tool — agents can use FK relationships to infer join paths between tables * FlightSQL `GetTables` — programmatic clients receive FK metadata in the schema response No configuration is required. If FK discovery fails for a schema (e.g., due to insufficient permissions on `information_schema`), tables are still registered without FK metadata and a warning is logged. ## Catalog Refresh[​](#catalog-refresh "Direct link to Catalog Refresh") The catalog is discovered once when the runtime starts, and then re-discovered every **60 seconds**. This interval is not currently configurable. Each cycle re-runs discovery from scratch, so schema changes at the source are picked up without restarting Spice: newly created tables and schemas appear, dropped ones disappear, and column additions, removals, and type changes are reflected when the table's schema is re-read. A source DDL change therefore becomes visible in Spice within roughly one refresh interval — up to about 60 seconds, plus the time the refresh itself takes. ### Source metadata load[​](#source-metadata-load "Direct link to Source metadata load") Because every cycle re-discovers the catalog, the metadata load on the source database scales with catalog size rather than with query volume. Each refresh issues: * One query to enumerate non-system schemas. * Per schema: one query for its relations (`pg_catalog.pg_class`), one for foreign-key constraints, and one for table and column comments. * Per selected table: one lookup to read its columns and build its Arrow schema. Use `include`/`exclude` to narrow the catalog on large databases. Filtered-out tables are skipped before their per-table schema lookup, so tighter patterns directly reduce the per-cycle query count. Note that `include`/`exclude` filter tables, not schemas — every non-system schema is still enumerated, so the per-schema queries are unaffected. ### Refresh failures[​](#refresh-failures "Direct link to Refresh failures") Failure is handled differently at initial load than on a later refresh cycle: * **Initial discovery** — the catalog does not register, its status is set to `Error`, and the load is retried with a fibonacci backoff until it succeeds. Two problems are treated as permanent misconfiguration and are *not* retried: no eligible tables for an accelerated catalog, and a replication slot already actively held by another consumer. * **Later refresh cycles** — a failure is logged at `ERROR` and the catalog keeps serving its last-known-good state until a subsequent cycle succeeds. The schema map is replaced atomically at the end of a cycle, so queries never observe a partially-refreshed catalog. Failures isolated to a single schema are handled per-schema rather than failing the whole cycle — see [Limitations](#limitations). ## Catalog-Level CDC Acceleration[​](#catalog-level-cdc-acceleration "Direct link to Catalog-Level CDC Acceleration") A PostgreSQL catalog can be accelerated as a whole. Adding an `acceleration` block bootstraps and CDC-accelerates every discovered table (subject to `include`/`exclude`) with no per-table dataset configuration. All accelerated tables share a single replication slot and publication — derived deterministically from the catalog `name`, see [Shared replication slot](#shared-replication-slot) — so the source's write-ahead log (WAL) is decoded once for the entire catalog instead of once per table. ``` catalogs: - from: pg name: my_pg include: - 'public.*' acceleration: engine: cayenne # optional; cayenne is the only supported engine refresh_mode: changes # required params: pg_connection_string: postgresql://${secrets:PG_USER}:${secrets:PG_PASS}@localhost:5432/my_database ``` ### `acceleration.engine`[​](#accelerationengine "Direct link to accelerationengine") Optional. The accelerator engine used for every table. Defaults to `cayenne`, which is currently the only supported value. ### `acceleration.refresh_mode`[​](#accelerationrefresh_mode "Direct link to accelerationrefresh_mode") Required — there is no catalog-level default. The only supported value is `changes` (CDC); there is no catalog-level `full` mode. An `acceleration` block without a `refresh_mode` is a configuration error. ### Requirements[​](#requirements "Direct link to Requirements") * **Logical replication must be enabled.** Before accelerating any table, Spice validates the PostgreSQL prerequisites CDC requires — `wal_level = logical` and the replication privilege — and fails fast with a specific, actionable error if either is missing. * **A replication slot must be available.** A CDC-accelerated catalog needs one [shared replication slot](#shared-replication-slot). When that slot does not yet exist, Spice compares the server's in-use slot count against `max_replication_slots` before trying to create it, and fails the catalog to load if the server is already at its limit: > Cannot start CDC catalog acceleration: PostgreSQL has no free replication slots (10 of 10 in use; `max_replication_slots` = 10). Drop an unused slot (inspect `pg_replication_slots`, then `SELECT pg_drop_replication_slot('');`), or raise `max_replication_slots` and restart PostgreSQL. The check is skipped when the catalog's slot **already** exists, because reusing it consumes no additional capacity — so a restart or reschedule still succeeds on a server whose slots are otherwise full. ### Shared replication slot[​](#shared-replication-slot "Direct link to Shared replication slot") The catalog's slot name is `spice_catalog_{catalog_name}_{hash}` — the sanitized catalog `name`, followed by a short hash of the full name so two long names that share a truncated prefix stay distinct, all within PostgreSQL's 63-byte identifier limit. The `spice_catalog_` prefix distinguishes it from the per-dataset `spice_` slots, so a catalog slot and a same-named dataset slot can never collide. The name is a pure function of the catalog `name`: it carries no instance, host, or process component, and Spice persists no slot identity of its own — the durable state is the PostgreSQL slot itself. Two consequences follow: * **Restarts and reschedules reuse the slot.** Restarting the runtime, or rescheduling the catalog onto a different node, recomputes the identical name and resumes from the existing slot instead of orphaning it and re-snapshotting the catalog from scratch. * **Two Spice instances cannot accelerate the same catalog.** PostgreSQL permits one consumer per replication slot, so before it starts streaming Spice checks whether the slot is already **actively** held. An absent slot is created by the per-table replication path; a present-but-inactive slot is reused; an actively-held slot fails the catalog to load with an error naming the slot and the consumer holding it. Because a slot can also read as active immediately after the runtime's *own* ungraceful exit — PostgreSQL keeps the walsender marked active until `wal_sender_timeout` elapses — Spice waits for it to free before concluding another consumer owns it. The wait is the server's `wal_sender_timeout` plus a 5-second grace, polled once a second, capped at 3 minutes; when `wal_sender_timeout` is `0` (disabled) a 90-second budget is used instead, because the server will not time out the dropped consumer on its own. warning Run only one Spice instance per accelerated PostgreSQL catalog. Because the slot name is instance-independent, a second instance configured with the same catalog `name` competes for the same slot and fails to load rather than silently splitting the change stream. ### Table eligibility[​](#table-eligibility "Direct link to Table eligibility") Each discovered table is accelerated according to its PostgreSQL [`REPLICA IDENTITY`](https://www.postgresql.org/docs/current/sql-altertable.html#SQL-ALTERTABLE-REPLICA-IDENTITY), which determines the row identity available for change data capture: * **`DEFAULT` with a primary key** — accelerated, keyed by the primary key. * **`USING INDEX`** — accelerated, keyed by the nominated unique index. * **`FULL`** — accelerated, but heavier: PostgreSQL logs the full old-row image on every `UPDATE`/`DELETE`, so a warning is logged. Prefer a primary key or `USING INDEX` where possible. * **No usable CDC key** (`NOTHING`, a keyless `DEFAULT` or `FULL`, or an unusable identity index) — **skipped with an actionable warning** and left out of the catalog's namespace. The rest of the catalog still replicates; a single ineligible table never fails the whole catalog. Use `include`/`exclude` to narrow scope and suppress the skip warning for tables you will handle another way (federation, or a per-dataset `refresh_mode: full`). Views, materialized views, and foreign tables are **not replicated**. They have no `REPLICA IDENTITY`, so they cannot be CDC-accelerated at all — unlike a table with an unusable replica identity, which is at least reported as skipped. Each one is named in a warning at load and is absent from the accelerated catalog's namespace. Query them through a non-accelerated catalog or an individual dataset instead, or exclude them via the catalog's `include`/`exclude` patterns to suppress the warning. This is the one way an accelerated PostgreSQL catalog's namespace differs from the [relation types discovered](#discovered-relations) by an un-accelerated one. The startup summary reports the accelerated tables broken down by the CDC key each one resolved to — primary key, `USING INDEX`, or `FULL` — alongside the skipped and excluded counts, and names the shared replication slot in use. If **no** table is eligible, the catalog fails to load with an `ERROR` status rather than registering an empty catalog. The error names the excluded and skipped counts and the fix. Because discovery happens at startup, an empty result is treated as a configuration problem: either every table lacks a usable CDC key, or the `include`/`exclude` patterns matched nothing. ### Acceleration metrics[​](#acceleration-metrics "Direct link to Acceleration metrics") Each catalog refresh records the current table dispositions as gauges — see [Available Metrics](/docs/next/features/observability#available-metrics): | Metric | Dimensions | Meaning | | ----------------------------------------- | --------------------- | ---------------------------------------------------------------------------------------------------------- | | `catalog_acceleration_tables` | `catalog`, `category` | Relations resolved into each disposition: `accelerated`, `skipped`, `excluded`, or `views_not_replicated`. | | `catalog_acceleration_accelerated_tables` | `catalog`, `kind` | Accelerated tables by the CDC key accelerating them: `primary_key`, `unique_index`, or `full`. | They are gauges rather than counters because each refresh re-plans the whole namespace, so a value can rise or fall. ### Behavior[​](#behavior "Direct link to Behavior") * Per-table-only acceleration concepts (`primary_key`, `on_conflict`, `indexes`, and other per-dataset overrides) are intentionally not configurable at the catalog level — they remain exclusively on an individual dataset's own `acceleration` block. * While a table's acceleration is still bootstrapping, that table is reported as not-yet-present rather than being served through the source, so queries do not transparently fall back to the un-accelerated PostgreSQL table. ## Limitations[​](#limitations "Direct link to Limitations") warning * **`include`/`exclude` filter tables, not schemas.** Both are matched against `schema.table`. All non-system schemas are still enumerated as (possibly empty) schemas even when no tables match. * **Partial discovery failures.** If discovery of a schema's tables fails, that schema is skipped with a warning rather than aborting the whole catalog load. On refresh, a transient per-schema failure falls back to the last-known-good state for that schema, so intermittent errors do not cause catalog flapping; a schema is only dropped for a cycle if it has never refreshed successfully. Total connectivity loss (a failed `list_schemas`) still fails hard. * **Amazon Redshift — datashare and external tables are not discovered.** Discovery reads Redshift's *local* catalog only (`information_schema.schemata` and `pg_catalog.pg_class`). Schemas and tables consumed from a datashare, and external schemas and tables (Redshift Spectrum), are absent from the local `pg_catalog` — Redshift exposes them only through its `svv_all_schemas` / `svv_all_tables` views, which the Catalog Connector does not query. Those relations do not appear in the catalog; register them as individual datasets with the [Redshift Data Connector](/docs/next/components/data-connectors/redshift) instead ([#12109](https://github.com/spiceai/spiceai/issues/12109)). * **Amazon Redshift — metadata coverage.** Redshift is supported over the PostgreSQL wire protocol, but its `pg_catalog` coverage is partial and it does not enforce foreign keys, so table/column comment and foreign-key metadata are often unavailable. Tables present in the local catalog are still registered. * **Read-only; per-table acceleration not configurable.** Catalog tables are read-only, and per-table `acceleration` blocks cannot be set on individually discovered tables. Catalog-wide CDC acceleration *is* available for PostgreSQL — see [Catalog-Level CDC Acceleration](#catalog-level-cdc-acceleration). ## Cookbook[​](#cookbook "Direct link to Cookbook") There is a [cookbook recipe](https://github.com/spiceai/cookbook/tree/trunk/catalogs/postgres) demonstrating the PostgreSQL Catalog Connector with the TPC-H dataset. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). --- # Snowflake Catalog Connector Connect to a [Snowflake](https://www.snowflake.com/) database as a catalog provider for federated SQL query. The Snowflake Catalog Connector automatically discovers schemas and tables within a Snowflake database and makes them available for querying in Spice. For connecting to individual Snowflake tables, see the [Snowflake Data Connector documentation](/docs/next/components/data-connectors/snowflake). ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: snowflake:MY_DATABASE name: my_snowflake include: - 'MY_SCHEMA.*' # include all tables from MY_SCHEMA params: snowflake_account: myaccount snowflake_username: ${secrets:SNOWFLAKE_USERNAME} snowflake_password: ${secrets:SNOWFLAKE_PASSWORD} snowflake_warehouse: COMPUTE_WH snowflake_role: ACCOUNTADMIN ``` ## `from`[​](#from "Direct link to from") The `from` field specifies the Snowflake database to use as a catalog. Use `snowflake:`, where `database_name` is the name of the Snowflake database. Hint Unquoted identifiers in Snowflake are stored as uppercase. Use uppercase database names in the `from` field. See [Snowflake Identifier Resolution](https://docs.snowflake.com/en/sql-reference/identifiers-syntax#label-identifier-casing). ## `name`[​](#name "Direct link to name") The `name` field specifies the name of the catalog in Spice. Tables from the Snowflake database will be available under this catalog name. The schema hierarchy of the Snowflake database is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` from any schema. Multiple `include` patterns are OR'ed together. ## `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ## `params`[​](#params "Direct link to params") | Parameter Name | Description | | ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `snowflake_account` | The Snowflake [account identifier](https://docs.snowflake.com/en/user-guide/admin-account-identifier). | | `snowflake_username` | The Snowflake username for authentication. | | `snowflake_password` | The Snowflake password for authentication. Use the [secret replacement syntax](/docs/next/components/secret-stores). | | `snowflake_warehouse` | The Snowflake [warehouse](https://docs.snowflake.com/en/user-guide/warehouses-tasks) to use for queries. | | `snowflake_role` | The Snowflake [role](https://docs.snowflake.com/en/user-guide/security-access-control-overview) to use. | | `snowflake_auth_type` | Authentication type. Accepts `password` (or `snowflake`) and `keypair` (or `snowflake_jwt`); matched case-insensitively. Defaults to password authentication unless only key-pair credentials (`snowflake_private_key` or `snowflake_private_key_path`) are provided, in which case key-pair is selected automatically. | | `snowflake_private_key` | The private key content for key pair authentication. Used when `auth_type` is `keypair`. | | `snowflake_private_key_path` | Path to a private key file for key pair authentication. Used when `auth_type` is `keypair`. | | `snowflake_private_key_passphrase` | Passphrase for the private key file, if encrypted. | ## Authentication[​](#authentication "Direct link to Authentication") ### Password authentication (default)[​](#password-authentication-default "Direct link to Password authentication (default)") ``` catalogs: - from: snowflake:MY_DATABASE name: my_snowflake params: snowflake_account: myaccount snowflake_username: ${secrets:SNOWFLAKE_USERNAME} snowflake_password: ${secrets:SNOWFLAKE_PASSWORD} ``` ### Key pair authentication[​](#key-pair-authentication "Direct link to Key pair authentication") ``` catalogs: - from: snowflake:MY_DATABASE name: my_snowflake params: snowflake_account: myaccount snowflake_username: ${secrets:SNOWFLAKE_USERNAME} snowflake_auth_type: keypair snowflake_private_key_path: /path/to/rsa_key.p8 snowflake_private_key_passphrase: ${secrets:SNOWFLAKE_KEY_PASSPHRASE} ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). --- # Spice.ai Catalog Connector Query datasets hosted in the [Spice.ai Cloud Platform](https://spice.ai). Discover available public datasets on [Spicerack](https://spicerack.org). ## Configuration[​](#configuration "Direct link to Configuration") Create a [Spice.ai Cloud Platform](https://spice.ai) account and login with the CLI using `spice login`. Example: ``` catalogs: - from: spice.ai:demo-org/tpch # Load tables from the `demo-org` organization's `tpch` app name: marketplace # Tables will be available in the "marketplace" catalog params: spiceai_api_key: ${secrets:SPICEAI_API_KEY} # Spice.ai API key spiceai_region: us-east-1 # Required: Spice Cloud region include: - "tpch.part*" # include only the tables from the "tpch" schema and that start with "part" - "tpch.supplier" # also include the "supplier" table ``` This configuration would load the following tables: ``` sql> show tables; +---------------+--------------+----------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------+----------------+------------+ | marketplace | tpch | part | BASE TABLE | | marketplace | tpch | partsupp | BASE TABLE | | marketplace | tpch | supplier | BASE TABLE | +---------------+--------------+----------------+------------+ ``` ## `from`[​](#from "Direct link to from") The `from` field specifies which organization and application to load tables from. The format is: ``` spice.ai//[/] ``` * `organization`: The Spice.ai organization that owns the application. * `application`: The specific application to load tables from. * `catalog_name` (optional): A specific catalog within the application. Defaults to `spice` if not specified. For example: 1. Load the default catalog from the "tpch" application in the "demo-org" organization ``` catalogs: - from: spice.ai/demo-org/tpch name: marketplace ``` This will create tables like: ``` SHOW TABLES; +---------------+--------------+----------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------+----------------+------------+ | marketplace | tpch | region | BASE TABLE | | marketplace | tpch | orders | BASE TABLE | | marketplace | tpch | part | BASE TABLE | | marketplace | tpch | supplier | BASE TABLE | | marketplace | tpch | customer | BASE TABLE | | marketplace | tpch | partsupp | BASE TABLE | | marketplace | tpch | lineitem | BASE TABLE | | marketplace | tpch | nation | BASE TABLE | +---------------+--------------+----------------+------------+ ``` 2. Load a specific catalog named "custom\_catalog" from an application ``` catalogs: - from: spice.ai/demo-org/tpch/custom_catalog name: marketplace ``` If the remote application has tables in `custom_catalog` like: ``` SHOW TABLES; +------------------+--------------+----------------+------------+ | table_catalog | table_schema | table_name | table_type | +------------------+--------------+----------------+------------+ | custom_catalog | ice_schema1 | table1 | BASE TABLE | | custom_catalog | ice_schema1 | table2 | BASE TABLE | | custom_catalog | ice_schema2 | table1 | BASE TABLE | +------------------+--------------+----------------+------------+ ``` They will be available locally as: ``` SHOW TABLES; +------------------+--------------+----------------+------------+ | table_catalog | table_schema | table_name | table_type | +------------------+--------------+----------------+------------+ | marketplace | ice_schema1 | table1 | BASE TABLE | | marketplace | ice_schema1 | table2 | BASE TABLE | | marketplace | ice_schema2 | table1 | BASE TABLE | +------------------+--------------+----------------+------------+ ``` ## `name`[​](#name "Direct link to name") The name field defines what catalog name the tables will be available under in the local Spice instance. For example, with the following configuration: ``` from: spice.ai/demo-org/tpch name: marketplace ``` Then tables that exist in the remote application as: ``` spice # Default catalog |- schema1 |- table1 |- table2 ``` Will be available locally as: ``` marketplace |- schema1 |- table1 |- table2 ``` Queries are run against the `marketplace` catalog, like `SELECT * FROM marketplace.schema1.table1`. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables: * `schema_name.*` - Include all tables from a specific schema * `*.table_name` - Include all tables with a specific name from any schema * `schema_name.table_name` - Include a specific table from a schema * `schema_name.table_prefix*` - Include all tables in a schema that start with a prefix Multiple include patterns can be specified and are OR'ed together. For example: ``` include: - "tpch.part*" # Include all tables from the "tpch" schema that start with "part" - "tpch.supplier" # Include the "supplier" table from the "tpch" schema ``` ## `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ## `params`[​](#params "Direct link to params") ### Authentication[​](#authentication "Direct link to Authentication") | Parameter Name | Description | | ----------------- | -------------------------------------------------------------------------------- | | `spiceai_api_key` | API key from the Spice.ai Cloud Platform. Takes precedence over `spiceai_token`. | | `spiceai_token` | Legacy alias for `spiceai_api_key`. Used only if `spiceai_api_key` is not set. | ### Region[​](#region "Direct link to Region") | Parameter Name | Description | | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `spiceai_region` | The Spice Cloud region to connect to (e.g. `us-east-1`). **Required** unless `spiceai_http_endpoint` is set. To list available regions, run `spice cloud regions`. | ### Endpoint Overrides[​](#endpoint-overrides "Direct link to Endpoint Overrides") These parameters override the default endpoints derived from `spiceai_region`. They are typically only needed for development or custom deployments. | Parameter Name | Description | | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `spiceai_http_endpoint` | Custom HTTP endpoint for the Spice Cloud catalog API (Iceberg REST catalog compatible). If set, `spiceai_region` is not required and region validation is skipped. | | `spiceai_endpoint` | Custom Arrow Flight endpoint for query execution. Takes precedence over `spiceai_flight_endpoint`. | | `spiceai_flight_endpoint` | Legacy alias for `spiceai_endpoint`. Used only if `spiceai_endpoint` is not set. | ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Spice Cloud as a catalog connector in Spice. [Spice Cloud Catalog Connector](https://github.com/spiceai/cookbook/tree/trunk/catalogs/spiceai#readme) --- # Unity Catalog Catalog Connector Connect to a [Unity Catalog](https://www.unitycatalog.io/) as a catalog provider for federated SQL query against [Delta Lake](https://delta.io/) tables. ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: unity_catalog:https://my_unity_catalog_host.com/api/2.1/unity-catalog/catalogs/my_catalog name: uc include: - "*.my_table" dataset_params: # delta_lake S3 parameters unity_catalog_aws_region: us-west-2 unity_catalog_aws_access_key_id: ${secrets:aws_access_key_id} unity_catalog_aws_secret_access_key: ${secrets:aws_secret_access_key} unity_catalog_aws_endpoint: s3.us-west-2.amazonaws.com ``` ## `from`[​](#from "Direct link to from") The `from` field is used to specify the catalog provider. For Unity Catalog, use `unity_catalog:`. The `catalog_path` is the URL to the [`getCatalog`](https://github.com/unitycatalog/unitycatalog/blob/main/api/Apis/CatalogsApi) endpoint of the Unity Catalog API. It should be formatted as `https:///api/2.1/unity-catalog/catalogs/`. ## `name`[​](#name "Direct link to name") The `name` field is used to specify the name of the catalog in Spice. The schema hierarchy of the external catalog is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. ## `exclude`[​](#exclude "Direct link to exclude") Optional. Use the `exclude` field to omit tables that would otherwise be included. It is matched against the same table name as `include`, using the same glob syntax, and multiple `exclude` patterns are OR'ed together. `exclude` takes precedence over `include`: a table is registered only when it matches `include` (or no `include` is set) **and** matches no `exclude` pattern. ## `params`[​](#params "Direct link to params") The `params` field is used to configure the connection to the Unity Catalog. The following parameters are supported: * `unity_catalog_token`: The [personal access token](https://docs.unitycatalog.io/server/auth/#use-admin-token-to-verify-admin-user-is-in-local-database) used to authenticate against the Unity Catalog API. * `unity_catalog_credential_vending`: When set to `enabled`, short-lived storage credentials for each table are fetched from the Unity Catalog [credential vending](https://docs.databricks.com/api/workspace/temporarytablecredentials) API instead of using the static storage credentials in `dataset_params`. Defaults to `disabled`. Works with both Databricks Unity Catalog and OSS Unity Catalog. When credential vending is enabled, the static object-store credentials below (`unity_catalog_aws_*`, `unity_catalog_azure_*`, `unity_catalog_google_*`) are not required: ``` catalogs: - from: unity_catalog:https:///api/2.1/unity-catalog/catalogs/my_catalog name: uc params: unity_catalog_token: ${secrets:UC_TOKEN} unity_catalog_credential_vending: enabled ``` ## `dataset_params`[​](#dataset_params "Direct link to dataset_params") The `dataset_params` field is used to configure the dataset-specific parameters for the catalog. ### Unity catalog object store parameters[​](#unity-catalog-object-store-parameters "Direct link to Unity catalog object store parameters") #### AWS S3[​](#aws-s3 "Direct link to AWS S3") * `unity_catalog_aws_region`: The AWS region for the S3 object store. E.g. `us-west-2`. * `unity_catalog_aws_access_key_id`: The access key ID for the S3 object store. * `unity_catalog_aws_secret_access_key`: The secret access key for the S3 object store. * `unity_catalog_aws_session_token`: Optional. The AWS session token for the S3 object store. Required with temporary (STS) credentials. * `unity_catalog_aws_endpoint`: The endpoint for the S3 object store. E.g. `s3.us-west-2.amazonaws.com`. * `unity_catalog_aws_allow_http`: Enables insecure HTTP connections to the AWS endpoint, useful for S3-compatible servers (e.g. MinIO). Defaults to `false`. #### Azure Blob[​](#azure-blob "Direct link to Azure Blob") Note One of the following auth values must be provided for Azure Blob: * `unity_catalog_azure_storage_account_key`, * `unity_catalog_azure_storage_client_id` and `unity_catalog_azure_storage_client_secret`, or * `unity_catalog_azure_storage_sas_key`. - `unity_catalog_azure_storage_account_name`: The Azure Storage account name. - `unity_catalog_azure_storage_account_key`: The Azure Storage master key for accessing the storage account. - `unity_catalog_azure_storage_client_id`: The service principal client id for accessing the storage account. - `unity_catalog_azure_storage_client_secret`: The service principal client secret for accessing the storage account. - `unity_catalog_azure_storage_sas_key`: The shared access signature key for accessing the storage account. - `unity_catalog_azure_storage_endpoint`: The endpoint for the Azure Blob storage account. #### Google Storage (GCS)[​](#google-storage-gcs "Direct link to Google Storage (GCS)") * `unity_catalog_google_service_account`: Filesystem path to the Google service account JSON key file. ## Limitations[​](#limitations "Direct link to Limitations") * Unity Catalog does not support reading Delta tables with the `V2Checkpoint` feature enabled. To use the Unity Catalog connector with such tables, drop the `V2Checkpoint` feature by executing the following command: ``` ALTER TABLE DROP FEATURE v2Checkpoint [TRUNCATE HISTORY]; ``` For more details on dropping Delta table features, refer to the official documentation: [Drop Delta table features](https://docs.delta.io/latest/delta-drop-feature.html) --- # Unity Catalog Catalog Connector Deployment Guide Production operating guide for the Unity Catalog catalog connector — discovering Databricks Unity Catalog tables and federating them through Spice. For Databricks-specific operational concerns (SQL Warehouse resilience, metrics, permissions flow as applied to Databricks workspaces), see the [Databricks Deployment Guide](/docs/next/components/data-connectors/databricks/deployment) — the Unity Catalog logic described there applies directly when the catalog connector targets a Databricks workspace. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") | Parameter | Description | | --------------------- | --------------------------------------------------------------------------------- | | `unity_catalog_token` | Bearer token for the Unity Catalog API. Use `${secrets:...}` from a secret store. | The catalog URL must match the pattern `https:///api/2.1/unity-catalog/catalogs/` and is parsed into the endpoint and catalog identifier at startup. Mismatched URLs are rejected as configuration errors. The token is optional — when unset, the catalog connector issues unauthenticated requests, suitable for locally-hosted Unity Catalog deployments (OSS UC) with permissive access. For Databricks workspaces, the token is always required. Secrets must be sourced from a [secret store](/docs/next/components/secret-stores) in production. Rotate tokens from the UC / Databricks console and update the secret store. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### HTTP Retry Policy[​](#http-retry-policy "Direct link to HTTP Retry Policy") The Unity Catalog client uses the shared `resilient_http` helper with these defaults: * Maximum retries: **3** * Backoff: fibonacci * Retriable conditions: HTTP `408`, `429`, `5xx`, and transient network errors (connect, timeout) * Respects `Retry-After`, `retry-after-ms`, `x-retry-after-ms` headers * Maximum backoff: 300 seconds These are not exposed as user-tunable parameters on the Unity Catalog connector itself. ### Discovery Concurrency[​](#discovery-concurrency "Direct link to Discovery Concurrency") The connector fans out schema and table enumeration with bounded concurrency to avoid thundering-herd on the UC API: * Schema refresh: up to **5** concurrent requests (`buffer_unordered(5)`) * Permission checks: up to **5** concurrent requests (`buffer_unordered(5)`) For catalogs with thousands of tables, initial discovery can take minutes while the connector respects these limits. ## Table Type and Permission Handling[​](#table-type-and-permission-handling "Direct link to Table Type and Permission Handling") ### Table Type Filtering[​](#table-type-filtering "Direct link to Table Type Filtering") | Table Type | Supported | Notes | | ------------------- | --------- | -------------------------------------- | | `MANAGED` | Yes | Standard Delta tables | | `EXTERNAL` | Yes | Tables with external storage locations | | `FOREIGN` | Yes | Lakehouse Federation foreign tables | | `MATERIALIZED_VIEW` | Yes | Materialized views | | `VIEW` | No | Skipped during discovery | | `STREAMING_TABLE` | No | Skipped during discovery | Unsupported table types are skipped during catalog discovery. When referenced directly, an error is returned. ### Effective Permissions[​](#effective-permissions "Direct link to Effective Permissions") Before creating a table provider, the connector checks permissions via `GET /api/2.1/unity-catalog/effective-permissions/table/{catalog.schema.table}`. The following privileges grant read access: * `SELECT` * `ALL_PRIVILEGES` / `ALL PRIVILEGES` * `OWNER` / `OWNERSHIP` **Behavior**: * **Discovery**: Tables without read permission are skipped. * **Direct reference**: An `InsufficientPermissions` error is returned. * **Foreign tables**: The precheck is skipped (`requires_read_permission_validation = false`) because Lakehouse Federation access can be valid when the UC effective-permissions endpoint does not report a table-level privilege. Access is still enforced by Databricks at query time. * **Graceful degradation**: If the UC API is unreachable or returns an error for the permissions endpoint, discovery proceeds with a warning — table providers are still created, and any per-query authorization failures surface at query time. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Initial discovery**: Scales with the number of schemas × tables. Bounded concurrency caps throughput; plan 5–30 minutes for catalogs with thousands of tables on a cold start. * **Refresh**: Catalog refresh re-enumerates schemas and tables at the configured interval. For very large catalogs, refresh less frequently (every few hours) unless schemas change rapidly. * **Permission-check cost**: One API call per table. The buffer of 5 caps concurrency. ## Metrics[​](#metrics "Direct link to Metrics") The Unity Catalog connector does not currently register UC-specific OpenTelemetry metric instruments. When used via the Databricks connector, the shared SQL Warehouse and UC spans produce task-history records that can be aggregated for operational insight. Monitor via: * Spice query execution metrics (`query_duration_ms`, `query_returned_rows`) from `runtime.metrics`. * Task-history spans listed below. * Databricks / UC workspace audit logs for API-level visibility. See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. ## Task History[​](#task-history "Direct link to Task History") Unity Catalog operations emit the following [task history](/docs/next/reference/task_history) spans: | Span | Input | Description | | ------------------------------ | -------------------------- | ---------------------------------------- | | `uc_get_table` | Fully-qualified table name | Fetch table metadata from Unity Catalog. | | `uc_get_catalog` | Catalog ID | Fetch catalog metadata. | | `uc_list_schemas` | Catalog ID | List schemas in a catalog. | | `uc_list_tables` | `catalog_id.schema_name` | List tables in a schema. | | `uc_get_effective_permissions` | Fully-qualified table name | Check effective permissions for a table. | ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **VIEW and STREAMING\_TABLE are skipped**: Only queryable table types are exposed. * **No UC write-back**: The connector is read-only; writes to UC are not supported through Spice. * **HTTP retry/concurrency parameters not exposed**: The resilient-HTTP defaults (3 retries, fibonacci backoff, concurrency 5) are not currently user-tunable on the UC connector. * **Graceful degradation on permission-endpoint failures**: If UC effective-permissions is unreachable, Spice proceeds; authorization errors surface at query time rather than discovery time. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ------------------------------------------------------ | ----------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- | | `401 Unauthorized` on catalog list | Missing, expired, or wrong-workspace token. | Regenerate token in UC / Databricks; update secret store. | | Table visible in UC but missing from the Spice catalog | Table type is VIEW / STREAMING\_TABLE or permissions were denied. | Confirm table type is supported and that the principal has `SELECT` (or equivalent). | | `InsufficientPermissions` on direct table reference | Role lacks read privilege on the table. | Grant `SELECT` on the table in UC. | | Slow catalog discovery on thousands of tables | Bounded concurrency + permission checks per table. | Expected behavior; schedule discovery during low-traffic windows and cache via accelerated datasets. | | Tables from a Lakehouse Federation source missing | FOREIGN precheck passed but Databricks denied at query time. | Verify the Databricks workspace has federation privileges granted to the principal. | --- # Data Accelerators Data sourced by Data Connectors can be locally materialized and accelerated using a Data Accelerator. A Data Accelerator queries/fetches data from a connected data source and stores/updates it locally in an embedded acceleration engine, such as Spice Cayenne, DuckDB, or SQLite. To set data refresh behavior, such as refreshing data on an interval, see [Data Refresh](/docs/next/features/data-acceleration/data-refresh). Dataset acceleration is enabled by setting the acceleration configuration: ``` datasets: - name: accelerated_dataset acceleration: enabled: true ``` For the complete reference specification, see [datasets](/docs/next/reference/spicepod/datasets). By default, datasets are locally materialized using in-memory Arrow records. ## Supported Data Accelerators[​](#supported-data-accelerators "Direct link to Supported Data Accelerators") | Name | Description | Status | Engine Modes | | ---------- | --------------------------------------------------------------------------------------------- | ----------------- | ---------------------------------------------- | | `cayenne` | [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) | Stable | `memory`, `file`, `file_create`, `file_update` | | `arrow` | In-Memory Arrow Records | Stable | `memory` | | `duckdb` | Embedded [DuckDB](/docs/next/components/data-accelerators/duckdb) | Stable | `memory`, `file`, `file_create`, `file_update` | | `postgres` | Attached [PostgreSQL](/docs/next/components/data-accelerators/postgres) (Spice.ai Enterprise) | Release Candidate | N/A | | `sqlite` | Embedded [SQLite](/docs/next/components/data-accelerators/sqlite) | Release Candidate | `memory`, `file`, `file_create`, `file_update` | | `turso` | Embedded [Turso](/docs/next/components/data-accelerators/turso) | Beta | `memory`, `file`, `file_create`, `file_update` | ## Choosing an Accelerator[​](#choosing-an-accelerator "Direct link to Choosing an Accelerator") Select the appropriate accelerator based on dataset size, query patterns, and resource constraints: | Use Case | Recommended Accelerator | Rationale | | ------------------------------------------ | ----------------------- | ------------------------------------------------------------------------------------------------ | | Small datasets (under 1 GB), maximum speed | `arrow` | In-memory storage provides lowest latency | | Small datasets (1-10 GB), complex SQL | `duckdb` | Mature SQL support with memory management | | Datasets 10 GB and above (up to 1+ TB) | `cayenne` | Needs 1/3 to 1/2 the memory of `duckdb`; Vortex columnar format scales beyond single-file limits | | Point lookups on large datasets | `cayenne` | Vortex provides 100x faster random access vs Parquet | | Simple queries, low resource usage | `sqlite` | Lightweight, minimal overhead | | Async operations, concurrent workloads | `turso` | Native async support, modern connection pooling | | External database integration | `postgres` | Use existing PostgreSQL infrastructure | ### Spice Cayenne vs DuckDB[​](#spice-cayenne-vs-duckdb "Direct link to Spice Cayenne vs DuckDB") Both [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) and [DuckDB](/docs/next/components/data-accelerators/duckdb) support file-based acceleration, but differ in architecture and performance characteristics. **Spice Cayenne is recommended for any dataset of 10 GB or larger**, because of DuckDB's memory requirements: Cayenne typically needs one-third to one-half the memory of the DuckDB accelerator for the same dataset. **Choose Spice Cayenne when:** * Datasets are 10 GB or larger * Memory headroom is constrained, or the deployment must run in a smaller container * Multi-file data ingestion is required (e.g., partitioned S3 data) * Workloads benefit from Vortex's [10-20x faster scans](https://bench.vortex.dev) * Point lookups and random access patterns are common ([100x faster than Parquet](https://bench.vortex.dev)) **Choose DuckDB when:** * Datasets are under 10 GB * Complex SQL features are required (window functions, CTEs) * Existing DuckDB tooling integration is beneficial * Explicit index control is required ## Data Types[​](#data-types "Direct link to Data Types") Data Accelerators may not support all possible Apache Arrow data types. For complete compatibility, see [specifications](/docs/next/reference/datatypes/accelerators). Memory Considerations When accelerating a dataset using `mode: memory` (the default), some or all of the dataset is loaded into memory. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. In-memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/next/components/data-accelerators/duckdb), [`sqlite`](/docs/next/components/data-accelerators/sqlite), and [`turso`](/docs/next/components/data-accelerators/turso) accelerators by specifying `mode: file`. ## Schema Handling[​](#schema-handling "Direct link to Schema Handling") Data accelerators store the schema that Spice infers from the data source at startup. This schema is fixed for the lifetime of the runtime process and defines the column names, data types, and nullability of the accelerated table. If the source schema changes while the runtime is running (for example, new columns are added or data types change), subsequent data refreshes into the accelerator will fail because the incoming data no longer matches the schema of the accelerated table. Restart the runtime to re-infer the schema and re-initialize the accelerated table. For details on how schema inference works per connector and recommendations for managing schema drift, see [Schema Inference](/docs/next/components/data-connectors#schema-inference). ## Data Accelerator Docs[​](#data-accelerator-docs "Direct link to Data Accelerator Docs") ## [🗃Spice Cayenne Data Accelerator](/docs/next/components/data-accelerators/cayenne) [2 items](/docs/next/components/data-accelerators/cayenne) ## [🗃In-Memory Arrow Data Accelerator](/docs/next/components/data-accelerators/arrow) [1 item](/docs/next/components/data-accelerators/arrow) ## [🗃DuckDB Data Accelerator](/docs/next/components/data-accelerators/duckdb) [1 item](/docs/next/components/data-accelerators/duckdb) ## [🗃SQLite Data Accelerator](/docs/next/components/data-accelerators/sqlite) [1 item](/docs/next/components/data-accelerators/sqlite) ## [🗃PostgreSQL Data Accelerator](/docs/next/components/data-accelerators/postgres) [1 item](/docs/next/components/data-accelerators/postgres) ## [📄️Turso Data Accelerator](/docs/next/components/data-accelerators/turso) [Turso (libSQL) Data Accelerator Documentation](/docs/next/components/data-accelerators/turso) ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") * [Performance Tuning](/docs/next/reference/performance-tuning) - Comprehensive optimization guide * [Managing Memory Usage](/docs/next/reference/memory) - Memory configuration reference * [Data Refresh](/docs/next/features/data-acceleration/data-refresh) - Refresh mode configuration * [Indexes](/docs/next/features/data-acceleration/indexes) - Index configuration for DuckDB, SQLite, and Turso --- # In-Memory Arrow Data Accelerator The In-Memory Arrow Data Accelerator is the default data accelerator in Spice. It uses Apache Arrow to store data in-memory for fast access and query performance. ## Configuration[​](#configuration "Direct link to Configuration") To use the In-Memory Arrow Data Accelerator, no additional configuration is required beyond enabling acceleration. Example: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: enabled: true ``` However Arrow can be specified explicitly using `arrow` as the `engine` for acceleration. ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: enabled: true engine: arrow ``` ## Hash Index[​](#hash-index "Direct link to Hash Index") Experimental Hash index is an experimental feature available in Spice v1.11.0-rc.2 and later. The In-Memory Arrow Data Accelerator supports an optional [hash index](/docs/next/features/data-acceleration/hash-index) for O(1) point lookups on primary key columns. Hash indexing activates automatically when a `primary_key` (or a secondary [`indexes`](/docs/next/features/data-acceleration/indexes) entry) is configured — no additional parameter is required: ``` datasets: - from: s3://bucket/orders.parquet name: orders acceleration: engine: arrow primary_key: order_id ``` The legacy `hash_index: enabled` parameter is accepted but no longer activates indexing on its own; when set, the runtime logs a warning and falls back to the automatic rules above. See [Hash Index](/docs/next/features/data-acceleration/hash-index) for configuration details, supported data types, and performance characteristics. ## Limitations[​](#limitations "Direct link to Limitations") * The In-Memory Arrow Data Accelerator does not support persistent storage. Data is stored in-memory and will be lost when the Spice runtime is stopped. * The In-Memory Arrow Data Accelerator does not support `Decimal256` (76 digits), as it exceeds Arrow's maximum Decimal width of 38 digits. * The In-Memory Arrow Data Accelerator does not support traditional [indexes](/docs/next/features/data-acceleration/indexes), but does support [hash indexes](/docs/next/features/data-acceleration/hash-index) (experimental) for point lookups. * The In-Memory Arrow Data Accelerator only supports primary-key [constraints](/docs/next/features/data-acceleration/constraints), not `unique` constraints. * With Arrow acceleration, mathematical operations like `value1 / value2` are treated as integer division if the values are integers. For example, `1 / 2` will result in 0 instead of the expected 0.5. Use casting to FLOAT to ensure conversion to a floating-point value: `CAST(1 AS FLOAT) / CAST(2 AS FLOAT)` (or `CAST(1 AS FLOAT) / 2`). Memory Considerations When accelerating a dataset using the In-Memory Arrow Data Accelerator, some or all of the dataset is loaded into memory. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. In-memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/next/components/data-accelerators/duckdb) and [`sqlite`](/docs/next/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure In-Memory Arrow as data accelerator in Spice. [In-Memory Arrow Data Accelerator](https://github.com/spiceai/cookbook/tree/trunk/arrow#readme) --- # Arrow Data Accelerator Deployment Guide Production operating guide for the Arrow in-memory data accelerator covering memory sizing, optional hash indexes, and observability. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The Arrow accelerator is an in-process, in-memory engine. There is no external storage and no authentication or secret management required. ## Resilience & Durability[​](#resilience--durability "Direct link to Resilience & Durability") The Arrow accelerator is **not durable**. Data is held in RAM and is lost on process restart; every restart re-materializes the dataset from the source connector. * **Crash recovery**: None — on restart, the dataset is refreshed from scratch. * **File modes**: File-mode acceleration is rejected at startup; Arrow is memory-only. Use [DuckDB](/docs/next/components/data-accelerators/duckdb/deployment), [SQLite](/docs/next/components/data-accelerators/sqlite/deployment), [PostgreSQL](/docs/next/components/data-accelerators/postgres/deployment), or [Cayenne](/docs/next/components/data-accelerators/cayenne/deployment) when durability or spill is required. * **Concurrency**: Arrow reads are lock-free. Refresh cadence is controlled by the runtime refresh semaphore, not by the accelerator itself. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Memory**: Plan for 1.0–1.5× the raw row-oriented size of the source data, plus overhead for string dictionaries. Use the source connector's schema and row count to estimate. * **Hash index**: Optional. Activated automatically when a `primary_key` (or secondary `indexes` entry) is configured, building a hash map over the indexed columns. Build time scales linearly with rows. The index stores no key bytes — only a 16-byte slot per entry (8-byte hash + packed 8-byte row location) plus roughly 1.25 bytes of bloom filter — so budget from the entry count, not the key width. Reported usage tracks allocated slot capacity, which exceeds the live entry count, so plan above the \~17 bytes/row floor. * **Startup cost**: Full-dataset materialization happens on startup. For tables larger than \~1 GB, consider a durable accelerator to avoid repeated full refresh on every restart. ## Metrics[​](#metrics "Direct link to Metrics") Generic acceleration metrics are available with the `dataset_acceleration_` prefix. Hash-index operations emit dedicated metrics when the index is enabled: | Metric | Type | Description | | ------------------------------ | --------- | ---------------------------------------------- | | `hash_index_builds` | Counter | Total hash-index builds (one per refresh). | | `hash_index_build_duration_ms` | Histogram | Time to build the hash index. | | `hash_index_entries` | Histogram | Number of entries in the index. | | `hash_index_memory_bytes` | Histogram | Approximate memory footprint of the index. | | `hash_index_lookups` | Counter | Total hash-index lookups performed by queries. | | `hash_index_lookup_rows` | Counter | Total rows returned via hash-index lookups. | See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting metrics. Refresh metrics are described in [Acceleration](/docs/next/features/data-acceleration). ## Task History[​](#task-history "Direct link to Task History") Arrow acceleration operations (refresh, query) participate in [task history](/docs/next/reference/task_history) through the shared acceleration spans (`accelerated_table_refresh`, `sql_query`). No Arrow-specific spans are emitted — the accelerator is a thin wrapper over Arrow memory. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **No persistence**: Every restart refreshes from the source. * **No traditional indexes**: Arrow does not support B-tree indexes. Hash index provides point-lookup acceleration but not range or sort-order optimization. * **Only primary-key hash index**: The hash index requires a `primary_key` constraint; `unique` constraints alone do not enable the index. * **Memory pressure**: If the dataset exceeds available RAM, the runtime will OOM; no spill-to-disk mechanism exists in the Arrow accelerator itself. * **`partition_by`**: Not applicable — Arrow accelerator holds a single in-memory representation. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ------------------------------------------- | ------------------------------------------ | ---------------------------------------------------------------------------------------------------------- | | OOM on refresh | Source dataset larger than RAM. | Switch to a durable accelerator (DuckDB / SQLite / Cayenne) that supports spill to disk. | | Long startup time | Full-dataset refresh runs on boot. | Switch to a durable accelerator so refresh is incremental, not full, on restart. | | `hash_index` ignored | No primary-key constraint on the dataset. | Add `primary_key:` to the dataset definition; hash index activates automatically. | | Query slow for point lookups | No primary key/index, or wrong key column. | Add a `primary_key:` (or secondary `indexes:` entry); ensure the query filter matches the indexed columns. | | Accelerator refuses to start with file mode | Arrow rejects file-mode acceleration. | Switch `engine:` to `duckdb`, `sqlite`, `postgres`, or `cayenne`. | --- # Spice Cayenne Data Accelerator Spice Cayenne is a data acceleration engine designed for high-performance, scalable query on large-scale datasets. Built on [Vortex](https://github.com/vortex-data/vortex), a high-performance columnar file format, Spice Cayenne combines columnar storage with in-process metadata management to provide fast query performance to scale to datasets beyond 1TB. ## Why Vortex?[​](#why-vortex "Direct link to Why Vortex?") Spice Cayenne uses Vortex as its storage format, providing significant performance advantages: * **100x faster random access reads** compared to modern Apache Parquet * **10-20x faster scans** for analytical queries * **5x faster writes** with similar compression ratios * **Zero-copy compatibility** with Apache Arrow for efficient data processing * **Extensible architecture** with pluggable encoding, compression, and layout strategies Vortex is a Linux Foundation (LF AI & Data) project under Apache-2.0 license with neutral governance. For performance benchmarks, see [bench.vortex.dev](https://bench.vortex.dev/). Spice Cayenne is recommended for any dataset of **10 GB or larger**. [DuckDB](/docs/next/components/data-accelerators/duckdb) suits smaller datasets, but its memory requirements grow faster: Cayenne typically needs one-third to one-half the memory of the DuckDB accelerator for the same dataset, and scales beyond the limits at which a single DuckDB file becomes impractical. ## Architecture[​](#architecture "Direct link to Architecture") Spice Cayenne follows a lakehouse architecture inspired by [DuckLake](https://ducklake.select/), separating metadata management from data storage: ![Spice Cayenne Architecture](/assets/images/cayenne-architecture-8444d170b17752fb61af914c8cdc9730.png) **Key Design Principles:** * **Virtual Files**: Each "file" is a Vortex `ListingTable` at a unique directory, enabling append operations and parallel reads * **Lazy Statistics**: Summary statistics are loaded on-demand for query optimization * **Sequence-based Ordering**: Iceberg-style sequence numbers enable upsert semantics without requiring separate tracking of "undeleted" records (rows that were deleted and then re-inserted) * **Pluggable Storage**: Data files can be stored locally or in S3 Express One Zone while metadata remains local ## Storage Recommendations[​](#storage-recommendations "Direct link to Storage Recommendations") For optimal performance, store Cayenne data files on NVMe storage. NVMe provides the lowest latency and highest throughput for the random access patterns that Vortex files require. Use [S3 Express One Zone](#aws-s3-express-one-zone-storage) when persistence of accelerations across restarts is required. S3 Express One Zone adds network latency compared to local NVMe but provides durability. Sharing accelerated data across multiple Spice instances is planned for a future release. ## Configuration[​](#configuration "Direct link to Configuration") To use Spice Cayenne as the data accelerator, specify `cayenne` as the `engine` for acceleration. Spice Cayenne supports two storage modes: * **`mode: file`** (durable) — data is written as Vortex files on local disk or S3 Express One Zone, with a SQLite/Turso metastore, and the acceleration survives restarts. This is the recommended mode for Cayenne and is used in the examples throughout this page. The `mode: file_create` and `mode: file_update` variants control how an existing on-disk acceleration is reused or rebuilt on startup. * **`mode: memory`** (ephemeral) — all data lives fully in RAM with an in-memory metastore; nothing is written to disk. The dataset is ephemeral and reloads from its source on restart (like the [Arrow](/docs/next/components/data-accelerators/arrow) accelerator). Memory mode works for all refresh modes (`full`/`append`/`changes`) and for both keyed and no-primary-key datasets, but does not support partitioned tables (`partition_by`), and it enforces a hard per-table RAM bound rather than spilling to disk (see [`cayenne_cdc_mem_tier_max_bytes`](#acceleration-parameters-accelerationparams)). ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: cayenne mode: file ``` ### Parameters[​](#parameters "Direct link to Parameters") Spice Cayenne is configured through two distinct parameter scopes: * **Acceleration parameters** are set per dataset under `acceleration.params` and control how that dataset's accelerated data is stored, compressed, written, and compacted. * **Runtime parameters** are set once per instance under `runtime.params` and control engine-global behavior — caches, optimizer rules, and dedicated memory pools — shared by every Cayenne-accelerated dataset. The two scopes are not interchangeable: setting a runtime parameter under `acceleration.params` (or a per-dataset parameter under `runtime.params`) has no effect — the value is ignored. #### Acceleration parameters (`acceleration.params`)[​](#acceleration-parameters-accelerationparams "Direct link to acceleration-parameters-accelerationparams") Set under a dataset's `acceleration.params`: | Parameter | Description | | ------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cayenne_tuning` | Auto-tuning mode. Accepts `auto` or `adaptive` (preview). `auto` derives the memory-, CPU-, and storage-sensitive knobs statically from the detected environment (cgroup-aware cores and memory, storage class) and the inferred schema (cardinality, row width, primary key) — no feedback loop. `adaptive` additionally runs a per-table closed-feedback controller that measures the live CDC ingest rate (and delete fraction and arrival burstiness) and the runtime's response (apply latency vs. offered load, read amplification, memory pressure) and adjusts the inline-memtable flush caps, the in-memory CDC tier byte cap, compaction cadence/trigger, and write concurrency over time, within the environment-derived bounds. **`adaptive` is reached only by setting it here — nothing else turns the closed loop on.** An unset value, an unrecognized value (warned about, then treated as `auto`), inferred schema, and a configured `cayenne_goal_*` SLO all resolve to `auto`. `adaptive` needs a non-zero `cayenne_compaction_background_interval_ms` (the controller runs on the background compaction tick); if it is `0`, Cayenne falls back to `auto`. Schema inference (always on) sharpens the `adaptive` warm-start but is not required — without inferred metadata the controller relearns the observed row width from live ingest. In both modes an explicit per-knob value overrides the derived value; under `adaptive` an explicitly-set knob is pinned (the loop will not move it). See [Self-Tuning](/docs/next/components/data-accelerators/cayenne/performance#self-tuning). | | `cayenne_goal_replication_lag` | Goal-driven adaptive tuning (preview): target end-to-end CDC replication lag, as a duration (e.g. `5s`). **Best set globally** under `runtime.params`; a value here overrides the global setpoint for this dataset. When set, the closed-loop controller converges toward this SLO in small, bounded steps within `cayenne_goal_convergence_window`. **Requires `cayenne_tuning: adaptive` on the dataset** — a goal never enables the loop by itself; on a table left at the `auto` default it is ignored, with a warning. Goal-seeking also requires a non-zero `cayenne_compaction_background_interval_ms`. See [Goal-driven tuning](/docs/next/components/data-accelerators/cayenne/performance#goal-driven-tuning). | | `cayenne_goal_freshness` | Goal-driven adaptive tuning (preview): target data freshness — the age of the newest queryable data — as a duration (e.g. `30s`). Best set globally under `runtime.params`; a value here overrides the global setpoint for this dataset. | | `cayenne_goal_query_latency` | Goal-driven adaptive tuning (preview): target p99 query latency on this table, as a duration (e.g. `250ms` or `10s`). Best set globally under `runtime.params`; a value here overrides the global setpoint for this dataset. | | `cayenne_goal_convergence_window` | Goal-driven adaptive tuning (preview): the time budget to converge toward the configured `cayenne_goal_*` SLOs, as a duration (e.g. `1m`). Defaults to `60s`. A per-dataset control-cadence knob — it paces how fast the loop chases the goals rather than declaring an outcome, so it is not part of the global SLO surface and has no `runtime.params` form. | | `cayenne_compression_strategy` | Compression algorithm for accelerated data. Defaults to `btrblocks`. Supports `btrblocks` or `zstd`. | | `cayenne_delta_encoding` | Encoding effort applied to delta (incremental) writes such as appends and inline-memtable flushes. Accepts `auto` (default) or a fixed level `0`–`10`. Higher levels search more encoding schemes for a better compression ratio at the cost of more write-time CPU; `auto` encodes every delta write at a light level regardless of size, because deltas are transient staged streams that compaction later re-encodes with the full cascade. Levels `7`–`10` all apply the full default cascade, so set `7` to encode every delta with the full cascade instead. Applies at write time only — changing it never re-encodes existing data or forces a table re-create. Invalid values fall back to `auto` with a warning. | | `cayenne_unsupported_type_action` | Action when an unsupported data type is encountered. Defaults to `error`. See [Data Type Support](#data-type-support). | | `cayenne_segment_cache_mb` | Size of the in-memory Vortex segment cache in megabytes, caching decompressed data segments for improved query performance. Accepts `auto` (default) or an explicit MB value. `auto` scales with machine memory (\~1/128 of RAM), but never below `256` MB and never above `1024` MB. | | `cayenne_force_view_types` | Whether scans emit Arrow view types (`Utf8View`/`BinaryView`) instead of native `Utf8`/`Binary`. Accepts `true` or `false`; defaults to `false`. View types are not compacted across `RepartitionExec`, so wide strings fanned out through a partitioned join can inflate query-pool memory reservations — the native types avoid that. Set to `true` to re-enable view types for a table. Any value other than `false` (case-insensitive) enables view types. | | `cayenne_file_path` | Custom path for storing Cayenne data files. Supports local paths or S3 Express One Zone URLs (e.g., `s3://bucket--usw2-az1--x-s3/prefix/`). | | `cayenne_target_file_size_mb` | Target size for individual Vortex files in MB. When writes exceed this size, a new Vortex file is created. Accepts `auto` (default) or an explicit MB value. `auto` is storage-aware: `256` MB on EBS-class network storage, `64` MB on RAM-backed (tmpfs) mounts, `512` MB on S3 Express (large immutable objects cut object count and per-request cost), and `256` MB on local SSD or unknown storage. Smaller files enable better parallelism and predicate pushdown. | | `cayenne_metadata_dir` | Custom directory for storing Cayenne metadata (SQLite catalog). Defaults to `{spice_data_path}/metadata`. Must resolve **outside** the dataset's data directory — see [Metastore location](#metastore-location). | | `cayenne_metastore` | Metastore backend type. Supports `sqlite` (default) or `turso` (requires `turso` feature flag). | | `cayenne_upload_concurrency` | Maximum number of concurrent file uploads when writing multiple Vortex files to S3 Express One Zone. Accepts `auto` (default) or an explicit value; `auto` uses the runtime's [CPU entitlement](/docs/next/reference/spicepod/runtime#runtimecpu) in whole cores. The aggregate encode concurrency across all Cayenne tables is separately bounded by a process-global budget sized from that same entitlement, less a query reserve of a quarter of the cores (at least 2). | | `cayenne_write_concurrency` | Writer partition override for unsorted ingests, controlling how many Vortex files are encoded in parallel during a write. Accepts `auto` (default) or an explicit value. `auto` encodes up to `min(4, session target_partitions)` files in parallel per write — an intentionally small per-table default (not the full CPU entitlement), so many independently-writing tables do not oversubscribe CPU under concurrent CDC. An explicitly-set value is capped at the session `target_partitions`, which defaults to the runtime's [CPU entitlement](/docs/next/reference/spicepod/runtime#runtimecpu) in whole cores; the aggregate encode concurrency across all Cayenne tables is separately bounded by a process-global budget sized from that entitlement. Values below `1` are clamped to `1`. The sort-and-rewrite compaction path always writes serially regardless of this setting. | | `cayenne_deletion_mode` | How primary-key deletions are recorded and applied. Accepts `auto`, `key`, or `position`; defaults to `auto`, which resolves to `position` (merge-on-read) for most tables, or to `key` for CDC datasets (`refresh_mode: changes`) that declare a `primary_key`. See [Deletion Strategies](#deletion-strategies). | | `cayenne_pk_conflict_detection` | Controls primary-key conflict detection on insert. Accepts `auto` or `none`; defaults to `auto`, which detects existing primary keys and resolves them as merge-on-read upserts. Set to `none` to skip conflict detection (blind append) for append-only CDC workloads where the source guarantees primary-key uniqueness. | | `cayenne_compaction_trigger_files` | Minimum number of small Vortex files in the current snapshot before tiered compaction runs. A "small" file is one whose size is below `cayenne_target_file_size_mb` / 4. Defaults to `4` for `refresh_mode: caching` / `changes`, or `append` with `refresh_check_interval` ≤ 5m; `8` otherwise. A value of `1` is clamped to a minimum of `2`. | | `cayenne_compaction_trigger_protected_snapshots` | Number of protected snapshots before snapshot-maintenance compaction runs. Separate from `cayenne_compaction_trigger_files` so small-file tuning does not silently change scan amplification behavior. Defaults to `4` for `refresh_mode: caching` / `changes`, or `append` with `refresh_check_interval` ≤ 5m; `8` otherwise. A value of `1` is clamped to a minimum of `2`. | | `cayenne_compaction_trigger_snapshot_age_ms` | Maximum age in milliseconds of the oldest protected snapshot before snapshot-maintenance compaction runs. Set to `0` to disable the age trigger. Defaults to `60000` for `refresh_mode: caching` / `changes`, or `append` with `refresh_check_interval` ≤ 5m; `300000` otherwise. | | `cayenne_compaction_max_levels` | Maximum number of consecutive compaction passes per trigger. Bounds write amplification when promotion keeps producing new candidates. Defaults to `3`. | | `cayenne_compaction_max_files_per_pick` | Maximum number of eligible file paths retained in one compaction candidate. For a table eligible for **subset** compaction — key-based deletes (`cayenne_deletion_mode: key`, including append-only tables), no configured `sort_columns`, and no protected snapshots — the runner rewrites only the picked candidate and carries the unpicked settled files into the new snapshot by hardlink (local filesystem) or copy (S3 or cross-device), so this value does bound rewrite IO and memory. A table using position deletes, a configured `sort_columns` (whose subset hardlinks would break the global sort attestation), or carrying protected snapshots (which the rewrite has to fold) still rewrites the whole current snapshot, and this value then only affects trigger selection and observability. Defaults to `32`. | | `cayenne_compaction_background_interval_ms` | Background compaction interval in milliseconds. The accelerator runs a per-table background task at this interval. Set to `0` to disable the background task — inline compaction on writes still runs. Defaults to `10000` for `refresh_mode: caching` / `changes`, or `append` with `refresh_check_interval` ≤ 5m; `0` for `refresh_mode: full`, where each refresh replaces the whole table and leaves nothing for the background task to consolidate (set the parameter explicitly to turn it back on for a table that also takes `INSERT`s); `30000` otherwise. When `refresh_mode` is left unset the connector fills it in — `debezium` and `cdc` resolve to `changes`, `sink` to `disabled`, every other connector to `full` — so an unannotated dataset takes the `full` default here. | | `cayenne_cdc_durability` | Durability mode for the inline CDC write path. Accepts `memory` (default) or `file`. `memory` appends batches to an in-RAM tier and defers the source-slot acknowledgement to a periodic or cap-triggered checkpoint, collapsing per-batch durability cost; on crash the un-checkpointed tail is replayed from the source slot, and because the apply is primary-key-idempotent this remains exactly-once. The in-memory path is eligibility-gated — it applies only to the small-write / CDC profile on non-partitioned tables, and other profiles always use `file`. `file` persists each CDC batch durably before advancing the source slot and is the conservative opt-out. | | `cayenne_cdc_mem_tier_max_bytes` | Per-table RAM-tier byte cap that forces a spill (checkpoint) and source-slot advance, in `cayenne_cdc_durability: memory` mode only. Auto-derived from host memory (\~1/64 of RAM, clamped to 256 MiB–1 GiB; 256 MiB on hosts at or under 16 GiB). A process-global byte budget also bounds aggregate resident memory across all tables; whichever cap is breached first triggers the spill. Set to `0` to disable the per-table cap (the global budget still applies). | | `cayenne_cdc_mem_tier_max_age_ms` | Maximum wall-clock milliseconds a RAM-tier epoch may age before a forced checkpoint, in `cayenne_cdc_durability: memory` mode only. Bounds the crash-replay window and the deferred source-slot acknowledgement for tables that never reach a byte threshold. Defaults to `10000` (10 s). Set to `0` to disable the age trigger. | | `cayenne_cdc_mem_tier_min_flush_bytes` | Minimum resident RAM-tier bytes before the periodic background checkpoint tick durably checkpoints, in `cayenne_cdc_durability: memory` mode only. Bounds snapshot / delete-file churn — below this size a tick is skipped unless the tier has reached `cayenne_cdc_mem_tier_max_age_ms`. Query freshness is unaffected (RAM rows are visible immediately); only the deferred slot acknowledgement waits. The write-path byte-cap spill is not gated by this value. Auto-derived as 1/8 of the resolved `cayenne_cdc_mem_tier_max_bytes` (clamped to 32–128 MiB; 32 MiB on hosts at or under 16 GiB). Set to `0` to flush on every tick. | | `cayenne_cdc_mem_tier_checkpoint_interval_ms` | Periodic background mem-tier checkpoint interval in milliseconds, in `cayenne_cdc_durability: memory` mode only. The accelerator spawns a per-table background task that checkpoints the RAM tier every interval, advancing the deferred source-slot acknowledgement on an idle or pure-upsert stream that never trips a write-path cap or event trigger. Defaults to `1000` (1 s). Set to `0` to disable the periodic task. | | `sort_columns` | Comma-separated list of columns to sort data by on refresh operations. Improves segment pruning for frequently filtered columns. | | `cayenne_sort_columns_origin` | Provenance of `cayenne_sort_columns`, which decides whether that sort order outranks the filter columns observed on scans. Accepts `user` (the default when absent) — the sort order is an explicit operator choice and is authoritative — or `inferred`, meaning schema inference filled it from the source's declared order (for most CDC datasets, the primary key). An `inferred` order is treated as a guess and ranks *below* the observed filter columns, so the default-on adaptive layout can cluster for the workload actually being queried. Schema inference sets this automatically whenever it populates `cayenne_sort_columns`; set it by hand only to reproduce an inferred configuration (for example in a benchmark or test). | | `unsupported_type_action` | Action when encountering unsupported data types. Options: `error` (default), `string`, `warn`, `ignore`. | ##### S3 Express One Zone parameters[​](#s3-express-one-zone-parameters "Direct link to S3 Express One Zone parameters") These are acceleration parameters (set under `acceleration.params`) used when storing Cayenne data files in S3 Express One Zone: | Parameter | Description | | ----------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | `cayenne_s3_zone_ids` | Comma-separated availability zone IDs (e.g., `usw2-az1,usw2-az2`). Auto-generates bucket names in format `spice-{app}-{dataset}--{zone}--x-s3`. | | `cayenne_s3_region` | AWS region (e.g., `us-west-2`). Auto-derived from zone ID if not specified. | | `cayenne_s3_auth` | Authentication method: `iam_role` (default) or `key`. | | `cayenne_s3_key` | AWS access key ID (required when `cayenne_s3_auth: key`). | | `cayenne_s3_secret` | AWS secret access key (required when `cayenne_s3_auth: key`). | | `cayenne_s3_session_token` | AWS session token (optional, for temporary credentials). | | `cayenne_s3_endpoint` | Custom S3 endpoint URL (optional, overrides auto-generated endpoint). | | `cayenne_s3_client_timeout` | Request timeout duration (e.g., `30s`, `5m`). Defaults to `120s`. | | `cayenne_s3_unsigned_payload` | Use unsigned payload for S3 Express One Zone requests. Defaults to `true`. | | `cayenne_s3_allow_http` | Set to `true` for testing with local S3-compatible storage. Defaults to `false`. | ##### Cold object-store tier parameters[​](#cold-object-store-tier-parameters "Direct link to Cold object-store tier parameters") These acceleration parameters (set under `acceleration.params`) configure the optional [cold object-store tier](#cold-object-store-tier). Setting `cayenne_datalake_location` enables the tier; the rest tune the clustering key, cold file size, the warm→cold promotion trigger, and the cold store's credentials. When `cayenne_datalake_location` is unset (the default), the cold tier is dormant and behaves byte-identically to a warm-only table. | Parameter | Description | | -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `cayenne_datalake_location` | Object-store URL prefix for the cold tier (the bottom tier of the storage cascade). Must be an `s3://` URL (e.g. `s3://bucket/prefix`) — a general-purpose S3 or S3-compatible bucket, distinct from the warm S3 Express One Zone store. When set, a background promotion stage graduates the warm local-disk tier to read-optimized, Z-order-clustered Vortex files on this store, and queries span the warm and cold tiers with per-tier pushdown. Unset (default) disables the cold tier. Requires key-based deletes (`cayenne_deletion_mode` auto-resolves to `key`; an explicit `position` is rejected at registration) and `refresh_mode: changes` or `append`. A `primary_key` is required to activate the tier — without one the tier registers but stays inactive. **v1 constraints:** local `file://` locations are not supported; partitioned and position-delete tables are not supported. | | `cayenne_datalake_clustering_columns` | Comma-separated liquid-clustering key columns for cold files (multi-column Z-order), e.g. `tenant_id,ts`. Clustering tightens each cold file's per-column zone maps so selective queries on any clustering dimension prune at the storage layer. When unset, falls back to an operator-configured `cayenne_sort_columns`; when that is also unset, cold-tier promotion clusters by the hottest columns observed in query pushdown filters (default-on adaptive layout), then by an inference-derived `cayenne_sort_columns` (see [`cayenne_sort_columns_origin`](#acceleration-parameters-accelerationparams)), and finally by the primary key when no usable observed column exists. An entry that does not exist in the schema is ignored with a warning. | | `cayenne_datalake_target_file_size_mb` | Target size for cold-tier Vortex files in MB. Larger than the warm `cayenne_target_file_size_mb` because object stores favor fewer, larger objects and cold scans are range reads. Accepts `auto` or an explicit MB value. Defaults to `512`. | | `cayenne_datalake_warm_max_bytes` | The warm tier graduates to cold once its total Vortex bytes reach this threshold. `0` (default) disables the byte trigger; set alongside `cayenne_datalake_warm_max_files` to bound warm-tier size. When the tier is enabled and neither trigger is set, Cayenne applies a default byte trigger (16× `cayenne_datalake_target_file_size_mb`) so promotion is not silently disabled. | | `cayenne_datalake_warm_max_files` | The warm tier graduates to cold once its Vortex file count reaches this threshold. `0` (default) disables the file-count trigger. | | `cayenne_datalake_tiering_check_interval_ms` | How often the background loop checks the warm→datalake tiering trigger, in milliseconds. Datalake tiering is not latency-critical, so this is coarser than compaction. Defaults to `60000` (60s). `0` disables the tiering and garbage-collection loop entirely and is rejected at registration when the datalake tier is enabled. | | `cayenne_datalake_gc_interval_ms` | Physical-GC cadence (and orphan grace) for superseded cold-tier objects: the background sweep runs about this often and deletes an object no longer referenced by the manifest only after it has been observed orphaned for at least one interval. Defaults to `300000` (5m). `0` collapses the grace period to zero (a superseded object could be deleted while a running query still reads it) and is rejected at registration when the cold tier is enabled. | The cold store is authenticated independently from the warm tier (which uses the `cayenne_s3_*` parameters) via the following `acceleration.params`: | Parameter | Description | | -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cayenne_datalake_s3_auth` | Authentication method for the cold S3 store: `iam_role` (default, uses environment/SDK credentials) or `key`. | | `cayenne_datalake_s3_key` | AWS access key ID for the cold store (required when `cayenne_datalake_s3_auth: key`). | | `cayenne_datalake_s3_secret` | AWS secret access key for the cold store (required when `cayenne_datalake_s3_auth: key`). | | `cayenne_datalake_s3_session_token` | AWS session token for the cold store (optional, for temporary credentials). | | `cayenne_datalake_s3_region` | AWS region of the cold bucket. Defaults to the environment region (`AWS_REGION` / `AWS_DEFAULT_REGION`), then `us-east-1`; inert for S3-compatible endpoints. | | `cayenne_datalake_s3_endpoint` | Custom S3 endpoint URL for the cold store (e.g. an S3-compatible store such as MinIO). `http://` endpoints implicitly allow HTTP. | | `cayenne_datalake_s3_allow_http` | Allow plain-HTTP connections to the cold S3 endpoint. Defaults to `false`. | | `cayenne_datalake_s3_client_timeout` | HTTP client timeout for cold-store requests, as a duration (e.g. `2m`). Defaults to `2m`. | | `cayenne_datalake_s3_unsigned_payload` | Use unsigned payloads for cold S3 uploads. Defaults to `true`. | #### Runtime parameters (`runtime.params`)[​](#runtime-parameters-runtimeparams "Direct link to runtime-parameters-runtimeparams") Set once under the top-level `runtime.params` and applied to every Cayenne-accelerated dataset in the instance. With the exception of the goal-driven SLO setpoints (`cayenne_goal_*`) — which set a global default that a dataset can override under its own `acceleration.params` (except `cayenne_goal_qph`, which is global-only) — these are **not** valid under a dataset's `acceleration.params`. A global `cayenne_goal_*` setpoint steers only the datasets that run the closed loop: it takes effect on a dataset that sets `cayenne_tuning: adaptive` and is ignored (with a warning) on every dataset left at the `auto` default. | Parameter | Description | | --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cayenne_footer_cache_mb` | Size of the engine-wide in-memory Vortex footer cache in megabytes. The footer cache stores Vortex file metadata (schemas, statistics, encoding information) and is shared across all Cayenne datasets. Larger values improve query performance for repeated scans. Optional; when unset, no explicit limit is applied and DataFusion's default file-metadata-cache limit of 50 MB applies (there is no fixed 128 MB default). | | `cayenne_filter_propagation` | Enables Cayenne's filter-propagation optimizer rules. Accepts `enabled` or `disabled`; defaults to `disabled`. | | `cayenne_optimizer_rules` | Selects which Cayenne optimizer rules run. Accepts `auto` (default — enables the recommended set, gated by `cayenne_filter_propagation`), `all`, `none` / `disabled`, or a comma-separated list of individual rule names. | | `cayenne_compaction_memory_fraction` | Fraction of the query memory pool carved out for a dedicated Cayenne compaction memory pool. Defaults to `0.2` and is clamped to a supported range. Only applied when at least one enabled Cayenne dataset can accumulate files to compact — a file acceleration `mode` on a refresh mode that is not a whole-table replace — and dedicated thread pools are not disabled. A pod whose Cayenne datasets are all `refresh_mode: full`, or all `mode: memory`, carves no compaction pool and brings up no compaction runtime. | | `cayenne_sort_merge_min_rows` | Advanced anti-join tuning: row-count threshold above which the filter-propagation optimizer switches to a sort-merge strategy. Defaults to an internally tuned value; override only when profiling indicates a need. | | `cayenne_sort_merge_memory_pool_fraction` | Advanced anti-join tuning: fraction of the memory pool the sort-merge anti-join strategy may use. Defaults to an internally tuned value. | | `cayenne_goal_replication_lag` | Goal-driven adaptive tuning (preview): global target end-to-end CDC replication lag, as a duration (e.g. `5s`), applied to every Cayenne dataset. Override per-dataset under a dataset's `acceleration.params`. See [Goal-driven tuning](/docs/next/components/data-accelerators/cayenne/performance#goal-driven-tuning). | | `cayenne_goal_freshness` | Goal-driven adaptive tuning (preview): global target data freshness — the age of the newest queryable data — as a duration (e.g. `30s`). Override per-dataset under a dataset's `acceleration.params`. | | `cayenne_goal_query_latency` | Goal-driven adaptive tuning (preview): global target p99 query latency, as a duration (e.g. `250ms` or `10s`). Override per-dataset under a dataset's `acceleration.params`. | | `cayenne_goal_qph` | Goal-driven adaptive tuning (preview): target query throughput in queries per hour (higher is better), e.g. `5000`. Must be a positive number. **Global-only** — query throughput is measured system-wide (a query spanning multiple datasets, such as a join, is counted once), so it has no per-dataset form; a value set under `acceleration.params` is ignored. | | `cayenne_metastore_cache_mb` | SQLite metastore page-cache size in megabytes. Applies to the default `sqlite` metastore backend (see `cayenne_metastore`); the `turso` backend uses MVCC and ignores the `cayenne_metastore_*` family. Defaults to `256`. | | `cayenne_metastore_mmap_mb` | SQLite metastore memory-mapped I/O size in megabytes. Defaults to `1024` (1 GiB). | | `cayenne_metastore_busy_timeout_ms` | SQLite metastore `busy_timeout` in milliseconds — how long a blocked connection waits for a lock before erroring. Defaults to `30000`. | | `cayenne_metastore_wal_autocheckpoint_pages` | SQLite metastore WAL auto-checkpoint threshold in pages. `0` disables the inline auto-checkpoint so the WAL is drained off the hot commit path by a dedicated background checkpoint instead. Defaults to `0`. | | `cayenne_metastore_wal_truncate_threshold_mb` | WAL size in megabytes above which the background checkpoint escalates to a TRUNCATE checkpoint to reclaim file space. Defaults to `160`. | | `cayenne_metastore_auto_vacuum` | SQLite metastore `auto_vacuum` mode: `none`, `incremental`, or `full`. Takes effect only on a fresh database (an existing database needs a full `VACUUM` to change it). Defaults to `none`. Under `incremental`, freed pages are marked reclaimable and reclaimed in bounded batches by the background maintenance pass — see `cayenne_metastore_incremental_vacuum_pages`. An unrecognized value logs a warning and falls back to `none`. | | `cayenne_metastore_incremental_vacuum_pages` | Freelist pages the background maintenance pass reclaims per tick when the metastore is in `incremental` `auto_vacuum` mode; ignored in every other mode. Defaults to `256`, which is 1 MiB at SQLite's 4 KiB default page size. Reclamation holds the write lock while it relocates pages, so the cap is what keeps each pause short; raise it to drain a large freelist faster at the cost of longer write-lock holds, or set `0` to stop reclaiming without changing the database's `auto_vacuum` mode. | ``` runtime: params: # Engine-global Cayenne tuning, shared by every Cayenne-accelerated dataset cayenne_footer_cache_mb: 512 cayenne_filter_propagation: enabled datasets: - from: s3://analytics-bucket/events/ name: events acceleration: engine: cayenne mode: file params: # Per-dataset Cayenne tuning cayenne_segment_cache_mb: 1024 ``` ## Performance Tuning[​](#performance-tuning "Direct link to Performance Tuning") Spice Cayenne is self-tuning by default and rarely needs manual tuning. For the full tuning reference — [self-tuning and goal-driven SLOs](/docs/next/components/data-accelerators/cayenne/performance#self-tuning), [cache sizing](/docs/next/components/data-accelerators/cayenne/performance#cache-tuning), [compression strategy](/docs/next/components/data-accelerators/cayenne/performance#compression-strategy), and [file-size tuning](/docs/next/components/data-accelerators/cayenne/performance#file-size-tuning) — see the dedicated **[Performance Tuning](/docs/next/components/data-accelerators/cayenne/performance)** page. ## Features[​](#features "Direct link to Features") ### DataFusion Query-Native Execution[​](#datafusion-query-native-execution "Direct link to DataFusion Query-Native Execution") Spice Cayenne is DataFusion query-native, meaning all query execution uses [Apache DataFusion](https://datafusion.apache.org/) and adheres to the `runtime.query.memory_limit` setting. This provides: * **Vectorized execution**: Multi-threaded, SIMD-optimized query processing * **Automatic memory management**: Query memory is tracked and spilled to disk when limits are exceeded * **Dynamic filter pushdown**: Filters from TopK, Join, and Aggregate operators push down to file scans DataFusion's [GreedyMemoryPool](https://docs.rs/datafusion/latest/datafusion/execution/memory_pool/struct.GreedyMemoryPool.html) allows memory reservations on a first-come, first-served basis, improving throughput for high-concurrency queries with many partitions. ### High-Performance Columnar Storage[​](#high-performance-columnar-storage "Direct link to High-Performance Columnar Storage") Spice Cayenne uses Vortex's advanced columnar format, which provides: * **Efficient Compression**: Cascading compression with nested encoding schemes including RLE, dictionary encoding, FastLanes, FSST, and ALP * **Rich Statistics**: Lazy-loaded summary statistics for query optimization * **Extensible Encodings**: Pluggable physical layouts optimized for different data patterns * **Wide Table Support**: Efficient handling of tables with many columns through zero-copy metadata access ### Point Lookups and Random Access[​](#point-lookups-and-random-access "Direct link to Point Lookups and Random Access") Vortex delivers [100x faster random access reads](https://bench.vortex.dev/) compared to Apache Parquet through several architectural features: **Segment Statistics (Zone-Map Equivalent):** Vortex's [ChunkedLayout](https://github.com/vortex-data/vortex) maintains per-segment statistics for each column, enabling segment pruning during query execution. Statistics include: | Statistic | Description | Use Case | | ------------- | -------------------------------- | -------------------------------- | | `min` | Minimum value in segment | Range predicate pruning | | `max` | Maximum value in segment | Range predicate pruning | | `null_count` | Count of null values | IS NULL/IS NOT NULL optimization | | `is_sorted` | Whether segment is sorted | Binary search for point lookups | | `is_constant` | Whether all values are identical | Immediate value return | When a query includes a `WHERE` clause, Spice Cayenne evaluates whether each segment could contain matching rows. Segments that cannot match based on min/max statistics are skipped entirely, similar to DuckDB's [zone-maps](https://duckdb.org/docs/stable/guides/performance/indexing#zonemaps) without requiring explicit index creation. **Example - Segment Pruning:** For a table with segments containing timestamp ranges `[2024-01-01, 2024-01-15]`, `[2024-01-16, 2024-01-31]`, `[2024-02-01, 2024-02-15]`, a query: ``` SELECT * FROM events WHERE timestamp > '2024-01-20'; ``` Prunes the first segment (max < 2024-01-20) and reads only the second and third segments. **Fast Random Access Encodings:** Vortex encodings support direct random access to compressed data: * [FSST](https://www.vldb.org/pvldb/vol13/p2649-boncz.pdf) (Fast Static Symbol Table): String compression with O(1) random access * [FastLanes](https://www.vldb.org/pvldb/vol16/p2132-afroozeh.pdf): High-performance integer encoding with vectorized decoding * [ALP](https://ir.cwi.nl/pub/33334/33334.pdf): Adaptive lossless floating-point compression with random access **Compute Push-Down:** Vortex supports executing filter and compute operations directly on compressed data, avoiding full decompression for predicate evaluation. This compute push-down reduces CPU and memory overhead by processing data in its compressed form: | Encoding | Data Type | Operations | | ---------- | --------- | ------------------------------------------------- | | FSST | Strings | Equality, prefix matching on compressed symbols | | FastLanes | Integers | SIMD-accelerated comparison on bit-packed data | | ALP | Floats | Range comparisons with minimal decompression | | Dictionary | Any | Lookup predicates evaluated on dictionary indices | | RLE | Any | Constant runs evaluated once per run | Array-level statistics (`is_sorted`, `is_constant`, `min`, `max`) enable additional optimizations beyond filtering. For example, `is_sorted` enables binary search for point lookups, and `is_constant` returns values immediately without scanning. **Performance Characteristics:** For point lookups and selective queries, Spice Cayenne with Vortex often matches or exceeds the performance of traditional B-tree indexes while consuming no additional memory for index structures. Performance scales with: * Data sorting (sorted columns benefit most from segment pruning) * Segment cache hit rate (hot data patterns) * Compression encoding match to data characteristics ## Change Data Capture (`refresh_mode: changes`)[​](#change-data-capture-refresh_mode-changes "Direct link to change-data-capture-refresh_mode-changes") Spice Cayenne is the recommended accelerator for [Change Data Capture (CDC)](/docs/next/features/cdc). A dataset configured with `refresh_mode: changes` is bootstrapped from a source snapshot and then kept continuously in sync by applying the source's row-level change stream — inserts, updates, and deletes — so the acceleration reflects the current state of the source row-for-row. CDC into Cayenne is available for every Spice CDC connector: * [PostgreSQL Logical Replication](/docs/next/features/cdc/postgres-replication) * [MySQL Binlog Replication](/docs/next/features/cdc/mysql-replication) * [MongoDB Change Streams](/docs/next/features/cdc/mongodb-streams) * [DynamoDB Streams](/docs/next/features/cdc/dynamodb-streams) * [Debezium](/docs/next/features/cdc/debezium) (over Kafka) See [Change Data Capture](/docs/next/features/cdc) for source-side setup and the [Changes refresh mode](/docs/next/features/data-acceleration/refresh-modes/changes) reference for the acceleration-side configuration. ``` datasets: - from: postgres:public.orders name: orders acceleration: engine: cayenne mode: file refresh_mode: changes ``` Cayenne is not the only engine that can apply a change stream — `arrow`, `duckdb`, and `sqlite` are also valid `changes` sinks — but Cayenne adds CDC-specific capabilities built for large-scale, continuously-updated datasets: * **Incrementally maintained aggregate views (IVM).** Aggregate views declared with `maintained_aggregates` are updated in place as each change event is applied, rather than recomputed, so `GROUP BY` rollups stay fresh at CDC speed. See [Maintained Aggregates](#maintained-aggregates). * **Key-based incremental deletes.** With a primary key, `cayenne_deletion_mode: auto` resolves to key-based deletion vectors so source `DELETE` and `UPDATE` events apply efficiently. See [Deletion Vectors](#deletion-vectors). * **In-memory CDC tier.** Change events can be staged through an in-memory tier (`cayenne_cdc_durability` and the `cayenne_cdc_mem_tier_*` parameters) for low apply latency, with exactly-once semantics via primary-key-idempotent replay. The [deployment guide](/docs/next/components/data-accelerators/cayenne/deployment#cdc-apply-metrics) documents the CDC apply metrics for monitoring lag and throughput. * **Replication-lag and freshness SLOs.** The adaptive self-tuner can target an end-to-end replication-lag or data-freshness goal (`cayenne_goal_replication_lag`, `cayenne_goal_freshness`) and adjust its apply behavior to meet it. See [Goal-driven tuning](/docs/next/components/data-accelerators/cayenne/performance#goal-driven-tuning). ### CDC Requirements[​](#cdc-requirements "Direct link to CDC Requirements") * **A primary key is required.** In `changes` mode Cayenne always applies an inferred or declared primary key and routes source updates through an upsert. Declare `primary_key` on the dataset (and, where the connector requires it, `on_conflict: upsert`); the per-connector CDC pages document the exact requirement. * **File or memory mode.** Durable resume across restarts requires `mode: file`. `mode: memory` is supported for ephemeral, in-RAM CDC where the accelerator is rebuilt from the source on restart. ## Maintained Aggregates[​](#maintained-aggregates "Direct link to Maintained Aggregates") Spice Cayenne can incrementally maintain aggregate views over a CDC-accelerated dataset so that `GROUP BY` aggregate queries are answered from continuously-updated summary state instead of re-scanning the full table. As change events are applied (`refresh_mode: changes`), each maintained view is updated with the incoming deltas; a matching query is then rewritten by the physical optimizer to read the maintained state directly. When a query does not match a declared view — or the view is not currently fresh — Cayenne transparently falls back to a full scan, so results are always correct. Maintained aggregates are declared per dataset under `acceleration.maintained_aggregates`: ``` datasets: - from: postgres:orders name: orders acceleration: enabled: true engine: cayenne refresh_mode: changes primary_key: id maintained_aggregates: - group_by: [customer_id] aggregates: - function: count - function: sum column: amount - function: avg column: amount - function: min column: amount - function: max column: amount ``` Each entry declares one maintained view: | Field | Required | Description | | ------------ | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | | `group_by` | Optional | Columns used as the `GROUP BY` key, in query output order. Omit for a single-group (grand total) view. | | `aggregates` | Required | Aggregate expressions maintained for each group, in query output order. | | `filter_sql` | Optional | A SQL row predicate (a `WHERE` expression over the dataset's columns, e.g. `ol_delivery_d > '2007-01-02'`) restricting which rows contribute to the view. | Each item in `aggregates` has: | Field | Required | Description | | ---------- | ----------- | ---------------------------------------------------------------------------------------------------- | | `function` | Required | One of `count`, `sum`, `avg`, `min`, or `max`. | | `column` | Conditional | The column to aggregate. Omit for `count` (`COUNT(*)`); required for `sum`, `avg`, `min`, and `max`. | Supported input column types by function: * `count` — any column (or `COUNT(*)` when `column` is omitted). * `sum` and `avg` — signed integers (`Int8`–`Int64`), unsigned integers (`UInt8`–`UInt64`), floats (`Float32`/`Float64`), and `Decimal128`. For integer and float inputs, `sum` widens to `BIGINT`/`Float64` and `avg` returns `Float64`. For a `Decimal128(p, s)` input, `sum` returns `Decimal128(min(38, p + 10), s)` and `avg` returns `Decimal128(min(38, p + 4), min(38, s + 4))`, matching DataFusion's exact decimal return types (`avg` requires a non-negative scale). `Decimal256` is not supported. * `min` and `max` — signed/unsigned integers, `Date32`/`Date64`, `Timestamp`, and `Decimal128`, preserving the input type (`MIN(Int32) -> Int32`). Float `min`/`max` is not yet supported. ### Query matching[​](#query-matching "Direct link to Query matching") A query is served from a maintained view only when it matches the view exactly: * The query's `GROUP BY` keys match the view's `group_by`, in order. * The query's `WHERE` predicate matches the view's `filter_sql` exactly — an unfiltered view (no `filter_sql`) answers only unfiltered queries, and a filtered view answers only a query carrying the identical predicate. A view whose `filter_sql` mirrors a common dashboard filter lets that filtered analytical query be served incrementally instead of by a full re-scan. Any query that does not match a declared view (or that reaches the view while it is stale) falls back to the base-table scan — correct, but not accelerated. ### Constraints[​](#constraints "Direct link to Constraints") * Maintained aggregates are a Spice Cayenne feature designed for CDC-accelerated datasets (`refresh_mode: changes`). * `min` and `max` are retraction-hard: they require a `primary_key` on the acceleration so that `UPDATE` and `DELETE` changes can retract a prior extremum. Set `acceleration.primary_key`, enable extended schema inference for a source primary key, or omit `min`/`max`. `count`, `sum`, and `avg` do not require a primary key. * Maintained aggregates are not supported on partitioned tables (`partition_by`). ### Retaining specs without maintaining them[​](#retaining-specs-without-maintaining-them "Direct link to Retaining specs without maintaining them") To keep a view's declaration in configuration without materializing or maintaining it, use the policy form with `mode: disabled`: ``` maintained_aggregates: mode: disabled views: - group_by: [customer_id] aggregates: - function: sum column: amount ``` When any view is declared with the list form above, maintenance is enabled by default; the policy form's `mode: enabled` is equivalent. ## Deletion Vectors[​](#deletion-vectors "Direct link to Deletion Vectors") Spice Cayenne implements efficient deletes without rewriting data files using deletion vectors. Deletion vectors track which rows have been logically deleted, and the information is applied transparently during query execution. ### Deletion Strategies[​](#deletion-strategies "Direct link to Deletion Strategies") How deletions are recorded and applied is controlled by the `cayenne_deletion_mode` parameter: | Mode | How deletes are applied | | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `auto` (default) | Resolves to `position` (merge-on-read) for most tables. For CDC datasets (`refresh_mode: changes`) that declare a `primary_key`, `auto` resolves to `key` instead, so deletes compact concurrently with the continuous writer. | | `position` | Per-file row-position `RoaringBitmap`s are pushed into the Vortex scan, skipping deleted rows at the storage layer with no per-row CPU cost. | | `key` | Deletes are applied above the Vortex scan via a per-row probe on the byte representation of the primary key columns. The explicit opt-out from merge-on-read for primary-key tables. | ``` datasets: - from: s3://bucket/events/ name: events acceleration: engine: cayenne mode: file primary_key: event_id params: cayenne_deletion_mode: auto # default; set to `key` to opt out of merge-on-read ``` Under `position` mode (the `auto` resolution for all tables except CDC datasets with a primary key): * **Tables without a primary key** record deletions by row position. Cayenne uses `RoaringBitmap` for memory-efficient storage of deleted row IDs, providing 50-90% memory savings compared to `HashSet` for sparse deletions. * **Tables with a primary key** capture row positions via a `row_idx()` read-back after each write, with a key-based fallback for any row whose position is not yet known. Pushing the deletes into the scan eliminates the per-row `RowConverter` deletion tax above it. **Key-based deletion** (`cayenne_deletion_mode: key`) uses the byte representation of primary key columns and applies deletes above the scan. This approach is position-independent and survives data reorganization. ### Orphaned deletion-vector cleanup[​](#orphaned-deletion-vector-cleanup "Direct link to Orphaned deletion-vector cleanup") On a primary-key upsert table, a deletion vector becomes **orphaned** once time-based retention empties the snapshot(s) it shadowed — the surviving-sequence floor rises above the DV's delete sequence, so the DV is a query-time no-op that still occupies disk and the in-memory deletion cache. Cayenne runs a background sweep that reclaims these orphaned `.arrow` files and their catalog rows. The sweep is always on and requires no configuration: * **Fixed per-table threshold** — the sweep runs per-table once at least 20 orphaned DVs have accumulated. This threshold is a fixed internal constant, not a tunable parameter. * **Retention-only** — a DV is orphaned only after retention (`retention_period` + `time_column`) empties the snapshots it shadowed. **Without a retention policy nothing is ever orphaned and the sweep is a no-op**, at zero ingest cost. * **No correctness impact** — orphaned DVs are already query-time no-ops; the benefit is reclaimed disk plus a smaller in-memory deletion cache and less file I/O at restart. * **Off the write path** — the sweep is lock-free and throttled, and runs on the compaction runtime; it never extends the write-lock window. The sweep bounds only the orphaned tail. The live (not-yet-orphaned) DV set is bounded separately by compaction. ### Primary Key Optimization[​](#primary-key-optimization "Direct link to Primary Key Optimization") For tables with a single-column `Int64` primary key, Cayenne uses an optimized direct lookup strategy that avoids serialization overhead: ``` datasets: - from: s3://bucket/events/ name: events acceleration: engine: cayenne mode: file primary_key: event_id # Int64 column - uses optimized deletion ``` ### Upsert Support[​](#upsert-support "Direct link to Upsert Support") When `on_conflict` is configured, Cayenne supports upsert semantics using sequence numbers (Iceberg-style ordering): ``` datasets: - from: kafka:events name: events acceleration: engine: cayenne mode: file primary_key: id on_conflict: id: upsert ``` When a primary key is deleted and then re-inserted: 1. The new insert gets a higher sequence number than the delete 2. During scan, the delete doesn't apply to data with higher sequence numbers 3. The new data is visible without requiring separate tracking of "undeleted" records ## AWS S3 Express One Zone Storage[​](#aws-s3-express-one-zone-storage "Direct link to AWS S3 Express One Zone Storage") Spice Cayenne supports storing data files in [AWS S3 Express One Zone](https://aws.amazon.com/s3/storage-classes/express-one-zone/) for single-digit millisecond latency, ideal for latency-sensitive query workloads that require persistence. Metadata remains on local disk for fast catalog operations while data files are stored in S3 Express One Zone. ### Why S3 Express One Zone?[​](#why-s3-express-one-zone "Direct link to Why S3 Express One Zone?") S3 Express One Zone directory buckets provide: * **Single-digit millisecond latency**: 10x faster than S3 Standard for first-byte latency * **High request throughput**: Up to 10x higher request rates than S3 Standard * **Cost efficiency**: Lower per-request costs for high-frequency access patterns * **Durability**: Same 99.999999999% (11 9s) durability as S3 Standard ### S3 Express Examples[​](#s3-express-examples "Direct link to S3 Express Examples") **Example 1 - Explicit bucket:** ``` datasets: - from: s3://source-bucket/events/ name: analytics_events acceleration: engine: cayenne enabled: true mode: file params: # Store data in S3 Express One Zone bucket cayenne_file_path: s3://my-bucket--usw2-az1--x-s3/cayenne/ cayenne_s3_region: us-west-2 ``` **Example 2 - Auto-generated bucket with IAM role:** ``` datasets: - from: postgresql://db/events name: fast_events acceleration: engine: cayenne enabled: true mode: file params: # Auto-generates bucket: spice-{spicepod-name}-fast_events--usw2-az1--x-s3 cayenne_s3_zone_ids: usw2-az1 ``` **Example 3 - Explicit credentials:** ``` datasets: - from: kafka:events name: realtime acceleration: engine: cayenne enabled: true mode: file params: cayenne_s3_zone_ids: use1-az4 cayenne_s3_region: us-east-1 cayenne_s3_auth: key cayenne_s3_key: ${secrets:AWS_ACCESS_KEY_ID} cayenne_s3_secret: ${secrets:AWS_SECRET_ACCESS_KEY} ``` ### Bucket Naming Conventions[​](#bucket-naming-conventions "Direct link to Bucket Naming Conventions") S3 Express One Zone buckets use a specific naming format: * **Format**: `{base-name}--{zone-id}--x-s3` * **Zone ID format**: `{region-code}-az{number}` (e.g., `usw2-az1`, `use1-az4`) * **Auto-generated names**: `spice-{app-name}-{dataset-name}--{zone-id}--x-s3` The zone ID is automatically extracted from the bucket name to configure the correct endpoint. ### Supported AWS Regions[​](#supported-aws-regions "Direct link to Supported AWS Regions") S3 Express One Zone is available in select regions. Spice automatically derives the region from zone IDs: | Zone ID Prefix | Region | | -------------- | -------------- | | `use1` | us-east-1 | | `use2` | us-east-2 | | `usw1` | us-west-1 | | `usw2` | us-west-2 | | `euw1` | eu-west-1 | | `euw2` | eu-west-2 | | `euw3` | eu-west-3 | | `euc1` | eu-central-1 | | `eun1` | eu-north-1 | | `eus1` | eu-south-1 | | `apne1` | ap-northeast-1 | | `apne2` | ap-northeast-2 | | `apse1` | ap-southeast-1 | | `apse2` | ap-southeast-2 | | `aps1` | ap-south-1 | | `sae1` | sa-east-1 | | `cac1` | ca-central-1 | | `afs1` | af-south-1 | | `mes1` | me-south-1 | See AWS documentation for the complete list of [S3 Express One Zone availability zones](https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-express-Regions-and-Zones.html). ### Important Considerations[​](#important-considerations "Direct link to Important Considerations") * **Standard S3 not supported**: Cayenne currently only supports S3 Express One Zone, not standard S3 buckets. * **Same-AZ optimization**: S3 Express One Zone is optimized for same-availability-zone access. For external access, Cayenne uses extended timeouts (a 2-minute per-request timeout by default, configurable via `cayenne_s3_client_timeout`) and retries. * **Bucket auto-creation**: When using `cayenne_s3_zone_ids`, Spice automatically creates the S3 Express directory bucket if it doesn't exist (requires appropriate IAM permissions). * **Metadata locality**: Cayenne metadata (SQLite catalog) remains on local disk. Only data files are stored in S3 Express. ## Cold Object-Store Tier[​](#cold-object-store-tier "Direct link to Cold Object-Store Tier") Cayenne can cascade data across three storage tiers — an in-RAM mem-tier, a local-disk **warm** tier, and an object-store **cold** tier — with each row living in exactly one tier. The cold tier is optional and disabled by default; it is enabled by setting [`cayenne_datalake_location`](#cold-object-store-tier-parameters). When unset, a table is warm-only and behaves byte-identically to before. The cold tier lets a table grow beyond local NVMe capacity while keeping recent, hot data on fast local storage and graduating older data to cheaper, durable object storage — without sacrificing pushdown on the cold data. ### How it works[​](#how-it-works "Direct link to How it works") * **Promotion (write path).** A dedicated background worker graduates the warm tier once the size or file-count threshold (`cayenne_datalake_warm_max_bytes` / `cayenne_datalake_warm_max_files`) is crossed, evaluated every `cayenne_datalake_tiering_check_interval_ms`. Promotion is **incremental carry-forward**: rather than re-materializing the whole table each cycle, it classifies the existing cold manifest into *dirty* files that may host a tombstoned key — using per-PK-column min/max rectangles from each file's persisted statistics, refined by a per-file PK bloom filter, and conservative in the safe direction (a false positive costs an extra rewrite; a missed tombstone is impossible) — and *clean* files that are provably untouched. Only the warm delta plus the dirty files are re-read (all deletes applied, one version per key), **Z-order clustered** for tight multi-column zone maps, and written as read-optimized Vortex at the larger `cayenne_datalake_target_file_size_mb` under a per-promotion prefix; clean files are carried forward by manifest reference and never re-read. The new files are atomically registered while the promoted warm files are cleared in a single transaction, so promotion cost tracks the *changed* data, not total table size. * **Cross-tier scan (read path).** Queries span all tiers and push filters, projection, and limits down to each. Cold files are pruned from statistics held in the metastore, so pruning requires **no object-store round-trip on the query path**. Each cold file's PK bloom lets an upsert's keyset rebuild answer cold-tier key existence without scanning the object store. A `DELETE` after promotion correctly hides a cold-resident row. * **Physical GC.** A periodic mark-and-sweep, rooted at the manifest, reclaims cold objects orphaned by carry-forward rewrites. It runs every `cayenne_datalake_gc_interval_ms` (default 5m), which doubles as the orphan grace period: an object no longer referenced by the manifest is deleted only after it has been observed orphaned for at least one interval. * **Clustering.** Cold files are clustered by `cayenne_datalake_clustering_columns` (multi-column Z-order / Morton order). When it is unset, clustering falls back to an operator-configured `cayenne_sort_columns`; when that is also unset, cold-tier promotion clusters by the hottest columns observed in query pushdown filters (default-on adaptive layout), then by an inference-derived `cayenne_sort_columns`, and finally the primary key. Clustering on more than one dimension prunes far better than a single-column sort for selective queries on any clustering column. Only an *operator-configured* sort order shadows the observed filter columns. A sort order that [schema inference](/docs/next/components/data-connectors#schema-inference) supplied — which for most CDC datasets resolves to the primary key — is tagged `cayenne_sort_columns_origin: inferred` and ranks below the observations, so the adaptive layout still clusters cold files for the queries the table actually receives rather than for inference's guess. It is used as the pre-observation fallback, before the primary key. ### Requirements and v1 constraints[​](#requirements-and-v1-constraints "Direct link to Requirements and v1 constraints") * **Key-based deletes required.** The cold tier requires `cayenne_deletion_mode: key`. When `cayenne_deletion_mode` is unset (or `auto`), Cayenne auto-resolves it to `key` for a datalake-enabled table. An explicit `cayenne_deletion_mode: position` conflicts with the cold tier and is **rejected at registration** — position deletes are file-path scoped and cannot survive the warm→cold rewrite. Set `cayenne_deletion_mode: key` (or remove it), or remove `cayenne_datalake_location`. * **Primary key required to activate.** Promotion classifies and rewrites cold files by primary key, so a table that sets `cayenne_datalake_location` but declares no `primary_key` registers successfully with the datalake tier **inactive** — data is never promoted and remains in the warm tier (Cayenne logs a loud warning). Add a `primary_key` to activate the tier, or remove `cayenne_datalake_location`. Leaving PK-less tables inactive (rather than failing registration) lets a fleet-wide datalake location coexist with datasets that have no primary key. * **Non-zero background intervals.** When the cold tier is enabled, `cayenne_datalake_tiering_check_interval_ms: 0` (which would disable the datalake tiering and garbage-collection loop) and `cayenne_datalake_gc_interval_ms: 0` (which would collapse the GC grace period, risking deletion of an object a running query still reads) are **rejected at registration**. Leave them at their defaults or set a positive value. * **S3-only location.** The cold location must be an `s3://` URL — a general-purpose S3 or S3-compatible bucket (for example MinIO via `cayenne_datalake_s3_endpoint`), authenticated independently through the `cayenne_datalake_s3_*` parameters. It is a separate store from the warm S3 Express One Zone tier and does not need to share its bucket. Local `file://` cold locations are not supported in v1. * **Continuous refresh only.** The cold tier requires `refresh_mode: changes` or `refresh_mode: append`; a `full` refresh re-materializes the whole table each cycle and is rejected. * **Unsupported in v1:** partitioned tables and position-delete tables. ``` datasets: - from: postgres:public.events name: events acceleration: engine: cayenne mode: file refresh_mode: changes primary_key: event_id params: # Enable the cold object-store tier (general-purpose S3, not S3 Express) cayenne_datalake_location: s3://my-cold-bucket/events/ cayenne_datalake_s3_region: us-west-2 cayenne_datalake_clustering_columns: tenant_id,created_at # Graduate the warm tier to cold once it reaches 8 GiB cayenne_datalake_warm_max_bytes: 8589934592 ``` ## Data Type Support[​](#data-type-support "Direct link to Data Type Support") Cayenne (via Vortex) supports most Arrow data types with the following considerations: ### Fully Supported Types[​](#fully-supported-types "Direct link to Fully Supported Types") * All integer types (`Int8`, `Int16`, `Int32`, `Int64`, `UInt*`) * Floating point (`Float32`, `Float64`) * Boolean * Utf8 and LargeUtf8 strings * Binary and LargeBinary * Timestamps (normalized to Microsecond precision) * Date32 and Date64 * Lists and FixedSizeLists * Maps * Structs ### Automatically Converted Types[​](#automatically-converted-types "Direct link to Automatically Converted Types") | Original Type | Converted To | Notes | | --------------------------- | ------------------------ | --------------------------------------------- | | `Float16` | `Float32` | Automatic conversion for Vortex compatibility | | `Timestamp(Nanosecond/...)` | `Timestamp(Microsecond)` | Precision normalized | ### Unsupported Types[​](#unsupported-types "Direct link to Unsupported Types") The following types require the `unsupported_type_action` parameter: * `Interval` types * `Duration` types * `FixedSizeBinary` **`unsupported_type_action` options:** | Value | Behavior | | -------- | ----------------------------------------------- | | `error` | Fail with error (default) | | `string` | Convert to Utf8 string | | `warn` | Include as-is with warning (may fail on insert) | | `ignore` | Skip the column entirely | ``` acceleration: engine: cayenne mode: file params: unsupported_type_action: string # Convert unsupported types to strings ``` ## Resource Considerations[​](#resource-considerations "Direct link to Resource Considerations") Resource requirements for Spice Cayenne depend on dataset size, query patterns, and cache configuration. ### Memory[​](#memory "Direct link to Memory") Spice Cayenne manages memory efficiently through columnar storage and selective caching. Memory allocation should account for: | Component | Default | Notes | | ---------------- | -------- | ------------------------------------------------------------------------------------------------------------------- | | Runtime overhead | \~500 MB | Fixed baseline for the Spice runtime | | Footer cache | 50 MB | Unset by default (DataFusion file-metadata-cache default); increase for datasets with many files (1-10 KB per file) | | Segment cache | 256 MB | Increase based on hot data volume | | Query execution | Variable | Depends on query complexity and concurrency | **Example - Memory-constrained environment:** ``` runtime: params: cayenne_footer_cache_mb: 64 datasets: - from: s3://my-bucket/data/ name: constrained_data acceleration: engine: cayenne mode: file params: cayenne_segment_cache_mb: 128 ``` ### Storage[​](#storage "Direct link to Storage") Spice Cayenne stores data in a columnar format optimized for analytical queries. Storage requirements include: * **Acceleration data**: Compressed Vortex files (typically 30-50% of raw data size with btrblocks) * **Metadata**: SQLite database for catalog and statistics (\~10 MB per 1000 files) * **Temporary files**: Query spill files during complex operations #### Metastore location[​](#metastore-location "Direct link to Metastore location") One metastore holds the catalog — manifests, snapshot pointers, and partition rows — for **every** Cayenne dataset sharing a `cayenne_metadata_dir`. Recreating a dataset deletes its data directory, so a metastore directory that resolves on or beneath that data directory would be deleted along with it, taking the catalog for unrelated datasets. A file-mode Cayenne dataset whose resolved metastore directory falls inside its resolved data directory therefore **fails to load**, naming the dataset and both paths. The stock defaults reach this without any operator error: the data directory defaults to `{spice_data_path}/{dataset_name}/` and the metastore to `{spice_data_path}/metadata`, so a dataset **named `metadata`** collides. An explicit `cayenne_metadata_dir` set beneath a dataset's data directory collides the same way. Resolve it by pointing `cayenne_metadata_dir` outside the data directory, or by renaming the dataset. Paths are compared after `.`/`..` are collapsed and symlinks are resolved, so neither hides an overlap, and a sibling that merely shares a name prefix (`…/meta` next to `…/metadata`) is not affected. Datasets whose data lives on object storage (for example an S3 Express `cayenne_file_path`) are exempt — the metastore is always local, so it cannot sit inside an object-store data path. ### CPU[​](#cpu "Direct link to CPU") Query performance scales with available CPU cores. Vortex's columnar format supports parallel decompression and scanning across multiple threads. Allocate sufficient CPU for: * Query execution parallelism * Data refresh and compression operations * Concurrent query workloads ## Transactions[​](#transactions "Direct link to Transactions") Cayenne supports serializable, gated transactions on accelerator-only Cayenne tables. A client submits a single `BEGIN … COMMIT` SQL body — over the HTTP `/v1/sql` endpoint or FlightSQL — and every statement in the body commits atomically, or not at all: ``` BEGIN; SELECT assert((SELECT balance FROM accounts WHERE id = 1) >= 100); -- optional gate UPDATE accounts SET balance = balance - 100 WHERE id = 1; UPDATE accounts SET balance = balance + 100 WHERE id = 2; COMMIT; ``` **How it works:** * **Detection is automatic.** A multi-statement body whose first statement is `BEGIN` (or `START TRANSACTION`) and whose last statement is `COMMIT` is executed as a transaction — there is no configuration flag to enable it. Every statement runs through the standard query path, so authorization, column masking, logging, and tracing apply to each. * **Atomicity and isolation.** Writes are staged off-lock and published together in a single metastore transaction at `COMMIT`. Each participant table's version is captured at `BEGIN` and re-validated at `COMMIT` using per-key optimistic concurrency control (OCC). If a participant changed since the transaction started, `COMMIT` fails with a retryable conflict (HTTP `409` on `/v1/sql`) — re-run the transaction against the latest committed state. Any statement error (including a failed gate) rolls back every staged write. * **Preconditions with `assert()`.** The `assert()` function evaluates its argument at execution time. If the expression is `false` or `NULL`, the transaction aborts with `assertion failed: gate expression was false or NULL`. Use it to enforce invariants (for example, a sufficient balance) as part of the transaction. Comparison gates such as `… >= cap` are NULL-safe. * **Return value.** On success, `/v1/sql` returns the final statement's result (for the canonical gate-plus-write shape, the last write's row-count summary), or `COMMIT` when the body produces no rows. **Durable federated write-back:** A Cayenne dataset configured with [`write_mode: write_back`](/docs/next/reference/spicepod/datasets#accelerationwrite_mode), `on_conflict`, and `refresh_mode: changes` (CDC) stages its committed writes to the accelerator and then reconciles them asynchronously back to the federated source by a per-table write-back worker. `write_mode: write_back` requires `replication.enabled: true` as an explicit opt-in to asynchronous source durability. Durable write-back is not currently available Reconciling a row to the source is delivered as a separate `DELETE` followed by an `INSERT`. Because the accelerator is CDC-fed from that same source, the standalone `DELETE` echoes back over the changes stream and removes the committed row from the accelerator; if the follow-up `INSERT` then fails, the write is silently gone from both sides. Closing that window needs a source that can apply a delivered row in **one atomic step** — a single transaction covering both legs, or a native conditional upsert. No connector advertises that capability yet, so rather than accept a configuration that can lose a committed write, Spice now **rejects the dataset at registration**: > Failed to register dataset `` (``): durable write-back needs a source that can apply a delivered row in one atomic step, and the `` connector cannot yet. Remove `on_conflict` to keep writes on the accelerator, or choose a different [`acceleration.write_mode`](/docs/next/reference/spicepod/datasets#accelerationwrite_mode). Per-connector atomic delivery is planned; until it lands, this combination cannot be loaded. **Requirements and v1 limitations:** * Write targets must be **accelerator-only, non-partitioned Cayenne datasets**. Other dataset modes route writes to the federated source — where the gate cannot govern them — and are rejected. Durable write-back datasets are not currently loadable (see the warning above). * Only **`INSERT` and `UPDATE`** writes are supported inside a transaction. `DELETE` and `MERGE` are rejected. * At most **one write per table** per transaction. Multiple tables may be written in the same transaction and are committed atomically together. * Reading a Cayenne table that is not a registered participant (for example, a partitioned table) fails the transaction closed. * Nullability-predicate gates (`IS NOT NULL`, `IS NULL`, `COALESCE`) are not yet reliable — prefer comparison gates. * A gated write whose superseded row is still in the in-memory/inline tier (recently CDC-streamed rows) is rejected; file-backed rows are supported. ## Limitations[​](#limitations "Direct link to Limitations") Consider the following limitations when using Spice Cayenne acceleration: * **Memory Mode Constraints**: `mode: memory` (fully in-RAM, ephemeral) is supported alongside `mode: file`, but it does not persist any data (the dataset reloads from its source on restart), does not support partitioned tables (`partition_by`), and enforces a hard per-table RAM bound instead of spilling to disk — a breach returns an error rather than growing without limit. Use `mode: file` when persistence across restarts is required. * **S3 Express Only**: Standard S3 buckets are not supported for remote storage. Only S3 Express One Zone directory buckets are supported. * **Unsupported Data Types**: `Interval`, `Duration`, and `FixedSizeBinary` types require `unsupported_type_action` configuration. * **No Traditional Indexes**: Spice Cayenne does not support explicit index creation via the `indexes` configuration. Vortex's segment statistics and fast random access encodings provide equivalent or better performance for most point lookup workloads. * **No MVCC**: Multi-version concurrency control is not yet implemented. Snapshots and time-travel queries are planned for future releases. * **Transaction Constraints**: [Transactions](#transactions) support gated `INSERT`/`UPDATE` writes on accelerator-only, non-partitioned Cayenne tables only (no `DELETE`/`MERGE`, one write per table). See [Transactions](#transactions) for the full list. ## Example Spicepod[​](#example-spicepod "Direct link to Example Spicepod") Complete example configuration using Spice Cayenne with performance tuning: ``` version: v1 kind: Spicepod name: cayenne-example runtime: query: memory_limit: 4GiB temp_directory: /tmp/spice params: # Engine-global Cayenne runtime tuning (shared by all Cayenne datasets) cayenne_footer_cache_mb: 256 datasets: # Local file storage example with upsert - from: s3://source-bucket/analytics/ name: analytics_data params: file_format: parquet time_column: created_at acceleration: engine: cayenne enabled: true mode: file primary_key: id on_conflict: id: upsert refresh_mode: append refresh_check_interval: 1h params: cayenne_compression_strategy: btrblocks cayenne_segment_cache_mb: 512 cayenne_target_file_size_mb: 64 sort_columns: created_at,id retention_sql: DELETE FROM analytics_data WHERE created_at < NOW() - INTERVAL '30 days' # S3 Express One Zone storage example - from: kafka:realtime-events name: realtime_events acceleration: engine: cayenne enabled: true mode: file primary_key: event_id refresh_mode: append params: # S3 Express One Zone for low-latency persistence cayenne_s3_zone_ids: usw2-az1 cayenne_s3_region: us-west-2 cayenne_compression_strategy: zstd # Fast writes for streaming cayenne_target_file_size_mb: 32 # Smaller files for faster ingestion ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Cayenne as a data accelerator in Spice. [Cayenne Data Accelerator](https://github.com/spiceai/cookbook/tree/trunk/cayenne#readme) ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") **Spice Documentation:** * [Performance Tuning](/docs/next/reference/performance-tuning) - Comprehensive performance optimization guide * [Managing Memory Usage](/docs/next/reference/memory) - Memory configuration reference * [Data Acceleration](/docs/next/features/data-acceleration) - Data acceleration overview **External References:** * [Apache DataFusion](https://datafusion.apache.org/) - Query execution engine * [DataFusion Configuration](https://datafusion.apache.org/user-guide/configs.html) - DataFusion settings and tuning * [Vortex Project](https://github.com/vortex-data/vortex) - Columnar file format * [Vortex Benchmarks](https://bench.vortex.dev/) - Performance benchmarks * [FSST Paper](https://www.vldb.org/pvldb/vol13/p2649-boncz.pdf) - Fast Static Symbol Table compression * [FastLanes Paper](https://www.vldb.org/pvldb/vol16/p2132-afroozeh.pdf) - High-performance integer encoding * [ALP Paper](https://ir.cwi.nl/pub/33334/33334.pdf) - Adaptive floating-point compression * [BtrBlocks Paper](https://www.cs.cit.tum.de/fileadmin/w00cfj/dis/papers/btrblocks.pdf) - Compression algorithm * [AWS S3 Express One Zone](https://aws.amazon.com/s3/storage-classes/express-one-zone/) - Low-latency object storage --- # Cayenne Data Accelerator Deployment Guide Production operating guide for [Spice Cayenne](https://spice.ai/cayenne) — a high-performance [Vortex](https://github.com/vortex-data/vortex)-based accelerator with file-mode storage. Covers storage layout, metastore durability, cache sizing, and observability. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") When Cayenne stores segments on S3 / S3 Express One Zone, authentication follows the same model as the [S3 connector](/docs/next/components/data-connectors/s3/deployment#authentication--secrets): the AWS credential chain with `iam_role_source` for explicit scoping. For local-disk Cayenne, no auth is required — the runtime process needs read/write on the storage path. ## Resilience & Durability[​](#resilience--durability "Direct link to Resilience & Durability") ### Storage Modes[​](#storage-modes "Direct link to Storage Modes") Cayenne supports two storage modes. In **`mode: file`** (durable, the recommended production mode), segments are written as Vortex files on local disk or S3 / S3 Express One Zone and the acceleration survives restarts — this guide is oriented to operating it. In **`mode: memory`** (ephemeral), all data lives fully in RAM with an in-memory metastore, nothing is written to disk, and the dataset reloads from its source on restart; it does not support partitioned tables and enforces a hard per-table RAM bound (no disk spill). Use `mode: file` when persistence across restarts is required. ### Metastore Durability[​](#metastore-durability "Direct link to Metastore Durability") Cayenne's metastore (table list, segment index, delete vectors) is backed by SQLite (default) or Turso. With the default **SQLite** backend, the metastore configures: * `journal_mode=WAL` for crash-safe writes. * `busy_timeout` to handle concurrent access. * `synchronous=NORMAL` for WAL-safe durability with acceptable write latency. The **Turso** backend (opt-in, requires the `turso` feature flag) uses its MVCC journal mode (`journal_mode='mvcc'`) instead of WAL. On shutdown, Cayenne performs a WAL checkpoint (SQLite) and runs `PRAGMA optimize` to minimize restart overhead. Graceful shutdown via `SIGTERM` is important — abrupt kills leave the WAL un-checkpointed (still recoverable, but restart is slower). ### Append WAL Crash Safety[​](#append-wal-crash-safety "Direct link to Append WAL Crash Safety") Staged appends use a crash-safe WAL. On startup Cayenne verifies each staged segment's checksum; corrupted or partially-uploaded segments are rejected and re-materialized from the source connector. ### Single-Writer Concurrency[​](#single-writer-concurrency "Direct link to Single-Writer Concurrency") Cayenne enforces single-writer-per-table concurrency via the metastore. Multiple Spice instances backed by the same Cayenne storage + metastore must not be configured as writers simultaneously; reader-only replicas are supported. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") ### Cache Tuning[​](#cache-tuning "Direct link to Cache Tuning") Two in-memory caches tune the random-read vs memory tradeoff: | Parameter | Scope | Description | | -------------------------- | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | `cayenne_footer_cache_mb` | `runtime.params` | Engine-global footer cache (Vortex file footers), shared by all Cayenne datasets. Low memory cost; enables fast plan-time decisions. | | `cayenne_segment_cache_mb` | `acceleration.params` | Per-dataset segment (data page) cache. Set proportional to your hot working set. | For point-lookup-heavy workloads, size `cayenne_segment_cache_mb` generously — Vortex random-access reads are \~100× faster for cached segments than cold S3 reads. ### Upload Concurrency[​](#upload-concurrency "Direct link to Upload Concurrency") | Parameter | Description | | ---------------------------- | --------------------------------------------------------- | | `cayenne_upload_concurrency` | Parallel segment uploads during refresh / append commits. | For S3 Express One Zone, 8–16 parallel uploads typically maximize throughput. For standard S3 across regions, higher concurrency helps hide per-request latency. ### Partitioning[​](#partitioning "Direct link to Partitioning") Cayenne supports `partition_by` (single and multi-expression). Partition on the column(s) that dominate query filters; this prunes segments at plan time. ### Storage Footprint[​](#storage-footprint "Direct link to Storage Footprint") Vortex compression typically delivers 2–4× better compression than Parquet Snappy for analytical datasets. Plan storage for 0.25–0.5× the raw data size as a starting estimate. ## Metrics[​](#metrics "Direct link to Metrics") Generic acceleration metrics are available with the `dataset_acceleration_` prefix. Cayenne also registers the following OpenTelemetry instruments for CDC ingestion, write/compaction, scan-path, and segment-cache observability, all tagged by `dataset`: ### CDC Apply Metrics[​](#cdc-apply-metrics "Direct link to CDC Apply Metrics") | Metric | Type | Unit | Description | | -------------------------------------------------- | --------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `dataset_acceleration_cdc_apply_burst_duration_ms` | Histogram | ms | Duration to apply one coalesced CDC burst. | | `dataset_acceleration_cdc_apply_burst_bytes` | Histogram | By | Arrow in-memory bytes in one coalesced CDC apply burst. | | `dataset_acceleration_cdc_apply_burst_envelopes` | Histogram | envelopes | Number of source envelopes in one coalesced CDC apply burst. | | `dataset_acceleration_cdc_apply_fixed_cost_ms` | Histogram | ms | Duration for fixed-cost phases of CDC apply (with `phase` label: `finalize_wait`, `commit_wait`, etc.). | | `dataset_acceleration_cdc_source_recv_wait_ms` | Histogram | ms | Duration the CDC apply loop waited to receive the next batch from the source-reader channel. High values indicate the apply loop is source-bound (slot read / WAL decode can't keep up); near-zero indicates it is apply-bound. | ### Scan-Path Metrics[​](#scan-path-metrics "Direct link to Scan-Path Metrics") | Metric | Type | Unit | Description | | ------------------------------------------ | --------- | ------- | ----------------------------------------------------------------------------------------------------------- | | `cayenne_scan_listing_table_cache_entries` | Gauge | entries | Number of entries in the scan `ListingTable` cache. Cleared on snapshot change (compaction/sort/overwrite). | | `cayenne_listing_fence_wait_duration_ms` | Histogram | ms | Time spent waiting on listing-fence reads during scans. | | `cayenne_listing_scan_duration_ms` | Histogram | ms | Duration of listing-table scans. | ### Write & Compaction Metrics[​](#write--compaction-metrics "Direct link to Write & Compaction Metrics") | Metric | Type | Unit | Description | | ------------------------------------------- | --------- | ------ | --------------------------------------------------------------------------------------------------------------------------- | | `cayenne_write_phase_duration_ms` | Histogram | ms | Time spent in Cayenne write-path phases. Labelled by `table` and `phase` (see [Write-phase labels](#write-phase-labels)). | | `cayenne_compaction_duration_ms` | Histogram | ms | Wall-clock time of Cayenne background compaction passes. The histogram's count doubles as the compaction-pass counter. | | `cayenne_compaction_memory_pool_bytes` | Gauge | By | Size of the dedicated compaction memory pool carved from the query memory limit (see `cayenne_compaction_memory_fraction`). | | `cayenne_compaction_memory_exhausted_total` | Counter | passes | Compaction passes that hit `ResourcesExhausted` on the dedicated compaction memory pool. | ### Memory Reconciliation Metrics[​](#memory-reconciliation-metrics "Direct link to Memory Reconciliation Metrics") Sampled every 2 seconds by the loop that resizes the in-memory CDC tier budget, so these are emitted whenever Cayenne acceleration is configured. Read them together: the pool gauges report what the memory accounting believes is reserved, `process_resident_memory_bytes` reports what the kernel will make its OOM decision on, and the gap between them is off-pool memory (encode buffers, caches, allocator retention) that no budget covers. | Metric | Type | Unit | Description | | ------------------------------------------- | ----- | ---- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `query_memory_pool_used_bytes` | Gauge | By | Live bytes reserved in the query memory pool (`runtime.query.memory_limit`), excluding the in-memory CDC tier's mirror account so the off-pool tier is not double-counted as query usage. | | `cayenne_compaction_memory_pool_used_bytes` | Gauge | By | Live bytes reserved in the dedicated compaction memory pool (whose size is reported by `cayenne_compaction_memory_pool_bytes`). | | `process_resident_memory_bytes` | Gauge | By | Resident set size of the `spiced` process. Read from `VmRSS` in `/proc/self/status` on Linux, and from the process's resident memory on other platforms. | #### Write-phase labels[​](#write-phase-labels "Direct link to Write-phase labels") `cayenne_write_phase_duration_ms` carries a `table` label (the accelerated dataset) and a `phase` label that attributes time across the write path. The `phase` values are: | `phase` | Description | | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cdc_path_synchronous` | Total latency of a synchronous CDC write, from slot-apply through publish completion. Also covers a staged inline-bearing upsert that could not be represented as a staged commit and fell back to the synchronous write path. | | `cdc_path_inlined` | A pipelined CDC append that completed as a small inlined write. | | `cdc_path_staged` | A staged (pipelined) CDC write: time to durable WAL and return. Publish/finalize is backgrounded, so this **excludes** publish. | | `cdc_path_inmemory` | An in-memory CDC append (`cayenne_cdc_durability: memory`, serial path): end-to-end latency from slot-apply through the RAM-tier append under the listing fence. The deferred source-slot acknowledgement is checkpointed separately. | | `cdc_path_inmemory_sharded` | An in-memory CDC append applied across PK-hash shards (intra-apply sharding) rather than the single serial index. | | `cdc_path_inmemory_fallback` | An in-memory CDC append that could not be admitted to the RAM tier (the process-global mem-tier byte budget was exhausted after waiting and spilling) and fell back to the durable write path. | | `cdc_path_inmemory_sharded_fallback` | A sharded in-memory apply that bailed under sustained overload before any tier mutation and re-streamed through the durable serial path. | | `inmemory_stream_drain` | Draining the prepared CDC stream into RAM and running deferred primary-key conflict validation — the upstream-bound produce-and-validate slice of `cdc_path_inmemory`. | | `inmemory_spill` | A synchronous RAM-tier checkpoint (spill) triggered when the per-table byte cap (`cayenne_cdc_mem_tier_max_bytes`) is breached, before the batch is appended. | | `inmemory_budget_wait` | Time spent waiting (bounded) for the process-global mem-tier byte budget to admit the batch, released by another table's checkpoint. | | `vortex_write` | Encoding and writing Vortex data files. | | `stage_wal_prepare` | Preparing the staged-append write-ahead log. | | `apply_on_conflict_deletions` | Applying merge-on-read deletions for on-conflict (upsert) writes. | | `publish` | Total publish/finalization of a new snapshot. | | `publish_lock_wait` | Waiting to acquire the visibility and listing-fence locks before publishing. | | `publish_seq` | Durably recording the new snapshot's sequence number before it becomes visible. | | `publish_cas` | The compare-and-swap that makes the new protected snapshot visible. | | `publish_wal_write` | Writing the staging WAL during backgrounded finalize. | | `publish_move_files` | Moving staged files into place during finalize. | | `publish_commit` | Committing the new snapshot during finalize. | The `cdc_path_*` phases are the mutually-exclusive terminal phase of a write — exactly one is recorded per write. The `cdc_path_inmemory*` phases and the `inmemory_*` sub-phases are emitted only under `cayenne_cdc_durability: memory`. The remaining phases (`vortex_write`, `stage_wal_prepare`, `apply_on_conflict_deletions`, `inmemory_*`, and `publish*`) are sub-components useful for attributing where write time is spent. ### Segment Cache Metrics[​](#segment-cache-metrics "Direct link to Segment Cache Metrics") The segment cache is the per-dataset Vortex decompressed-segment cache (`cayenne_segment_cache_mb`). All five instruments are observable — sampled on every collection — and caches sharing a dataset label are aggregated into one series. `accesses` and `hits` are monotonic counters and keep counting across a cache being recreated, so query them with counter operations such as `rate()` or `increase()` (hit rate over a window = `rate(cayenne_segment_cache_hits[5m]) / rate(cayenne_segment_cache_accesses[5m])`). | Metric | Type | Unit | Description | | -------------------------------------- | ------- | -------- | -------------------------------------------------- | | `cayenne_segment_cache_accesses` | Counter | accesses | Cumulative Vortex segment cache `get()` calls. | | `cayenne_segment_cache_hits` | Counter | hits | Cumulative Vortex segment cache hits. | | `cayenne_segment_cache_entries` | Gauge | entries | Live Vortex segment cache entry count. | | `cayenne_segment_cache_weighted_bytes` | Gauge | By | Live Vortex segment cache size in bytes. | | `cayenne_segment_cache_capacity_bytes` | Gauge | By | Configured Vortex segment cache capacity in bytes. | See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting metrics. ## Task History[​](#task-history "Direct link to Task History") Cayenne refresh, append, and query operations participate in [task history](/docs/next/reference/task_history) through the shared acceleration spans (`accelerated_table_refresh`, `sql_query`) plus Cayenne's own internal spans for segment uploads and metastore commits. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Memory mode is ephemeral**: `mode: memory` keeps all data in RAM with no durable storage — the dataset reloads from its source on restart and enforces a hard RAM bound (no disk spill). Use `mode: file` when persistence across restarts is required; for a non-Cayenne pure in-memory accelerator, see [Arrow](/docs/next/components/data-accelerators/arrow/deployment). * **Single-writer per table**: Two Spice instances cannot write the same Cayenne table concurrently. * **Vortex version compatibility**: Cayenne files are tied to the Vortex binary version shipped with Spice. Cross-version reads may be supported but not cross-version writes. * **Object-store write atomicity**: Standard S3 is eventually consistent for multipart uploads. S3 Express One Zone provides strong read-after-write consistency and is recommended for latency-sensitive workloads. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Slow restart after a crash | WAL not checkpointed due to ungraceful shutdown. | Use graceful shutdown (`SIGTERM`); first restart will catch up the WAL automatically. | | `database is locked` metastore errors | Two writers sharing one metastore path. | Ensure only one writer; use distinct metastore paths per instance. | | Dataset fails to load naming a data directory that contains the metastore directory | The resolved metastore sits inside the dataset's data directory — commonly a dataset named `metadata` under the stock defaults. | Set `cayenne_metadata_dir` outside the data directory, or rename the dataset. See [Metastore location](/docs/next/components/data-accelerators/cayenne#metastore-location). | | Query slower than expected for cold data | Segment cache too small for working set. | Increase `cayenne_segment_cache_mb`. | | High S3 request cost | Segment cache misses on every query. | Increase segment cache; consider `partition_by` aligned with query filters. | | Upload throughput does not scale with concurrency | Network or S3 Express One Zone TPS limit. | Use S3 Express One Zone in the same AZ; benchmark with `upload_concurrency` to find the right setting. | | Corrupted segment refused on startup | Crash mid-upload; checksum mismatch. | Segments are re-materialized on refresh. Check storage for partial uploads and remove if orphaned. | --- # Spice Cayenne Performance Tuning Spice Cayenne performance can be optimized through cache configuration, compression strategy selection, and resource allocation. By default Cayenne is self-tuning — see [Self-Tuning](#self-tuning) — so manual tuning is only needed to override a specific knob. ## Self-Tuning[​](#self-tuning "Direct link to Self-Tuning") Cayenne sizes its memory-, CPU-, and storage-sensitive knobs automatically so it runs well on any host without hand-tuning, from a small container to a large multi-core box across different storage classes. Every numeric `cayenne_*` knob also accepts the literal `auto`, which lets the runtime derive that knob's value while leaving the rest of the configuration untouched. The mode is controlled by the `cayenne_tuning` acceleration parameter: * **`auto`** — derive the correct configuration values statically from the detected environment (cgroup-aware cores and memory, storage class) and the inferred schema (cardinality, row width, primary key). No feedback loop runs. * **`adaptive`** (preview) — in addition to the static derivation, run a per-table closed-feedback controller that measures the live CDC ingest rate and the runtime's response (apply latency vs. offered load, read amplification, memory pressure) and adjusts the inline-memtable flush caps, compaction cadence/trigger, and write concurrency over time. Adjustments are bounded by the same environment-derived `[floor, ceiling]` the static tier uses, so the loop can only ever pick a value the static tier could have picked. Environment detection feeds both modes. On AWS EC2, the runtime additionally probes the [Instance Metadata Service (IMDSv2)](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html) once at first table registration to detect T-family burstable CPU and the instance's EBS baseline bandwidth, and measures real storage write throughput with a cloud-agnostic calibration probe — both refine the environment-derived `[floor, ceiling]` bounds. The IMDS probe is non-blocking and fail-open: it uses a tight timeout and a single attempt, so off-AWS (or when the metadata service is unreachable) it is a fast no-op and detection falls back to the calibration probe alone. To skip the IMDS probe entirely, set `SPICE_DISABLE_IMDS=1` (the standard AWS `AWS_EC2_METADATA_DISABLED` environment variable is also honored). `adaptive` is reached **only** by setting `cayenne_tuning: adaptive` on the dataset — nothing else turns the closed loop on. An unset value, an unrecognized value (logged as a warning, then treated as `auto`), schema inference, and a configured `cayenne_goal_*` SLO all resolve to `auto`. Schema inference is always attempted and seeds the controller's warm-start, but is not what selects the mode. ``` acceleration: engine: cayenne params: cayenne_tuning: adaptive ``` Preview `cayenne_tuning: adaptive` is in preview. Cayenne logs a startup warning when it is enabled; verify query correctness and performance before using it for production workloads. The static `auto` mode is recommended for production. `adaptive` needs a non-zero `cayenne_compaction_background_interval_ms`, since the controller runs on the background compaction tick; if it is `0`, Cayenne logs a warning and falls back to `auto`. A `refresh_mode: full` dataset defaults that interval to `0` — a whole-table replace has nothing for background compaction to consolidate — so `adaptive` falls back to `auto` on such a table unless you also set `cayenne_compaction_background_interval_ms` explicitly. Schema inference (always on) sharpens the `adaptive` warm-start — the loop's data-aware warm-start uses the inferred cardinality and size — but is not required for an explicit `cayenne_tuning: adaptive`; without inferred metadata the controller relearns the observed row width from live ingest and converges from the hardware-derived warm-start. In both modes, setting any `cayenne_*` knob to an explicit value overrides the derived value. Under `adaptive`, an explicitly-set knob is **pinned** — the controller will not move it. ### Goal-driven tuning[​](#goal-driven-tuning "Direct link to Goal-driven tuning") Under `adaptive`, you can give the controller high-level service-level objectives (SLOs) instead of leaving it to optimize from its built-in signals alone. Set one or more `cayenne_goal_*` SLOs — **globally under `runtime.params`**, where they apply to every Cayenne-accelerated dataset — and the closed loop converges toward each target in small, bounded steps: A goal does not enable the loop A goal declares a target, not a choice of controller. `cayenne_goal_*` is read on every Cayenne dataset, but it steers the loop only where the loop is already running — set on a dataset left at the `auto` default it is **ignored**, and Cayenne logs a warning naming the dataset once at registration. Because the SLOs are designed to be set once under `runtime.params`, enable the loop per dataset with `cayenne_tuning: adaptive` on each table you want it to steer. | Goal | Parameter | Example | | ----------------------------------------------------- | ------------------------------ | ------- | | End-to-end CDC replication lag | `cayenne_goal_replication_lag` | `5s` | | Data freshness (age of the newest queryable data) | `cayenne_goal_freshness` | `30s` | | p99 query latency | `cayenne_goal_query_latency` | `250ms` | | Query throughput (queries per hour, higher is better) | `cayenne_goal_qph` | `5000` | The time-based goals accept a duration string (e.g. `5s`, `1m`, `250ms`); `cayenne_goal_qph` is a positive number. Scope each goal where the runtime reads it: * **Time-based goals** (`cayenne_goal_replication_lag`, `cayenne_goal_freshness`, `cayenne_goal_query_latency`) are set globally under `runtime.params` and may be overridden per-dataset under a dataset's `acceleration.params`. * **`cayenne_goal_qph`** is **global-only**: query throughput is measured system-wide — a query spanning multiple datasets (such as a join) is counted once — so it is read only from `runtime.params`, and a value set under `acceleration.params` is ignored. * **`cayenne_goal_convergence_window`** (duration, default `60s`) sets the time budget the controller targets for convergence. It is a per-dataset control-cadence knob (set under `acceleration.params`) and is not part of the global SLO surface. Goal-seeking runs only where the closed loop runs, so it needs `cayenne_tuning: adaptive` on the dataset and — like `adaptive` itself — a non-zero `cayenne_compaction_background_interval_ms`. ``` runtime: params: # Global SLOs — apply to every Cayenne dataset that runs the closed loop cayenne_goal_replication_lag: 5s cayenne_goal_query_latency: 250ms cayenne_goal_qph: 5000 # global-only — no per-dataset form datasets: - from: postgres:public.orders name: orders acceleration: engine: cayenne params: # Required — without it this table runs static `auto` and the goals are ignored cayenne_tuning: adaptive # Per-dataset override of the global query-latency SLO cayenne_goal_query_latency: 100ms cayenne_goal_convergence_window: 1m ``` ## Cache Tuning[​](#cache-tuning "Direct link to Cache Tuning") Spice Cayenne uses two in-memory caches to accelerate query performance: **Footer Cache (`cayenne_footer_cache_mb`) — runtime parameter:** The footer cache stores Vortex file metadata, including schemas, statistics, and encoding information. It is engine-global and shared across every Cayenne-accelerated dataset, so it is set under `runtime.params`, not per dataset. Larger cache sizes benefit workloads with many files. * Default: unset — when omitted, DataFusion's default file-metadata-cache limit of 50 MB applies * Increase for datasets with many small files * Each file requires approximately 1-10 KB of footer cache **Segment Cache (`cayenne_segment_cache_mb`) — acceleration parameter:** The segment cache stores decompressed data segments. It is configured per dataset under `acceleration.params`. Larger cache sizes benefit workloads with repeated queries on the same data. * Default: 256 MB * Increase for workloads with hot data patterns * Size based on frequently accessed data volume **Example - High-throughput configuration:** ``` runtime: params: # Engine-global footer cache, shared by all Cayenne datasets cayenne_footer_cache_mb: 512 datasets: - from: s3://analytics-bucket/events/ name: events acceleration: engine: cayenne mode: file params: # Per-dataset segment cache cayenne_segment_cache_mb: 1024 ``` ## Compression Strategy[​](#compression-strategy "Direct link to Compression Strategy") Spice Cayenne supports two compression strategies, each with different performance characteristics. The [BtrBlocks](https://www.cs.cit.tum.de/fileadmin/w00cfj/dis/papers/btrblocks.pdf) compression algorithm is designed for fast analytical queries, while [zstd](https://facebook.github.io/zstd/) provides fast write performance. Additionally, `zstd` achieves better compression ratios when data contains large chunks of binary or text. | Strategy | Compression | Read Speed | Write Speed | Best For | | ----------- | ----------- | ---------- | ----------- | ------------------------------------------------ | | `btrblocks` | Higher | Faster | Moderate | Read-heavy analytics (default) | | `zstd` | High | Moderate | Faster | Write-heavy workloads, large binary or text data | **Example - Write-optimized configuration:** ``` datasets: - from: kafka:events name: realtime_events acceleration: engine: cayenne mode: file refresh_mode: append params: cayenne_compression_strategy: zstd ``` ## File Size Tuning[​](#file-size-tuning "Direct link to File Size Tuning") The `cayenne_target_file_size_mb` parameter controls when new Vortex files are created during writes: * **Smaller files (32-64 MB)**: Better parallelism, finer-grained statistics, faster ingestion * **Larger files (128-256 MB)**: Fewer files to manage, reduced metadata overhead ``` params: cayenne_target_file_size_mb: 64 # More parallelism for high-concurrency workloads ``` --- # DuckDB Data Accelerator The DuckDB Data Accelerator helps improve query performance by using [DuckDB](https://duckdb.org/), an embedded analytical database engine optimized for efficient data processing. It supports in-memory and file-based operation modes, enabling workloads that exceed available memory and optionally providing persistent storage for datasets. To enable DuckDB acceleration, set the dataset's `acceleration.engine` to `duckdb`: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: duckdb mode: file ``` ## Modes[​](#modes "Direct link to Modes") ### Memory Mode[​](#memory-mode "Direct link to Memory Mode") By default, DuckDB acceleration uses `mode: memory`, loading datasets into memory. ### File Mode[​](#file-mode "Direct link to File Mode") When using `mode: file`, datasets are stored by default in a DuckDB file on disk in the `.spice/data` directory relative to the spicepod.yaml. Specify the `duckdb_file` parameter to store the DuckDB file in a different location. For datasets intended to be joined, set the same `duckdb_file` path for all related datasets. ## Configuration Parameters[​](#configuration-parameters "Direct link to Configuration Parameters") DuckDB acceleration supports the following optional parameters under `acceleration.params`: * `duckdb_file` (string, default:`.spice/data/accelerated_duckdb.db`): Path to the DuckDB database file. Applies if `mode` is set to `file`. If the file does not exist, Spice creates it automatically. * `duckdb_data_dir` (string, default:`.spice/data/`): Path to the directory the DuckDB database file(s) will be placed in. If both `duckdb_data_dir` and `duckdb_file` are specified, `duckdb_file` will be used and `duckdb_data_dir` will be ignored. * `duckdb_memory_limit` (string, default: none — the runtime computes a [coordinated memory budget](#coordinated-memory-budget) when this is unset): Limits DuckDB's memory usage for instance. Acceptable units are KB, MB, GB, TB (decimal: 1000^i) or KiB, MiB, GiB, TiB (binary: 1024^i). See [DuckDB memory limit documentation](https://duckdb.org/docs/stable/configuration/overview). * `duckdb_preserve_insertion_order` (boolean, default: `true`): Controls whether DuckDB preserves the insertion order of rows in tables. When set to `true`, rows are returned in the order they were inserted. See [DuckDB preserve insertion order documentation](https://duckdb.org/docs/stable/guides/performance/how_to_tune_workloads#the-preserve_insertion_order-option) and [order preservation documentation](https://duckdb.org/docs/stable/sql/dialect/order_preservation). * `connection_pool_size` (integer, default: `10` for local SSD / tmpfs / unspecified storage profiles, or `4` for `ebs`; whichever is larger between that floor and the number of datasets sharing the same DuckDB file): Controls the maximum number of connections to keep open in the connection pool for concurrent query execution. See [`acceleration.storage_profile`](/docs/next/reference/spicepod/datasets#accelerationstorage_profile) for how the storage profile is selected. * `on_refresh_recompute_statistics` (string, default: `enabled`, `disabled` when `refresh_mode` is `changes`): Triggers automatic `ANALYZE` execution after data refreshes. This keeps DuckDB optimizer statistics up-to-date for efficient query plans and performance. Set to `disabled` to turn automatic statistics recomputation off. See [DuckDB ANALYZE statement documentation](https://duckdb.org/docs/stable/sql/statements/analyze). * `duckdb_index_scan_percentage` (float, default: `0.001`): Sets the threshold percentage for performing an index scan instead of a table scan. An index scan is used when the number of matching rows is less than the maximum of `duckdb_index_scan_max_count` and `duckdb_index_scan_percentage` multiplied by total row count. Must be between `0.0` and `1.0`. * `duckdb_index_scan_max_count` (integer, default: `2048`): Sets the maximum row count threshold for performing an index scan instead of a table scan. An index scan is used when the number of matching rows is less than the maximum of `duckdb_index_scan_max_count` and `duckdb_index_scan_percentage` multiplied by total row count. Must be a non-negative integer. * `on_refresh_sort_columns` (string, default: none): Sorts data after each refresh by the specified columns, improving DuckDB [zone map](https://duckdb.org/2025/05/14/sorting-for-fast-selective-queries) (min/max) statistics for query pruning and significantly faster lookup queries. Format: `column1 ASC, column2 DESC` or `column1, column2` (defaults to ASC). Specified columns must exist in the dataset schema, and sort direction must be `ASC` or `DESC`. * `on_full_refresh` (string, default: `reuse_file`): How a full refresh writes into a file-mode acceleration, and whether the space held by the previous copy of the data is reclaimed. One of `reuse_file`, `replace_file`, or `checkpoint_file` — see [Bounding acceleration file growth](#bounding-acceleration-file-growth). `replace_file` and `checkpoint_file` require `mode: file`; configuring either with `mode: memory` is rejected at load time. * `optimizer_duckdb_aggregate_pushdown` (string, default: `disabled`): Enables aggregate pushdown optimization to execute supported aggregate queries directly in DuckDB. Set to `enabled` to push down aggregations for improved query performance on supported functions like `count`, `sum`, `avg`, `min`, and `max`. Requires `query_federation` to be `disabled`. Refer to the [datasets configuration reference](/docs/next/reference/spicepod/datasets#acceleration) for additional supported fields. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: duckdb mode: file params: duckdb_file: /my/chosen/location/duckdb.db duckdb_memory_limit: '2GB' ``` ## Bounding Acceleration File Growth[​](#bounding-acceleration-file-growth "Direct link to Bounding Acceleration File Growth") A full refresh (`refresh_mode: full`) bulk-loads a fresh copy of the data into the DuckDB file. Bulk loads write row groups directly to the database file and send only block pointers to the WAL, so DuckDB's WAL-growth-based automatic checkpoint never fires — and the blocks freed by dropping the previous copy of the table are only returned to the free list at a checkpoint. On a repeatedly full-refreshed file-mode acceleration, the DuckDB file therefore **grows without bound** even though the data it holds does not. Set the `on_full_refresh` parameter to reclaim that space after each refresh: | `on_full_refresh` | Behavior | Query impact | File size | | ----------------- | ------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | | `reuse_file` | Default. Keeps writing into the current database file. No space is reclaimed. | None. | Grows with every refresh. | | `replace_file` | Writes a new database file, checkpoints it, and atomically replaces the live file with it. | Readers are never interrupted; writers pause briefly during the replacement. | Reclaimed on every refresh; the file can shrink. | | `checkpoint_file` | Keeps the current file and checkpoints it after each refresh commits. | A checkpoint that has to escalate stalls queries on that file for a bounded window (see below). | Plateaus near the working set, but never shrinks below the file's high-water mark. | Both `replace_file` and `checkpoint_file` require file-mode acceleration. Configuring either alongside `mode: memory` is rejected at load time. ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: duckdb mode: file refresh_mode: full params: duckdb_file: /data/shared.duckdb on_full_refresh: replace_file # default: reuse_file ``` ### `replace_file`[​](#replace_file "Direct link to replace_file") A full refresh streams into a fresh staging database file while the live file keeps serving queries, then: 1. copies every other object sharing the file into the staging file — other datasets' tables, views, and indexes, the `spice_sys_*` metadata tables (dataset checkpoints, CDC offsets), and HNSW indexes, 2. checkpoints and cleanly closes the staging file, so it is compact and WAL-free, and 3. atomically renames it over the configured path and repoints the shared connection pool at it. In-flight queries drain against the old file through their already-open descriptors and new queries see the new file, so readers never block; only writers pause, for the duration of the copy-and-replace window. Because the replacement always produces a checkpointed file, [acceleration snapshots](/docs/next/features/data-acceleration/snapshots) taken from it are exact. Several datasets can share one DuckDB file: replacements serialize on a per-file write gate and each one carries the other datasets' current data forward. A dataset accelerated into a *different* DuckDB file that reads this one needs no special handling — its attachment re-resolves when the file underneath it is replaced. If the process is interrupted mid-replacement, the leftover files are cleaned up on the next startup: incomplete staging files are deleted, and the newest completed replacement is adopted when the configured file itself is missing. `replace_file` cannot be combined with `refresh_mode: snapshot` on the same DuckDB file — whether on the same dataset or on another dataset sharing the file — and the combination is rejected at load time. Both mechanisms replace the file out-of-band on their own schedules, so refreshes would fail intermittently as each retires the other's file. Give one of them its own `duckdb_file`, or set `on_full_refresh: reuse_file`. ### `checkpoint_file`[​](#checkpoint_file "Direct link to checkpoint_file") After each full-refresh overwrite commits, the runtime runs `CHECKPOINT` on the live database. A plain `CHECKPOINT` fails fast while other transactions are open, in which case it escalates to `FORCE CHECKPOINT`, which waits for the in-flight transactions to finish while blocking new ones from starting — a stall bounded by the slowest in-flight query plus the checkpoint write itself, paid only when the escalation is needed. A checkpoint that fails is logged and never fails the refresh; the refreshed data is already durably committed. This is lighter than `replace_file` — there is no staging copy of cohabiting objects — at the cost of that stall, and of the file plateauing at its high-water mark instead of shrinking. Prefer `replace_file` where in-flight queries must not be interrupted. ## Limitations[​](#limitations "Direct link to Limitations") Consider the following limitations when using DuckDB acceleration: * DuckDB does not support [enum and dictionary field types](https://duckdb.org/docs/sql/data_types/overview). * DuckDB's maximum decimal precision is 38 digits. `Decimal256` (76 digits) is unsupported. * Timezone-aware timestamp columns (e.g. a PostgreSQL `timestamptz` source) are stored at microsecond precision. DuckDB's `TIMESTAMP WITH TIME ZONE` type has no nanosecond variant, so a nanosecond-precision timezone-aware column is normalized to microsecond when accelerated, and sub-microsecond precision is not preserved. Timezone-naive timestamp columns are unaffected (DuckDB has a native nanosecond `TIMESTAMP_NS` type). * Queries using `on_zero_results: use_source` cannot filter binary columns directly (e.g., `WHERE col_blob <> ''`). Instead, cast binary columns to another type (e.g., `WHERE CAST(col_blob AS TEXT) <> ''`). * DuckDB indexes currently do not support spilling to disk. * Hot-reloading dataset configurations while the Spice Runtime is active disables DuckDB query federation until the runtime restarts. * `on_refresh_sort_columns` is not currently supported with primary keys or indexes. * DuckDB acceleration does not support [`partition_by`](/docs/features/data-acceleration/partitioning). Configuring it is rejected at load time. Use the `arrow` or `cayenne` engine for partitioned acceleration. ## Resource Considerations[​](#resource-considerations "Direct link to Resource Considerations") Resource requirements depend on workload, dataset size, query complexity, and refresh modes. ### Memory[​](#memory "Direct link to Memory") Datasets 10 GB or larger For any dataset of **10 GB or larger**, [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) is recommended over DuckDB, because of DuckDB's memory requirements. Cayenne typically needs **one-third to one-half** the memory of the DuckDB accelerator for the same dataset, and its query execution is governed by [`runtime.query.memory_limit`](/docs/reference/spicepod/runtime#runtimequerymemory_limit) with spill-to-disk rather than a separate per-instance pool. DuckDB manages memory through streaming execution, intermediate spilling, and buffer management. Left to itself, each DuckDB instance (one per distinct DuckDB file, plus one shared instance for all `mode: memory` datasets) sizes its own `memory_limit` at roughly 80% of **host** RAM — independently of every other instance and of the Spice query engine. To control memory usage explicitly, set the `duckdb_memory_limit` parameter: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: duckdb mode: file params: duckdb_file: '/data/shared_duckdb_instance.db' duckdb_memory_limit: '4GB' ``` Note that `duckdb_memory_limit` only limits the DuckDB instance it is set on, not the entire runtime process. Additionally, it does not cover all DuckDB operations, such as some insert operations. Index creation and scans are limited by the `duckdb_memory_limit` so ensure adequate memory is provisioned. Allocate at least 30% more container/machine memory for the runtime process. #### Coordinated memory budget[​](#coordinated-memory-budget "Direct link to Coordinated memory budget") Because those per-instance ceilings do not know about each other, a Spicepod with several DuckDB files declares several independent 80%-of-RAM ceilings, stacked on top of the [`runtime.query.memory_limit`](/docs/next/reference/spicepod/runtime#runtimequerymemory_limit) pool (90% of RAM by default, 70% when Cayenne acceleration is also active) — an over-commit that risks an OOM kill under load. At startup, and again on hot-reload, Spice computes a coordinated budget so the **sum** of those ceilings fits within the memory the process can actually use — its cgroup memory limit when one binds, otherwise host RAM. The limit is read from the process's own cgroup path, taking the smallest limit at any level, so a container limit, a `systemd` unit's `MemoryMax=`, a capped parent slice, and a Kubernetes pod cgroup are all honored. Coordination is always on and has no configuration parameter: * Each distinct DuckDB instance with **no** `duckdb_memory_limit` is capped at an equal share of what the query pool and any explicit ceilings leave, with a floor of 128 MiB per instance. * The query pool is reduced by the same amount, taking roughly half of the contested region and never dropping below a quarter of its uncoordinated default (or 256 MiB when every instance has an explicit ceiling). * An explicit `runtime.query.memory_limit` is honored verbatim, and an explicit `duckdb_memory_limit` remains that instance's ceiling — the coordination only sizes what you have not. * If the floors above cannot fit the projection, the ceilings are still applied and the residual over-commit is reported. Coordination is skipped entirely when no DuckDB accelerator is configured, or when the uncoordinated ceilings already fit. Whenever it engages, the runtime logs a warning naming the un-limited instances, the projected uncoordinated ceiling, and the caps it applied — set `duckdb_memory_limit` (and, if needed, `runtime.query.memory_limit`) to replace the automatic split with a deliberate one. note Because `memory_limit` is a per-instance DuckDB setting, an automatic cap is not applied to an instance where any dataset sharing the same DuckDB file sets `duckdb_memory_limit` explicitly — that would clobber the explicit value. ### Indexes and Memory[​](#indexes-and-memory "Direct link to Indexes and Memory") DuckDB indexes currently do not support spilling to disk. While index memory usage is registered through the buffer manager, index buffers are not managed by the buffer eviction mechanism. As a result, indexes may consume significant memory, impacting memory-intensive query performance. Indexes are serialized to disk and loaded lazily upon database reopening, ensuring they do not affect database opening performance. Also consider index serialization when allocating disk storage. For more details, see DuckDB's [Indexes and Memory documentation](https://duckdb.org/docs/stable/guides/performance/indexing.html#indexes-and-memory). ### CPU[​](#cpu "Direct link to CPU") Query performance, data load, and refresh operations scale with available CPU resources. Allocate sufficient CPU cores based on query complexity and concurrency. ### Storage[​](#storage "Direct link to Storage") Ensure adequate disk space for temporary files, swap files, WAL files, and intermediate spilling. Monitor disk usage regularly and adjust storage capacity based on dataset growth and query patterns. ## Temporary Directory[​](#temporary-directory "Direct link to Temporary Directory") The Spice runtime supports configuring a temporary directory for query and acceleration operations that spill to disk. By default, this is the directory of the `duckdb_file`. Set the `runtime.query.temp_directory` parameter to specify a custom temporary directory. This can help distribute I/O operations across multiple volumes for improved throughput. For example, setting `runtime.query.temp_directory` to a high-IOPS volume separate from the DuckDB data file can improve performance for workloads exceeding available memory. Example configuration: ``` runtime: query: temp_directory: /tmp/spice ``` Use this parameter when: * Handling workloads that frequently spill to disk. * Distributing swap and data I/O operations across multiple storage volumes. For more details, refer to the [runtime parameters documentation](/docs/next/reference/spicepod/runtime#runtimequerytemp_directory). For detailed DuckDB limits, see the [DuckDB Memory Management Guide](https://duckdb.org/docs/operations_manual/limits.html). ## Cookbook[​](#cookbook "Direct link to Cookbook") For practical examples, see the [DuckDB Data Accelerator Cookbook Recipe](https://github.com/spiceai/cookbook/tree/trunk/duckdb/accelerator#readme). ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") * [Performance Tuning](/docs/next/reference/performance-tuning) - Zone-maps, indexes, and optimization patterns * [Managing Memory Usage](/docs/next/reference/memory) - Memory configuration reference * [Data Refresh](/docs/next/features/data-acceleration/data-refresh) - Refresh mode configuration --- # DuckDB Data Accelerator Deployment Guide Production operating guide for the DuckDB data accelerator covering memory vs file mode, checkpointing, spill, and observability. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") DuckDB is an embedded, in-process engine. No external authentication is required. For file-mode, protect the DuckDB database file with filesystem permissions and encrypt at rest (LUKS/dm-crypt, EBS encryption, etc.). ## Resilience & Durability[​](#resilience--durability "Direct link to Resilience & Durability") ### Memory vs File Mode[​](#memory-vs-file-mode "Direct link to Memory vs File Mode") | Mode | Durability | Spill-to-disk | Restart behavior | | -------- | -------------------------- | -------------------------------- | ---------------------------- | | `memory` | None — lost on restart. | Via configured `temp_directory`. | Full refresh on startup. | | `file` | Crash-safe via DuckDB WAL. | Via configured `temp_directory`. | Incremental refresh resumes. | Use `mode: file` for any dataset larger than a few hundred MB or where restart speed matters. ### Checkpointing[​](#checkpointing "Direct link to Checkpointing") The DuckDB accelerator enables `PRAGMA enable_checkpoint_on_shutdown` once per DuckDB instance, when the instance is set up. Graceful shutdown writes a clean checkpoint, making restart near-instantaneous. Ungraceful shutdowns leave a WAL to replay, slowing the first subsequent startup. Full-refresh bulk loads bypass the WAL, so DuckDB's WAL-growth-based automatic checkpoint never fires on a repeatedly full-refreshed acceleration and the freed blocks are never returned to the free list. Set [`on_full_refresh`](/docs/next/components/data-accelerators/duckdb#bounding-acceleration-file-growth) to `replace_file` or `checkpoint_file` to reclaim that space on every refresh. ### Spill Directory[​](#spill-directory "Direct link to Spill Directory") Large queries (sort, aggregate, join) can spill to disk. The spill directory is controlled by `runtime.query.temp_directory`. Point this at a fast local volume (NVMe SSD) and ensure adequate free space (2-4× the largest join input is a safe starting point). ### Vacuum[​](#vacuum "Direct link to Vacuum") DuckDB does not require explicit `VACUUM`; its storage layout compacts on checkpoint. For file-mode accelerations on `refresh_mode: full`, the [`on_full_refresh`](/docs/next/components/data-accelerators/duckdb#bounding-acceleration-file-growth) parameter is the Spice-level control over that reclamation: `replace_file` rebuilds and atomically swaps in a compact file on every refresh, `checkpoint_file` checkpoints the live file in place, and the default `reuse_file` reclaims nothing. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") ### Connection Pool[​](#connection-pool "Direct link to Connection Pool") | Parameter | Default | Description | | ---------------------- | -------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `connection_pool_size` | `max(floor, number of datasets on the same instance)`, where `floor` is `4` for `ebs` and `10` otherwise | Maximum connections in the shared DuckDB pool. Floor depends on the resolved [`acceleration.storage_profile`](/docs/next/reference/spicepod/datasets#accelerationstorage_profile). | | *(pool min idle)* | Same as the `floor` above (`4` for `ebs`, `10` otherwise), capped at `connection_pool_size` | Minimum idle connections. | Datasets sharing a DuckDB instance share the pool. For write-heavy refresh plus read-heavy query workloads, size the pool to cover expected concurrency plus a small headroom; DuckDB's serializable concurrency model limits benefit beyond the point of write contention. ### Memory[​](#memory "Direct link to Memory") Datasets 10 GB or larger For any dataset of **10 GB or larger**, deploy [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) instead of DuckDB, because of DuckDB's memory requirements. Cayenne typically needs **one-third to one-half** the memory of the DuckDB accelerator for the same dataset, which lowers the memory request a container needs to run the workload safely. DuckDB self-tunes its memory limit from **host** memory, not the cgroup limit, so under any cgroup memory cap — a container, a `systemd` unit's `MemoryMax=`, a capped parent slice, or a Kubernetes pod cgroup — each instance's own default over-states what the process may use. When `duckdb_memory_limit` is unset, Spice caps each un-limited DuckDB instance from a cgroup-aware [coordinated memory budget](/docs/next/components/data-accelerators/duckdb#coordinated-memory-budget) shared with the query pool, and warns when it does so. Set the `duckdb_memory_limit` acceleration parameter to replace that automatic split with a deliberate ceiling. Plan for the DuckDB working set plus \~2× for query execution headroom. ### Index Parameters[​](#index-parameters "Direct link to Index Parameters") | Parameter | Description | | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------- | | `duckdb_index_scan_percentage` | Optimizer hint: fraction of rows below which index scan is preferred over table scan. | | `duckdb_index_scan_max_count` | Optimizer hint: maximum rows for which index scan is preferred. | | `on_refresh_sort_columns` | Columns to sort by during refresh. **Caution**: current implementation uses `CREATE OR REPLACE`, which drops constraints and indexes. | DuckDB supports traditional B-tree / ART indexes via SQL `CREATE INDEX` against the accelerated table. Define them once the dataset schema is stable. ## Metrics[​](#metrics "Direct link to Metrics") Generic acceleration metrics are available with the `dataset_acceleration_` prefix. DuckDB-specific OpenTelemetry instruments are not currently registered at the runtime layer. For DuckDB-internal telemetry, query DuckDB directly via Spice: ``` SELECT * FROM duckdb_memory(); PRAGMA database_size; ``` See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting runtime metrics. ## Task History[​](#task-history "Direct link to Task History") DuckDB acceleration operations participate in [task history](/docs/next/reference/task_history) through the shared acceleration spans (`accelerated_table_refresh`, `sql_query`) plus DuckDB's SQL execution wrapped in DataFusion plan nodes. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **`on_refresh_sort_columns` drops indexes**: The current implementation issues `CREATE OR REPLACE TABLE ... ORDER BY ...`, which drops pre-existing indexes and constraints. Re-run `CREATE INDEX` statements after sort-column refreshes or pin DDL changes via startup scripts. * **Single writer**: A DuckDB file has one writer at a time. Two Spice instances must not share the same file in write mode. * **Version pinning**: DuckDB database files are tied to the DuckDB binary version. Upgrading Spice to a version with a newer embedded DuckDB may require re-materialization. * **No built-in remote replication**: Cross-host replication is not provided; use file-level replication or a cloud block-store snapshot. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Slow first startup after restart | WAL replay due to ungraceful shutdown. | Use graceful shutdown (`SIGTERM`). Subsequent starts will be fast once the checkpoint is clean. | | OOM on refresh | DuckDB memory limit too high for container cgroup. | Set the `duckdb_memory_limit` acceleration parameter. Check the startup log for the coordinated-budget warning to see what the runtime capped each un-limited instance to. | | Disk fills during large queries | Spill directory on undersized volume. | Point `runtime.query.temp_directory` at a larger volume; monitor free space. | | Query uses table scan when an index exists | `duckdb_index_scan_percentage` / `duckdb_index_scan_max_count` too low. | Tune thresholds; `EXPLAIN` to confirm. | | Indexes disappear after refresh | `on_refresh_sort_columns` triggers `CREATE OR REPLACE`. | Re-create indexes post-refresh, or avoid sort-column refreshes until the underlying behavior is updated. | | `IO Error: Could not set lock on file` | Another process holds a write lock. | Ensure single-writer semantics; verify no other Spice instance is using the same file. | | DuckDB file grows on every refresh | `refresh_mode: full` bulk loads bypass the WAL, so no automatic checkpoint reclaims the previous copy of the data. | Set [`on_full_refresh`](/docs/next/components/data-accelerators/duckdb#bounding-acceleration-file-growth) to `replace_file` (or `checkpoint_file` for a lighter, in-place checkpoint). | --- # PostgreSQL Data Accelerator Spice.ai Enterprise The PostgreSQL Data Accelerator is available only in [Spice.ai Enterprise](https://docs.spice.ai/docs/enterprise). To use PostgreSQL as Data Accelerator, specify `postgres` as the `engine` for acceleration. ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: postgres ``` ## Configuration[​](#configuration "Direct link to Configuration") The connection to PostgreSQL can be configured by providing the following `params`: * `pg_host`: The hostname of the PostgreSQL server. * `pg_port`: The port of the PostgreSQL server. * `pg_db`: The name of the database to connect to. * `pg_user`: The username to connect with. * `pg_pass`: The password to connect with. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load the password from a secret store, e.g. `${secrets:my_pg_pass}`. * `pg_sslmode`: Optional. Specifies the SSL/TLS behavior for the connection, supported values: * `verify-full`: (default) This mode requires an SSL connection, a valid root certificate, and the server host name to match the one specified in the certificate. * `verify-ca`: This mode requires a TLS connection and a valid root certificate. * `require`: This mode requires a TLS connection. * `prefer`: This mode will try to establish a secure TLS connection if possible, but will connect insecurely if the server does not support TLS. * `disable`: This mode will not attempt to use a TLS connection, even if the server supports it. * `pg_sslrootcert`: Optional. Path to a custom PEM certificate file that the connector will trust. * `pg_connection_pool_min`: Optional. The minimum number of connections to keep open in the pool, lazily created when requested. Default is `5`. * `connection_pool_size`: Optional. The maximum number of connections created in the connection pool. Default is `10`. Configuration `params` are provided in the `acceleration` section of a dataset. ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: postgres params: pg_host: localhost pg_port: 5432 pg_db: my_database pg_user: my_user pg_pass: ${secrets:my_pg_pass} pg_sslmode: require ``` Specify different secrets for a PostgreSQL source and acceleration: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset params: pg_host: localhost pg_port: 5432 pg_db: data_store pg_user: my_user pg_pass: ${secrets:pg1_pass} acceleration: engine: postgres params: pg_host: localhost pg_port: 5433 pg_db: acceleration pg_user: two_user_two_furious pg_pass: ${secrets:pg2_pass} ``` Limitations * The Postgres accelerator does not support `Map` types. * The Postgres federated queries may result in unexpected result types due to the difference in DataFusion and Postgres size increase rules. Explicitly specify the expected output type of aggregation functions when writing queries involving Postgres tables in Spice. For example, rewrite `SUM(int_col)` into `CAST (SUM(int_col) as BIGINT)`. ## Arrow to PostgreSQL Type Mapping[​](#arrow-to-postgresql-type-mapping "Direct link to Arrow to PostgreSQL Type Mapping") The table below lists the supported [Apache Arrow data types](https://arrow.apache.org/rust/arrow/datatypes/enum.DataType.html) and their mappings to [PostgreSQL types](https://www.postgresql.org/docs/current/datatype.html) when stored | Arrow Type | sea\_query ColumnType | PostgreSQL Type | | -------------------------------------- | ----------------------- | ----------------------------- | | `Int8` | `TinyInteger` | `smallint` | | `Int16` | `SmallInteger` | `smallint` | | `Int32` | `Integer` | `integer` | | `Int64` | `BigInteger` | `bigint` | | `UInt8` | `TinyUnsigned` | `smallint` | | `UInt16` | `SmallUnsigned` | `smallint` | | `UInt32` | `Unsigned` | `bigint` | | `UInt64` | `BigUnsigned` | `numeric` | | `Decimal128` / `Decimal256` | `Decimal` | `decimal` | | `Float32` | `Float` | `real` | | `Float64` | `Double` | `double precision` | | `Utf8 / LargeUtf8` | `Text` | `text` | | `Boolean` | `Boolean` | `bool` | | `Binary / LargeBinary` | `VarBinary` | `bytea` | | `FixedSizeBinary` | `Binary` | `bytea` | | `Timestamp` (no Timezone) | `Timestamp` | `timestamp` without time zone | | `Timestamp` (with Timezone) | `TimestampWithTimeZone` | `timestamp` with time zone | | `Date32` / `Date64` | `Date` | `date` | | `Time32` / `Time64` | `Time` | `time` | | `Interval` | `Interval` | `interval` | | `Duration` | `BigInteger` | `bigint` | | `List` / `LargeList` / `FixedSizeList` | `Array` | `array` | | `Struct` | `N/A` | `Composite` (Custom type) | ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure PostgreSQL as a data accelerator in Spice. [PostgreSQL Data Accelerator](https://github.com/spiceai/cookbook/tree/trunk/postgres/accelerator#readme) --- # PostgreSQL Data Accelerator Deployment Guide Spice.ai Enterprise The PostgreSQL Data Accelerator is available only in [Spice.ai Enterprise](https://docs.spice.ai/docs/enterprise). Production operating guide for the PostgreSQL data accelerator — materializing source data into a dedicated PostgreSQL database or schema for durable, SQL-native acceleration. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The accelerator uses the same Postgres wire-protocol authentication as the [PostgreSQL data connector](/docs/next/components/data-connectors/postgres/deployment#authentication--secrets). | Parameter | Description | | ---------------- | ----------------------------------------------------------------------------------------------- | | `pg_host` | Postgres server hostname. | | `pg_port` | TCP port (default `5432`). | | `pg_db` | Database name used for acceleration storage. | | `pg_user` | Postgres user. Must have `CREATE`, `INSERT`, `UPDATE`, `DELETE`, `SELECT` on the target schema. | | `pg_pass` | Password. Use `${secrets:...}` to resolve from a configured secret store. | | `pg_sslmode` | TLS mode: `disable` / `prefer` / `require` / `verify-ca` / `verify-full`. | | `pg_sslrootcert` | CA bundle file path for `verify-ca` / `verify-full`. | For production, use `pg_sslmode: verify-full` and source passwords from a [secret store](/docs/next/components/secret-stores). The accelerator sets `application_name` on each connection to the Spice.ai version, which surfaces in `pg_stat_activity` for attribution. ### Permissions[​](#permissions "Direct link to Permissions") The accelerator creates and writes tables in the configured database. Grant the role the minimum privileges on the target schema: `CREATE`, `INSERT`, `UPDATE`, `DELETE`, `SELECT`, and `TRUNCATE`. For refresh modes that use `CREATE OR REPLACE`, the role must also be able to `DROP` its own tables. ## Resilience & Durability[​](#resilience--durability "Direct link to Resilience & Durability") ### Connection Pool[​](#connection-pool "Direct link to Connection Pool") | Parameter | Default | Description | | ------------------------ | ------- | ------------------------------------------ | | `pg_connection_pool_min` | `5` | Minimum idle connections held by the pool. | | `connection_pool_size` | `10` | Maximum connections the pool will open. | `connection_pool_min <= connection_pool_size` is enforced at startup; mismatched values are rejected as configuration errors. ### Durability[​](#durability "Direct link to Durability") Durability is delegated to the PostgreSQL server. Configure Postgres WAL, `synchronous_commit`, and backup policy according to the RPO/RTO requirements of your deployment. For multi-AZ durability, use a Postgres HA setup (Patroni, RDS Multi-AZ, Cloud SQL HA, etc.) — Spice does not replicate across Postgres instances. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Server sizing**: Size the Postgres server for the sum of all accelerated datasets' working-set size plus WAL overhead. Plan for the peak during refresh, which may double-buffer rows in `UPDATE`-heavy paths. * **Connection budget**: `connection_pool_size` across all Spice datasets + all other Postgres clients must not exceed the server's `max_connections`. Use PgBouncer in front of a shared Postgres if many datasets share the server. * **Index management**: Create indexes via SQL on the accelerated tables for query performance. The accelerator does not automatically infer indexes. * **Partitioning**: `partition_by` is **not supported** by the PostgreSQL accelerator and is rejected at configuration validation. Use native Postgres table partitioning managed out-of-band if required. ## Metrics[​](#metrics "Direct link to Metrics") Generic acceleration metrics are available with the `dataset_acceleration_` prefix. The accelerator does not currently register Postgres-specific dataset-level OpenTelemetry instruments. Monitor via: * Spice acceleration metrics (`dataset_acceleration_refresh_duration_ms`, `dataset_acceleration_refresh_errors`). * Postgres server metrics: `pg_stat_activity`, `pg_stat_bgwriter`, `pg_stat_user_tables`, `pg_stat_statements`. * Infrastructure metrics on the Postgres host (CPU, I/O wait, WAL throughput). See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. ## Task History[​](#task-history "Direct link to Task History") PostgreSQL accelerator operations participate in [task history](/docs/next/reference/task_history) through the shared acceleration spans (`accelerated_table_refresh`, `sql_query`). ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **`partition_by` is rejected**: Use native Postgres partitioning if required. * **No cross-instance replication via Spice**: Durability and HA are the Postgres server's responsibility. * **Schema migrations on refresh**: When the source schema changes, the accelerator re-creates the table — existing indexes and foreign keys on the accelerated table are dropped and must be reapplied. * **Write amplification**: `UPSERT`-style refresh modes generate dead tuples; monitor and tune `autovacuum` accordingly. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ---------------------------------------------------------------- | ------------------------------------------------------------ | ----------------------------------------------------------------------------------------------- | | `connection_pool_min must be <= connection_pool_size` at startup | Misconfiguration. | Correct the values so `min <= size`. | | `FATAL: too many clients already` | Sum of pool sizes + other clients exceeds `max_connections`. | Reduce `connection_pool_size`, raise `max_connections`, or front with PgBouncer. | | Refresh fails with `permission denied for table` | Role lacks write/drop privileges on the target schema. | Grant `CREATE`, `INSERT`, `UPDATE`, `DELETE`, `SELECT`, `TRUNCATE`, `DROP` on the schema. | | Indexes disappear after refresh | Accelerator re-created the table. | Reapply indexes post-refresh, or use a refresh mode that preserves the table structure. | | Bloat / slow queries over time | `UPDATE`-heavy refresh without autovacuum tuning. | Tune autovacuum thresholds on the accelerated tables; consider scheduled `VACUUM FULL` windows. | | `partition_by` rejected at startup | Feature not supported. | Remove `partition_by`; use native Postgres partitioning if necessary. | --- # SQLite Data Accelerator To use SQLite as Data Accelerator, specify `sqlite` as the `engine` for acceleration. ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: sqlite ``` ## Configuration[​](#configuration "Direct link to Configuration") The connection to SQLite can be configured by providing the following `params`: * `sqlite_file`: The filename for the file to back the SQLite database. Only applies if `mode` is `file`. * `busy_timeout`: Optional. Specifies the duration for the SQLite [busy timeout](https://www.sqlite.org/c3ref/busy_timeout.html) when connecting to the database file. Default: 5000 ms, or 15000 ms when the acceleration `storage_profile` resolves to EBS-class network storage (where fsync latency spikes are more frequent). Configuration `params` are provided in the `acceleration` section of a dataset. Other common `acceleration` fields can be configured for sqlite, see see [datasets](/docs/next/reference/spicepod/datasets). ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: sqlite mode: file params: sqlite_file: /my/chosen/location/sqlite.db ``` Limitations * The SQLite accelerator doesn't support arrow `Interval` types, as [SQLite](https://www.sqlite.org/lang_datefunc.html) doesn't have a native interval type. * The SQLite accelerator only supports arrow `List` types of primitive data types; lists with structs are not supported. * The SQLite accelerator doesn't support `Dictionary` or `Map` types. * SQLite may not be suitable for high row count use cases with complex join queries. Use [DuckDB](/docs/next/components/data-accelerators/duckdb) instead. * The SQLite accelerator doesn't support advanced grouping features such as `ROLLUP` and `GROUPING`. * In SQLite, `CAST(value AS DECIMAL)` doesn't convert an integer to a floating-point value if the casted value is an integer. Operations like `CAST(1 AS DECIMAL) / CAST(2 AS DECIMAL)` will be treated as integer division, resulting in 0 instead of the expected 0.5. Use `FLOAT` to ensure conversion to a floating-point value: `CAST(1 AS FLOAT) / CAST(2 AS FLOAT)`. * Updating a dataset with SQLite acceleration while the Spice Runtime is running (hot-reload) will cause SQLite accelerator query federation to disable until the Runtime is restarted. Memory Considerations When accelerating a dataset using `mode: memory` (the default), some or all of the dataset is loaded into memory. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. In-memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/next/components/data-accelerators/duckdb) and [`sqlite`](/docs/next/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure SQLite as a data accelerator in Spice. [SQLite Data Accelerator](https://github.com/spiceai/cookbook/tree/trunk/sqlite/accelerator#readme) --- # SQLite Data Accelerator Deployment Guide Production operating guide for the SQLite data accelerator covering file vs memory mode, busy-timeout handling, and observability. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") SQLite is an embedded, in-process engine. No external authentication is required. For file-mode, protect the SQLite database file with filesystem permissions and encrypt at rest if the data is sensitive. ## Resilience & Durability[​](#resilience--durability "Direct link to Resilience & Durability") ### Memory vs File Mode[​](#memory-vs-file-mode "Direct link to Memory vs File Mode") | Mode | Durability | Restart behavior | | -------- | ------------------------------------------ | ---------------------------- | | `memory` | None — lost on restart. | Full refresh on startup. | | `file` | Durable; persisted to the configured path. | Incremental refresh resumes. | Use `mode: file` for any dataset larger than a few hundred MB or where restart speed matters. ### Busy Timeout[​](#busy-timeout "Direct link to Busy Timeout") | Parameter | Default | Description | | -------------- | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `busy_timeout` | `5000` | Milliseconds SQLite will wait for a table lock before returning `SQLITE_BUSY`. Defaults to `15000` when the acceleration `storage_profile` resolves to EBS-class network storage. | Raise this when you observe `database is locked` errors under sustained concurrent refresh + read load. ### Journal Mode[​](#journal-mode "Direct link to Journal Mode") For file-mode databases, the connection pool automatically sets the following pragmas on each connection: | Pragma | Value | Purpose | | -------------- | -------- | ---------------------------------------------- | | `journal_mode` | `WAL` | Enables concurrent readers during writes. | | `synchronous` | `NORMAL` | Balances durability with write performance. | | `cache_size` | `-20000` | Sets the page cache to \~20 MB. | | `foreign_keys` | `true` | Enables foreign key constraint enforcement. | | `temp_store` | `memory` | Stores temporary tables and indices in memory. | These are no-ops for in-memory databases and are set by the runtime; the SQLite accelerator does not expose a connection string for overriding them. Use the `busy_timeout` parameter to tune concurrent-writer handling. ### Federation Across Files[​](#federation-across-files "Direct link to Federation Across Files") File-mode SQLite datasets on the same runtime can be federated using SQLite's `ATTACH DATABASE` mechanism; the accelerator wires up peer attachments automatically for co-located file-mode accelerators. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Single writer**: SQLite serializes writes globally per file. High-concurrency write workloads (e.g., very short refresh intervals on many datasets) hit the write mutex — prefer [DuckDB](/docs/next/components/data-accelerators/duckdb/deployment) or [PostgreSQL](/docs/next/components/data-accelerators/postgres/deployment) for those cases. * **Memory**: The default page cache is \~20 MB (`cache_size = -20000`) and is managed by the runtime; it is not directly configurable. For large read-heavy workloads, prefer [DuckDB](/docs/next/components/data-accelerators/duckdb/deployment). * **Disk**: Plan for 1.2–1.5× the raw data size (SQLite uses row-oriented storage with no strong compression by default). ## Metrics[​](#metrics "Direct link to Metrics") Generic acceleration metrics are available with the `dataset_acceleration_` prefix. SQLite-specific OpenTelemetry instruments are not currently registered at the runtime layer. See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting metrics. ## Task History[​](#task-history "Direct link to Task History") SQLite acceleration operations participate in [task history](/docs/next/reference/task_history) through the shared acceleration spans (`accelerated_table_refresh`, `sql_query`). ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **`partition_by` is rejected**: SQLite accelerator does not support partitioning; use [DuckDB](/docs/next/components/data-accelerators/duckdb/deployment), [PostgreSQL](/docs/next/components/data-accelerators/postgres/deployment), or [Cayenne](/docs/next/components/data-accelerators/cayenne/deployment) when partitioning is required. * **Single writer**: Only one write transaction at a time per file. * **Column store advantages absent**: For wide analytical scans, DuckDB and Cayenne will outperform SQLite materially. * **No built-in remote replication**: Cross-host replication is not provided; use file-level replication, `VACUUM INTO`, or a cloud block-store snapshot. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ---------------------------------------- | ----------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | `database is locked` | Concurrent writer contention exceeds `busy_timeout`. | Raise `busy_timeout`; reduce concurrent refreshes; or switch to DuckDB/Postgres. | | Slow reads on a large file-mode database | Default page cache is small for the working set. | The page cache is managed by the runtime and is not directly configurable; consider DuckDB for large-scan workloads. | | Acceleration rejects `partition_by` | Feature not supported. | Remove `partition_by` or switch engines. | | Queries return stale data after refresh | Readers using long-lived transactions hold an old snapshot. | Ensure read paths do not keep connections open across refresh boundaries (runtime handles this, but custom SQL in pre/post refresh hooks can affect it). | --- # Turso Data Accelerator Beta The Turso Data Accelerator is in Beta. Features and configuration may change. The Turso Data Accelerator uses [Turso](https://turso.tech/) (built on [libSQL](https://github.com/tursodatabase/libsql), a fork of SQLite) as the acceleration engine for caching and accelerating query performance. Turso provides native async support and modern improvements over traditional SQLite while maintaining SQLite compatibility. To enable Turso acceleration, set the dataset's `acceleration.engine` to `turso`: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: turso ``` ## Modes[​](#modes "Direct link to Modes") ### Memory Mode[​](#memory-mode "Direct link to Memory Mode") By default, Turso acceleration uses `mode: memory`, creating an in-memory database for fast temporary caching: ``` datasets: - from: postgres:transactions name: transactions acceleration: enabled: true engine: turso ``` ### File Mode[​](#file-mode "Direct link to File Mode") When using `mode: file`, datasets are stored by default in a Turso database file on disk in the `.spice/data` directory relative to the spicepod.yaml. Specify the `turso_file` parameter to store the database file in a different location: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: turso mode: file params: turso_file: /my/chosen/location/data.turso ``` ## Configuration Parameters[​](#configuration-parameters "Direct link to Configuration Parameters") Turso acceleration supports the following optional parameters under `acceleration.params`: * `turso_file` (string, default: `.spice/data/{dataset_name}.turso`): Path to the Turso database file. Only applies if `mode` is `file`. If the file does not exist, Spice creates it automatically. * `internal_timestamp_format` (string, default: `rfc3339`): Internal timestamp storage format. See [Timestamp Storage](#timestamp-storage) section. Values: `rfc3339`, `integer_millis`. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - from: mysql:orders name: orders acceleration: enabled: true engine: turso mode: file refresh_mode: full refresh_check_interval: 10m params: turso_file: ./orders.turso ``` Refer to the [datasets configuration reference](/docs/next/reference/spicepod/datasets#acceleration) for additional supported fields. ## Timestamp Storage[​](#timestamp-storage "Direct link to Timestamp Storage") Turso supports two timestamp storage formats to accommodate different use cases: ### RFC3339 Format (Default)[​](#rfc3339-format-default "Direct link to RFC3339 Format (Default)") By default, timestamps are stored as **RFC3339 TEXT strings** (e.g., `"2024-01-01T00:00:00.123456789Z"`). This format preserves all timestamp information: * ✅ Full nanosecond precision preserved * ✅ Timezone information preserved * ✅ All Arrow timestamp types supported (Second, Millisecond, Microsecond, Nanosecond) * ✅ Human-readable in database tools * ⚠️ Nanosecond-precision timestamps outside the range \~1677–2262 return `NULL` (i64 nanosecond overflow). Second, Millisecond, and Microsecond precisions are not affected. ### Integer Milliseconds Format[​](#integer-milliseconds-format "Direct link to Integer Milliseconds Format") For performance-critical use cases, set `internal_timestamp_format: "integer_millis"` to store timestamps as integers (milliseconds since Unix epoch): ``` acceleration: engine: turso params: internal_timestamp_format: integer_millis ``` * ✅ Higher performance (direct integer operations) * ✅ Compact storage (8 bytes vs \~30 bytes for RFC3339) * ⚠️ Millisecond precision only (sub-millisecond data is truncated) * ⚠️ No timezone preservation (UTC assumed) * ⚠️ Limited to Second and Millisecond timestamp types ## Features[​](#features "Direct link to Features") ### MVCC Support[​](#mvcc-support "Direct link to MVCC Support") Multi-Version Concurrency Control (MVCC) is automatically enabled for all Turso-accelerated datasets. The Turso accelerator sets `PRAGMA journal_mode = 'mvcc'` during connection pool initialization, enabling concurrent transactions via `BEGIN CONCURRENT`. When MVCC is active: * Concurrent reads and writes perform better * Standard (`enabled`) and unique (`unique`) indexes are created and enforced normally ### Indexes[​](#indexes "Direct link to Indexes") Turso supports index creation for improved query performance: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: turso indexes: id: enabled '(created_at, status)': unique ``` See [Indexes](/docs/next/features/data-acceleration/indexes) for more details. ### Connection Pooling[​](#connection-pooling "Direct link to Connection Pooling") Turso uses connection pooling for efficient database access. Connection pools are shared across datasets using the same database file. ### Query Federation[​](#query-federation "Direct link to Query Federation") Turso supports query federation, where queries can span multiple data sources. The accelerator pushes down filters, projections, and limits when possible for improved performance. ## Limitations[​](#limitations "Direct link to Limitations") * **Remote databases not supported**: Only local Turso databases (file-based or in-memory) are supported as accelerators. Remote Turso databases using `turso_url` and `turso_auth_token` are not supported in this accelerator context. Remote Turso support will be available when Turso is implemented as a data connector. * **Arrow Interval types**: Not supported, as SQLite/libSQL doesn't have a native interval type. * **Complex List types**: Only Arrow `List` types of primitive data types are supported; lists with structs are not supported. * **Dictionary and Map types**: Not supported. * **Hot-reload federation**: Updating a dataset with Turso acceleration while the Spice Runtime is running (hot-reload) may disable query federation until the runtime is restarted. * **ROLLUP and GROUPING**: Advanced grouping features are not supported. ## Comparison with SQLite[​](#comparison-with-sqlite "Direct link to Comparison with SQLite") Turso is built on libSQL (a fork of SQLite) and offers several advantages: | Feature | Turso | SQLite | | ---------------- | -------------------------------------- | -------------------------------- | | Async Support | Native async via libSQL | Synchronous with spawn\_blocking | | Connection Model | Modern async pooling | Traditional sync connections | | Performance | Optimized for modern workloads | Traditional SQLite performance | | File Extensions | `.turso`, `.db`, `.sqlite`, `.sqlite3` | `.db`, `.sqlite`, `.sqlite3` | Choose Turso when you need: * Native async database operations * Modern connection pooling * Better performance for concurrent workloads Choose SQLite when you need: * Maximum compatibility with existing SQLite tools * Simpler deployment without additional dependencies Memory Considerations When accelerating a dataset using `mode: memory` (the default), the dataset is loaded into memory. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. In-memory limitations can be mitigated by storing acceleration data on disk using `mode: file`. ## Building Spice with Turso Support[​](#building-spice-with-turso-support "Direct link to Building Spice with Turso Support") Turso is an optional feature in Spice. To build with Turso support: ``` # Build spiced with Turso support cargo build --release --features turso # Or for development cargo build --features turso ``` ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") * [Data Acceleration](/docs/next/features/data-acceleration) - Data acceleration overview * [Data Refresh](/docs/next/features/data-acceleration/data-refresh) - Refresh mode configuration * [Indexes](/docs/next/features/data-acceleration/indexes) - Index configuration * [SQLite Data Accelerator](/docs/next/components/data-accelerators/sqlite) - Alternative SQLite-based acceleration --- # Data Connectors Data Connectors provide connections to databases, data warehouses, and data lakes for federated SQL queries and data replication. Each connector is configured using the `from` field in a dataset definition. For example: ``` datasets: - from: postgres:public.orders # Database connector name: orders params: pg_host: localhost pg_db: mydb pg_user: reader pg_pass: ${secrets:PG_PASS} - from: s3://my-bucket/events/ # Object storage connector name: events params: file_format: parquet s3_auth: iam_role ``` Supported Data Connectors include: | Name | Description | Status | Protocol/Format | | ---------------------------------- | ----------------------------------------------------------------------------------------------------- | ----------------- | --------------------------------------------------------------------------------- | | `databricks (mode: delta_lake)` | [Databricks](https://github.com/spiceai/cookbook/tree/trunk/databricks#readme) | Stable | S3/Delta Lake | | `delta_lake` | Delta Lake | Stable | Delta Lake | | `dremio` | [Dremio](https://github.com/spiceai/cookbook/tree/trunk/dremio#readme) | Stable | Arrow Flight | | `duckdb` | DuckDB | Stable | Embedded | | `file` | File | Stable | Parquet, CSV | | `github` | GitHub | Stable | GitHub API | | `postgres` | PostgreSQL (with native WAL CDC) | Stable | | | `s3` | [S3](https://github.com/spiceai/cookbook/tree/trunk/s3#readme) | Stable | Parquet, CSV | | `mysql` | MySQL (with native binlog CDC) | Stable | | | `spice.ai` | [Spice.ai](https://github.com/spiceai/cookbook/tree/trunk/spiceai#readme) | Stable | Arrow Flight | | `dynamodb` | Amazon DynamoDB (with Streams) | Stable | | | `graphql` | GraphQL | Release Candidate | JSON | | `cosmosdb` | Azure Cosmos DB (NoSQL) | Release Candidate | | | `git` | Git repositories | Release Candidate | | | `snowflake` | Snowflake | Release Candidate | Arrow | | `adbc` | ADBC | Release Candidate | Arrow | | `iceberg` | [Apache Iceberg](https://github.com/spiceai/cookbook/tree/trunk/catalogs/iceberg#readme) (read+write) | Release Candidate | Parquet | | `databricks (mode: spark_connect)` | [Databricks](https://github.com/spiceai/cookbook/tree/trunk/databricks#readme) | Beta | [Spark Connect](https://spark.apache.org/docs/latest/spark-connect-overview.html) | | `ducklake` | \[DuckLake]\[ducklake] | Beta | Parquet | | `flightsql` | FlightSQL | Beta | Arrow Flight SQL | | `mssql` | Microsoft SQL Server | Beta | Tabular Data Stream (TDS) | | `odbc` | ODBC (Spice.ai Enterprise) | Beta | ODBC | | `spark` | Spark | Beta | [Spark Connect](https://spark.apache.org/docs/latest/spark-connect-overview.html) | | `sharepoint` | Microsoft SharePoint | Beta | Object-store listing | | `oracle` | Oracle | Alpha | [Oracle ODPI-C](https://oracle.github.io/odpi/) | | `abfs` | Azure BlobFS | Alpha | Parquet, CSV | | `clickhouse` | ClickHouse | Alpha | | | `debezium` | Debezium CDC | Alpha | Kafka + JSON | | `elasticsearch` | Elasticsearch (BM25 + kNN + RRF) (Spice.ai Enterprise) | Alpha | | | `gcs`, `gs` | \[Google Cloud Storage]\[gcs] | Alpha | Parquet, CSV, JSON | | `kafka` | Kafka | Alpha | Kafka + JSON | | `ftp`, `sftp` | FTP/SFTP | Alpha | Parquet, CSV | | `glue` | [AWS Glue](https://github.com/spiceai/cookbook/tree/trunk/glue#readme) | Alpha | Iceberg, Parquet, CSV | | `http`, `https` | HTTP(s) (dynamic headers, pagination) | Alpha | Parquet, CSV, JSON | | `imap` | IMAP | Alpha | IMAP Emails | | `localpod` | [Local dataset replication](https://github.com/spiceai/cookbook/blob/trunk/localpod/README.md) | Alpha | | | `mongodb` | MongoDB (with native Change Streams CDC) | Alpha | | | `scylladb` | ScyllaDB | Alpha | | | `smb` | SMB 3.1.1 | Alpha | SMB | | `nfs` | NFS (Spice.ai Enterprise) | Alpha | Parquet, CSV, JSON | ## File Formats[​](#file-formats "Direct link to File Formats") Data connectors that read files from object stores (S3, Azure Blob, GCS) or network-attached storage (FTP, SFTP, SMB, NFS) support a variety of file formats. These connectors work with both structured data formats (Parquet, CSV) and document formats (Markdown, PDF). ### Specifying File Format[​](#specifying-file-format "Direct link to Specifying File Format") When connecting to a **directory**, specify the file format using `params.file_format`: ``` datasets: - from: s3://bucket/data/sales/ name: sales params: file_format: parquet ``` When connecting to a **specific file**, the format is inferred from the file extension: ``` datasets: - from: sftp://files.example.com/reports/quarterly.parquet name: quarterly_report ``` ### Supported Formats[​](#supported-formats "Direct link to Supported Formats") | Name | Parameter | Status | Description | | --------------------------------------------- | ---------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------ | | [Apache Parquet](https://parquet.apache.org/) | `file_format: parquet` | Stable | Columnar format optimized for analytics | | [CSV](/docs/next/reference/file_format#csv) | `file_format: csv` | Stable | Comma-separated values | | JSON | `file_format: json` | Stable | JavaScript Object Notation | | [Delta Lake](https://delta.io/) | `file_format: delta` | Stable | Open table format with ACID transactions. Object stores only. | | [Apache Iceberg](https://iceberg.apache.org/) | `file_format: iceberg` | Stable | Open table format for large analytic datasets. Object stores only. Requires a [catalog](/docs/next/components/catalogs). | | Microsoft Excel | `file_format: xlsx` | Roadmap | Excel spreadsheet format | | Markdown | `file_format: md` | Stable | Plain text with formatting (document format) | | Text | `file_format: txt` | Stable | Plain text files (document format) | | PDF | `file_format: pdf` | Beta | Portable Document Format (document format) | | Microsoft Word | `file_format: docx` | Alpha | Word document format (document format) | ### Format-Specific Parameters[​](#format-specific-parameters "Direct link to Format-Specific Parameters") File formats support additional parameters for fine-grained control. Common examples include: | Parameter | Applies To | Description | | ---------------- | ---------- | ------------------------------------------------ | | `csv_has_header` | CSV | Whether the first row contains column headers | | `csv_delimiter` | CSV | Field delimiter character (default: `,`) | | `csv_quote` | CSV | Quote character for fields containing delimiters | For complete format options, see [File Formats Reference](/docs/next/reference/file_format). ### Applicable Connectors[​](#object-store-file-formats "Direct link to Applicable Connectors") The following data connectors support file format configuration: | Connector Type | Connectors | | ---------------------------- | -------------------------------------- | | **Object Stores** | S3, Azure Blob (ABFS), GCS, HTTP/HTTPS | | **Network-Attached Storage** | FTP, SFTP, SMB, NFS | | **Local Storage** | File | ### Hive Partitioning[​](#hive-partitioning "Direct link to Hive Partitioning") File-based connectors support Hive-style partitioning, which extracts partition columns from folder names. Enable with `hive_partitioning_enabled: true`. Given a folder structure: ``` /data/ year=2024/ month=01/ data.parquet month=02/ data.parquet ``` Configure the dataset: ``` datasets: - from: s3://bucket/data/ name: partitioned_data params: file_format: parquet hive_partitioning_enabled: true ``` Query with partition filters: ``` SELECT * FROM partitioned_data WHERE year = '2024' AND month = '01'; ``` Partition pruning improves query performance by reading only the relevant files. ### Metadata Columns[​](#metadata-columns "Direct link to Metadata Columns") File-based connectors can expose per-file object store metadata as virtual columns in the dataset schema. These columns are not stored in the data files — they are derived from object store file metadata at query time. #### Available Columns[​](#available-columns "Direct link to Available Columns") | Column | Type | Description | | ---------------- | ---------------------- | ------------------------------- | | `_location` | `Utf8` | Full URI of the source file | | `_last_modified` | `Timestamp(µs, "UTC")` | When the file was last modified | | `_size` | `UInt64` | File size in bytes | #### Enabling Metadata Columns[​](#enabling-metadata-columns "Direct link to Enabling Metadata Columns") Metadata columns are enabled by adding a `metadata` section to the dataset definition with each desired column set to `enabled`: ``` datasets: - from: s3://bucket/data/ name: my_data params: file_format: parquet metadata: _location: enabled _last_modified: enabled _size: enabled ``` Each column can be individually enabled or omitted: ``` metadata: _location: enabled # Only add the _location column ``` note If the data files already contain a column with the same name as a metadata column (e.g., a Parquet file with a `_size` column), the metadata column is not added to avoid conflicts. #### Querying Metadata Columns[​](#querying-metadata-columns "Direct link to Querying Metadata Columns") Once enabled, metadata columns appear alongside the regular data columns: ``` SELECT * FROM my_data LIMIT 3; ``` ``` +----+---------+------+-------+-----+----------------------+--------------------------------------------------------------+-------+ | id | value | year | month | day | _last_modified | _location | _size | +----+---------+------+-------+-----+----------------------+--------------------------------------------------------------+-------+ | 0 | value_0 | 2022 | 1 | 1 | 2024-10-10T05:36:59Z | s3://bucket/data/year=2022/month=1/day=1/data_0.parquet | 2317 | | 1 | value_1 | 2022 | 1 | 1 | 2024-10-10T05:36:59Z | s3://bucket/data/year=2022/month=1/day=1/data_0.parquet | 2317 | | 2 | value_2 | 2022 | 1 | 1 | 2024-10-10T05:36:59Z | s3://bucket/data/year=2022/month=1/day=1/data_0.parquet | 2317 | +----+---------+------+-------+-----+----------------------+--------------------------------------------------------------+-------+ ``` Metadata columns can be used in filters, projections, aggregations, and joins like any other column: ``` -- Filter by file location SELECT id, value FROM my_data WHERE _location = 's3://bucket/data/year=2022/month=1/day=1/data_0.parquet'; -- Find recently modified files SELECT DISTINCT _location, _last_modified FROM my_data WHERE _last_modified > '2024-01-01T00:00:00Z'; -- Aggregate by file SELECT _location, COUNT(*) AS row_count, _size FROM my_data GROUP BY _location, _size ORDER BY _location; ``` #### Applicable Connectors[​](#applicable-connectors "Direct link to Applicable Connectors") Metadata columns are supported by all file-based connectors: | Connector Type | Connectors | | ---------------------------- | --------------------------------- | | **Object Stores** | S3, Azure Blob (ABFS), HTTP/HTTPS | | **Network-Attached Storage** | FTP, SFTP, SMB, NFS | | **Local Storage** | File | ## Schema Inference[​](#schema-inference "Direct link to Schema Inference") Spice infers the schema for each dataset from its data source at startup. The inferred schema defines the column names, data types, and nullability used by the dataset for the lifetime of that runtime process. Schema inference happens once, when the dataset is first registered. Some connectors support tuning the inference behavior with connector-specific parameters: | Connector | Parameter | Default | Description | | ---------------------------------------------------------- | ---------------------------------- | ------- | --------------------------------------------------- | | [Kafka](/docs/next/components/data-connectors/kafka) | `schema_infer_max_records` | 1 | Number of messages sampled to infer the JSON schema | | [DynamoDB](/docs/next/components/data-connectors/dynamodb) | `schema_infer_max_records` | 10 | Number of items sampled to infer the schema | | [MongoDB](/docs/next/components/data-connectors/mongodb) | `mongodb_schema_infer_max_records` | 400 | Number of documents sampled to infer the schema | | [CSV files](/docs/next/reference/file_format) | `csv_schema_infer_max_records` | 1000 | Number of rows sampled to infer the CSV schema | For connectors that read self-describing formats (Parquet, Arrow, Avro), the schema is read directly from file metadata and does not require sampling. ### Runtime Schema Changes[​](#runtime-schema-changes "Direct link to Runtime Schema Changes") Spice does not apply schema changes at runtime. If the source schema changes while the runtime is running — for example, new columns are added, columns are removed, or data types change — subsequent data refreshes will fail with an error such as: ``` Failed to load data for dataset : Cannot cast struct field ... ``` This behavior is by design. Blocking runtime schema evolution protects accelerated tables from unintentional or breaking schema changes that could corrupt data or produce unexpected query results. To apply a new source schema, restart the Spice runtime. On startup, Spice re-infers the schema from the source and re-initializes the dataset with the updated column definitions. Recommendation Pin a known-good schema version in the data source or use the [`columns`](/docs/next/reference/spicepod/datasets#columns) configuration to explicitly define the expected columns. This makes schema expectations explicit and produces clear errors if the source drifts. note Runtime schema evolution controls are planned for a future release. When available, schema evolution will remain off by default. | Name | Parameter | Supported | Is Document Format | | --------------------------------------------- | ---------------------- | --------- | ------------------ | | [Apache Parquet](https://parquet.apache.org/) | `file_format: parquet` | ✅ | ❌ | | [CSV](/docs/next/reference/file_format#csv) | `file_format: csv` | ✅ | ❌ | | [Delta Lake](https://delta.io/) | `file_format: delta` | ✅ | ❌ | | [Apache Iceberg](https://iceberg.apache.org/) | `file_format: iceberg` | ✅ | ❌ | | JSON | `file_format: json` | ✅ | ❌ | | Microsoft Excel | `file_format: xlsx` | Roadmap | ❌ | | Markdown | `file_format: md` | ✅ | ✅ | | Text | `file_format: txt` | ✅ | ✅ | | PDF | `file_format: pdf` | Beta | ✅ | | Microsoft Word | `file_format: docx` | Alpha | ✅ | ### Document Formats[​](#document-formats "Direct link to Document Formats") []() Document formats (Markdown, Text, PDF, Word) are handled differently from structured data formats. Each file becomes a row in the resulting table, with the file contents stored in a `content` column. Note Document formats in Alpha (DOCX) may not parse all structure or text from the underlying documents correctly. #### Document Table Schema[​](#document-table-schema "Direct link to Document Table Schema") | Column | Type | Description | | ---------- | ------ | --------------------------------- | | `location` | String | Path to the source file | | `content` | String | Full text content of the document | #### Example[​](#example "Direct link to Example") Consider a local filesystem: ``` >>> ls -la total 232 drwxr-sr-x@ 22 jeadie staff 704 30 Jul 13:12 . drwxr-sr-x@ 18 jeadie staff 576 30 Jul 13:12 .. -rw-r--r--@ 1 jeadie staff 1329 15 Jan 2024 DR-000-Template.md -rw-r--r--@ 1 jeadie staff 4966 11 Aug 2023 DR-001-Dremio-Architecture.md -rw-r--r--@ 1 jeadie staff 2307 28 Jul 2023 DR-002-Data-Completeness.md ``` And the spicepod ``` datasets: - name: my_documents from: file:docs/decisions/ params: file_format: md ``` A Document table will be created. ``` >>> SELECT * FROM my_documents LIMIT 3 +----------------------------------------------------+--------------------------------------------------+ | location | content | +----------------------------------------------------+--------------------------------------------------+ | Users/docs/decisions/DR-000-Template.md | # DR-000: DR Template | | | **Date:** <> | | | **Decision Makers:** | | | - @<> | | | - @<> | | | ... | | Users/docs/decisions/DR-001-Dremio-Architecture.md | # DR-001: Add "Cached" Dremio Dataset | | | | | | ## Context | | | | | | We use [Dremio](https://www.dremio.com/) to p... | | Users/docs/decisions/DR-002-Data-Completeness.md | # DR-002: Append-Only Data Completeness | | | | | | ## Context | | | | | | Our Ethereum append-only dataset is incomple... | +----------------------------------------------------+--------------------------------------------------+ ``` ## Identifier Case Sensitivity and Quoting[​](#identifier-case-sensitivity-and-quoting "Direct link to Identifier Case Sensitivity and Quoting") Spice follows [PostgreSQL conventions](/docs/next/reference/sql/select) for identifier handling: **unquoted identifiers are normalized to lowercase**. This applies to both the `from` field in dataset definitions and the `name` field used for SQL queries. ### Quoting in the `from` field[​](#quoting-in-the-from-field "Direct link to quoting-in-the-from-field") To reference a table or schema with mixed-case or uppercase characters in the `from` field, wrap each case-sensitive part in double quotes: ``` datasets: # Without quoting — "ActionExecutions" is lowercased to "actionexecutions" - from: postgres:my_schema.ActionExecutions name: action_executions # With quoting — case is preserved for the table name - from: postgres:my_schema."ActionExecutions" name: action_executions # Quote each part individually as needed - from: postgres:"MySchema"."ActionExecutions" name: action_executions ``` Each dotted part of the identifier is treated independently — quote only the parts that require case preservation. For example, `postgres:my_schema."ActionExecutions"` preserves the case of `ActionExecutions` while `my_schema` is normalized to lowercase. This applies to all federated database connectors where the `from` field references a table identifier (e.g. `postgres`, `mysql`, `snowflake`, `databricks`, `clickhouse`, `mssql`, `duckdb`, `dremio`, `flightsql`, `spark`, `mongodb`, `oracle`, `adbc`). Connectors that interpret `from` as a file path (e.g. `s3`, `delta_lake`, `ftp`, `abfs`) do not apply identifier normalization. ### Quoting in the `name` field[​](#quoting-in-the-name-field "Direct link to quoting-in-the-name-field") The `name` field controls the table name used in Spice SQL queries and follows the same lowercase normalization. To preserve case in the dataset name, wrap the value in double quotes. In YAML, use single quotes around the double-quoted value: ``` datasets: - from: postgres:my_schema."ActionExecutions" name: '"ActionExecutions"' ``` ``` -- Query using the preserved-case name SELECT * FROM "ActionExecutions"; ``` If you don't need to preserve case in queries, a lowercase `name` works without quoting: ``` datasets: - from: postgres:my_schema."ActionExecutions" name: action_executions ``` ``` SELECT * FROM action_executions; ``` Dataset `name` quoting works regardless of connector type. See the [datasets `name` reference](/docs/next/reference/spicepod/datasets#name) for more details. ## Data Connector Docs[​](#data-connector-docs "Direct link to Data Connector Docs") ## [📄️Redshift Data Connector](/docs/next/components/data-connectors/redshift) [Connect to Amazon Redshift using the PostgreSQL connector in Spice.](/docs/next/components/data-connectors/redshift) ## [📄️Azure BlobFS Data Connector](/docs/next/components/data-connectors/abfs) [Azure BlobFS Data Connector Documentation](/docs/next/components/data-connectors/abfs) ## [📄️ADBC Data Connector](/docs/next/components/data-connectors/adbc) [ADBC Data Connector Documentation](/docs/next/components/data-connectors/adbc) ## [📄️ClickHouse Data Connector](/docs/next/components/data-connectors/clickhouse) [ClickHouse Data Connector Documentation](/docs/next/components/data-connectors/clickhouse) ## [🗃Azure Cosmos DB Data Connector](/docs/next/components/data-connectors/cosmosdb) [1 item](/docs/next/components/data-connectors/cosmosdb) ## [🗃Databricks Data Connector](/docs/next/components/data-connectors/databricks) [1 item](/docs/next/components/data-connectors/databricks) ## [📄️Debezium Data Connector](/docs/next/components/data-connectors/debezium) [Debezium Data Connector Documentation](/docs/next/components/data-connectors/debezium) ## [🗃Delta Lake Data Connector](/docs/next/components/data-connectors/delta-lake) [1 item](/docs/next/components/data-connectors/delta-lake) ## [🗃Dremio Data Connector](/docs/next/components/data-connectors/dremio) [1 item](/docs/next/components/data-connectors/dremio) ## [🗃DuckDB Data Connector](/docs/next/components/data-connectors/duckdb) [1 item](/docs/next/components/data-connectors/duckdb) ## [📄️DuckLake Data Connector](/docs/next/components/data-connectors/ducklake) [DuckLake Data Connector Documentation](/docs/next/components/data-connectors/ducklake) ## [🗃DynamoDB Data Connector](/docs/next/components/data-connectors/dynamodb) [1 item](/docs/next/components/data-connectors/dynamodb) ## [🗃Elasticsearch Data Connector](/docs/next/components/data-connectors/elasticsearch) [1 item](/docs/next/components/data-connectors/elasticsearch) ## [🗃File Data Connector](/docs/next/components/data-connectors/file) [1 item](/docs/next/components/data-connectors/file) ## [📄️Flight SQL Data Connector](/docs/next/components/data-connectors/flightsql) [Flight SQL Data Connector Documentation](/docs/next/components/data-connectors/flightsql) ## [📄️FTP/SFTP Data Connector](/docs/next/components/data-connectors/ftp) [FTP/SFTP Data Connector Documentation](/docs/next/components/data-connectors/ftp) ## [📄️GCS Data Connector](/docs/next/components/data-connectors/gcs) [GCS (Google Cloud Storage) Data Connector Documentation](/docs/next/components/data-connectors/gcs) ## [🗃GitHub Data Connector](/docs/next/components/data-connectors/github) [1 item](/docs/next/components/data-connectors/github) ## [📄️Glue Data Connector](/docs/next/components/data-connectors/glue) [Connect to and query tables in an AWS Glue Data Catalog](/docs/next/components/data-connectors/glue) ## [🗃GraphQL Data Connector](/docs/next/components/data-connectors/graphql) [1 item](/docs/next/components/data-connectors/graphql) ## [🗃HTTP(s) Data Connector](/docs/next/components/data-connectors/https) [1 item](/docs/next/components/data-connectors/https) ## [📄️Iceberg Data Connector](/docs/next/components/data-connectors/iceberg) [Connect to and query Apache Iceberg tables](/docs/next/components/data-connectors/iceberg) ## [📄️IMAP Data Connector](/docs/next/components/data-connectors/imap) [IMAP Data Connector Documentation](/docs/next/components/data-connectors/imap) ## [📄️Kafka Data Connector](/docs/next/components/data-connectors/kafka) [Kafka Data Connector Documentation](/docs/next/components/data-connectors/kafka) ## [📄️Localpod Data Connector](/docs/next/components/data-connectors/localpod) [Localpod Data Connector Documentation](/docs/next/components/data-connectors/localpod) ## [📄️Memory Data Connector](/docs/next/components/data-connectors/memory) [Memory Data Connector Documentation](/docs/next/components/data-connectors/memory) ## [📄️MongoDB Data Connector](/docs/next/components/data-connectors/mongodb) [MongoDB Data Connector Documentation](/docs/next/components/data-connectors/mongodb) ## [🗃Microsoft SQL Server](/docs/next/components/data-connectors/mssql) [1 item](/docs/next/components/data-connectors/mssql) ## [🗃MySQL Data Connector](/docs/next/components/data-connectors/mysql) [1 item](/docs/next/components/data-connectors/mysql) ## [📄️NFS Data Connector](/docs/next/components/data-connectors/nfs) [NFS Data Connector Documentation](/docs/next/components/data-connectors/nfs) ## [📄️ODBC Data Connector](/docs/next/components/data-connectors/odbc) [ODBC Data Connector Documentation](/docs/next/components/data-connectors/odbc) ## [📄️Oracle Data Connector](/docs/next/components/data-connectors/oracle) [Oracle Data Connector Documentation](/docs/next/components/data-connectors/oracle) ## [🗃PostgreSQL Data Connector](/docs/next/components/data-connectors/postgres) [1 item](/docs/next/components/data-connectors/postgres) ## [🗃S3 Data Connector](/docs/next/components/data-connectors/s3) [1 item](/docs/next/components/data-connectors/s3) ## [🗃ScyllaDB Data Connector](/docs/next/components/data-connectors/scylladb) [1 item](/docs/next/components/data-connectors/scylladb) ## [📄️SharePoint Data Connector](/docs/next/components/data-connectors/sharepoint) [SharePoint Data Connector Documentation](/docs/next/components/data-connectors/sharepoint) ## [📄️SMB Data Connector](/docs/next/components/data-connectors/smb) [SMB Data Connector Documentation](/docs/next/components/data-connectors/smb) ## [📄️Snowflake Data Connector](/docs/next/components/data-connectors/snowflake) [Snowflake Data Connector Documentation](/docs/next/components/data-connectors/snowflake) ## [📄️Apache Spark Connector](/docs/next/components/data-connectors/spark) [Apache Spark Connector Documentation](/docs/next/components/data-connectors/spark) ## [🗃Spice.ai Data Connector](/docs/next/components/data-connectors/spiceai) [1 item](/docs/next/components/data-connectors/spiceai) --- # Azure BlobFS Data Connector The Azure BlobFS (ABFS) Data Connector enables federated SQL queries on files stored in Azure Blob-compatible endpoints. This includes Azure Data Lake Storage Gen2 endpoints accessed via the `abfs://` and `abfss://` schemes. When a folder path is provided, all the contained files will be loaded. File formats are specified using the `file_format` parameter, as described in [File Formats](/docs/next/components/data-connectors/#file-formats). ``` datasets: - from: abfs://foocontainer/taxi_sample.csv name: azure_test params: abfs_account: spiceadls abfs_access_key: ${ secrets:access_key } file_format: csv ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Defines the ABFS-compatible URI to a folder or object: * `from: abfs:///` with the account name configured using `abfs_account` parameter, or * `from: abfs://@.dfs.core.windows.net/` ### `name`[​](#name "Direct link to name") Defines the dataset name, which is used as the table name within Spice. Example: ``` datasets: - from: abfs://foocontainer/taxi_sample.csv name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") #### Basic parameters[​](#basic-parameters "Direct link to Basic parameters") | Parameter name | Description | | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `file_format` | Specifies the data format. Required if not inferrable from `from`. Options: `parquet`, `csv`, `json`. Refer to [File Formats](/docs/next/components/data-connectors/#file-formats) for details. | | `abfs_account` | Azure storage account name | | `abfs_container_name` | Azure container name | | `abfs_sas_string` | SAS (Shared Access Signature) Token to use for authorization | | `abfs_endpoint` | Storage endpoint, default: `https://{account}.blob.core.windows.net` | | `abfs_use_emulator` | Use `true` or `false` to connect to a local emulator. Default: `false` | | `abfs_use_fabric_endpoint` | Use Microsoft Fabric endpoint. Default: `false` | | `abfs_authority_host` | Alternative authority host, default: `https://login.microsoftonline.com` | | `abfs_proxy_url` | Proxy URL | | `abfs_proxy_ca_certificate` | CA certificate for the proxy | | `abfs_proxy_excludes` | A list of hosts to exclude from proxy connections | | `abfs_disable_tagging` | Disable tagging objects. Use this if your backing store doesn't support tags | | `allow_http` | Allow insecure HTTP connections | | `client_timeout` | Optional. Timeout for Azure client operations. | | `abfs_versioning` | Enable Azure blob versioning. Default: `disabled` | | `hive_partitioning_enabled` | Enable partitioning using hive-style partitioning from the folder structure. Defaults to `false` | | `schema_source_path` | Specifies the URL used to infer the dataset schema. Default to the most recently modified file | #### Authentication parameters[​](#authentication-parameters "Direct link to Authentication parameters") The following authentication methods are mutually exclusive — only one can be used at a time: * `abfs_access_key` * `abfs_bearer_token` * `abfs_sas_string` * Client credentials (`abfs_client_id` + `abfs_client_secret` + `abfs_tenant_id`) * `abfs_use_cli` * `abfs_msi_endpoint` * `abfs_federated_token_file` * `abfs_skip_signature` If none of these are set the connector will default to using a [managed identity](https://learn.microsoft.com/en-us/entra/identity/managed-identities-azure-resources/overview) | Parameter name | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `abfs_access_key` | Secret access key | | `abfs_bearer_token` | `BEARER` access token for user authentication. The token can be obtained from the OAuth2 flow (see [access token authentication](#access-token-authentication)). | | `abfs_client_id` | Client ID for client authentication flow | | `abfs_client_secret` | Client Secret to use for client authentication flow | | `abfs_tenant_id` | Tenant ID to use for client authentication flow | | `abfs_skip_signature` | Skip credentials and request signing for public containers | | `abfs_msi_endpoint` | Endpoint for managed identity tokens | | `abfs_federated_token_file` | File path for federated identity token in Kubernetes | | `abfs_use_cli` | Set to `true` to use the Azure CLI to acquire access tokens | #### Retry parameters[​](#retry-parameters "Direct link to Retry parameters") | Parameter name | Description | | ------------------------------- | -------------------------------------------- | | `abfs_max_retries` | Maximum retries. Default: `3` | | `abfs_retry_timeout` | Total timeout for retries (e.g., `5s`, `1m`) | | `abfs_backoff_initial_duration` | Initial retry delay (e.g., `5s`) | | `abfs_backoff_max_duration` | Maximum retry delay (e.g., `1m`) | | `abfs_backoff_base` | Exponential backoff base (e.g., `0.1`) | ## Authentication[​](#authentication "Direct link to Authentication") ABFS connector supports three types of authentication, as detailed in the [authentication parameters](#authentication-parameters) ### Service principal authentication[​](#service-principal-authentication "Direct link to Service principal authentication") Configure service principal authentication by setting the `abfs_client_secret` parameter. 1. Create a new Azure AD application in the [Azure portal](https://portal.azure.com/#view/Microsoft_AAD_IAM/ActiveDirectoryMenuBlade/~/Overview) and generate a `client secret` under `Certificates & secrets`. 2. Grant the Azure AD application read access to the storage account under `Access Control (IAM)`, this can typically be done using the `Storage Blob Data Reader` built-in role. ### Access key authentication[​](#access-key-authentication "Direct link to Access key authentication") Configure service principal authentication by setting the `abfs_access_key` parameter to [Azure Storage Account Access Key](https://learn.microsoft.com/en-us/azure/storage/common/storage-account-keys-manage?tabs=azure-portal) ### Access token authentication[​](#access-token-authentication "Direct link to Access token authentication") Configure access token authentication by setting the `abfs_bearer_token` parameter, typically obtained through the following the OAuth2 flow with `spice login abfs`. 1. Create a new Azure AD application in the [Azure portal](https://portal.azure.com/#view/Microsoft_AAD_IAM/ActiveDirectoryMenuBlade/~/Overview). 2. Under the application's `API permissions`, add the permission: `Azure Storage - user_impersonation`. 3. Under the applications's `Authentication`, add `http://localhost` as Mobile and desktop applications redirect URI. 4. Grant the user read access to the storage account under `Access Control (IAM)`, this can typically be done using the `Storage Blob Data Reader` built-in role. 5. Obtain the `abfs_bearer_token` using the following command. The `abfs_bearer_token`, `abfs_client_id`, `abfs_tenant_id` will be automatically filled in environment secret after login. Refere to [`spice login`](/docs/next/cli/reference/login) documentation for more details. ``` spice login abfs --tenant-id $TENANT_ID --client-id $CLIENT_ID ``` ## Supported file formats[​](#supported-file-formats "Direct link to Supported file formats") Specify the file format using `file_format` parameter. More details in [File Formats](/docs/next/components/data-connectors/#file-formats). ## Examples[​](#examples "Direct link to Examples") ### Reading a CSV file with an Access Key[​](#reading-a-csv-file-with-an-access-key "Direct link to Reading a CSV file with an Access Key") ``` datasets: - from: abfs://foocontainer/taxi_sample.csv name: azure_test params: abfs_account: spiceadls abfs_access_key: ${ secrets:ACCESS_KEY } file_format: csv ``` ### Using Public Containers[​](#using-public-containers "Direct link to Using Public Containers") ``` datasets: - from: abfs://pubcontainer/taxi_sample.csv name: pub_data params: abfs_account: spiceadls abfs_skip_signature: true file_format: csv ``` ### Connecting to the Storage Emulator[​](#connecting-to-the-storage-emulator "Direct link to Connecting to the Storage Emulator") ``` datasets: - from: abfs://test_container/test_csv.csv name: test_data params: abfs_use_emulator: true file_format: csv ``` ### Using secrets for Account name[​](#using-secrets-for-account-name "Direct link to Using secrets for Account name") ``` datasets: - from: abfs://my_container/my_csv.csv name: prod_data params: abfs_account: ${ secrets:PROD_ACCOUNT } file_format: csv ``` ### Authenticating using Client Authentication[​](#authenticating-using-client-authentication "Direct link to Authenticating using Client Authentication") ``` datasets: - from: abfs://my_data/input.parquet name: my_data params: abfs_tenant_id: ${ secrets:MY_TENANT_ID } abfs_client_id: ${ secrets:MY_CLIENT_ID } abfs_client_secret: ${ secrets:MY_CLIENT_SECRET } ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). --- # ADBC Data Connector [ADBC](https://arrow.apache.org/adbc/) (Arrow Database Connectivity) is a columnar, minimal-overhead alternative to JDBC/ODBC for analytical data access. It transfers data using [Apache Arrow](https://arrow.apache.org/), avoiding serialization overhead between the database driver and Spice. The ADBC data connector dynamically loads any ADBC-compatible driver at runtime and provides federated SQL query access through a managed connection pool. It supports both read and write operations, and pushes filters, projections, limits, aggregations, sorting, and joins down to the source database. Drivers are available for [BigQuery](https://docs.adbc-drivers.org/drivers/bigquery/index.html), [Trino](https://docs.adbc-drivers.org/drivers/trino/index.html), [Snowflake](https://docs.adbc-drivers.org/drivers/snowflake/index.html), [Amazon Redshift](https://docs.adbc-drivers.org/drivers/redshift/index.html), [Databricks](https://docs.adbc-drivers.org/drivers/databricks/index.html), and more. See [ADBC Driver Foundry](https://docs.adbc-drivers.org/) for the full list. ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: bigquery adbc_uri: "bigquery:///my-gcp-project" adbc_driver_options: >- adbc.bigquery.sql.dataset_id=my_dataset ``` ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") An ADBC-compatible driver must be installed on the system where Spice runs. Spice loads the driver shared library by name (e.g., `bigquery`, `trino`, `snowflake`, `redshift`) or by an explicit file path. The recommended way to install drivers is with [`dbc`](https://docs.columnar.tech/dbc/), the command-line tool for installing and managing ADBC drivers: ``` # Install dbc curl -LsSf https://dbc.columnar.tech/install.sh | sh # Install drivers dbc install bigquery dbc install trino dbc install redshift ``` See the [`dbc` documentation](https://docs.columnar.tech/dbc/) for other installation methods (pip, Homebrew, Windows MSI) and platform support. Spice includes built-in SQL dialect support for BigQuery, translating federated queries into BigQuery-compatible SQL automatically. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `adbc:table_name`, where `table_name` is the name of the table to read from the connected database. Spice normalizes unquoted identifiers to lowercase. To preserve mixed-case or uppercase table names, wrap each case-sensitive part in double quotes: ``` datasets: # BigQuery table with mixed-case name - from: adbc:my_dataset."ActionExecutions" name: action_executions params: adbc_driver: bigquery adbc_uri: "bigquery:///my-gcp-project" # Quote each part individually as needed - from: adbc:"MyDataset"."ActionExecutions" name: action_executions params: adbc_driver: bigquery adbc_uri: "bigquery:///my-gcp-project" ``` Table name casing conventions vary by source database — Snowflake uses uppercase by default, while BigQuery and Trino use lowercase. When `adbc_catalog` and `adbc_schema` are not set, fully qualified names can be specified in the `from` field. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting) for details. ### `name`[​](#name "Direct link to name") The dataset name, used as the table name within Spice. ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: bigquery adbc_uri: "bigquery:///my-gcp-project" adbc_driver_options: >- adbc.bigquery.sql.dataset_id=my_dataset ``` ``` SELECT COUNT(*) FROM my_table; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | `adbc_driver` | Required. The ADBC driver name (e.g., `bigquery`, `trino`, `snowflake`, `redshift`). | | `adbc_uri` | Required. Database URI or connection string for the ADBC driver. In-memory URIs (e.g., `:memory:`) are not supported. | | `adbc_driver_path` | Optional. Absolute path to the ADBC driver shared library. When omitted, the driver is loaded by name from the system library path. | | `adbc_username` | Optional. Username for database authentication. Supports [Secrets Stores](/docs/next/components/secret-stores). | | `adbc_password` | Optional. Password for database authentication. Supports [Secrets Stores](/docs/next/components/secret-stores). | | `adbc_driver_options` | Optional. Semicolon-delimited key-value pairs of driver-specific options. See [Driver Options](#driver-options-adbc_driver_options). | | `adbc_catalog` | Optional. Sets the default catalog for the connection. | | `adbc_schema` | Optional. Sets the default schema for the connection. | | `connection_pool_size` | Optional. Maximum number of connections in the connection pool. Default: `5`. | | `connection_pool_min_idle` | Optional. Minimum number of idle connections in the pool. Default: `1`. | | `query_federation` | Optional. Controls whether queries are federated to the ADBC source. Values: `enabled`, `disabled`. Default: `enabled`. | In-memory databases In-memory database URIs (e.g., `:memory:` or URIs containing `mode=memory`) are not supported. ### Driver Options (`adbc_driver_options`)[​](#driver-options-adbc_driver_options "Direct link to driver-options-adbc_driver_options") The `adbc_driver_options` parameter passes driver-specific configuration as semicolon-delimited `key=value` pairs. Each key is automatically prefixed with `adbc.` if it does not already start with that prefix. For example, the following two configurations are equivalent: ``` # Without adbc. prefix (prefix is added automatically) adbc_driver_options: bigquery.sql.project_id=my-project;bigquery.sql.dataset_id=my_dataset # With explicit adbc. prefix adbc_driver_options: adbc.bigquery.sql.project_id=my-project;adbc.bigquery.sql.dataset_id=my_dataset ``` Trailing semicolons are permitted. Entries without an `=` sign or with an empty key are ignored. For multi-line readability, use YAML's `>-` folded block scalar: ``` adbc_driver_options: >- adbc.bigquery.sql.project_id=my-project; adbc.bigquery.sql.dataset_id=my_dataset ``` Driver-specific options are documented at [ADBC Driver Foundry](https://docs.adbc-drivers.org/). #### BigQuery Driver Options[​](#bigquery-driver-options "Direct link to BigQuery Driver Options") The [BigQuery ADBC driver](https://docs.adbc-drivers.org/drivers/bigquery/index.html) accepts the following options through `adbc_driver_options`. **Connection** | Option Key | Description | | ------------------------------ | ----------------------------------- | | `adbc.bigquery.sql.project_id` | The Google Cloud project ID. | | `adbc.bigquery.sql.dataset_id` | The BigQuery dataset ID. | | `adbc.bigquery.sql.table_id` | The BigQuery table ID. | | `adbc.bigquery.sql.location` | The BigQuery location (e.g., `US`). | **Authentication** | Option Key | Description | | -------------------------------------- | --------------------------------------------------------------------------------------------------- | | `adbc.bigquery.sql.auth_type` | Authentication type. One of the values below. Default: `adbc.bigquery.sql.auth_type.auth_bigquery`. | | `adbc.bigquery.sql.auth_credentials` | Credential data: file path or JSON string, depending on `auth_type`. | | `adbc.bigquery.sql.auth.client_id` | OAuth client ID (for `oauth_client_ids` auth type). | | `adbc.bigquery.sql.auth.client_secret` | OAuth client secret (for `oauth_client_ids` auth type). | | `adbc.bigquery.sql.auth.refresh_token` | OAuth refresh token (for `user_authentication` auth type). | | `adbc.bigquery.sql.auth.quota_project` | Quota project for billing. | Supported `auth_type` values: | Value | Description | | ----------------------------------------------------- | --------------------------------------------- | | `adbc.bigquery.sql.auth_type.auth_bigquery` | Default BigQuery authentication. | | `adbc.bigquery.sql.auth_type.json_credential_file` | Service account JSON key file. | | `adbc.bigquery.sql.auth_type.json_credential_string` | Service account JSON as a string. | | `adbc.bigquery.sql.auth_type.json_credentials` | JSON credentials as a byte array. | | `adbc.bigquery.sql.auth_type.user_authentication` | User-based OAuth authentication. | | `adbc.bigquery.sql.auth_type.app_default_credentials` | Google Application Default Credentials (ADC). | | `adbc.bigquery.sql.auth_type.oauth_client_ids` | OAuth client ID credentials. | **Service Account Impersonation** | Option Key | Description | | ------------------------------------------------ | ---------------------------------------------------------- | | `adbc.bigquery.sql.impersonate.target_principal` | Service account email to impersonate. | | `adbc.bigquery.sql.impersonate.delegates` | Delegation chain (comma-separated service account emails). | | `adbc.bigquery.sql.impersonate.scopes` | OAuth 2.0 scopes (comma-separated). | | `adbc.bigquery.sql.impersonate.lifetime` | Token lifetime duration (e.g., `3600s`). | **Query** | Option Key | Type | Description | Default | | --------------------------------------------------- | ------ | ----------------------------------------------------------- | ------------ | | `adbc.bigquery.sql.query.parameter_mode` | string | Query parameter mode: `positional` (`?`) or `named` (`@p`). | `positional` | | `adbc.bigquery.sql.query.destination_table` | string | Destination table for query results. | | | `adbc.bigquery.sql.query.default_project_id` | string | Default project ID for queries. | | | `adbc.bigquery.sql.query.default_dataset_id` | string | Default dataset ID for queries. | | | `adbc.bigquery.sql.query.create_disposition` | string | Table creation behavior (e.g., `CREATE_IF_NEEDED`). | | | `adbc.bigquery.sql.query.write_disposition` | string | Table write behavior (e.g., `WRITE_TRUNCATE`). | | | `adbc.bigquery.sql.query.disable_query_cache` | bool | Disable query cache. | `false` | | `adbc.bigquery.sql.query.disable_flattened_results` | bool | Disable flattened results. | `false` | | `adbc.bigquery.sql.query.allow_large_results` | bool | Allow large query results. | `false` | | `adbc.bigquery.sql.query.priority` | string | Query priority (`BATCH` or `INTERACTIVE`). | | | `adbc.bigquery.sql.query.max_billing_tier` | int | Maximum billing tier. | | | `adbc.bigquery.sql.query.max_bytes_billed` | int | Maximum bytes billed. | | | `adbc.bigquery.sql.query.use_legacy_sql` | bool | Use legacy SQL syntax. | `false` | | `adbc.bigquery.sql.query.dry_run` | bool | Execute a dry run (no data returned). | `false` | | `adbc.bigquery.sql.query.create_session` | bool | Create a session for the query. | `false` | | `adbc.bigquery.sql.query.job_timeout` | int | Job timeout in milliseconds. | | | `adbc.bigquery.sql.query.result_buffer_size` | int | Result buffer size. | `200` | | `adbc.bigquery.sql.query.prefetch_concurrency` | int | Number of concurrent prefetch operations. | `10` | BigQuery also supports connection string URIs as the `adbc_uri` value: ``` bigquery:///my-project-123 bigquery:///my-project-123?OAuthType=1&AuthCredentials=/path/to/key.json bigquery:///my-project-123?DatasetId=analytics&Location=US ``` See the [BigQuery ADBC driver documentation](https://docs.adbc-drivers.org/drivers/bigquery/index.html) for the full reference. #### Trino Driver Options[​](#trino-driver-options "Direct link to Trino Driver Options") The [Trino ADBC driver](https://docs.adbc-drivers.org/drivers/trino/index.html) uses a connection string URI as the `adbc_uri` value: ``` trino://[user[:password]@]host[:port][/catalog[/schema]][?param=value&...] ``` Examples: ``` trino://trino.example.com:8080/hive/default trino://user:pass@trino.example.com:8443/postgresql/public?SSL=true trino://user@localhost:8443/memory/default?SSLVerification=NONE ``` By default, connections use HTTPS. Add `SSL=false` to use HTTP. For self-signed certificates, add `SSLVerification=NONE`. See the [Trino ADBC driver documentation](https://docs.adbc-drivers.org/drivers/trino/index.html) for the full reference. ### Catalog and Schema[​](#catalog-and-schema "Direct link to Catalog and Schema") The `adbc_catalog` and `adbc_schema` parameters set connection-level defaults that apply to all queries on the connection. These map to the ADBC standard connection options `adbc.connection.catalog` and `adbc.connection.db_schema`. How these map to source database concepts varies by driver: | Driver | `adbc_catalog` | `adbc_schema` | | --------- | ------------------ | --------------------- | | BigQuery | GCP project ID | BigQuery dataset | | Trino | Trino catalog | Schema within catalog | | Redshift | Redshift database | Redshift schema | | Snowflake | Snowflake database | Snowflake schema | ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: bigquery adbc_uri: "bigquery:///my-gcp-project" adbc_catalog: my-gcp-project adbc_schema: my_dataset ``` ### Connection Pooling[​](#connection-pooling "Direct link to Connection Pooling") The ADBC connector maintains a pool of database connections for concurrent query execution. The pool is configured with: * `connection_pool_size`: The maximum number of connections. Increase for workloads with many concurrent queries. * `connection_pool_min_idle`: The minimum number of idle connections kept open to reduce connection setup latency. Both values must be positive integers, and `connection_pool_min_idle` must not exceed `connection_pool_size` — a larger value causes connection pool initialization to fail. ### Query Pushdown[​](#query-pushdown "Direct link to Query Pushdown") The ADBC connector pushes SQL operations down to the source database when possible, reducing the amount of data transferred: * **Filter pushdown**: `WHERE` clauses are pushed to the source. * **Projection pushdown**: Only the columns referenced in the query are fetched. * **Limit pushdown**: `LIMIT` clauses are applied at the source. * **Aggregation pushdown**: `GROUP BY`, `SUM`, `COUNT`, `AVG`, and other aggregations are executed on the source. * **Sort pushdown**: `ORDER BY` clauses are applied at the source. * **Join pushdown**: Joins between datasets from the same ADBC URI are executed on the remote database. No special configuration is required. Pushdown happens automatically when the source database supports the operation. ## Auth[​](#auth "Direct link to Auth") Authentication varies by driver. Credentials can be provided through `adbc_username`, `adbc_password`, `adbc_driver_options`, the connection URI, or through [Secrets Stores](/docs/next/components/secret-stores). For BigQuery, authentication typically uses Google Cloud Application Default Credentials or a service account JSON file passed through `adbc_driver_options`. For Trino, credentials are embedded in the connection URI (`trino://user:pass@host/catalog`). For drivers that accept `adbc_username` and `adbc_password` (e.g., Redshift, Snowflake), these can be set through environment variables: ``` SPICE_SECRET_ADBC_USERNAME=myuser \ SPICE_SECRET_ADBC_PASSWORD=mypassword \ spice run ``` ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: redshift adbc_uri: "redshift://my-cluster.region.redshift.amazonaws.com:5439/dev" adbc_username: ${secrets:ADBC_USERNAME} adbc_password: ${secrets:ADBC_PASSWORD} ``` ## Examples[​](#examples "Direct link to Examples") ### BigQuery[​](#bigquery "Direct link to BigQuery") Connect to a BigQuery table using Application Default Credentials. Spice translates federated queries into BigQuery-compatible SQL automatically. ``` # Install the driver dbc install bigquery # Authenticate with Google Cloud gcloud auth application-default login ``` ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: bigquery adbc_uri: "bigquery:///my-gcp-project?DatasetId=my_dataset" ``` #### BigQuery with Service Account[​](#bigquery-with-service-account "Direct link to BigQuery with Service Account") Authenticate with a service account JSON key file: ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: bigquery adbc_uri: "bigquery:///my-gcp-project" adbc_catalog: my-gcp-project adbc_schema: my_dataset adbc_driver_options: >- adbc.bigquery.sql.auth_type=adbc.bigquery.sql.auth_type.json_credential_file; adbc.bigquery.sql.auth_credentials=/path/to/service-account.json ``` #### BigQuery with Service Account JSON Secret[​](#bigquery-with-service-account-json-secret "Direct link to BigQuery with Service Account JSON Secret") Authenticate using a service account JSON string stored in a [Secrets Store](/docs/next/components/secret-stores): ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: bigquery adbc_uri: "bigquery:///my-gcp-project" adbc_driver_options: >- adbc.bigquery.sql.auth_type=adbc.bigquery.sql.auth_type.json_credential_string; adbc.bigquery.sql.auth_credentials=${secrets:BIGQUERY_SERVICE_ACCOUNT_JSON}; adbc.bigquery.sql.dataset_id=my_dataset ``` ### Trino[​](#trino "Direct link to Trino") Connect to a Trino cluster: ``` dbc install trino ``` ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: trino adbc_uri: "trino://trino.example.com:8080/hive/default?SSL=false" ``` #### Trino with Authentication[​](#trino-with-authentication "Direct link to Trino with Authentication") Connect over HTTPS with HTTP Basic authentication: ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: trino adbc_uri: "trino://user:${secrets:TRINO_PASSWORD}@trino.example.com:8443/hive/default" connection_pool_size: 10 ``` For self-signed certificates, add `SSLVerification=NONE`: ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: trino adbc_uri: "trino://user@trino.internal:8443/hive/default?SSLVerification=NONE" ``` ### Amazon Redshift[​](#amazon-redshift "Direct link to Amazon Redshift") Connect to a Redshift Provisioned cluster: ``` dbc install redshift ``` ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: redshift adbc_uri: "redshift://admin:${secrets:REDSHIFT_PASSWORD}@my-cluster.region.redshift.amazonaws.com:5439/dev" ``` #### Redshift Serverless with IAM[​](#redshift-serverless-with-iam "Direct link to Redshift Serverless with IAM") ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: redshift adbc_uri: "redshift:///dev?cluster_type=redshift-serverless&workgroup_name=my-workgroup" ``` ### Custom Driver Path[​](#custom-driver-path "Direct link to Custom Driver Path") When the ADBC driver shared library is not on the system library path (e.g., not installed via `dbc`), specify its location with `adbc_driver_path`: ``` datasets: - from: adbc:my_table name: my_table params: adbc_driver: bigquery adbc_driver_path: /opt/drivers/libadbc_driver_bigquery.so adbc_uri: "bigquery:///my-gcp-project?DatasetId=my_dataset" ``` --- # ClickHouse Data Connector ClickHouse is a fast, open-source columnar database management system designed for online analytical processing (OLAP) and real-time analytics. This connector enables federated SQL queries from a ClickHouse server. ``` datasets: - from: clickhouse:my.dataset name: my_dataset ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field for the ClickHouse connector takes the form of `from:db.dataset` where `db.dataset` is the path to the Dataset within ClickHouse. In the example above it would be `my.dataset`. The `clickhouse_db` parameter is required when not using `clickhouse_connection_string`. When using a connection string without a database path, it defaults to the `default` database. info Unquoted identifiers are normalized to lowercase. To reference a table or database with mixed-case characters, wrap each case-sensitive part in double quotes: `clickhouse:my_db."MixedCaseTable"`. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. ``` datasets: - from: clickhouse:my.dataset name: cool_dataset ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords) or any of the following keywords that are reserved by ClickHouse: * `PREWHERE` * `SETTINGS` * `FORMAT` ### `params`[​](#params "Direct link to params") The ClickHouse data connector can be configured by providing the following `params`: | Parameter Name | Definition | | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `clickhouse_connection_string` | The connection string to use to connect to the ClickHouse server. This can be used instead of providing individual connection parameters. | | `clickhouse_host` | The hostname of the ClickHouse server. | | `clickhouse_tcp_port` | The port of the ClickHouse server. | | `clickhouse_db` | The name of the database to connect to. | | `clickhouse_user` | The username to connect with. | | `clickhouse_pass` | The password to connect with. | | `clickhouse_secure` | Optional. Specifies the SSL/TLS behavior for the connection, supported values:
- `true`: (default) This mode requires an SSL connection. If a secure connection cannot be established, server will not connect.
- `false`: This mode will not attempt to use an SSL connection, even if the server supports it. | | `connection_timeout` | Optional. Specifies the connection timeout in milliseconds. Default is `10000` (10 seconds). | ## Types[​](#types "Direct link to Types") The table below shows the ClickHouse data types supported, along with the type mapping to Apache Arrow types in Spice. | ClickHouse Type | Arrow Type | | --------------- | ------------------------- | | `Bool` | `Boolean` | | `Int8` | `Int8` | | `Int16` | `Int16` | | `Int32` | `Int32` | | `Int64` | `Int64` | | `UInt8` | `UInt8` | | `UInt16` | `UInt16` | | `UInt32` | `UInt32` | | `UInt64` | `UInt64` | | `Float32` | `Float32` | | `Float64` | `Float64` | | `Decimal` | `Decimal128` | | `String` | `Utf8` | | `FixedString` | `Utf8` | | `UUID` | `Utf8` | | `Date` | `Date32` | | `Date32` | `Date32` | | `DateTime` | `Timestamp(Second, None)` | | `Nullable(T)` | Mapped inner type `T` | ## Examples[​](#examples "Direct link to Examples") ### Connecting to localhost[​](#connecting-to-localhost "Direct link to Connecting to localhost") ``` datasets: - from: clickhouse:my.dataset name: my_dataset params: clickhouse_host: localhost clickhouse_tcp_port: 9000 clickhouse_db: my_database clickhouse_user: my_user clickhouse_pass: ${secrets:my_clickhouse_pass} connection_timeout: 10000 clickhouse_secure: false ``` ### Specifying a connection timeout[​](#specifying-a-connection-timeout "Direct link to Specifying a connection timeout") ``` datasets: - from: clickhouse:my.dataset name: my_dataset params: clickhouse_connection_string: tcp://my_user:${secrets:my_clickhouse_pass}@localhost:9000/my_database connection_timeout: 10000 clickhouse_secure: true ``` ### Using a connection string[​](#using-a-connection-string "Direct link to Using a connection string") ``` datasets: - from: clickhouse:my.dataset name: my_dataset params: clickhouse_connection_string: tcp://my_user:${secrets:my_clickhouse_pass}@localhost:9000/my_database?connection_timeout=10000&secure=true ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure ClickHouse as data connector in Spice. [Clickhouse Data Connector](https://github.com/spiceai/cookbook/tree/trunk/clickhouse#readme) --- # Azure Cosmos DB Data Connector The Azure Cosmos DB Data Connector exposes Cosmos DB containers (NoSQL / Core SQL API) as SQL tables in Spice. The connector samples a configurable number of documents at startup, infers an Arrow schema, and streams documents into DataFusion for federated SQL queries alongside data from other connectors. ``` datasets: - from: cosmosdb:store.products name: products params: cosmosdb_connection_string: ${secrets:COSMOSDB_CONNECTION_STRING} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `cosmosdb:{database}.{container}` or `cosmosdb:{database}/{container}`. The connector also accepts a bare `{container}` when `cosmosdb_database` is provided in `params`. ``` datasets: - from: cosmosdb:store.orders name: orders ``` ### `name`[​](#name "Direct link to name") The dataset name used as the table name within Spice. The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") #### Authentication[​](#authentication "Direct link to Authentication") Provide either a full Cosmos DB connection string (preferred) or the discrete `account_endpoint` + `account_key` pair. Secrets must be sourced from a [secret store](/docs/next/components/secret-stores) in production. | Parameter Name | Description | Required | | ---------------------------- | -------------------------------------------------------------------------------------------------------------- | -------------------------------- | | `cosmosdb_connection_string` | Full connection string copied from the Azure portal. Takes precedence over `account_endpoint` / `account_key`. | Either this or both endpoint+key | | `cosmosdb_account_endpoint` | Account endpoint URL, e.g. `https://my-account.documents.azure.com:443/`. | When connection string isn't set | | `cosmosdb_account_key` | Primary or secondary account key. | When connection string isn't set | Microsoft Entra ID and managed-identity authentication are tracked as a post-RC enhancement and are not supported in the current release. #### Data Shape[​](#data-shape "Direct link to Data Shape") | Parameter Name | Description | Default | | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------- | | `cosmosdb_database` | Database name. When unset, parsed from the `from:` path (`database.container`). | - | | `query` | Cosmos SQL query used to scan the container. Useful when the container is large and only a subset should be surfaced as a dataset. | `SELECT * FROM c` | | `schema_infer_max_records` | Number of documents sampled during schema inference at dataset registration. Larger samples produce a more precise schema at the cost of more RU consumption. | `100` | #### Resilience[​](#resilience "Direct link to Resilience") The connector applies per-account concurrency limits, bounded retries with backoff, and a permanent-error latch that disables the connector account-wide on 401/403/404 responses. | Parameter Name | Description | Default | | ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------- | | `max_concurrent_requests` | Maximum number of concurrent Cosmos DB requests per account endpoint, shared across all datasets pointing at the same account. | `4` | | `http_max_retries` | Maximum number of retries for transient errors (HTTP 429, 5xx, network) during the schema-inference pass at dataset registration. Retries honor `Retry-After` and `x-ms-retry-after-ms` headers. | `3` | | `backoff_method` | Backoff strategy between retries. `exponential` doubles the delay each attempt (capped at 30s); `fibonacci` follows the Fibonacci sequence (capped at 30s). | `exponential` | | `disable_on_permanent_error` | When `true`, a permanent error (401/403/404) latches the connector into a disabled state and short-circuits subsequent requests until Spice is restarted. | `true` | See the [deployment guide](/docs/next/components/data-connectors/cosmosdb/deployment) for sizing, troubleshooting, and observability details. ## Authentication[​](#authentication-1 "Direct link to Authentication") ### Connection String (recommended)[​](#connection-string-recommended "Direct link to Connection String (recommended)") Copy the connection string from the Azure portal under **Settings → Keys** for the Cosmos DB account, then reference it from a secret store: ``` datasets: - from: cosmosdb:store.products name: products params: cosmosdb_connection_string: ${secrets:COSMOSDB_CONNECTION_STRING} ``` ### Explicit Endpoint and Key[​](#explicit-endpoint-and-key "Direct link to Explicit Endpoint and Key") When the endpoint and key are stored separately (for example in Key Vault), provide both: ``` datasets: - from: cosmosdb:store.products name: products params: cosmosdb_account_endpoint: https://my-account.documents.azure.com:443/ cosmosdb_account_key: ${secrets:COSMOSDB_ACCOUNT_KEY} ``` If both styles are supplied, the connection string takes precedence. ## Schema Inference[​](#schema-inference "Direct link to Schema Inference") Cosmos DB has no native schema. At dataset registration the connector runs the configured `query` (default `SELECT * FROM c`) limited to `schema_infer_max_records` documents and hands the result to Arrow's JSON inference. The inferred schema is locked for the lifetime of the runtime process. ### JSON → Arrow type mapping[​](#json--arrow-type-mapping "Direct link to JSON → Arrow type mapping") | Cosmos / JSON value | Arrow type | Notes | | --------------------------- | ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | | `"abc"` | `Utf8` | | | Integer (`42`, `-7`) | `Int64` | Widens to `Float64` if any sampled document contains a decimal value for the same field. | | Floating (`3.14`, `1.0e9`) | `Float64` | | | `true` / `false` | `Boolean` | | | Object `{ ... }` | `Struct` | Nested objects are preserved as Arrow structs. | | Array `[ ... ]` | `List` | The element type is inferred from the first non-null item; heterogeneous arrays may surface as `Utf8` or require a wider sample to disambiguate. | | All-null in sample | `Null` | Warn-dropped by default. Set `unsupported_type_action: string` to coerce to `Utf8`, or widen the sample so real values appear. | | System fields (`_rid`, ...) | stripped | The system fields `_rid`, `_self`, `_etag`, `_attachments`, and `_ts` are stripped and never appear in the dataset schema. | Cosmos does not emit `Date`, `Time`, `Timestamp`, `Decimal`, or `Binary` natively — they round-trip as strings and should be handled with `CAST` at query time. ### Tuning Schema Inference[​](#tuning-schema-inference "Direct link to Tuning Schema Inference") When optional fields are sparse in the first 100 documents but present in production data, increase `schema_infer_max_records`: ``` datasets: - from: cosmosdb:store.orders name: orders params: cosmosdb_connection_string: ${secrets:COSMOSDB_CONNECTION_STRING} schema_infer_max_records: '500' ``` Each unit increase costs additional Request Units (RUs) at startup. Pin a schema explicitly via `columns:` for the most precise control. ## `unsupported_type_action`[​](#unsupported_type_action "Direct link to unsupported_type_action") Optional. Controls behavior for fields that infer as `DataType::Null` (every sampled document had `null` for the field). Defaults to `warn`. * `error` — Fail dataset registration. * `warn` — Log a warning and drop the column. (Default.) * `ignore` — Silently drop the column. * `string` — Coerce the column to `Utf8`. ``` datasets: - from: cosmosdb:store.products name: products unsupported_type_action: string params: cosmosdb_connection_string: ${secrets:COSMOSDB_CONNECTION_STRING} ``` ## Querying[​](#querying "Direct link to Querying") After registering a dataset, query it like any other Spice table: ``` SELECT id, name, price FROM products WHERE price > 100 ORDER BY price DESC LIMIT 10; ``` ``` SELECT category, COUNT(*) AS count, AVG(price) AS avg_price FROM products GROUP BY category ORDER BY count DESC; ``` ### Custom Cosmos SQL Queries[​](#custom-cosmos-sql-queries "Direct link to Custom Cosmos SQL Queries") When the container is large and only a subset should be surfaced as a dataset, push the predicate to Cosmos with a custom `query`: ``` datasets: - from: cosmosdb:store.orders name: active_orders params: cosmosdb_connection_string: ${secrets:COSMOSDB_CONNECTION_STRING} query: "SELECT * FROM c WHERE c.status = 'active'" ``` ### Joins Across Containers[​](#joins-across-containers "Direct link to Joins Across Containers") Cosmos DB does not support joins across containers. Spice federates joins between Cosmos-backed datasets (and any other connector) in the local DataFusion engine: ``` SELECT o.id AS order_id, p.name AS product_name, p.price AS unit_price FROM active_orders o JOIN products p ON o.product_id = p.id LIMIT 50; ``` ### Acceleration[​](#acceleration "Direct link to Acceleration") Standard Spice acceleration (DuckDB, SQLite, Arrow in-memory, Cayenne) works on top of the Cosmos DB connector. Acceleration is recommended when the container is large or when query latency matters — it avoids per-query RU consumption against the Cosmos account. ``` datasets: - from: cosmosdb:store.products name: products params: cosmosdb_connection_string: ${secrets:COSMOSDB_CONNECTION_STRING} acceleration: enabled: true engine: duckdb mode: file refresh_check_interval: 1h ``` ## Limitations[​](#limitations "Direct link to Limitations") * **Read-only**: Only `SELECT` scans are supported. Writes (`INSERT` / `UPDATE` / `DELETE`) are not implemented. * **No filter / projection / limit pushdown into Cosmos**: SQL predicates are evaluated locally by DataFusion after the documents have been streamed in. Use the `query` parameter to narrow at the Cosmos side. * **Schema is frozen at registration**: New fields added in Cosmos after startup are not picked up until the runtime restarts. * **No change feed**: `acceleration.refresh_mode: changes` is not supported. * **No native temporal / decimal / binary types**: Cosmos stores everything as JSON. Round-trip these as strings and `CAST` in SQL. * **Microsoft Entra ID / managed identity not supported**: Use account keys via connection string or `account_endpoint` + `account_key`. ## Cookbook[​](#cookbook "Direct link to Cookbook") A copy-pasteable example Spicepod is in the runtime repo at [`examples/cosmosdb-connector/`](https://github.com/spiceai/spiceai/tree/trunk/examples/cosmosdb-connector). --- # Azure Cosmos DB Data Connector Deployment Guide Production operating guide for the Azure Cosmos DB (NoSQL / Core SQL API) data connector covering authentication, Request Unit (RU) cost, resilience tuning, observability, and troubleshooting. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The connector currently supports key-based authentication only. Microsoft Entra ID and managed identity are tracked as a post-RC enhancement. | Parameter | Description | | ---------------------------- | ------------------------------------------------------------------------------------------------------ | | `cosmosdb_connection_string` | Full connection string from the Azure portal (`AccountEndpoint=...;AccountKey=...`). Takes precedence. | | `cosmosdb_account_endpoint` | Account endpoint URL when storing endpoint and key separately. | | `cosmosdb_account_key` | Primary or secondary account key. | Credentials must be sourced from a [secret store](/docs/next/components/secret-stores) in production. Prefer the **secondary** account key for Spice and rotate keys via the Azure portal — this lets you revoke access without taking the primary down. Because the connector authenticates with account keys, access cannot be narrowed with data-plane RBAC roles such as **Cosmos DB Built-in Data Reader** — those apply only to Microsoft Entra ID identities, which this connector does not yet support (see above). ### TLS[​](#tls "Direct link to TLS") Cosmos DB endpoints are HTTPS-only. The Azure-issued certificate is signed by a public CA, so no extra trust-store configuration is required. Self-hosted gateways or proxies in front of Cosmos must be trusted by the runtime's host OS / container. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Per-Account Concurrency Budget[​](#per-account-concurrency-budget "Direct link to Per-Account Concurrency Budget") The connector enforces a per-account concurrency semaphore that is shared across every dataset targeting the same Cosmos endpoint. This matches Cosmos DB's per-account RU model — multiple datasets pointing at the same account compete for the same backend budget. | Parameter | Default | Notes | | ------------------------- | ------- | --------------------------------------------------------------------------------------------------------------------- | | `max_concurrent_requests` | `4` | Per-account upper bound. Datasets configured with conflicting values keep the **first-seen** value and log a warning. | For workloads that fan out across many datasets, raise the budget (e.g. `8`–`16`) only after observing how it affects the account's provisioned RU/s consumption. Datasets that rarely query against the same account can each set their own value. ### Retries and Backoff[​](#retries-and-backoff "Direct link to Retries and Backoff") Retries apply to the **schema-inference sampling pass** at dataset registration. Errors surfaced *during* a streaming scan propagate immediately — a `FeedPager` cannot be safely rewound after rows have been emitted. Spice's dataset refresh layer handles retry at the query boundary. | Parameter | Default | Behavior | | ------------------ | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `http_max_retries` | `3` | Retries for HTTP 429, 5xx, and transient network failures. The connector honors `Retry-After` and `x-ms-retry-after-ms` headers; the effective sleep is `max(retry_after, backoff)`. | | `backoff_method` | `exponential` | `exponential`: 500ms × 2ⁿ, capped at 30s. `fibonacci`: 500ms × Fₙ, capped at 30s. | For accounts at provisioned RU limits, prefer `fibonacci` — it grows slower than exponential between attempts 3 and 5 and reduces head-of-line stalls for downstream datasets sharing the budget. ### Permanent-Error Latch[​](#permanent-error-latch "Direct link to Permanent-Error Latch") A 401 (unauthorized), 403 (forbidden), or 404 (not found) from any request flips a per-account flag that short-circuits subsequent requests. This avoids a thundering herd of failed calls when credentials are wrong or the database/container has been deleted. | Parameter | Default | Behavior | | ---------------------------- | ------- | ---------------------------------------------------------------------------------------- | | `disable_on_permanent_error` | `true` | When `true`, latches the connector account-wide on 401/403/404 until Spice is restarted. | The latch is per-account-endpoint, not per-dataset — fixing the credentials and restarting clears the state. Set to `false` only in development when you want to see every failure surface immediately. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") ### Request Units (RU)[​](#request-units-ru "Direct link to Request Units (RU)") Every Cosmos DB read consumes RUs from the account's provisioned (or autoscale) budget. The Spice connector contributes RU consumption in three phases: 1. **Dataset registration** — schema inference samples up to `schema_infer_max_records` documents (default `100`). 2. **Per-query scans** — non-accelerated datasets stream the entire `query` result set on each query. 3. **Acceleration refresh** — accelerated datasets stream the entire query at every `refresh_check_interval` (or `refresh_cron`). For accounts close to their RU ceiling: * **Accelerate the dataset** (`acceleration.enabled: true`) to amortize RU cost across queries. * **Narrow the dataset** with a custom `query: SELECT * FROM c WHERE ...` to push the predicate to Cosmos. * **Tune `refresh_check_interval`** to control how often the connector replays the scan against the account. * **Lower `schema_infer_max_records`** if the schema is stable and the default sample is an avoidable RU cost on dataset registration. Cosmos DB exposes RU consumption per query in the response headers. Monitor account-level RU/s in the Azure portal under **Insights → Throughput** — sustained 429 retries indicate the account is undersized for the Spice workload. ### Schema Inference Cost[​](#schema-inference-cost "Direct link to Schema Inference Cost") Each dataset's schema inference samples documents once at registration. The cost is roughly `schema_infer_max_records × per-document RU cost`. For containers with large documents (multi-KB JSON), prefer a smaller sample (e.g. `50`) and pin the schema explicitly via `columns:` if needed. ### Connection Pool[​](#connection-pool "Direct link to Connection Pool") The connector uses a single shared HTTP/2 connection pool to each account endpoint. Cosmos DB's gateway tolerates many concurrent streams over a single connection — the bottleneck is RU/s, not TCP sockets. ## Metrics[​](#metrics "Direct link to Metrics") The Cosmos DB connector exposes one observable gauge, registered automatically for every dataset: | Metric Name | Description | | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `inflight_operations` | Number of Cosmos DB operations currently holding a concurrency permit. Incremented per operation and held across retry-backoff sleeps. **Per-dataset**, not per-account. | This metric is auto-registered — no configuration is required to export it. To disable it for a dataset, set `enabled: false` in the dataset's `metrics` section: ``` datasets: - from: cosmosdb:store.products name: products params: cosmosdb_connection_string: ${secrets:COSMOSDB_CONNECTION_STRING} metrics: - name: inflight_operations enabled: false ``` See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. For broader observability, also monitor: * Spice query execution metrics (`query_duration_ms`, `query_returned_rows`, `query_failures`) from `runtime.metrics`. * Azure portal **Cosmos DB account → Insights → Throughput** for RU/s consumption and 429 rates. * Account-level Azure Monitor metrics: `TotalRequestUnits`, `TotalRequests`, `MetadataRequests`. ## Task History[​](#task-history "Direct link to Task History") Cosmos DB requests participate in [task history](/docs/next/reference/task_history) through the connector span. Each query is captured as a child of the enclosing `sql_query` or `accelerated_table_refresh` task. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Read-only**: Writes (`INSERT` / `UPDATE` / `DELETE`) are not supported. * **No filter / projection / limit pushdown**: SQL predicates are evaluated locally by DataFusion. Use a custom `query:` to narrow at the Cosmos side. * **Schema is frozen at registration**: Mapping changes after startup require a runtime restart. * **No change feed**: `RefreshMode::Changes` is not wired. * **Mid-stream retries are not safe**: Retries apply to the schema-inference pass only. Errors during a streaming scan propagate immediately; rely on dataset refresh-level retry instead. * **No fine-grained partition-key routing**: All scans are cross-partition. * **Microsoft Entra ID / managed identity unsupported**: Key-based auth only. * **No native temporal / decimal / binary types**: Round-trip as strings; cast in SQL. * **Cosmos emulator is not used in CI**: Tested against a live account; emulator coverage is tracked as a future enhancement. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | -------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | `EmptyContainer` error at dataset load | The container has no documents, or the custom `query` returns zero rows. | Populate the container, broaden the `query`, or pin a schema via the dataset `columns:` configuration. | | Connector latched disabled — every query fails immediately | A 401/403/404 was observed, and `disable_on_permanent_error` is `true` (the default). | Fix the credential or restore the missing database/container, then restart `spice run`. Or set `disable_on_permanent_error: 'false'` during development. | | 429 retries dominate the request budget | Account RU/s is undersized for the Spice workload. | Increase RU/s in Azure, accelerate the dataset, or lower `max_concurrent_requests` to back off. | | RU consumption spikes on every restart | `schema_infer_max_records` × document size. | Lower the sample size or pin a schema via `columns:`. | | Schema doesn't include a field that exists in production | The first `schema_infer_max_records` documents had `null` for that field. | Increase `schema_infer_max_records`, or pin the schema explicitly. | | `Invalid Azure Cosmos DB connection string` | Connection string was edited or trimmed. | Re-copy the full string from the Azure portal — `AccountEndpoint=...;AccountKey=...;` (note the trailing `;`). | | `Could not determine Cosmos DB database from dataset path` error at registration | `from:` does not match `cosmosdb:database.container`. | Use `cosmosdb:database.container` or `cosmosdb:database/container`, or set `cosmosdb_database` and use `cosmosdb:container`. | | Multiple datasets, one with a different `max_concurrent_requests` | Spice keeps the **first-seen** value across datasets sharing an endpoint. | Set the same value on every dataset that targets the same account, or accept the warning logged at startup. | | Mid-stream scan failure leaves dataset partially loaded | Cosmos returned an error after some rows had been emitted; mid-stream retry is not safe. | The dataset refresh policy retries at the query boundary. For incidental failures, lower the `query` row count or accelerate. | --- # Databricks Data Connector Databricks as a connector for federated SQL query against Databricks using [Spark Connect](https://www.databricks.com/blog/2022/07/07/introducing-spark-connect-the-power-of-apache-spark-everywhere.html), directly from [Delta Lake](https://delta.io/) tables, or using the [SQL Statement Execution API](https://docs.databricks.com/aws/en/dev-tools/sql-execution-tutorial). ``` datasets: - from: databricks:spiceai.datasets.my_awesome_table # A reference to a table in the Databricks unity catalog name: my_delta_lake_table params: mode: delta_lake databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_token: ${secrets:my_token} databricks_aws_access_key_id: ${secrets:aws_access_key_id} databricks_aws_secret_access_key: ${secrets:aws_secret_access_key} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field for the Databricks connector takes the form `databricks:catalog.schema.table` where `catalog.schema.table` is the fully-qualified path to the table to read from. info Unquoted identifiers are normalized to lowercase. To reference a table, schema, or catalog with mixed-case characters, wrap each case-sensitive part in double quotes: `databricks:my_catalog."MySchema"."MyTable"`. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: databricks:spiceai.datasets.my_awesome_table name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") Use the [secret replacement syntax](/docs/next/components/secret-stores) to reference a secret, e.g. `${secrets:my_token}`. | Parameter Name | Description | | ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mode` | The execution mode for querying against Databricks. The default is `spark_connect`. Possible values:
- `spark_connect`: Use Spark Connect to query against Databricks. Requires a Spark cluster to be available.
- `delta_lake`: Query directly from Delta Tables. Requires the object store credentials to be provided.
- `sql_warehouse`: Use the SQL Statement Execution API to query against a Databricks SQL Warehouse. | | `databricks_endpoint` | The endpoint of the Databricks instance. Required for all modes. | | `databricks_sql_warehouse_id` | The ID of the SQL Warehouse in Databricks to use for the query. Only valid when `mode` is `sql_warehouse`. | | `databricks_cluster_id` | The ID of the compute cluster in Databricks to use for the query. Only valid when `mode` is `spark_connect`. | | `databricks_credential_vending` | When set to `enabled` (requires `mode` to be `delta_lake`), short-lived storage credentials for each table are fetched from the Unity Catalog credential vending API instead of using static object store credentials. Possible values: `enabled`, `disabled`. Defaults to `disabled`. | | `databricks_use_ssl` | If true, use a TLS connection to connect to the Databricks Spark Connect endpoint. Only valid when `mode` is `spark_connect`. Default is `true`. | | `client_timeout` | Optional. Specifies timeout for HTTP operations. In `delta_lake` mode, applies to object store operations. In `sql_warehouse` mode, applies per-HTTP-call (statement submit, status poll, chunk fetch) — not total query duration. Default: `30s`. E.g. `client_timeout: 2m` | | `connect_timeout` | Optional. Timeout for establishing TCP/TLS connections. Applies in `sql_warehouse` mode. Default: `10s`. E.g. `connect_timeout: 15s` | | `databricks_token` | The Databricks API token to authenticate with the Unity Catalog API. Can't be used with `databricks_client_id` and `databricks_client_secret`. | | `databricks_client_id` | The Databricks OAuth client ID. Used with `databricks_client_secret` for service-principal (M2M) auth, or alone for interactive User-to-Machine (U2M) auth. Can't be used with `databricks_token`. | | `databricks_client_secret` | The Databricks Service Principal Client Secret. Required for M2M auth; omit for U2M auth. Can't be used with `databricks_token`. | | `databricks_auth_mode` | Optional. Pins the authentication flow instead of inferring it from the credentials that are set. One of `auto`, `token`, `m2m`, `u2m` (case-insensitive). Defaults to `auto`. See [Pinning the authentication mode](#pinning-the-authentication-mode). | #### SQL Warehouse tuning[​](#sql-warehouse-tuning "Direct link to SQL Warehouse tuning") The following parameters apply only when `mode` is `sql_warehouse` and control connection resilience and concurrency: | Parameter Name | Description | | ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `connect_timeout` | Optional. Timeout for establishing TCP/TLS connections to the Databricks API. Default: `10s`. E.g. `connect_timeout: 15s` | | `client_timeout` | Optional. Per-HTTP-call timeout (statement submit, status poll, chunk fetch) — not total query duration. Overall query time is bounded by `statement_max_retries` × backoff. Default: `30s`. E.g. `client_timeout: 2m` | | `max_concurrent_requests` | Optional. Maximum number of concurrent HTTP requests to the SQL Warehouse API. Default: `8`. | | `http_max_retries` | Optional. Maximum number of HTTP-level retries for transient failures (429, 5xx). Default: `3`. | | `backoff_method` | Optional. Backoff strategy for transient HTTP retries. Options: `fibonacci`, `exponential`. Default: `fibonacci`. | | `statement_max_retries` | Optional. Maximum number of poll retries when waiting for async statement completion. Default: `14`. | | `disable_on_permanent_error` | Optional. When `true`, non-retryable errors (401, 403, 404) permanently disable the connector. Default: `true`. | #### Rate control[​](#rate-control "Direct link to Rate control") The Databricks connector supports per-dataset rate control parameters when `mode` is `spark_connect` or `sql_warehouse`. These override [`runtime.params`](/docs/next/reference/spicepod/runtime#runtimeparams) HTTP rate control defaults. When [`runtime.source_rate_control.state_location`](/docs/next/reference/spicepod/runtime#runtimesource_rate_control) is configured, rate limits are coordinated across the cluster. | Parameter Name | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `requests_per_second_limit` | Optional. Maximum HTTP requests per second to the Databricks endpoint. Overrides `runtime.params.http_requests_per_second_limit`. | | `requests_per_minute_limit` | Optional. Maximum HTTP requests per minute to the Databricks endpoint. Overrides `runtime.params.http_requests_per_minute_limit`. | | `rate_control_jitter_min` | Optional. Minimum random delay before HTTP requests when rate control is active. Defaults to `5ms` when a rate limit is configured. Accepts durations like `5ms`. | | `rate_control_jitter_max` | Optional. Maximum random delay before HTTP requests when rate control is active. Defaults to `10ms` when a rate limit is configured. Accepts durations like `10ms`. | ## Authentication[​](#authentication "Direct link to Authentication") ### Personal access token[​](#personal-access-token "Direct link to Personal access token") To learn more about how to set up personal access tokens, see [Databricks PAT docs](https://docs.databricks.com/aws/en/dev-tools/auth/pat). ``` datasets: - from: databricks:spiceai.datasets.my_awesome_table name: my_awesome_table params: databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_cluster_id: 1234-567890-abcde123 databricks_token: ${secrets:DATABRICKS_TOKEN} # PAT ``` ### Databricks service principal[​](#databricks-service-principal "Direct link to Databricks service principal") Spice supports the Machine-to-Machine (M2M) OAuth flow with service principal credentials by utilizing the `databricks_client_id` and `databricks_client_secret` parameters. The runtime will automatically refresh the token. Ensure that you grant your service principal the "Data Reader" privilege preset for the catalog and "Can Attach" cluster permissions when using Spark Connect mode. To Learn more about how to set up the service principal, see [Databricks M2M OAuth docs](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m). ``` datasets: - from: databricks:spiceai.datasets.my_awesome_table name: my_awesome_table params: databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_cluster_id: 1234-567890-abcde123 databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} # service principal client id databricks_client_secret: ${secrets:DATABRICKS_CLIENT_SECRET} # service principal client secret ``` ### User-to-Machine (U2M) OAuth[​](#user-to-machine-u2m-oauth "Direct link to User-to-Machine (U2M) OAuth") Spice supports the User-to-Machine (U2M) OAuth flow for interactive sign-in against Databricks. To use U2M auth, supply only `databricks_client_id` (without `databricks_token` or `databricks_client_secret`), or set `databricks_auth_mode: u2m` to pin the flow — see [Pinning the authentication mode](#pinning-the-authentication-mode). When U2M auth is configured, the connector defers initialization until first use. On the first query the runtime opens a browser to complete the Databricks OAuth sign-in, then caches and refreshes the resulting token for subsequent requests. To learn more about how to set up U2M OAuth, see the [Databricks U2M OAuth docs](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-u2m). note U2M auth is supported with `mode: delta_lake` and `mode: sql_warehouse`. It is not supported with `mode: spark_connect` — use a personal access token or service principal credentials when querying through Spark Connect. ``` datasets: - from: databricks:spiceai.datasets.my_awesome_table name: my_awesome_table params: mode: sql_warehouse databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_sql_warehouse_id: 2b4e24cff378fb24 databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} # OAuth app client id ``` ### Pinning the authentication mode[​](#pinning-the-authentication-mode "Direct link to Pinning the authentication mode") By default (`databricks_auth_mode: auto`) the connector infers the flow from whichever credentials resolve: a `databricks_token` alone selects personal access token, `databricks_client_id` alone selects U2M, and `databricks_client_id` together with `databricks_client_secret` selects M2M. Credentials do not only come from the Spicepod. `databricks_token`, `databricks_client_secret`, and the other secret parameters are auto-loaded from the [secret stores](/docs/next/components/secret-stores) and the environment (e.g. `DATABRICKS_CLIENT_SECRET`, `SPICE_DATABRICKS_CLIENT_SECRET`) when the Spicepod omits them — so an ambient client secret on the host is enough to switch a U2M dataset to machine-to-machine, which then fails with a service principal `401`. Set `databricks_auth_mode` to pin the flow. A pinned mode ignores the credentials the flow does not use: | Value | Flow | Requires | | ---------------- | --------------------------------- | -------------------------------------------------- | | `auto` (default) | Inferred from the credentials set | — | | `token` | Personal access token | `databricks_token` | | `m2m` | Service principal (M2M) OAuth | `databricks_client_id`, `databricks_client_secret` | | `u2m` | User-to-machine OAuth | `databricks_client_id` | Values are matched case-insensitively, so `M2M` and `U2M` are also accepted. A required parameter that is missing for the pinned mode fails at load with an error naming it; an unrecognized value is rejected. ``` datasets: - from: databricks:spiceai.datasets.my_awesome_table name: my_awesome_table params: mode: sql_warehouse databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_sql_warehouse_id: 2b4e24cff378fb24 databricks_auth_mode: u2m # ignore any ambient client secret or token databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} # OAuth app client id ``` ## Delta Lake object store parameters[​](#delta-lake-object-store-parameters "Direct link to Delta Lake object store parameters") Configure the connection to the object store when using `mode: delta_lake`. Use the [secret replacement syntax](/docs/next/components/secret-stores) to reference a secret, e.g. `${secrets:aws_access_key_id}`. ### AWS S3[​](#aws-s3 "Direct link to AWS S3") | Parameter Name | Description | | ---------------------------------- | --------------------------------------------------------------------------------------------------- | | `databricks_aws_region` | Optional. The AWS region for the S3 object store. E.g. `us-west-2`. | | `databricks_aws_access_key_id` | The access key ID for the S3 object store. | | `databricks_aws_secret_access_key` | The secret access key for the S3 object store. | | `databricks_aws_session_token` | Optional. The AWS session token for the S3 object store. Required with temporary (STS) credentials. | | `databricks_aws_endpoint` | Optional. The endpoint for the S3 object store. E.g. `s3.us-west-2.amazonaws.com`. | | `databricks_aws_allow_http` | Optional. Enables insecure HTTP connections to `databricks_aws_endpoint`. Defaults to `false`. | ### Azure Blob[​](#azure-blob "Direct link to Azure Blob") Note **One** of the following auth values must be provided for Azure Blob: * `databricks_azure_storage_account_key`, * `databricks_azure_storage_client_id` and `databricks_azure_storage_client_secret`, or * `databricks_azure_storage_sas_key`. | Parameter Name | Description | | ---------------------------------------- | ---------------------------------------------------------------------- | | `databricks_azure_storage_account_name` | The Azure Storage account name. | | `databricks_azure_storage_account_key` | The Azure Storage key for accessing the storage account. | | `databricks_azure_storage_client_id` | The Service Principal client ID for accessing the storage account. | | `databricks_azure_storage_client_secret` | The Service Principal client secret for accessing the storage account. | | `databricks_azure_storage_sas_key` | The shared access signature key for accessing the storage account. | | `databricks_azure_storage_endpoint` | Optional. The endpoint for the Azure Blob storage account. | ### Google Storage (GCS)[​](#google-storage-gcs "Direct link to Google Storage (GCS)") | Parameter Name | Description | | ----------------------------------- | ------------------------------------------------------------ | | `databricks_google_service_account` | Filesystem path to the Google service account JSON key file. | ## Examples[​](#examples "Direct link to Examples") ### Spark Connect[​](#spark-connect "Direct link to Spark Connect") ``` - from: databricks:spiceai.datasets.my_spark_table # A reference to a table in the Databricks unity catalog name: my_delta_lake_table params: mode: spark_connect databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_cluster_id: 1234-567890-abcde123 databricks_token: ${secrets:my_token} ``` ### SQL Warehouse[​](#sql-warehouse "Direct link to SQL Warehouse") ``` - from: databricks:spiceai.datasets.my_table # A reference to a table in the Databricks unity catalog name: my_table params: mode: sql_warehouse databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_sql_warehouse_id: 2b4e24cff378fb24 databricks_token: ${secrets:my_token} ``` ### Delta Lake (S3)[​](#delta-lake-s3 "Direct link to Delta Lake (S3)") ``` - from: databricks:spiceai.datasets.my_delta_table # A reference to a table in the Databricks unity catalog name: my_delta_lake_table params: mode: delta_lake databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_token: ${secrets:my_token} databricks_aws_region: us-west-2 # Optional databricks_aws_access_key_id: ${secrets:aws_access_key_id} databricks_aws_secret_access_key: ${secrets:aws_secret_access_key} databricks_aws_endpoint: s3.us-west-2.amazonaws.com # Optional ``` ### Delta Lake (Azure Blobs)[​](#delta-lake-azure-blobs "Direct link to Delta Lake (Azure Blobs)") ``` - from: databricks:spiceai.datasets.my_adls_table # A reference to a table in the Databricks unity catalog name: my_delta_lake_table params: mode: delta_lake databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_token: ${secrets:my_token} # Account Name + Key databricks_azure_storage_account_name: my_account databricks_azure_storage_account_key: ${secrets:my_key} # OR Service Principal + Secret databricks_azure_storage_client_id: my_client_id databricks_azure_storage_client_secret: ${secrets:my_secret} # OR SAS Key databricks_azure_storage_sas_key: my_sas_key ``` ### Delta Lake (GCP)[​](#delta-lake-gcp "Direct link to Delta Lake (GCP)") ``` - from: databricks:spiceai.datasets.my_gcp_table # A reference to a table in the Databricks unity catalog name: my_delta_lake_table params: mode: delta_lake databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_token: ${secrets:my_token} databricks_google_service_account: /path/to/service-account.json ``` ## Types[​](#types "Direct link to Types") ### mode: delta\_lake[​](#mode-delta_lake "Direct link to mode: delta_lake") The table below shows the Databricks (mode: delta\_lake) data types supported, along with the type mapping to Apache Arrow types in Spice. | Databricks SQL Type | Arrow Type | | ------------------- | ------------------------------------- | | `STRING` | `Utf8` | | `BIGINT` | `Int64` | | `INT` | `Int32` | | `SMALLINT` | `Int16` | | `TINYINT` | `Int8` | | `FLOAT` | `Float32` | | `DOUBLE` | `Float64` | | `BOOLEAN` | `Boolean` | | `BINARY` | `Binary` | | `DATE` | `Date32` | | `TIMESTAMP` | `Timestamp(Microsecond, Some("UTC"))` | | `TIMESTAMP_NTZ` | `Timestamp(Microsecond, None)` | | `DECIMAL` | `Decimal128` | | `ARRAY` | `List` | | `STRUCT` | `Struct` | | `MAP` | `Map` | ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Limitations[​](#limitations "Direct link to Limitations") * Databricks connector (mode: delta\_lake) does not support reading Delta tables with the `V2Checkpoint` feature enabled. To use the Databricks connector (mode: delta\_lake) with such tables, drop the `V2Checkpoint` feature by executing the following command: ``` ALTER TABLE DROP FEATURE v2Checkpoint [TRUNCATE HISTORY]; ``` For more details on dropping Delta table features, refer to the official documentation: [Drop Delta table features](https://docs.databricks.com/en/delta/drop-feature.html#:~:text=Databricks%20provides%20limited%20support%20for,data%20files%20backing%20the%20table.) * When using `mode: spark_connect`, correlated scalar subqueries can only be used in filters, aggregations, projections, and UPDATE/MERGE/DELETE commands. [Spark Docs](https://spark.apache.org/docs/latest/sql-error-conditions-unsupported-subquery-expression-category-error-class.html#unsupported_correlated_scalar_subquery) Memory Considerations When using the Databricks (mode: delta\_lake) Data connector without acceleration, data is loaded into memory during query execution. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. Memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/next/components/data-accelerators/duckdb) and [`sqlite`](/docs/next/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. * The Databricks Connector (`mode: spark_connect`) does not yet support streaming query results from Spark. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Databricks as a data connector in Spice. [Spice on Databricks](https://github.com/spiceai/cookbook/tree/trunk/databricks) --- # Databricks Deployment Guide Production operating guide for the Databricks connector covering resilience tuning, Unity Catalog awareness, metrics, and observability. These features apply primarily to `sql_warehouse` mode unless noted otherwise. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Retry and Concurrency Parameters[​](#retry-and-concurrency-parameters "Direct link to Retry and Concurrency Parameters") When using `mode: sql_warehouse`, the following parameters control HTTP retry behavior and concurrency limits for the Databricks SQL Statements API. | Parameter | Type | Default | Description | | ---------------------------- | -------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------- | | `connect_timeout` | duration | `10s` | Timeout for establishing TCP/TLS connections to the Databricks API. | | `client_timeout` | duration | `30s` | Per-HTTP-call timeout (statement submit, status poll, chunk fetch). Set to the longest expected single call, not total query duration. | | `max_concurrent_requests` | integer | `8` | Maximum concurrent HTTP requests to the SQL Warehouse API. | | `http_max_retries` | integer | `3` | Maximum HTTP-level retries for transient failures (429, 5xx). | | `backoff_method` | string | `fibonacci` | Backoff strategy for transient HTTP retries: `fibonacci` or `exponential`. | | `statement_max_retries` | integer | `14` | Maximum poll retries when waiting for an async SQL statement to complete. | | `disable_on_permanent_error` | boolean | `true` | Permanently disable the connector on non-retryable errors (401, 403, 404). | #### Example[​](#example "Direct link to Example") ``` catalogs: - from: databricks:my_catalog name: my_catalog params: databricks_endpoint: my-workspace.cloud.databricks.com mode: sql_warehouse databricks_sql_warehouse_id: abc123def456 databricks_client_id: ${env:DBX_CLIENT_ID} databricks_client_secret: ${env:DBX_CLIENT_SECRET} connect_timeout: 10s client_timeout: 2m max_concurrent_requests: '4' http_max_retries: '5' backoff_method: exponential statement_max_retries: '20' disable_on_permanent_error: 'true' ``` ### Shared Concurrency Semaphore[​](#shared-concurrency-semaphore "Direct link to Shared Concurrency Semaphore") When multiple datasets or catalog-discovery paths target the same SQL Warehouse (same `endpoint` + `sql_warehouse_id`), a single concurrency semaphore is shared across all of them. The `max_concurrent_requests` limit is enforced globally for that warehouse, not per dataset or per catalog. Every dataset and catalog entry that targets the same warehouse must resolve to the **same** `max_concurrent_requests` value. A component that omits the parameter resolves to the default `8`, so setting a non-default limit on one entry and leaving it off another is itself a conflict — set the same value on every component that shares the warehouse. A mismatch fails at load with an error naming the requested and existing limits (`Conflicting Databricks SQL Warehouse concurrency limits for endpoint ... requested 4, existing 8`). A value of `0` or a non-numeric value is not applied: the runtime logs a warning and uses the default `8`. ### Permanent-Disable Behavior[​](#permanent-disable-behavior "Direct link to Permanent-Disable Behavior") When `disable_on_permanent_error` is `true` (default), non-retryable HTTP status codes on statement-execution requests permanently disable the connector. Subsequent queries immediately return a `PermanentlyDisabled` error instead of issuing further HTTP requests. The following errors trigger permanent disable: * **401 Unauthorized** — expired or invalid credentials. * **403 Forbidden** — the service principal or token lacks permission to execute statements on the warehouse. * **404 Not Found** — the SQL Warehouse has been deleted or the endpoint is incorrect. This prevents cascading failures (e.g., every dataset refresh hammering a warehouse that will never accept the request). info Permanent-disable detection is **not** applied to statement-poll or result-fetch requests. Transient 403/404 responses on those paths (e.g., expired pre-signed URLs or purged statement results) do not indicate a configuration problem. To recover from a permanent-disable state, fix the underlying issue (e.g., renew credentials, restore the warehouse) and restart the Spice runtime. ### Retry Behavior[​](#retry-behavior "Direct link to Retry Behavior") The SQL Warehouse connector has two retry layers: 1. **HTTP-level retries** retries on 408 (request timeout), 429 (rate-limit), and 5xx (server error) responses, as well as transient network and connection errors. Respects `Retry-After`, `retry-after-ms`, and `x-retry-after-ms` headers. Uses the configured `backoff_method` with a maximum backoff of 300 seconds. 2. **Statement poll retries** when a SQL statement enters PENDING or RUNNING state, the connector polls for completion using fibonacci backoff up to `statement_max_retries` times. If the statement does not reach a terminal state within the retry budget, a `QueryStillRunning` or `InvalidWarehouseState` error is returned. ## Unity Catalog Awareness[​](#unity-catalog-awareness "Direct link to Unity Catalog Awareness") ### Table Type Filtering[​](#table-type-filtering "Direct link to Table Type Filtering") The connector checks each table's type against Unity Catalog metadata before creating a table provider. The following table types are supported: | Table Type | Supported | Notes | | ------------------- | --------- | -------------------------------------- | | `MANAGED` | Yes | Standard Delta tables | | `EXTERNAL` | Yes | Tables with external storage locations | | `FOREIGN` | Yes | Lakehouse Federation foreign tables | | `MATERIALIZED_VIEW` | Yes | Materialized views | | `VIEW` | No | Skipped during discovery | | `STREAMING_TABLE` | No | Skipped during discovery | Unsupported table types are silently skipped during catalog discovery. When referenced directly (e.g., `databricks:catalog.schema.view_name`), an error is returned. ### Permission Checking[​](#permission-checking "Direct link to Permission Checking") Before creating a table provider, the connector verifies the current principal has a read-compatible privilege on the table using the Unity Catalog Effective Permissions API. The following privileges grant read access: `SELECT`, `ALL_PRIVILEGES`, `ALL PRIVILEGES`, `OWNER`, and `OWNERSHIP`. **Catalog discovery**: Tables without read permissions are skipped. **Direct table references**: An `InsufficientPermissions` error is returned. **Foreign tables**: `FOREIGN` tables skip the table-level permission precheck because Lakehouse Federation access can be valid even when the effective-permissions endpoint does not report a table-level read privilege. Access is still enforced by Databricks at query time. **Graceful degradation**: If the Unity Catalog API is unreachable or the table is not found in UC, the connector logs a warning and proceeds without validation. ## Metrics[​](#metrics "Direct link to Metrics") The SQL Warehouse connector exposes per-dataset operational metrics. Most metrics must be explicitly enabled in the dataset's `metrics` section. The `inflight_operations` metric is auto-registered and always available. For general information about component metrics, see [Component Metrics](/docs/features/observability/component_metrics). ### Available Metrics[​](#available-metrics "Direct link to Available Metrics") | Metric Name | Type | Category | Description | | ----------------------------- | ------- | --------------- | ---------------------------------------------------------------------------------------------------------------------------------- | | `requests_total` | Counter | Requests | Total HTTP requests issued (excluding retries). | | `retries_total` | Counter | Requests | Total HTTP retries for transient failures. | | `permanent_errors_total` | Counter | Requests | Total non-retryable errors (401, 403, 404). | | `inflight_operations` | Gauge | Requests | Current in-flight operations holding a concurrency permit. Global across datasets sharing the same warehouse. **Auto-registered.** | | `statements_executed_total` | Counter | Statements | Total SQL statements submitted. | | `statement_polls_total` | Counter | Statements | Total polls for async statement completion. | | `statements_failed_total` | Counter | Statements | Total SQL statements that completed with FAILED status. | | `pool_connections_total` | Counter | Connection Pool | Total pool `connect()` calls. | | `pool_active_connections` | Gauge | Connection Pool | Current active connection handles. | | `semaphore_available_permits` | Gauge | Concurrency | Available permits in the request concurrency semaphore. | | `chunks_fetched_total` | Counter | Data Transfer | Total Arrow result chunks fetched. | | `connector_disabled` | Gauge | Connector State | Whether the connector is permanently disabled (`1` = yes, `0` = no). | ### Enabling Metrics[​](#enabling-metrics "Direct link to Enabling Metrics") Add a `metrics` list to the dataset definition in your spicepod: ``` datasets: - from: databricks:my_catalog.my_schema.my_table name: my_table params: mode: sql_warehouse databricks_sql_warehouse_id: abc123def456 databricks_endpoint: my-workspace.cloud.databricks.com databricks_client_id: ${env:DBX_CLIENT_ID} databricks_client_secret: ${env:DBX_CLIENT_SECRET} metrics: - name: requests_total - name: retries_total - name: permanent_errors_total - name: statements_executed_total - name: statements_failed_total - name: pool_active_connections - name: semaphore_available_permits - name: chunks_fetched_total - name: connector_disabled ``` Individual metrics can be disabled by setting `enabled: false`. This includes auto-registered metrics: ``` metrics: - name: inflight_operations enabled: false ``` ### Metric Naming[​](#metric-naming "Direct link to Metric Naming") Metrics are exposed as OpenTelemetry instruments with the naming convention: ``` dataset_databricks_{metric_name} ``` For example, `requests_total` becomes `dataset_databricks_requests_total`. Each instrument carries a `name` attribute set to the dataset instance name, so metrics from multiple datasets sharing the same warehouse can be distinguished. ### Shared Warehouse Attribution[​](#shared-warehouse-attribution "Direct link to Shared Warehouse Attribution") When multiple datasets share the same SQL Warehouse, compare `dataset_databricks_*` metrics by their `name` attribute to understand per-dataset load. The `semaphore_available_permits` metric reflects the shared semaphore, so all datasets targeting the same warehouse observe the same underlying concurrency budget. ### Accessing Metrics[​](#accessing-metrics "Direct link to Accessing Metrics") Registered metrics are available through: * **Prometheus endpoint**: `GET /metrics` when the metrics server is enabled. * **`runtime.metrics` SQL table**: `SELECT * FROM runtime.metrics WHERE name LIKE 'dataset_databricks_%'`. * **OTLP push exporter**: Pushed to any configured OpenTelemetry collector. ## Task History[​](#task-history "Direct link to Task History") All major Databricks operations are instrumented with tracing spans for the Spice [task history](/docs/reference/task_history) system. This applies to both `sql_warehouse` and `delta_lake` modes. ### SQL Warehouse Spans[​](#sql-warehouse-spans "Direct link to SQL Warehouse Spans") | Span Name | Input Field | Description | | ------------------------------ | ------------ | ------------------------------------------------------- | | `databricks_get_schema` | Table name | Schema inference via `information_schema` or `DESCRIBE` | | `databricks_execute_statement` | SQL text | SQL statement execution via the Statements API | | `databricks_poll_statement` | Statement ID | Polling for async statement completion | ### Unity Catalog Spans[​](#unity-catalog-spans "Direct link to Unity Catalog Spans") | Span Name | Input Field | Description | | ------------------------------ | -------------------------- | --------------------------------------- | | `uc_get_table` | Fully-qualified table name | Fetch table metadata from Unity Catalog | | `uc_get_catalog` | Catalog ID | Fetch catalog metadata | | `uc_list_schemas` | Catalog ID | List schemas in a catalog | | `uc_list_tables` | `catalog_id.schema_name` | List tables in a schema | | `uc_get_effective_permissions` | Fully-qualified table name | Check effective permissions for a table | All SQL Warehouse spans include a `warehouse_id` field. Unity Catalog spans include the table or catalog identifier as the `input` field. ## Token Management[​](#token-management "Direct link to Token Management") How authentication tokens are managed depends on the authentication method: * **Service Principal (M2M OAuth)**: A background task refreshes the OAuth2 token 5 minutes before expiry. Refresh failures use fibonacci backoff capped at 5 minutes. * **Personal Access Token**: Used as-is with no automatic refresh. --- # Debezium Data Connector [Debezium](https://debezium.io/) is an open-source platform that enables [Change Data Capture (CDC)](/docs/next/features/cdc) for efficient real-time updates of locally accelerated datasets. Spice supports connecting to a Kafka topic managed by Debezium to keep datasets up-to-date with the source data. ``` datasets: - from: debezium:my_kafka_topic_with_debezium_changes name: my_dataset params: debezium_transport: kafka # Optional. Only `kafka` is currently supported. debezium_message_format: json # Optional. Only `json` is currently supported. kafka_bootstrap_servers: broker1:9092,broker2:9092,broker3:9092 # Required. A comma separated list of Kafka broker servers. kafka_security_protocol: sasl_ssl # Default is `sasl_ssl`. Valid values are `plaintext`, `ssl`, `sasl_plaintext`, `sasl_ssl`. kafka_sasl_mechanism: SCRAM-SHA-512 # Default is `SCRAM-SHA-512`. Valid values are `PLAIN`, `SCRAM-SHA-256`, `SCRAM-SHA-512`. kafka_sasl_username: kafka # Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. kafka_sasl_password: ${secrets:kafka_sasl_password} # Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. kafka_ssl_ca_location: ./certs/kafka_ca_cert.pem # Optional. Used to verify the SSL/TLS certificate of the Kafka broker. kafka_ssl_certificate_location: ./certs/client_cert.pem # Optional. Client SSL/TLS certificate for mTLS authentication. kafka_ssl_key_location: ./certs/client_key.pem # Optional. Client SSL/TLS private key for mTLS authentication. kafka_ssl_key_password: ${secrets:kafka_ssl_key_password} # Optional. Password for the client SSL/TLS private key, if encrypted. kafka_enable_ssl_certificate_verification: true # Default is `true`. Set to `false` to disable SSL/TLS certificate verification. kafka_ssl_endpoint_identification_algorithm: https # Default is `https`. Valid values are `none` and `https`. batch_max_size: 10000 # Default is `10000`. Maximum number of change events to batch together before processing. batch_max_duration: 1s # Default is `1s`. Maximum time to wait for a batch to fill before processing. acceleration: enabled: true # Acceleration is required for the debezium connector. engine: duckdb # `duckdb`, `sqlite` and `postgres` are supported acceleration engines for Debezium. refresh_mode: changes # Optional. If specified, this is required to be set to `changes` - any other value is an error. mode: file # Persistence is recommended to not have to rebuild the table each time Spice starts. ``` ## Overview[​](#overview "Direct link to Overview") Upon startup, Spice subscribes to the specified Debezium-managed Kafka topic using either a uniquely generated consumer group or a custom one specified via `kafka_consumer_group_id`. If a persistent acceleration engine is used (with `mode: file`), data is fetched starting from the last processed record, so Spice can resume without reprocessing all historical change events. ## Consumer Group Management[​](#consumer-group-management "Direct link to Consumer Group Management") The Debezium connector manages consumer groups to ensure data consistency across restarts. Offsets are committed to Kafka, so Spice can track consumption progress. **Default behavior:** When no `kafka_consumer_group_id` is specified, Spice automatically generates a unique consumer group ID and stores it in the acceleration metadata. On subsequent restarts, Spice retrieves and reuses this stored consumer group ID to maintain offset tracking and resume consumption from where it left off. **Custom consumer group:** If you specify a custom `kafka_consumer_group_id`, Spice stores this ID in the acceleration metadata. The same consumer group must be used on subsequent restarts. If no acceleration data exists and a custom consumer group is provided, Spice will reset its position to the oldest available offset and begin consuming from the start of the topic. **Consumer group mismatch error:** Spice will return an error if a restart is attempted with a different consumer group than what is stored in the acceleration metadata. This applies to both auto-generated and custom consumer group IDs. This safeguard prevents data inconsistency that could occur from mixing offsets between different consumer groups. To resolve a consumer group mismatch, either: * Use the same consumer group ID as stored in the acceleration * Reset the acceleration data to start fresh with a new consumer group ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form of `debezium:kafka_topic` where `kafka_topic` is the name of the Kafka topic where Debezium is notifying consumers about any upstream changes. In the example above it would listen to the `my_kafka_topic_with_debezium_changes` topic. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. ``` datasets: - from: debezium:my_kafka_topic_with_debezium_changes name: cool_dataset ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `debezium_transport` | Optional. The message broker transport to use. The default is `kafka`. Possible values:- `kafka`: Use Kafka as the message broker transport. Spice may support additional transports in the future. | | `debezium_message_format` | Optional. The message format to use. The default is `json`. Possible values: - `json`: Use JSON as the message format. Spice is expected to support additional message formats in the future, like `avro`. | | `kafka_bootstrap_servers` | **Required**. A list of host/port pairs for establishing the initial Kafka cluster connection. The client will use all servers, regardless of the bootstrapping servers specified here. This list only affects the initial hosts used to discover the full server set and should be formatted as `host1:port1,host2:port2,...`. | | `kafka_security_protocol` | Security protocol for Kafka connections. Default: `sasl_ssl`. Options: - `plaintext`
- `ssl`
- `sasl_plaintext`
- `sasl_ssl` | | `kafka_sasl_mechanism` | SASL (Simple Authentication and Security Layer) authentication mechanism. Default: `SCRAM-SHA-512`. Options: - `PLAIN`
- `SCRAM-SHA-256`
- `SCRAM-SHA-512` | | `kafka_sasl_username` | SASL username. | | `kafka_sasl_password` | SASL password. | | `kafka_ssl_ca_location` | Path to the SSL/TLS CA certificate file for server verification. | | `kafka_ssl_certificate_location` | Path to the client SSL/TLS certificate file for mTLS authentication. | | `kafka_ssl_key_location` | Path to the client SSL/TLS private key file for mTLS authentication. | | `kafka_ssl_key_password` | Password for the client SSL/TLS private key, if encrypted. | | `kafka_enable_ssl_certificate_verification` | Enable SSL/TLS certificate verification. Default: `true`. | | `kafka_ssl_endpoint_identification_algorithm` | SSL/TLS endpoint identification algorithm. Default: `https`. Options: - `none`
- `https` | | `kafka_consumer_group_id` | Kafka consumer group ID to use. If not set, a unique ID will be generated automatically. The consumer group ID (whether auto-generated or custom) is stored in the acceleration metadata and must remain consistent across restarts. See [Consumer Group Management](#consumer-group-management) for details. | | `batch_max_size` | Maximum number of change events to batch together before processing. Default: `10000`. | | `batch_max_duration` | Maximum time to wait for a batch to fill before processing. Default: `1s`. | | `schema_evolution` | Enable automatic schema evolution detection on reload. When `true`, the connector peeks at the latest Kafka message to detect schema changes. Default: `false`. | ### `metrics`[​](#metrics "Direct link to metrics") The connector supports the following optional [component metrics](/docs/next/features/observability/component_metrics): | Metric Name | Type | Description | | ------------------------ | ------- | ------------------------------------------------------------------------------------ | | `bytes_consumed_total` | Counter | Total number of bytes consumed from the Kafka topic | | `records_consumed_total` | Counter | Total number of records (messages) consumed from Kafka topics | | `records_lag` | Gauge | Total consumer lag across all topic partitions (number of messages not yet consumed) | These metrics are not enabled by default, enable them by setting the `metrics` parameter: ``` datasets: - from: debezium:my_kafka_topic_with_debezium_changes name: cool_dataset metrics: - name: records_lag - name: records_consumed_total - name: bytes_consumed_total params: ... ``` ### Acceleration Settings[​](#acceleration-settings "Direct link to Acceleration Settings") warning Using the Debezium connector **requires** [acceleration](/docs/next/components/data-accelerators) to be enabled. The following settings are required: | Parameter Name | Description | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `enabled` | Required. Must be set to `true` to enable acceleration. | | `engine` | Required. The acceleration engine to use. Possible valid values: - `duckdb`: Use [DuckDB](/docs/next/components/data-accelerators/duckdb) as the acceleration engine.
- `sqlite`: Use [SQLite](/docs/next/components/data-accelerators/sqlite) as the acceleration engine.
- `postgres`: Use [PostgreSQL](/docs/next/components/data-accelerators/postgres) as the acceleration engine. | | `refresh_mode` | Optional. The refresh mode to use. If specified, this must be set to `changes`. Any other value is an error. | | `mode` | Optional. The persistence mode to use. When using the `duckdb` and `sqlite` engines, it is recommended to set this to `file` to persist the data across restarts. Spice also persists metadata about the dataset, so it can resume from the last known state of the dataset instead of re-fetching the entire dataset. | ## JSON Nesting[​](#json-nesting "Direct link to JSON Nesting") When a change event carries many fields but you only need a few as discrete columns, you can consolidate the rest into a single JSON column using the `json_object` metadata option. Declare the fields you want as top-level columns explicitly in the `columns` list, then add a "catch-all" column with `json_object: "*"` metadata — every field not otherwise listed is nested into it as a JSON object. This applies to the change-stream (CDC) events the connector decomposes into the accelerator. ``` datasets: - from: debezium:my_kafka_topic_with_debezium_changes name: orders columns: - name: id - name: status - name: data_json metadata: json_object: '*' acceleration: enabled: true engine: duckdb refresh_mode: changes ``` Any field other than `id`, `status`, and `data_json` is folded into `data_json` as a JSON object. Primary-key columns must be declared explicitly and cannot be folded into the catch-all column. Limitations * The `json_object` metadata only accepts `"*"` as its value, which captures all unspecified fields. * Only one column can carry the `json_object` metadata; declaring more than one is an error. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * See an example of configuring a dataset to use CDC with Debezium by following the sample [Streaming changes in real-time with Debezium CDC](https://github.com/spiceai/cookbook/tree/trunk/cdc-debezium#readme). * An example of [Streaming changes in real-time with Debezium CDC and SASL/SCRAM authentication](https://github.com/spiceai/cookbook/tree/trunk/cdc-debezium/sasl-scram#readme) is available as well. --- # Delta Lake Data Connector Delta Lake data connector enables SQL queries from [Delta Lake](https://delta.io/) tables. ``` datasets: - from: delta_lake:s3://my_bucket/path/to/s3/delta/table/ name: my_delta_lake_table params: delta_lake_aws_access_key_id: ${secrets:aws_access_key_id} delta_lake_aws_secret_access_key: ${secrets:aws_secret_access_key} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field for the Delta Lake connector takes the form of `delta_lake:path` where `path` is any supported path, either local or to a cloud storage location. See the [examples](#examples) section below. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: delta_lake:s3://my_bucket/path/to/s3/delta/table/ name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") Use the [secret replacement syntax](/docs/next/components/secret-stores) to reference a secret, e.g. `${secrets:aws_access_key_id}`. | Parameter Name | Description | | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | | `client_timeout` | Optional. Specifies timeout for object store operations. No default; uses the underlying HTTP client timeout when not set. E.g. `client_timeout: 60s` | ## Delta Lake object store parameters[​](#delta-lake-object-store-parameters "Direct link to Delta Lake object store parameters") ### AWS S3[​](#aws-s3 "Direct link to AWS S3") | Parameter Name | Description | | ---------------------------------- | ---------------------------------------------------------------------------------------------- | | `delta_lake_aws_region` | Optional. The AWS region for the S3 object store. E.g. `us-west-2`. | | `delta_lake_aws_access_key_id` | The access key ID for the S3 object store. | | `delta_lake_aws_secret_access_key` | The secret access key for the S3 object store. | | `delta_lake_aws_session_token` | Optional. The AWS session token for S3 object store. | | `delta_lake_aws_endpoint` | Optional. The endpoint for the S3 object store. E.g. `s3.us-west-2.amazonaws.com`. | | `delta_lake_aws_allow_http` | Optional. Enables insecure HTTP connections to `delta_lake_aws_endpoint`. Defaults to `false`. | ### Azure Blob[​](#azure-blob "Direct link to Azure Blob") Note **One** of the following auth values must be provided for Azure Blob: * `delta_lake_azure_storage_account_key`, * `delta_lake_azure_storage_client_id`, `delta_lake_azure_storage_client_secret`, and `delta_lake_azure_storage_tenant_id`, or * `delta_lake_azure_storage_sas_key`. | Parameter Name | Description | | ---------------------------------------- | ---------------------------------------------------------------------- | | `delta_lake_azure_storage_account_name` | The Azure Storage account name. | | `delta_lake_azure_storage_account_key` | The Azure Storage master key for accessing the storage account. | | `delta_lake_azure_storage_client_id` | The service principal client id for accessing the storage account. | | `delta_lake_azure_storage_client_secret` | The service principal client secret for accessing the storage account. | | `delta_lake_azure_storage_tenant_id` | The service principal tenant id for accessing the storage account. | | `delta_lake_azure_storage_sas_key` | The shared access signature key for accessing the storage account. | | `delta_lake_azure_storage_endpoint` | Optional. The endpoint for the Azure Blob storage account. | ### Google Storage (GCS)[​](#google-storage-gcs "Direct link to Google Storage (GCS)") | Parameter Name | Description | | ----------------------------------- | ------------------------------------------------------------ | | `delta_lake_google_service_account` | Filesystem path to the Google service account JSON key file. | ## Examples[​](#examples "Direct link to Examples") ### Delta Lake + Local[​](#delta-lake--local "Direct link to Delta Lake + Local") ``` - from: delta_lake:/path/to/local/delta/table # A local filesystem path to a Delta Lake table name: my_delta_lake_table ``` ### Delta Lake + S3[​](#delta-lake--s3 "Direct link to Delta Lake + S3") ``` - from: delta_lake:s3://my_bucket/path/to/s3/delta/table/ # A reference to a table in S3 name: my_delta_lake_table params: delta_lake_aws_region: us-west-2 # Optional delta_lake_aws_access_key_id: ${secrets:aws_access_key_id} delta_lake_aws_secret_access_key: ${secrets:aws_secret_access_key} delta_lake_aws_endpoint: s3.us-west-2.amazonaws.com # Optional ``` ### Delta Lake + MinIO[​](#delta-lake--minio "Direct link to Delta Lake + MinIO") ``` - from: delta_lake:s3://my_bucket/path/to/s3/delta/table/ # A reference to a table in MinIO name: my_delta_lake_table params: delta_lake_aws_region: us-east-1 # Best practice for MinIO delta_lake_aws_access_key_id: ${secrets:aws_access_key_id} delta_lake_aws_secret_access_key: ${secrets:aws_secret_access_key} delta_lake_aws_endpoint: http://localhost:9000 # MinIO Endpoint delta_lake_aws_allow_http: true ``` ### Delta Lake + Azure Blob[​](#delta-lake--azure-blob "Direct link to Delta Lake + Azure Blob") ``` - from: delta_lake:abfss://my_container@my_account.dfs.core.windows.net/path/to/azure/delta/table/ # A reference to a table in Azure Blob name: my_delta_lake_table params: # Account Name + Key delta_lake_azure_storage_account_name: my_account delta_lake_azure_storage_account_key: ${secrets:my_key} # OR Service Principal + Secret delta_lake_azure_storage_client_id: my_client_id delta_lake_azure_storage_client_secret: ${secrets:my_secret} delta_lake_azure_storage_tenant_id: ${secrets:my_tenant_id} # OR SAS Key delta_lake_azure_storage_sas_key: my_sas_key ``` ### Delta Lake + Google Storage[​](#delta-lake--google-storage "Direct link to Delta Lake + Google Storage") ``` params: delta_lake_google_service_account: /path/to/service-account.json ``` ## Types[​](#types "Direct link to Types") The table below shows the Delta Lake data types supported, along with the type mapping to Apache Arrow types in Spice. | Delta Lake Type | Arrow Type | | --------------- | ------------------------------------- | | `String` | `Utf8` | | `Long` | `Int64` | | `Integer` | `Int32` | | `Short` | `Int16` | | `Byte` | `Int8` | | `Float` | `Float32` | | `Double` | `Float64` | | `Boolean` | `Boolean` | | `Binary` | `Binary` | | `Date` | `Date32` | | `Timestamp` | `Timestamp(Microsecond, Some("UTC"))` | | `TimestampNtz` | `Timestamp(Microsecond, None)` | | `Decimal` | `Decimal128` | | `Array` | `List` | | `Struct` | `Struct` | | `Variant` | `Struct` | | `Map` | `Map` | ## Column Mapping[​](#column-mapping "Direct link to Column Mapping") Delta Lake tables that use [column mapping](https://github.com/delta-io/delta/blob/master/PROTOCOL.md#column-mapping) with `columnMapping.mode = name` or `columnMapping.mode = id` are supported. Tables in either mode store data in Parquet files under opaque physical column names (e.g., `col-b70f7585-…`) that differ from the logical names exposed by the table schema. Spice resolves logical-to-physical names at read time, including for partitions, predicate pushdown, and nested `Struct`, `List`, and `Map` field renames. ## Limitations[​](#limitations "Direct link to Limitations") * Delta Lake connector does not support reading Delta tables with the `V2Checkpoint` feature enabled. To use the Delta Lake connector with such tables, drop the `V2Checkpoint` feature by executing the following command: ``` ALTER TABLE DROP FEATURE v2Checkpoint [TRUNCATE HISTORY]; ``` For more details on dropping Delta table features, refer to the official documentation: [Drop Delta table features](https://docs.delta.io/latest/delta-drop-feature.html) ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Delta Lake as a data connector in Spice. [Delta Lake Data Connector](https://github.com/spiceai/cookbook/tree/trunk/delta-lake#readme) --- # Delta Lake Data Connector Deployment Guide Production operating guide for the Delta Lake data connector covering object-store authentication, metadata handling, and operational tuning. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The Delta Lake connector reads Delta tables directly from an underlying object store (S3, ABFS/Azure, GCS, or the local filesystem). Authentication parameters depend on the object store: | Object store | Auth parameters | | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **S3** / S3-compatible | `delta_lake_aws_region`, `delta_lake_aws_access_key_id`, `delta_lake_aws_secret_access_key`, `delta_lake_aws_session_token`, `delta_lake_aws_endpoint`, `delta_lake_aws_allow_http`. Defaults to the AWS credential chain when unset. | | **Azure ADLS** | `delta_lake_azure_storage_account_name`, `delta_lake_azure_storage_account_key`, `delta_lake_azure_storage_client_id`, `delta_lake_azure_storage_client_secret`, `delta_lake_azure_storage_sas_key`, `delta_lake_azure_storage_endpoint`, `delta_lake_azure_storage_tenant_id`. | | **Google Cloud Storage** | `delta_lake_google_service_account`. | | **Local filesystem** | No auth. Ensure the Spice process has read permission on the Delta table directory. | Credentials must be sourced from a [secret store](/docs/next/components/secret-stores) in production. For AWS deployments, prefer instance-profile or IRSA-based auth (leave `aws_access_key_id` unset). ### Databricks Unity Catalog[​](#databricks-unity-catalog "Direct link to Databricks Unity Catalog") To access Delta tables registered in Unity Catalog, use the [Databricks connector](/docs/next/components/data-connectors/databricks) in `mode: delta_lake`. It handles UC metadata resolution automatically, and can optionally fetch short-lived storage credentials via Unity Catalog credential vending when `databricks_credential_vending: enabled` is set (defaults to `disabled`) — see the [Databricks Deployment Guide](/docs/next/components/data-connectors/databricks/deployment) for details. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Object Store Retry[​](#object-store-retry "Direct link to Object Store Retry") Object-store I/O uses the AWS/Azure/GCS SDK default retry strategies (adaptive backoff on throttling and transient 5xx responses). Per-operation retry parameters are not exposed at the Spice layer. ### Transaction Log Caching[​](#transaction-log-caching "Direct link to Transaction Log Caching") The Delta Lake connector reads the `_delta_log` transaction log on each query plan. For high-query-rate workloads, accelerate the dataset (see [Acceleration](/docs/next/features/data-acceleration)) to avoid repeatedly scanning the log on the hot path. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Metadata cost**: Reading a Delta table requires reading its `_delta_log`, which grows with every commit. Checkpointing (handled by the writer) bounds this cost. Ensure the writer is issuing checkpoints every 10-100 commits for large, high-churn tables. * **Partition pruning**: The connector prunes partitions using Delta's partition metadata during query planning. Partition datasets by the dominant filter columns (date, tenant) for best performance. * **Column pruning and predicate pushdown**: Filters and column projections are pushed down into the underlying Parquet readers. Queries over a narrow column set over a large table are efficient without requiring acceleration. * **Acceleration sizing**: When materializing into DuckDB/SQLite/Postgres, size the target engine to hold the refreshed snapshot plus WAL overhead (typically 1.3–2× the raw Parquet size for row-oriented accelerators). ## Metrics[​](#metrics "Direct link to Metrics") The `runtime-object-store` layer that performs Delta Lake object-store I/O is not instrumented, so Spice does not emit object-store transport metrics. See [Component Metrics](/docs/next/features/observability/component_metrics) for the metrics that are available. The Delta Lake connector does not currently register connector-specific dataset-level instruments. Monitor Delta operations via: * Query execution metrics (`query_duration_ms`, `query_returned_rows`) from `runtime.metrics`. * Upstream cloud metrics (S3 request count, Azure blob throughput). * Acceleration refresh metrics when the dataset is accelerated. ## Task History[​](#task-history "Direct link to Task History") Delta Lake reads participate in Spice [task history](/docs/next/reference/task_history) through DataFusion's execution-plan spans. Individual object reads are attributed to their enclosing `sql_query` or `accelerated_table_refresh` task. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Read-only**: The Delta Lake connector cannot write or update Delta tables. * **Time travel**: Querying by version or timestamp is not exposed through the connector; only the latest snapshot is read. * **Drop table features**: Some Delta features (deletion vectors, column mapping, row tracking) may require specific writer protocol versions. See [Databricks docs on dropping Delta table features](https://docs.databricks.com/en/delta/drop-feature.html) when compatibility issues arise. * **Change Data Feed (CDF)**: CDF is not exposed through the Delta Lake connector. For CDC, use the [Debezium](/docs/next/components/data-connectors/debezium) connector against the source database, or use [Databricks](/docs/next/components/data-connectors/databricks) mode to federate a CDF-enabled view. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ------------------------------------------------ | ------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- | | `Access Denied` on `_delta_log/` GET | Role lacks read on the `_delta_log/` prefix. | Grant `s3:GetObject` / equivalent on the table root and `_delta_log/` prefix. | | Query returns an empty result after a new commit | Stale transaction-log cache. | Trigger a dataset refresh (acceleration) or re-plan the query. | | `Protocol version unsupported` | Writer committed a newer Delta protocol version than the reader supports. | Upgrade Spice to a version with the required Delta protocol reader, or drop the writer feature. | | Slow queries on very high-commit tables | Transaction log grown very large without checkpoints. | Ensure the writer is issuing checkpoints; consider acceleration. | | `No such file or directory` for `_delta_log/...` | Table is uninitialized or path is incorrect. | Confirm the `from:` path points at the Delta table root, not a data subdirectory. | --- # Dremio Data Connector [Dremio](https://www.dremio.com/) is a data lake engine that enables high-performance SQL queries directly on data lake storage. It provides a unified interface for querying and analyzing data from various sources without the need for complex data movement or transformation. This connector enables using Dremio as a data source for federated SQL queries. ``` - from: dremio:datasets.dremio_dataset name: dremio_dataset params: dremio_endpoint: grpc://127.0.0.1:32010 dremio_username: demo dremio_password: ${secrets:my_dremio_pass} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `dremio:dataset` where `dataset` is the fully qualified name of the dataset to read from. info Unquoted identifiers are normalized to lowercase. To reference a dataset with mixed-case characters, wrap each case-sensitive part in double quotes: `dremio:my_source."MixedCaseDataset"`. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). \[Limitations] Currently, only up to three levels of nesting are supported for dataset names (e.g., a.b.c). Additional levels are not supported at this time. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: dremio:datasets.dremio_dataset name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `dremio_endpoint` | Required. The endpoint used to connect to the Dremio server. | | `dremio_username` | Optional. The username used to connect to the Dremio endpoint. | | `dremio_password` | Optional. The password used to connect to the Dremio endpoint. Use the [secret replacement syntax](#secrets) to load the password from a secret store, e.g. `${secrets:my_dremio_pass}`. | ## Examples[​](#examples "Direct link to Examples") ### Connecting to a GRPC endpoint[​](#connecting-to-a-grpc-endpoint "Direct link to Connecting to a GRPC endpoint") ``` - from: dremio:datasets.dremio_dataset name: dremio_dataset params: dremio_endpoint: grpc://127.0.0.1:32010 dremio_username: demo dremio_password: ${secrets:my_dremio_pass} ``` ## Types[​](#types "Direct link to Types") The table below shows the Dremio data types supported, along with the type mapping to Apache Arrow types in Spice. | Dremio Type | Arrow Type | | ----------- | ------------------------------ | | `INT` | `Int32` | | `BIGINT` | `Int64` | | `FLOAT` | `Float32` | | `DOUBLE` | `Float64` | | `DECIMAL` | `Decimal128` | | `VARCHAR` | `Utf8` | | `VARBINARY` | `Binary` | | `BOOL` | `Boolean` | | `DATE` | `Date64` | | `TIME` | `Time32` | | `TIMESTAMP` | `Timestamp(Millisecond, None)` | | `INTERVAL` | `Interval` | | `LIST` | `List` | | `STRUCT` | `Struct` | | `MAP` | `Map` | ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Limitations[​](#limitations "Direct link to Limitations") * Dremio connector does not support queries with the EXCEPT and INTERSECT keywords in Spice REPL. Use DISTINCT and IN/NOT IN instead. See the example below. ``` # fail SELECT ws_item_sk FROM web_sales INTERSECT SELECT ss_item_sk FROM store_sales; # success SELECT DISTINCT ws_item_sk FROM web_sales WHERE ws_item_sk IN ( SELECT DISTINCT ss_item_sk FROM store_sales ); # fail SELECT ws_item_sk FROM web_sales EXCEPT SELECT ss_item_sk FROM store_sales; # success SELECT DISTINCT ws_item_sk FROM web_sales WHERE ws_item_sk NOT IN ( SELECT DISTINCT ss_item_sk FROM store_sales ); ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Dremio as data connector in Spice. [Dremio Data Connector](https://github.com/spiceai/cookbook/tree/trunk/dremio#readme) --- # Dremio Data Connector Deployment Guide Production operating guide for the Dremio data connector covering authentication, Flight SQL transport, and operational tuning. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The Dremio connector connects over [Arrow Flight SQL](https://arrow.apache.org/docs/format/FlightSql.html) with username/password authentication. | Parameter | Description | | ----------------- | ------------------------------------------------------------- | | `dremio_endpoint` | Flight SQL endpoint, e.g. `grpc+tls://dremio.internal:32010`. | | `dremio_username` | Dremio user. | | `dremio_password` | Dremio password. Use `${secrets:...}` from a secret store. | Use TLS endpoints (`grpc+tls://`) in production. Credentials must be sourced from a [secret store](/docs/next/components/secret-stores). ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Flight SQL Transport[​](#flight-sql-transport "Direct link to Flight SQL Transport") Data transfer uses gRPC. Transient `UNAVAILABLE` / `DEADLINE_EXCEEDED` errors surface to the caller; the connector does not perform automatic retries, and no per-operation retry parameters are exposed at the Spice layer. Build retry/backoff into the calling application, or rely on Spice acceleration to reduce direct coordinator queries. ### Query Pushdown[​](#query-pushdown "Direct link to Query Pushdown") Dremio is a federated engine; the connector pushes SQL predicates and projections into Dremio where possible. For heavy analytical joins, execute the join in Dremio (via a Dremio view) rather than pulling raw rows into Spice. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Network**: Flight SQL is gRPC over HTTP/2. Co-locate Spice with Dremio for best latency. * **Coordinator load**: Every Spice query opens a Flight ticket against a Dremio coordinator. For dashboard-heavy workloads, accelerate the dataset in Spice (`acceleration: enabled`) to offload repeat queries from the coordinator. * **Result streaming**: Large result sets are streamed as Arrow record batches; memory footprint scales with the configured DataFusion batch size, not the result-set total. ## Metrics[​](#metrics "Direct link to Metrics") The Flight SQL client is not instrumented, so this connector emits no transport-level metrics, and it does not currently register Dremio-specific dataset-level instruments. Monitor via: * Spice query execution metrics (`query_duration_ms`, `query_returned_rows`, `query_failures`) from `runtime.metrics`. * Dremio's own job metrics exposed via the Dremio UI / API (`job_id` correlation). See [Component Metrics](/docs/next/features/observability/component_metrics) for configuration. ## Task History[​](#task-history "Direct link to Task History") Dremio queries participate in [task history](/docs/next/reference/task_history) via Flight client spans. Each Flight request is captured as a child of the enclosing `sql_query` or `accelerated_table_refresh` task. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Temporary tables**: Dremio temporary objects are not visible to Spice; use Dremio views for shared logic. * **Reflection-aware routing**: The connector does not explicitly hint Dremio reflections; they are still applied by the coordinator transparently. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ----------------------------------------- | -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | | `UNAUTHENTICATED` on handshake | Wrong or expired username/password credentials. | Verify `dremio_username` and `dremio_password`; reset the credentials via the Dremio UI if needed. | | `UNAVAILABLE` intermittent errors | Network partition or coordinator restart. | Errors surface to the caller; retry the request from the application and check coordinator health if persistent. | | `PERMISSION_DENIED` on a specific dataset | Dremio role lacks SELECT on the underlying source. | Grant access in Dremio via the user/role management UI. | | Slow queries for repeated dashboards | Coordinator overloaded by repeat queries. | Enable Spice acceleration for the dataset to cache results locally. | | TLS handshake failures | Self-signed cert or missing CA. | Configure TLS at the Flight client; ensure the CA bundle is trusted by the Spice runtime. | --- # DuckDB Data Connector DuckDB is an in-process SQL OLAP (Online Analytical Processing) database management system designed for analytical query workloads. It is optimized for fast execution and can be embedded directly into applications, providing efficient data processing without the need for a separate database server. This connector supports DuckDB [persistent databases](https://duckdb.org/docs/connect/overview#persistent-database) as a data source for federated SQL queries. ``` datasets: - from: duckdb:database.schema.table name: my_dataset params: duckdb_open: path/to/duckdb_file.duckdb ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field supports one of two forms: | `from` | Description | | ------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `duckdb:database.schema.table` | Read data from a table named `database.schema.table` in the DuckDB file | | `duckdb:*` | Read data using any DuckDB function that produces a table. For example one of the [data import](https://duckdb.org/docs/data/overview) functions such as `read_json`, `read_parquet` or `read_csv`. | info Unquoted identifiers are normalized to lowercase. To reference a table or schema with mixed-case characters, wrap each case-sensitive part in double quotes: `duckdb:my_database."MySchema"."MyTable"`. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: duckdb:database.schema.table name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The DuckDB data connector can be configured by providing the following `params`: | Parameter Name | Description | | -------------- | ----------------------------------------- | | `duckdb_open` | Path to the DuckDB database file to open. | Configuration `params` are provided either in the top level `dataset` for a dataset source, or in the `acceleration` section for a data store. ## Examples[​](#examples "Direct link to Examples") ### Reading from a relative path[​](#reading-from-a-relative-path "Direct link to Reading from a relative path") A generic example of DuckDB data connector configuration. ``` datasets: - from: duckdb:database.schema.table name: my_dataset params: duckdb_open: path/to/duckdb_file.duckdb ``` ### Reading from an absolute path[​](#reading-from-an-absolute-path "Direct link to Reading from an absolute path") ``` datasets: - from: duckdb:sample_data.nyc.rideshare name: nyc_rideshare params: duckdb_open: /my/path/my_database.db ``` ### DuckDB Functions[​](#duckdb-functions "Direct link to DuckDB Functions") Common [data import](https://duckdb.org/docs/data/overview) DuckDB functions can also define datasets. Instead of a fixed table reference (e.g. `database.schema.table`), a DuckDB function is provided in the `from:` key. For example ``` datasets: - from: duckdb:database.schema.table name: my_dataset params: duckdb_open: path/to/duckdb_file.duckdb - from: duckdb:read_csv('test.csv', header = false) name: from_function ``` Datasets created from DuckDB functions are similar to a standard `SELECT` query. For example: ``` datasets: - from: duckdb:read_csv('test.csv', header = false) ``` is equivalent to: ``` -- from_function SELECT * FROM read_csv('test.csv', header = false); ``` Many DuckDB data imports can be rewritten as DuckDB functions, making them usable as Spice datasets. For example: ``` SELECT * FROM 'todos.json'; -- As a DuckDB function SELECT * FROM read_json('todos.json'); ``` Limitations * The DuckDB connector does not support enum, dictionary, or map [field types](https://duckdb.org/docs/sql/data_types/overview). For example: * Unsupported: * `SELECT MAP(['key1', 'key2', 'key3'], [10, 20, 30])` * The DuckDB connector does not support `Decimal256` (76 digits), as it exceeds DuckDB's maximum Decimal width of 38 digits. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure DuckDB as a data connector in Spice. [DuckDB Data Connector](https://github.com/spiceai/cookbook/tree/trunk/duckdb/connector#readme) --- # DuckDB Data Connector Deployment Guide Production operating guide for the DuckDB data connector (used to federate queries against an existing DuckDB database file). ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") DuckDB is an embedded engine; the connector reads a local DuckDB database file. No network authentication is involved. | Parameter | Description | | ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `duckdb_open` | Path to the DuckDB database file. Required when reading from a table reference (`from: duckdb:database.schema.table`); it may be omitted only when the dataset's `from:` uses a DuckDB table function (e.g. `duckdb:read_csv(...)`), which runs against an in-memory database. Omitting it for a table-reference dataset fails with a `MissingDuckDBFile` error. | Protect the DuckDB file with filesystem permissions. Store it on encrypted storage (LUKS/dm-crypt, EBS encryption, etc.) for data-at-rest protection. For data loaded from cloud object stores inside DuckDB, configure AWS/Azure/GCS credentials via DuckDB extensions rather than Spice parameters. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### File Concurrency[​](#file-concurrency "Direct link to File Concurrency") DuckDB supports a single writer with many readers per database file. The Spice DuckDB data connector always opens the database in read-only mode, so it will not conflict with other readers. However, if another process holds a write lock, the connector may return an I/O error on open. Co-locate the writer and the Spice reader on the same host and ensure the writer releases its lock before the connector opens the file. ### Crash Recovery[​](#crash-recovery "Direct link to Crash Recovery") DuckDB's WAL provides crash recovery for any process that wrote to the file. The Spice connector does not itself write (the data connector is read-only; the DuckDB *accelerator* is distinct and handles write paths). ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Memory**: DuckDB manages its own memory based on available system memory when executing queries against the file. The connector opens the database read-only and exposes no memory-limit or connection-string setting — its only parameter is `duckdb_open`. For a Spice-configurable memory limit, use the [DuckDB accelerator](/docs/next/components/data-accelerators/duckdb), which supports the `duckdb_memory_limit` acceleration parameter. * **Disk**: Plan for 1.5–2× the raw data size to accommodate DuckDB's internal compression, WAL, and temporary spill files during query execution. * **Temporary spill**: Large queries can spill to DuckDB's temp directory (by default, alongside the database file). The connector exposes no spill-directory setting, so ensure the volume holding the database file has adequate free space. ## Metrics[​](#metrics "Direct link to Metrics") The DuckDB connector does not register connector-specific instruments. Monitor via Spice's query metrics (`query_duration_ms`, `query_returned_rows`). See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. For DuckDB-internal metrics, use DuckDB's `duckdb_memory()` and `pragma database_size` via a SQL query against the connector. ## Task History[​](#task-history "Direct link to Task History") DuckDB queries participate in [task history](/docs/next/reference/task_history) through DataFusion's execution-plan spans. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Read-only via the data connector**: For a writable, Spice-managed DuckDB, use the [DuckDB accelerator](/docs/next/components/data-accelerators/duckdb) instead. * **Single-writer**: A DuckDB file cannot be written by two processes concurrently. Coordinate writers out-of-band. * **Version compatibility**: DuckDB files are tied to the DuckDB binary version. Upgrading DuckDB in Spice may require regenerating older database files. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ---------------------------------------------------------------------------------- | --------------------------------------------------------- | ------------------------------------------------------------------------------------- | | `IO Error: Could not set lock on file` | Another process holds the DuckDB write lock. | Ensure only one writer; open in read-only mode if Spice should not hold a write lock. | | `Catalog Error: Table ... does not exist` | Table name mismatch or database not at the expected path. | Query `SELECT * FROM information_schema.tables` via the connector to list tables. | | Queries spill aggressively, slow performance | Working set exceeds memory. | Increase system memory or set a smaller batch size; direct temp to faster storage. | | `Serialization Error: Failed to deserialize ... database ... not a valid database` | DuckDB version mismatch. | Upgrade/downgrade Spice's DuckDB version to match the file producer. | --- # DuckLake Data Connector [DuckLake](https://ducklake.select/) is an open lakehouse format that stores metadata in a SQLite-compatible database (or PostgreSQL) and data in Parquet files. This connector enables querying individual DuckLake tables as datasets in Spice. For automatic discovery of all schemas and tables in a DuckLake catalog, use the [DuckLake Catalog Connector](/docs/next/components/catalogs/ducklake) instead. ``` datasets: - from: ducklake:my_table name: my_table params: ducklake_connection_string: s3://my-bucket/path/metadata.ducklake ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field specifies the DuckLake table to connect to. Use `ducklake:`, where `table_path` is the table name or a schema-qualified table name. | `from` | Description | | ----------------------------- | ------------------------------------------------- | | `ducklake:my_table` | Read from `my_table` in the default `main` schema | | `ducklake:my_schema.my_table` | Read from `my_table` in the `my_schema` schema | ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. ``` datasets: - from: ducklake:customer name: tpch_customer params: ducklake_connection_string: s3://my-bucket/metadata.ducklake ``` ``` SELECT COUNT(*) FROM tpch_customer; ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `ducklake_connection_string` | **Required**. The DuckLake metadata location (e.g., `s3://bucket/path/metadata.ducklake`). | | `ducklake_name` | The name to attach the DuckLake catalog as in DuckDB. Default: `ducklake`. | | `ducklake_open` | Path to an existing DuckDB file for persistent storage. If not provided, an in-memory DuckDB instance is used. | | `ducklake_aws_region` | Optional. The AWS region for S3 storage. Default: `us-east-1` when explicit credentials are provided. | | `ducklake_aws_access_key_id` | Optional. The AWS access key ID for S3 storage. Must be set together with `ducklake_aws_secret_access_key`. | | `ducklake_aws_secret_access_key` | Optional. The AWS secret access key for S3 storage. Must be set together with `ducklake_aws_access_key_id`. | | `ducklake_aws_session_token` | Optional. The AWS session token for S3 storage. Required with temporary (STS) credentials. Ignored, with a warning, unless `ducklake_aws_access_key_id` is also set. | | `ducklake_aws_endpoint` | Optional. Custom S3-compatible endpoint URL (e.g., for MinIO). | | `ducklake_aws_allow_http` | Optional. Set to `true` to allow HTTP (non-TLS) connections to S3. Default: `false`. | | `ducklake_automatic_migration` | Optional. Set to `true` to automatically migrate an older DuckLake catalog schema to the version required by the DuckLake extension on attach. Default: `false`. Migration rewrites catalog metadata and **cannot be undone**. | ### Connection string formats[​](#connection-string-formats "Direct link to Connection string formats") | Backend | Example | | ---------- | ------------------------------------------------------------------- | | Local file | `/path/to/metadata.ducklake` | | AWS S3 | `s3://bucket/path/metadata.ducklake` | | PostgreSQL | `postgres:dbname=mydb host=localhost user=postgres password=secret` | ## Authentication[​](#authentication "Direct link to Authentication") ### AWS S3[​](#aws-s3 "Direct link to AWS S3") When no explicit S3 credentials are configured, DuckDB falls back to its built-in credential chain provider: 1. Environment variables (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN`) 2. Shared credentials file (`~/.aws/credentials`) 3. IAM instance profiles (on EC2/ECS) To provide explicit S3 credentials, use the `ducklake_aws_*` parameters: ``` datasets: - from: ducklake:customer name: customer params: ducklake_connection_string: s3://my-bucket/metadata.ducklake ducklake_aws_region: us-west-2 ducklake_aws_access_key_id: ${secrets:AWS_ACCESS_KEY_ID} ducklake_aws_secret_access_key: ${secrets:AWS_SECRET_ACCESS_KEY} ``` For S3-compatible storage (e.g., MinIO), use `ducklake_aws_endpoint`: ``` datasets: - from: ducklake:customer name: customer params: ducklake_connection_string: s3://my-bucket/metadata.ducklake ducklake_aws_endpoint: http://minio:9000 ducklake_aws_access_key_id: ${secrets:MINIO_ACCESS_KEY} ducklake_aws_secret_access_key: ${secrets:MINIO_SECRET_KEY} ducklake_aws_allow_http: true ``` ## Write Support[​](#write-support "Direct link to Write Support") This connector supports writing data to DuckLake tables using SQL [`INSERT INTO`](/docs/next/reference/sql/dml#insert) statements when `access` is set to `read_write`: ``` datasets: - from: ducklake:customer name: customer access: read_write params: ducklake_connection_string: s3://my-bucket/metadata.ducklake ``` ``` INSERT INTO customer (c_custkey, c_name) VALUES (1, 'Acme Corp'); ``` `UPDATE` and `DELETE FROM` are not supported. For DDL operations (`CREATE TABLE`, `DROP TABLE`), use the [DuckLake Catalog Connector](/docs/next/components/catalogs/ducklake) with `access: read_write_create`. ## Examples[​](#examples "Direct link to Examples") ### Reading from a local DuckLake catalog[​](#reading-from-a-local-ducklake-catalog "Direct link to Reading from a local DuckLake catalog") ``` datasets: - from: ducklake:customer name: customer params: ducklake_connection_string: /path/to/metadata.ducklake ``` ### Reading from S3[​](#reading-from-s3 "Direct link to Reading from S3") ``` datasets: - from: ducklake:customer name: customer params: ducklake_connection_string: s3://my-bucket/lakehouse/metadata.ducklake ``` ### Reading from a specific schema[​](#reading-from-a-specific-schema "Direct link to Reading from a specific schema") ``` datasets: - from: ducklake:analytics.events name: events params: ducklake_connection_string: s3://my-bucket/metadata.ducklake ``` ### PostgreSQL metadata backend[​](#postgresql-metadata-backend "Direct link to PostgreSQL metadata backend") ``` datasets: - from: ducklake:customer name: customer params: ducklake_connection_string: "postgres:dbname=ducklake_catalog host=localhost user=postgres password=postgres" ``` ### Multiple tables with YAML anchors[​](#multiple-tables-with-yaml-anchors "Direct link to Multiple tables with YAML anchors") ``` datasets: - from: ducklake:customer name: customer params: &ducklake_params ducklake_connection_string: s3://my-bucket/metadata.ducklake - from: ducklake:orders name: orders params: *ducklake_params - from: ducklake:lineitem name: lineitem params: *ducklake_params ``` ### With data acceleration[​](#with-data-acceleration "Direct link to With data acceleration") ``` datasets: - from: ducklake:customer name: customer params: ducklake_connection_string: s3://my-bucket/metadata.ducklake acceleration: enabled: true engine: duckdb mode: file refresh_interval: 1h ``` Limitations * Spice uses DuckDB 1.5.3, which supports DuckLake 1.0. Older DuckLake catalogs require a metadata migration before use — set `ducklake_automatic_migration: true` to perform it on attach (this rewrites catalog metadata and cannot be undone). See [DuckLake migration guide](https://ducklake.select/docs/stable/duckdb/guides/troubleshooting#connecting-to-an-older-ducklake). * The DuckLake DuckDB extension is downloaded at runtime on first use, requiring network connectivity. * The `ducklake_connection_string` parameter is required — unlike the catalog connector, it cannot be omitted. * Each dataset creates its own DuckDB connection pool. For querying many tables from the same catalog, consider using the [DuckLake Catalog Connector](/docs/next/components/catalogs/ducklake) instead, which shares a single connection pool. * Writes are limited to `INSERT INTO`. `UPDATE`, `DELETE FROM`, and DDL (`CREATE TABLE`, `DROP TABLE`) are not supported on the data connector — use the [DuckLake Catalog Connector](/docs/next/components/catalogs/ducklake) for schema operations. --- # DynamoDB Data Connector Amazon DynamoDB is a fully managed NoSQL database service that provides fast and predictable performance with seamless scalability. This connector enables using DynamoDB tables as data sources for federated SQL queries in Spice. ``` datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 dynamodb_aws_access_key_id: ${secrets:aws_access_key_id} # Optional dynamodb_aws_secret_access_key: ${secrets:aws_secret_access_key} # Optional dynamodb_aws_session_token: ${secrets:aws_session_token} # Optional ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field should specify the DynamoDB table name: | `from` | Description | | ---------------- | --------------------------------------------- | | `dynamodb:table` | Read data from a DynamoDB table named `table` | note If an expected table is not found, verify the `dynamodb_aws_region` parameter. DynamoDB tables are region-specific. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: dynamodb:users name: my_users params: ... ``` ``` SELECT COUNT(*) FROM my_users; ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The DynamoDB data connector supports the following configuration parameters: | Parameter Name | Description | | -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `dynamodb_aws_region` | Required. The AWS region containing the DynamoDB table | | `dynamodb_aws_access_key_id` | Optional. AWS access key ID for authentication. If not provided, credentials will be loaded from environment variables or IAM roles | | `dynamodb_aws_secret_access_key` | Optional. AWS secret access key for authentication. If not provided, credentials will be loaded from environment variables or IAM roles | | `dynamodb_aws_session_token` | Optional. AWS session token for authentication | | `dynamodb_aws_auth` | Optional. Authentication method. Use `iam_role` (default) for IAM role-based authentication or `key` for explicit access key credentials. | | `dynamodb_aws_iam_role_source` | Optional. IAM role credential source (only used when `dynamodb_aws_auth: iam_role`). `auto` (default) uses the default AWS credential chain, `metadata` uses only instance/container metadata (IMDS, ECS, EKS/IRSA), `env` uses only environment variables | | `unnest_depth` | Optional. Maximum nesting depth for unnesting embedded documents into a flattened structure. Higher values expand deeper nested fields. | | `schema_infer_max_records` | Optional. The number of documents to use to infer the schema. Defaults to 10 | | `scan_segments` | Optional. Number of segments for `Scan` request. 'auto' by default, which will calculate number of segments based on number of the records in a table | | `scan_interval` | Optional. Interval between polling for new records in a DynamoDB stream. Default: `0s`. See [Streams](#streams). | | `dynamodb_replication_ready_lag` | Optional. For `refresh_mode: changes`, the dataset is marked Ready once its replication lag (now minus the newest applied source-commit time) falls below this. It stays not-ready while snapshotting or draining a backlog, so it never serves stale data. Default: `2s`. See [Streams](#streams). | | `dynamodb_replication_initial_snapshot` | Optional. When `refresh_mode: changes` first loads the table's existing items: `auto` (default) scans when no resumable stream checkpoint exists and resumes without a scan when one does; `disabled` streams changes only, from the current stream tip; `always` scans on every start, discarding any persisted checkpoint. Default: `auto`. | | `dynamodb_replication_invalid_checkpoint_behavior` | Optional. Behavior when the persisted stream checkpoint can no longer be honored (past the \~24h shard retention). One of `error` (default — marks dataset as Error) or `restart` (re-bootstraps from a fresh scan). Default: `error`. | | `endpoint_url` | Optional. Custom endpoint URL for DynamoDB-compatible services (e.g., DynamoDB Local, ScyllaDB Alternator). | | `ready_lag` | *Deprecated.* Renamed to `dynamodb_replication_ready_lag`; still accepted as an alias. | | `lag_exceeds_shard_retention_behavior` | *Deprecated.* Renamed to `dynamodb_replication_invalid_checkpoint_behavior` (`error` \| `restart`). `ready_after_load` maps to `restart`; `ready_before_load` has been removed and also maps to `restart`. | | `time_format` | Optional. Go-style time format used for parsing/formatting timestamps. See [Time Format](#time-format) | | `write_parallelism` | Optional. Number of parallel operations for writing and deleting data to DynamoDB. Default: `10` | ### Authentication[​](#authentication "Direct link to Authentication") The DynamoDB connector supports two authentication methods controlled by the `dynamodb_aws_auth` parameter: ``` datasets: - from: dynamodb:my_table params: dynamodb_aws_auth: iam_role | key # iam_role is default dynamodb_aws_iam_role_source: auto | metadata | env # auto is default (only used with iam_role) ``` #### IAM Role Authentication (`dynamodb_aws_auth: iam_role`)[​](#iam-role-authentication-dynamodb_aws_auth-iam_role "Direct link to iam-role-authentication-dynamodb_aws_auth-iam_role") This is the default authentication method. When using IAM role authentication, the `dynamodb_aws_iam_role_source` parameter controls which credential sources are used: | Source Value | Description | Credential Sources | | ---------------- | ------------------------------------- | ----------------------------------------------------------------------------- | | `auto` (default) | Uses the default AWS credential chain | All sources listed below, in order | | `metadata` | Uses only instance/container metadata | Web Identity Token, ECS Container Credentials, EC2 Instance Metadata (IMDSv2) | | `env` | Uses only environment variables | `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN` | note When using `iam_role` authentication, any explicitly provided access keys (`dynamodb_aws_access_key_id`, `dynamodb_aws_secret_access_key`) are ignored. #### Key-Based Authentication (`dynamodb_aws_auth: key`)[​](#key-based-authentication-dynamodb_aws_auth-key "Direct link to key-based-authentication-dynamodb_aws_auth-key") When `dynamodb_aws_auth` is set to `key`, credentials must be provided explicitly: ``` datasets: - from: dynamodb:my_table params: dynamodb_aws_auth: key dynamodb_aws_access_key_id: ${secrets:aws_access_key_id} dynamodb_aws_secret_access_key: ${secrets:aws_secret_access_key} dynamodb_aws_session_token: ${secrets:aws_session_token} # Optional, for temporary credentials ``` #### Default Credential Chain (`auto`)[​](#default-credential-chain-auto "Direct link to default-credential-chain-auto") When using `dynamodb_aws_auth: iam_role` with `dynamodb_aws_iam_role_source: auto` (or when both parameters are omitted), credentials are loaded from the following sources in order: 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate DynamoDB permissions (e.g., `dynamodb:Scan`, `dynamodb:Query`, `dynamodb:DescribeTable`) to access the tables. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access DynamoDB tables: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "dynamodb:Scan", "dynamodb:Query", "dynamodb:DescribeTable" ], "Resource": [ "arn:aws:dynamodb:*:*:table/YOUR_TABLE_NAME" ] } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------------- | -------------------------------------------------------------------------------- | | `dynamodb:Scan` | Required. Allows reading all items from the table | | `dynamodb:Query` | Required. Allows reading items from the table using partition key | | `dynamodb:DescribeTable` | Required. Allows fetching table metadata and schema information | | `dynamodb:BatchWriteItem` | Required for INSERT and DELETE. Both write paths issue `BatchWriteItem` requests | | `dynamodb:UpdateItem` | Required for UPDATE. Allows modifying existing items in the table | Streams / CDC permissions When using `refresh_mode: changes` (DynamoDB Streams), the IAM role or user additionally needs `dynamodb:DescribeStream`, `dynamodb:GetShardIterator`, and `dynamodb:GetRecords`, scoped to the table's stream ARN (`arn:aws:dynamodb:*:*:table/YOUR_TABLE_NAME/stream/*`). ### Example IAM Policies[​](#example-iam-policies "Direct link to Example IAM Policies") #### Minimal Policy (Read-only access to specific table)[​](#minimal-policy-read-only-access-to-specific-table "Direct link to Minimal Policy (Read-only access to specific table)") ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "dynamodb:Scan", "dynamodb:Query", "dynamodb:DescribeTable" ], "Resource": "arn:aws:dynamodb:us-west-2:123456789012:table/users" } ] } ``` #### Access to Multiple Tables[​](#access-to-multiple-tables "Direct link to Access to Multiple Tables") ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "dynamodb:Scan", "dynamodb:Query", "dynamodb:DescribeTable" ], "Resource": [ "arn:aws:dynamodb:us-west-2:123456789012:table/users", "arn:aws:dynamodb:us-west-2:123456789012:table/orders" ] } ] } ``` #### Access to All Tables in a Region[​](#access-to-all-tables-in-a-region "Direct link to Access to All Tables in a Region") ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "dynamodb:Scan", "dynamodb:Query", "dynamodb:DescribeTable" ], "Resource": "arn:aws:dynamodb:us-west-2:123456789012:table/*" } ] } ``` Security Considerations * Avoid using `dynamodb:*` permissions as it grants more access than necessary. * Consider using more restrictive policies in production environments. * When using IAM roles with EKS, ensure the [service account is properly configured with IRSA](https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html). ## Data Types[​](#data-types "Direct link to Data Types") The table below shows the DynamoDB data types supported, along with the type mapping to Apache Arrow types in Spice. | DynamoDB Type | Description | Arrow Type | Notes | | ------------- | ----------- | ------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------- | | `Bool` | Boolean | `Boolean` | | | `S` | String | `Utf8` | | | `S` | String | `Timestamp(Millisecond)` | Naive timestamp if it matches `time_format` without timezone | | `S` | String | `Timestamp(Millisecond, )` | Timezone-aware timestamp if it matches `time_format` with timezone | | `S` | String | `Date32` | Date-only string in `YYYY-MM-DD` format (when it does not match `time_format`) | | `Ss` | String Set | `List` | | | `N` | Number | `Int64` \| `Float64` | | | `Ns` | Number Set | `List` | | | `B` | Binary | `Binary` | | | `Bs` | Binary Set | `List` | | | `L` | List | `List` | DynamoDB arrays can be heterogeneous e.g. `[1, "foo", true]`, Arrow arrays must be homogeneous - use strings to preserve all data | | `M` | Map | `Utf8` or Unflattened | Depending on `unnest_depth` value | ## Time format[​](#time-format "Direct link to Time format") Since DynamoDB stores timestamps as strings, Spice supports parsing timestamps using a customizable format. By default, Spice will try to parse timestamps using ISO8601 format, but you can provide a custom format using the `time_format` parameter. Once Spice is able to parse a timestamp, it will convert it to a `Timestamp(Millisecond)` Arrow type, and will use the same format to serialize it back to DynamoDB for filter pushdown. This parameter uses Go-style time formatting, which uses a reference time of `Mon Jan 2 15:04:05 MST 2006`. | Format Pattern | Example Value | Description | | ------------------------------- | ------------------------------- | ---------------------------------------------------------- | | `2006-01-02T15:04:05.000Z07:00` | `2024-03-15T14:30:00.000Z` | ISO8601 / RFC3339 with milliseconds and timezone (default) | | `2006-01-02T15:04:05.999Z07:00` | `2024-03-15T14:30:00.123-07:00` | ISO8601 with milliseconds and timezone | | `2006-01-02T15:04:05` | `2024-03-15T14:30:00` | ISO8601 without timezone (naive timestamp) | | `2006-01-02 15:04:05` | `2024-03-15 14:30:00` | Date and time with space separator | | `01/02/2006 15:04:05` | `03/15/2024 14:30:00` | US-style date with time | | `02/01/2006 15:04:05` | `15/03/2024 14:30:00` | European-style date with time | | `Jan 2, 2006 3:04:05 PM` | `Mar 15, 2024 2:30:00 PM` | Human-readable with 12-hour clock | | `20060102150405` | `20240315143000` | Compact format (no separators) | Go's format uses specific reference values that must appear exactly as shown: | Component | Reference Value | Alternatives | | ------------ | --------------- | ------------------------------------- | | Year | `2006` | `06` (2-digit) | | Month | `01` | `1`, `Jan`, `January` | | Day | `02` | `2` | | Hour (24h) | `15` | — | | Hour (12h) | `03` | `3` | | Minute | `04` | `4` | | Second | `05` | `5` | | AM/PM | `PM` | `pm` | | Timezone | `Z07:00` | `-0700`, `MST` | | Milliseconds | `.000` | `.999` (trailing zeros trimmed) | | Microseconds | `.000000` | `.999999` (trailing zeros trimmed) | | Nanoseconds | `.000000000` | `.999999999` (trailing zeros trimmed) | ## Unnesting[​](#unnesting "Direct link to Unnesting") Consider the following document: ``` { "a": 1, "b": { "x": 2, "y": { "z": 3 } } } ``` Using `unnest_depth` you can control the unnesting behavior. Here are the examples: ### unnest\_depth: 0[​](#unnest_depth-0 "Direct link to unnest_depth: 0") ``` sql> select * from test_table; +-----------+---------------------+ | a (Int32) | b (Utf8) | +-----------+---------------------+ | 1 | {"x":2,"y":{"z":3}} | +---+-----------------------------+ ``` ### unnest\_depth: 1[​](#unnest_depth-1 "Direct link to unnest_depth: 1") ``` sql> select * from test_table; +-----------+-------------+------------+ | a (Int32) | b.x (Int32) | b.y (Utf8) | +-----------+-------------+------------+ | 1 | 2 | {"z":3} | +-----------+-------------+------------+ ``` ### unnest\_depth: 2[​](#unnest_depth-2 "Direct link to unnest_depth: 2") ``` sql> select * from test_table; +-----------+-------------+---------------+ | a (Int32) | b.x (Int32) | b.y.z (Int32) | +-----------+-------------+---------------+ | 1 | 2 | 3 | +-----------+-------------+---------------+ ``` ## JSON Nesting[​](#json-nesting "Direct link to JSON Nesting") When working with DynamoDB tables that have many columns, you can consolidate unspecified columns into a single JSON column using the `json_object` metadata option. This is useful when you only need a few columns as discrete fields and want to bundle the remaining columns into a single JSON structure. ### Configuration[​](#configuration-1 "Direct link to Configuration") To use JSON nesting, define your desired columns explicitly in the `columns` list and add a "catch-all" column with `json_object: "*"` metadata. Any columns from the source table that are not explicitly listed will be nested into this JSON column. ``` datasets: - from: dynamodb:my_table name: my_table params: dynamodb_aws_region: us-west-2 columns: - name: PK - name: SK - name: Baz - name: data_json metadata: json_object: "*" ``` ### Example[​](#example "Direct link to Example") Given a DynamoDB table with this schema: | Column | Type | | ------ | ------ | | PK | String | | SK | String | | Foo | Map | | Bar | List | | Baz | String | The configuration above produces: | Column | Type | | ---------- | -------------------------------------- | | PK | String | | SK | String | | Baz | String | | data\_json | JSON (`{"Foo": , "Bar": }`) | The `Foo` and `Bar` columns, which were not explicitly listed, are automatically nested into the `data_json` column as a JSON object. ### Interaction with `unnest_depth`[​](#interaction-with-unnest_depth "Direct link to interaction-with-unnest_depth") When both `unnest_depth` and `json_object` are specified, the operations are applied in this order: 1. **Unnesting first**: Nested structures are flattened according to the `unnest_depth` value 2. **JSON nesting second**: Unspecified columns are then consolidated into the `json_object` column Consider this DynamoDB dataset: ``` +----+------+-------------------+------------------------------------------------------+ | PK | SK | Baz | Foo | +----+------+-------------------+------------------------------------------------------+ | 1 | 200 | some_string_value | { "Age" : { "N" : "35" }, "Name" : { "S" : "Joe" } } | +----+------+-------------------+------------------------------------------------------+ ``` And this configuration ``` datasets: - from: dynamodb:my_table name: my_table params: dynamodb_aws_region: us-west-2 unnest_depth: 1 columns: - name: PK - name: SK metadata: json_object: "*" ``` Will produce the following Spice dataset: ``` +----+-----+---------------------------------------------------------+ | PK | SK | json_data | +----+-----+---------------------------------------------------------+ | 1 | 200 | {"Baz":"some_string", "Foo.Age":35.0, "Foo.Name":"Joe"} | +----+-----+---------------------------------------------------------+ ``` Limitations * The `json_object` metadata only accepts `"*"` as its value, which captures all unspecified columns * Only one column can have the `json_object` metadata. Specifying multiple columns with `json_object` will result in an error ## Data Manipulation (DML)[​](#data-manipulation-dml "Direct link to Data Manipulation (DML)") The DynamoDB connector supports `INSERT`, `UPDATE`, and `DELETE` operations. ``` -- Insert a new item INSERT INTO users (id, email, name) VALUES (42, 'user@example.com', 'Jane'); -- Update existing items UPDATE users SET name = 'Jane Doe' WHERE id = 42; -- Delete items DELETE FROM users WHERE id = 42; ``` Write operations use the table's primary key (partition key and optional sort key) to identify items. The `write_parallelism` parameter controls how many DynamoDB API calls are issued concurrently for batch operations (default: `10`). note INSERT and DELETE both use `BatchWriteItem` (with `PutRequest` and `DeleteRequest` items respectively) for efficiency. UPDATE uses per-item `UpdateItem` calls. All write operations are issued in parallel chunks sized by `write_parallelism`. ## Examples[​](#examples "Direct link to Examples") ### Basic Configuration with Environment Credentials[​](#basic-configuration-with-environment-credentials "Direct link to Basic Configuration with Environment Credentials") ``` version: v1 kind: Spicepod name: dynamodb datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 acceleration: enabled: true ``` ### Configuration with Explicit Credentials[​](#configuration-with-explicit-credentials "Direct link to Configuration with Explicit Credentials") ``` version: v1 kind: Spicepod name: dynamodb datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 dynamodb_aws_auth: key dynamodb_aws_access_key_id: ${secrets:aws_access_key_id} dynamodb_aws_secret_access_key: ${secrets:aws_secret_access_key} acceleration: enabled: true ``` ### Configuration with Metadata-Only Credentials (ECS/EKS)[​](#configuration-with-metadata-only-credentials-ecseks "Direct link to Configuration with Metadata-Only Credentials (ECS/EKS)") ``` version: v1 kind: Spicepod name: dynamodb datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 dynamodb_aws_auth: iam_role dynamodb_aws_iam_role_source: metadata acceleration: enabled: true ``` ### Configuration with time\_format[​](#configuration-with-time_format "Direct link to Configuration with time_format") ``` version: v1 kind: Spicepod name: dynamodb datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 time_format: 2006-01-02 15:04:05 acceleration: enabled: true ``` ### Querying Nested Structures[​](#querying-nested-structures "Direct link to Querying Nested Structures") DynamoDB supports complex nested JSON structures. These fields can be queried using SQL: ``` -- Query nested structs SELECT metadata.registration_ip, metadata.user_agent FROM users LIMIT 5; -- Query nested structs in arrays SELECT address.city FROM ( SELECT unnest(addresses) AS address FROM users ) WHERE address.city = 'San Francisco'; ``` ### Limitations[​](#limitations "Direct link to Limitations") Limitations * The DynamoDB connector will scan the first 10 items to determine the schema of the table. This may miss columns that are not present in the first 10 items. * The DynamoDB connector does not support Decimal type. Example schema from a users table: ``` describe users; ``` ``` +----------------+------------------+-------------+ | column_name | data_type | is_nullable | +----------------+------------------+-------------+ | email | Utf8 | YES | | id | Int64 | YES | | metadata | Struct | YES | | addresses | List(Struct) | YES | | preferences | Struct | YES | | created_at | Utf8 | YES | ... +----------------+------------------+-------------+ ``` ## Streams[​](#streams "Direct link to Streams") The DynamoDB Data Connector integrates with [DynamoDB Streams](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.html) to enable real-time streaming of table changes. This feature supports both initial table bootstrapping and continuous change data capture (CDC), so Spice can automatically detect and stream inserts, updates, and deletes from DynamoDB tables. warning Using DynamoDB Streams **requires** [acceleration](/docs/next/components/data-accelerators) with `refresh_mode: changes`. ### Basic Configuration[​](#basic-configuration "Direct link to Basic Configuration") To enable streaming from DynamoDB, enable acceleration and set the `refresh_mode` to `changes` in your dataset configuration. ``` datasets: - from: dynamodb:my_table name: orders_stream acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes ``` ### Configuration Parameters[​](#configuration-parameters "Direct link to Configuration Parameters") #### Dataset Parameters[​](#dataset-parameters "Direct link to Dataset Parameters") * **`dynamodb_replication_ready_lag`** - Defines the maximum lag threshold before the dataset is reported as "Ready". Once the stream lag falls below this value, queries can be executed against the dataset. Default: `2s`. (Previously `ready_lag`, still accepted as a deprecated alias.) * **`scan_interval`** - Controls the polling frequency for checking new records in the DynamoDB stream. Lower values provide more real-time updates but increase API calls. Higher values reduce API usage but may introduce additional latency. #### Acceleration Parameters[​](#acceleration-parameters "Direct link to Acceleration Parameters") * **`snapshots`** - Optional. Controls snapshots behavior. Supported values are `disabled` (default), `enabled`, `create_only`, `bootstrap_only`. * **`snapshots_trigger`** - Optional. Determines type of trigger for creating snapshots. Supported values are `time_interval` (default) and `stream_batches`. * **`snapshots_trigger_threshold`** - Optional. Threshold value for snapshot creation. The format depends on the `snapshots_trigger` type: * When `snapshots_trigger` is `stream_batches`: a raw integer specifying the number of batches (e.g., `100`, `1000`). * When `snapshots_trigger` is `time_interval`: an integer with a time unit suffix (e.g., `10m`, `30s`, `1h`). See [Acceleration snapshots](/docs/next/features/data-acceleration/snapshots) for more details. ### Metrics[​](#metrics "Direct link to Metrics") The following [Component Metrics](/docs/next/features/observability/component_metrics) are provided for monitoring streaming performance and health: | Metric | Type | Description | | -------------------------------------------------------- | ------- | -------------------------------------------------------------------------- | | `shards_active` | Gauge | Current number of active shards in the stream | | `records_consumed_total` | Counter | Total number of records consumed from the stream | | `lag_ms` | Gauge | Current lag in milliseconds between stream watermark and the current time | | `errors_transient_total` | Counter | Total number of transient errors encountered while polling from the stream | | `reinitializations_on_lag_exceeds_shard_retention_total` | Counter | Total rebootstrap operations triggered due to expired shards | These metrics are not enabled by default, enable them by setting the metrics parameter: ``` datasets: - from: dynamodb:user_events name: events acceleration: enabled: true refresh_mode: changes metrics: - name: shards_active - name: records_consumed_total - name: lag_ms - name: errors_transient_total - name: reinitializations_on_lag_exceeds_shard_retention_total ``` You can find an example dashboard for DynamoDB Streams in [monitoring/grafana-dashboard.json](https://github.com/spiceai/spiceai/blob/trunk/monitoring/grafana-dashboard.json). ## Advanced Configuration[​](#advanced-configuration "Direct link to Advanced Configuration") For production workloads requiring fine-tuned control over streaming behavior and performance characteristics: ``` datasets: - from: dynamodb:my_table name: orders_stream params: dynamodb_replication_ready_lag: 1s # Dataset reports as Ready when lag is below 1 second scan_interval: 100ms # Poll for new stream records every 100 milliseconds acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes snapshots: enabled snapshots_trigger: stream_batches snapshots_trigger_threshold: 5 # Create snapshots every 5 batch updates metrics: - name: shards_active - name: records_consumed_total - name: lag_ms - name: errors_transient_total ``` Limitations * DynamoDB Streams connector does not support `refresh_sql`. ## Cookbooks[​](#cookbooks "Direct link to Cookbooks") * A cookbook recipe to configure DynamoDB as a data connector in Spice. [DynamoDB Data Connector](https://github.com/spiceai/cookbook/tree/trunk/dynamodb#readme) * A cookbook recipe to configure DynamoDB Streams as a data connector in Spice. [DynamoDB Streams Data Connector](https://github.com/spiceai/cookbook/tree/trunk/dynamodb/streams#readme) --- # DynamoDB Data Connector Deployment Guide Production operating guide for the DynamoDB data connector covering IAM, DynamoDB Streams CDC, checkpointing, and lag handling. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") DynamoDB authentication uses the standard AWS credential chain. Configure via the same parameters as the [S3 connector](/docs/next/components/data-connectors/s3/deployment#authentication--secrets): | Parameter | Description | | -------------------------------- | ------------------------------------------------------------------------------ | | `dynamodb_aws_region` | AWS region of the DynamoDB table. | | `dynamodb_aws_access_key_id` | Explicit access key (optional; falls back to the credential chain when unset). | | `dynamodb_aws_secret_access_key` | Explicit secret key (optional). | | `dynamodb_aws_session_token` | Session token for temporary credentials (optional). | For production on EKS/ECS, leave access-key parameters unset and rely on instance-profile, IRSA, or ECS task-role credentials. Grant the role `dynamodb:Scan`, `dynamodb:Query`, and `dynamodb:DescribeTable` on the table; for streams, additionally grant `dynamodb:DescribeStream`, `dynamodb:GetShardIterator`, and `dynamodb:GetRecords`. Secrets should be sourced from a [secret store](/docs/next/components/secret-stores) when not using IAM role auth. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Streams and Checkpointing[​](#streams-and-checkpointing "Direct link to Streams and Checkpointing") The DynamoDB connector supports CDC via DynamoDB Streams with an accelerated dataset as the sink. Stream state is persisted as a checkpoint alongside the accelerator, allowing resumption after a restart. | Parameter | Default | Description | | -------------------------------------------------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `scan_interval` | `0s` | Interval between polls for new records in a DynamoDB stream. | | `dynamodb_replication_ready_lag` | `2s` | Once replication lag falls below this threshold, the dataset is reported as `Ready`. (Previously `ready_lag`, still accepted as a deprecated alias.) | | `dynamodb_replication_initial_snapshot` | `auto` | When `refresh_mode: changes` first loads existing items: `auto` scans only when no resumable checkpoint exists, `disabled` streams changes only, `always` scans on every start. | | `dynamodb_replication_invalid_checkpoint_behavior` | `error` | Behavior when the persisted checkpoint can no longer be honored (past \~24h shard retention): `error` or `restart`. (Previously `lag_exceeds_shard_retention_behavior`, still accepted as a deprecated alias.) | ### Shard Retention and Lag[​](#shard-retention-and-lag "Direct link to Shard Retention and Lag") DynamoDB Streams retain records for 24 hours. If Spice is offline longer than the retention window, the checkpoint becomes stale and the next stream open returns `ShardNotFound`. Behavior is controlled by `dynamodb_replication_invalid_checkpoint_behavior`: * **`error`** (default): Mark the dataset `Error`. Requires operator intervention to re-bootstrap. * **`restart`**: Drop the stale checkpoint and re-bootstrap the accelerated dataset from a fresh scan. A checkpoint older than **18 hours** is treated as near-expired and triggers the same recovery path even if the shard has not yet been dropped by DynamoDB. The deprecated `lag_exceeds_shard_retention_behavior` alias is still accepted: `ready_after_load` maps to `restart`, and `ready_before_load` (which briefly served stale data before reloading) has been removed and also maps to `restart`. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Read capacity**: Full-table scans consume provisioned read capacity. Use on-demand billing or reserve sufficient RCU for refresh windows to avoid throttling. * **Stream throughput**: DynamoDB Streams shards cap at 1000 records/sec and 2 MB/sec each. Wide or high-write tables automatically partition into more shards. * **Checkpoint storage**: Checkpoint records live in the acceleration engine and add roughly one row per stream shard. Negligible for sizing. ## Metrics[​](#metrics "Direct link to Metrics") The DynamoDB connector registers the following metrics: | Metric | Type | Description | | -------------------------------------------------------- | ------- | --------------------------------------------------------------------------------------- | | `shards_active` | Gauge | Number of active DynamoDB Streams shards being consumed. | | `records_consumed_total` | Counter | Total stream records consumed. | | `lag_ms` | Gauge | Current lag behind the stream head in milliseconds (approximate, summed across shards). | | `errors_transient_total` | Counter | Transient stream read errors (retried automatically). | | `reinitializations_on_lag_exceeds_shard_retention_total` | Counter | Number of times the stream was reinitialized due to lag exceeding shard retention. | Metrics are exposed with the `dataset_dynamodb_` prefix. Monitor `lag_ms` together with `errors_transient_total` — a climbing lag with rising transient errors indicates the connector is falling behind retention. See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting metrics. ## Task History[​](#task-history "Direct link to Task History") Stream polling and bootstrap operations emit spans that participate in [task history](/docs/next/reference/task_history) under the enclosing `accelerated_table_refresh` and changes-stream tasks. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Global Secondary Indexes**: Not exposed as separate datasets. Query the base table and let DataFusion filter. * **Conditional writes**: DynamoDB conditional expressions (e.g., `attribute_exists`) are not supported in DML operations. * **Cross-region streams**: Must configure `dynamodb_aws_region` to match the region of the source table; cross-region access requires resource policies and is not recommended. * **Table with `StreamSpecification` disabled**: CDC mode is unavailable; fall back to full-table refresh. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ---------------------------------------------------------- | --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | | Dataset stuck in `Error` after restart with stream enabled | Checkpoint older than 18h or exceeded 24h retention. | Set `dynamodb_replication_invalid_checkpoint_behavior: restart` to auto-recover, or trigger a manual refresh. | | `ProvisionedThroughputExceededException` | RCU exhausted during initial scan. | Switch to on-demand billing, raise RCU for the refresh window, or slow the refresh via acceleration settings. | | `TrimmedDataAccessException` | Records trimmed from the stream before they could be processed. | Same recovery path as `ShardNotFound` — re-bootstrap. Reduce bootstrap duration via parallel segments if supported. | | `AccessDeniedException` on `DescribeStream` | IAM role lacks stream permissions. | Add `dynamodb:DescribeStream`, `GetShardIterator`, `GetRecords` to the role. | | `ResourceNotFoundException` on stream start | Stream not enabled on the table. | Enable streams on the DynamoDB table (`NEW_AND_OLD_IMAGES` recommended). | --- # Elasticsearch Data Connector The Elasticsearch Data Connector exposes Elasticsearch indexes as SQL tables in Spice. Index mappings are translated to Arrow schemas so that documents can be queried with federated SQL alongside data from other connectors. To run [vector, full-text, or hybrid search](/docs/next/features/search) (the `vector_search`, `text_search`, and `rrf` UDTFs) against an Elasticsearch index, the dataset must additionally be configured for search — as an [Elasticsearch Vector Engine](/docs/next/components/vectors/elasticsearch) with an embedding model for vector search, and/or with full-text search columns for `text_search`. See [Vector and Full-Text Search](#vector-and-full-text-search) below. Registering an index through the data connector alone exposes it for federated SQL but does not make it searchable through those UDTFs. ``` datasets: - from: elasticsearch:products name: products params: elasticsearch_endpoint: https://localhost:9200 elasticsearch_user: ${secrets:es_user} elasticsearch_pass: ${secrets:es_pass} ``` Enterprise edition The Elasticsearch connector is available in the Spice [Enterprise edition](https://docs.spice.ai/docs/enterprise/getting-started/distributions). ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `elasticsearch:{index_name}` where `index_name` is the Elasticsearch index to query. ``` datasets: - from: elasticsearch:products name: products ``` Dot-separated paths may be used to refer to nested fields in query results (e.g. `address.city`); the connector flattens object mappings into Arrow columns using that convention. ### `name`[​](#name "Direct link to name") The dataset name used as the table name within Spice. The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The Elasticsearch connector accepts the following `params`. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load credentials from a secret store. | Parameter Name | Description | Required | Default | | ------------------------ | --------------------------------------------- | -------- | ------- | | `elasticsearch_endpoint` | Cluster URL (e.g., `https://localhost:9200`). | Yes | - | | `elasticsearch_user` | Username for HTTP basic authentication. | No | - | | `elasticsearch_pass` | Password for HTTP basic authentication. | No | - | ## Types[​](#types "Direct link to Types") The connector derives an Arrow schema from each index's mapping via `GET //_mapping`. Elasticsearch field types map to Arrow as follows: | Elasticsearch Field Type | Arrow Type | Notes | | -------------------------------------------------------------------- | ------------------------------ | -------------------------------------------------------------------------------------------------------- | | `text`, `keyword`, `wildcard`, `constant_keyword`, `match_only_text` | `Utf8` | | | `long` | `Int64` | | | `unsigned_long` | `UInt64` | Accepts both numeric values and digit strings (JS clients commonly serialize values > 253-1 as strings). | | `integer` | `Int32` | | | `short` | `Int16` | | | `byte` | `Int8` | | | `double` | `Float64` | | | `float`, `half_float`, `scaled_float` | `Float32` | | | `boolean` | `Boolean` | | | `date`, `date_nanos` | `Utf8` | ES dates are flexibly formatted; preserved as strings. | | `binary` | `Utf8` | Base64-encoded in the JSON response. | | `ip` | `Utf8` | | | `dense_vector` (with `dims`) | `FixedSizeList` | Required `dims` field must fit in `i32`. | | `dense_vector` (missing `dims`) | `Utf8` | Falls back to raw JSON when dims cannot be resolved. | | `object` (with sub-fields) | *(flattened)* | Expanded into dot-separated columns (e.g. `address.city`). | | `object` (no sub-fields), `nested` | `Utf8` | Serialized JSON. | | Any other mapping type | `Utf8` | Fallback — the raw JSON value is preserved as a string. | Nested `object` fields are flattened by concatenating field names with dots (e.g. `address.city`). `nested` fields are preserved as JSON strings because per-document ordering must be retained. ## Querying[​](#querying "Direct link to Querying") After registering a dataset, query it like any other Spice table: ``` SELECT name, price FROM products WHERE price > 100 ORDER BY price DESC LIMIT 10; ``` ### Vector and Full-Text Search[​](#vector-and-full-text-search "Direct link to Vector and Full-Text Search") An Elasticsearch dataset is **not** searchable through the search UDTFs by virtue of being registered with the data connector. To enable search against an Elasticsearch index, configure the dataset for search: * For **vector** and **hybrid** search, configure the dataset as an [Elasticsearch Vector Engine](/docs/next/components/vectors/elasticsearch) (`vectors: { engine: elasticsearch, enabled: true }`) with a column-level `embeddings` entry naming an embedding model. The embedding model is required — it is used to embed the query text at search time. * For **full-text** search, enable `full_text_search` on the column(s) to search. Once configured, the following UDTFs are available against the dataset: * **Vector similarity search** via [`vector_search`](/docs/next/reference/sql/search#vector-search-vector_search) — executed natively as an Elasticsearch kNN query. * **Full-text search** via [`text_search`](/docs/next/reference/sql/search#full-text-search-text_search) — executed using Elasticsearch's native BM25 ranking. * **Hybrid search** via [`rrf`](/docs/next/reference/sql/search#reciprocal-rank-fusion-rrf) — combining both with Reciprocal Rank Fusion. These operations run against the Elasticsearch cluster directly rather than ingesting vectors into an accelerator, keeping indexing and search colocated in Elasticsearch. Example: ``` -- kNN vector search against Elasticsearch SELECT product_id, name, score FROM vector_search(products, 'wireless noise cancelling headphones') ORDER BY score DESC LIMIT 10; -- BM25 full-text search SELECT product_id, name, score FROM text_search(products, 'headphones waterproof', description) ORDER BY score DESC LIMIT 10; -- Hybrid search via RRF SELECT product_id, name, fused_score FROM rrf( vector_search(products, 'wireless noise cancelling headphones'), text_search(products, 'headphones waterproof', description), join_key => 'product_id' ) ORDER BY fused_score DESC LIMIT 10; ``` See [Search Functionality](/docs/next/features/search) for the full search feature guide. ## Authentication[​](#authentication "Direct link to Authentication") The connector uses HTTP basic authentication when `elasticsearch_user` and `elasticsearch_pass` are provided. For production deployments, store credentials in a [secret store](/docs/next/components/secret-stores) and reference them with `${secrets:...}` rather than hard-coding them in `spicepod.yaml`. TLS is enabled automatically for `https://` endpoints. ## Limitations[​](#limitations "Direct link to Limitations") * Nested object fields are exposed as JSON strings rather than structured columns. * `date` and `date_nanos` fields are preserved as strings because Elasticsearch accepts heterogeneous date formats; cast to a timestamp in SQL when numeric comparison is required. * `dense_vector` fields without a declared `dims` value fall back to `Utf8` and are not usable as a vector column. * For queries with `LIMIT N` where N ≤ 10,000, the connector issues a single `_search` request. For larger result sets or queries without `LIMIT`, the connector automatically paginates using Point-In-Time (PIT) + `search_after`, fetching all matching documents in 10,000-hit batches. * SQL `WHERE` predicates are not pushed down to the Elasticsearch query DSL; all filter expressions are evaluated locally by DataFusion after fetching results (only `LIMIT` is pushed down, as the query `size`). Elasticsearch can also be configured as a [Vector Engine](/docs/next/components/vectors/elasticsearch) for datasets sourced from other connectors (storing Spice-managed embeddings in Elasticsearch rather than querying an existing index). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Elasticsearch as a data connector in Spice. [Elasticsearch Data Connector](https://github.com/spiceai/cookbook/tree/trunk/elasticsearch/connector#readme) --- # Elasticsearch Data Connector Deployment Guide Production operating guide for the Elasticsearch data connector covering authentication, TLS, resilience, capacity planning, and search routing. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The connector uses HTTP Basic authentication. Credentials must be sourced from a [secret store](/docs/next/components/secret-stores) in production. | Parameter | Description | | ------------------------ | ------------------------------------------------------------- | | `elasticsearch_endpoint` | Cluster URL. Required. Use `https://...` to enable TLS. | | `elasticsearch_user` | Username for HTTP Basic authentication. Use `${secrets:...}`. | | `elasticsearch_pass` | Password for HTTP Basic authentication. Use `${secrets:...}`. | Scope the user to the minimum required permissions: * **Read-only access** to the indexes the connector will query (`read` privilege). * **`monitor`** cluster privilege if you intend to inspect mappings programmatically. For Elastic Cloud and self-managed deployments protected by API keys, generate a dedicated user (or service account) for Spice rather than reusing administrative credentials. ### TLS[​](#tls "Direct link to TLS") Use `https://` endpoints in production. TLS is enabled automatically when the endpoint scheme is HTTPS. Self-signed certificates require a trusted CA bundle in the container or host OS trust store. The connector does not currently expose certificate-pinning or custom CA-bundle parameters — rely on the system trust store, or front the cluster with a TLS-terminating proxy you trust. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Retries[​](#retries "Direct link to Retries") The Elasticsearch client library includes a retry mechanism with exponential backoff for transient errors (HTTP 429 and 5xx). However, retries are currently only active on the **write path** used by the [Elasticsearch Vector Engine](/docs/next/components/vectors/elasticsearch) (`bulk_index` operations). The data connector's read operations (`_search`, `_mapping`) do **not** retry transient errors — failures are surfaced immediately. Retry tuning is exposed only on the [Elasticsearch Vector Engine](/docs/next/components/vectors/elasticsearch) (`elasticsearch_max_retries`, `elasticsearch_retry_initial_backoff`). ### Timeouts[​](#timeouts "Direct link to Timeouts") | Setting | Default | Behavior | | --------------- | ------- | -------------------------------------------------------------- | | Connect timeout | `10s` | Maximum time to establish a TCP/TLS connection to the cluster. | | Request timeout | `30s` | Maximum time for each individual HTTP request. | Long-running search responses (very large `LIMIT`, deep pagination, or expensive aggregations) may exceed the default request timeout. Either narrow the query, accelerate the dataset, or use the [vector engine](/docs/next/components/vectors/elasticsearch) `client_timeout` parameter when running the workload through the embedding-write path. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Throughput**: Bounded by the Elasticsearch cluster's request handling and (for kNN) HNSW search cost. Plan refresh intervals and concurrent query load to stay within the cluster's tested capacity. * **Result size**: For queries with `LIMIT N` where N ≤ 10,000, the connector issues a single `_search` request. For larger result sets or queries without `LIMIT`, the connector automatically paginates using Point-In-Time (PIT) + `search_after`, fetching all matching documents in 10,000-hit batches (bounded by the Elasticsearch `index.max_result_window` setting per batch). * **Mapping fetches**: At dataset registration the connector fetches the index mapping once via `GET //_mapping`. Mapping changes after registration are not picked up until the runtime restarts. ## Search Routing[​](#search-routing "Direct link to Search Routing") When an index has a `dense_vector` field, Spice's search UDTFs compile to native Elasticsearch queries: * `vector_search(...)` → kNN query against the `dense_vector` field. By default the candidate pool (`num_candidates`) is twice the requested `k`. * `text_search(...)` → BM25 `match` query on the specified text field. * `rrf(...)` → both queries issued in parallel and fused using Reciprocal Rank Fusion. RRF tuning (per-query `rank_weight`, recency decay, smoothing `k`) is evaluated by Spice rather than Elasticsearch. For more, see [Search Functionality](/docs/next/features/search) and the [SQL search reference](/docs/next/reference/sql/search). ## Pushdown Behavior[​](#pushdown-behavior "Direct link to Pushdown Behavior") | Predicate | Pushdown to ES Query DSL | | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | | `WHERE` filters (any field) | None — the connector pushes down no `WHERE` predicates; all are evaluated locally by DataFusion after fetch (the scan issues `match_all`). | | `LIMIT N` | Translated to `size: N`. | | `ORDER BY` | Evaluated locally unless paired with a search UDTF. | | `vector_search` / `text_search` / `rrf` | Native — issued as kNN / BM25 query bodies. | For workloads dominated by selective filters, accelerate the dataset (`acceleration.enabled: true`) into DuckDB / SQLite / Cayenne so DataFusion can apply filters at acceleration time rather than fetching unfiltered hits. ## Schema Stability[​](#schema-stability "Direct link to Schema Stability") The connector derives an Arrow schema from `GET //_mapping` at registration time. Once registered, the schema is locked for the lifetime of the runtime process — adding fields or changing types in Elasticsearch does **not** re-trigger schema inference. Restart the runtime to pick up mapping changes. For schema-evolution-friendly workloads, prefer accelerating the dataset and refreshing on a schedule against a stable subset of fields. ## Metrics[​](#metrics "Direct link to Metrics") The Elasticsearch connector does not register connector-specific instruments in the current release. Monitor via: * Spice query execution metrics (`query_duration_ms`, `query_returned_rows`, `query_failures`) from `runtime.metrics`. * Elasticsearch's own `/_nodes/stats` endpoint and Kibana dashboards for cluster-side request latency, CPU, JVM heap, and shard health. See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. ## Task History[​](#task-history "Direct link to Task History") Elasticsearch requests participate in [task history](/docs/next/reference/task_history) through the HTTP client's span. Each `_search` and `_mapping` call is a child of the enclosing `sql_query` or `accelerated_table_refresh` task. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Read-only**: The connector is read-only. Writes (indexing documents, updating mappings) are not supported. Use the [Elasticsearch Vector Engine](/docs/next/components/vectors/elasticsearch) when Spice should manage an index. * **Schema is frozen at registration**: Mapping changes after startup are not picked up. Restart the runtime to refresh the schema. * **`date` and `date_nanos` are strings**: Elasticsearch accepts heterogeneous date formats. The connector preserves them as `Utf8` — cast to `TIMESTAMP` in SQL when comparison is needed. * **`nested` and `object` are JSON strings**: Nested objects are exposed as `Utf8` JSON, not structured Arrow types. * **`dense_vector` without `dims`**: Falls back to `Utf8` and is not usable as a vector column. Declare `dims` in the index mapping. * **No filter pushdown**: The connector pushes down no SQL `WHERE` predicates — every filter is evaluated locally by DataFusion after fetch (the scan issues `match_all`; only `LIMIT` is pushed down, as the query `size`). For selective filters, accelerate the dataset. * **Tested against Elasticsearch 8.17**: Other major versions (7.x, 9.x) may work but are not part of the integration test matrix. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ------------------------------------------------------- | --------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | | `401 Unauthorized` on dataset registration | Wrong/expired credentials or insufficient privileges. | Verify `elasticsearch_user`/`elasticsearch_pass`; confirm the user has `read` on the target index. | | `Elasticsearch index 'X' not found in mapping response` | The index does not exist or the user lacks read access. | Create the index, or grant `view_index_metadata` privilege. | | `dense_vector` column missing from query results | The mapping omits `dims` for that field. | Add `dims` to the index mapping; reconfirm with `GET //_mapping`. | | `vector_search` / `text_search` returns nothing | Wrong vector field name, or the index has no documents. | Verify the field is a populated `dense_vector` / `text` field; check via `GET //_count`. | | Schema drift after deploying mapping changes | Schema is frozen at registration time. | Restart the runtime to re-infer the schema. | | Refresh exceeds `request_timeout` | Large response or slow cluster. | Narrow the query, accelerate the dataset, or front Elasticsearch with a cache. | | TLS handshake fails with self-signed certificate | The certificate's CA is not in the runtime's trust store. | Install the CA bundle in the container/host trust store; do not disable TLS verification in production. | --- # File Data Connector The File Data Connector enables federated SQL queries on files stored by locally accessible filesystems. It supports querying individual files or entire directories, where all child files within the directory will be loaded and queried. File formats are specified using the `file_format` parameter, as described in [File Formats](/docs/next/components/data-connectors/#file-formats). Example `spicepod.yml` ``` datasets: - from: file://path/to/customer.parquet name: customer params: file_format: parquet ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field for the File connector takes the form `file://path` where `path` is the path to the file to read from. See the [examples](#examples) below for examples of relative and absolute paths ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: file://path/to/customer.parquet name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter name | Description | | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `file_format` | Specifies the data file format. Required if the format cannot be inferred from the `from` path. Refer to [File Formats](/docs/next/components/data-connectors/#file-formats) for details. | | `hive_partitioning_enabled` | Enable partitioning using hive-style partitioning from the folder structure. Defaults to `false` | | `schema_source_path` | Specifies the path used to infer the dataset schema. Default to the most recently modified file | For additional CSV, JSON, and Parquet specific parameters, see [File Formats](/docs/next/reference/file_format). ## Trigger data refresh on file change[​](#trigger-data-refresh-on-file-change "Direct link to Trigger data refresh on file change") In addition to standard [Data Refresh](/docs/next/features/data-acceleration/data-refresh), a data refresh can also be triggered when the source file is modified. The File Data Connector uses a file system watcher to be notified the file has changed. The file watcher is disabled by default and can be enabled by setting the `file_watcher` parameter to `enabled` in the acceleration parameters. ``` datasets: - from: file://path/to/my_file.csv name: my_file acceleration: enabled: true refresh_mode: full params: file_watcher: enabled ``` When the file is modified, the acceleration will be refreshed and will include the latest data. ## Types[​](#types "Direct link to Types") Refer to [Object Store Data Types](/docs/next/reference/datatypes/object_store) for data type mapping from object store files to arrow data type. ## Examples[​](#examples "Direct link to Examples") ### Absolute path[​](#absolute-path "Direct link to Absolute path") In this example, `path` is an absolute path to the file on the filesystem. ``` datasets: - from: file:///path/to/customer.parquet name: customer params: file_format: parquet ``` ### Relative path[​](#relative-path "Direct link to Relative path") In this example, the path is relative to the directory where the `spicepod.yaml` is located. ``` ├── foo │   └── yellow_tripdata_2024-01.parquet └── spicepod.yaml ``` ``` datasets: - from: file://foo/yellow_tripdata_2024-01.parquet name: trip_data params: file_format: parquet ``` Performance Considerations When using the File Data connector without acceleration, data is loaded into memory during query execution. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. Memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/next/components/data-accelerators/duckdb) and [`sqlite`](/docs/next/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. ## Cookbook[​](#cookbook "Direct link to Cookbook") Refer to the [File cookbook recipe](https://github.com/spiceai/cookbook/tree/trunk/file) to see an example of the File connector in use. --- # File Data Connector Deployment Guide Production operating guide for the File data connector (reading files from the local or mounted filesystem). ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The File connector has no authentication layer. Access control is enforced by the operating system: * The Spice runtime process must have read permission on the target file or directory. * For containers, mount source files as read-only volumes. * For Kubernetes, prefer `ConfigMap` / `Secret` / `PersistentVolumeClaim` mounts over host paths. For secrets embedded in data files (credentials, tokens), encrypt at rest and restrict filesystem ACLs to the Spice process user. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") The File connector reads local files synchronously; there is no network layer, retry, or concurrency semaphore. Failures are filesystem errors (`ENOENT`, `EACCES`, `EIO`) and surface directly to the caller. Filesystem issues (e.g., an NFS mount going stale) must be handled at the infrastructure layer. For hot-reloading of updated data files, accelerate the dataset and configure a `refresh_interval` — the connector re-reads the file on each refresh. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Throughput**: Bounded by local disk bandwidth. For NVMe and locally-attached SSDs, expect single-threaded reads of hundreds of MB/s for uncompressed formats and proportionally less for compressed formats due to CPU cost. * **Memory**: File reads are streamed; memory footprint is bounded by DataFusion's 8192-row record batch size. * **Directory listings**: Glob patterns and directory paths list the full matching set at plan time. For directories with tens of thousands of files, expect multi-second planning overhead. * **Hive partitioning**: Enable `hive_partitioning_enabled: true` when reading partitioned directories to prune at plan time. ### File Formats[​](#file-formats "Direct link to File Formats") See [File Formats](/docs/next/reference/file_format) for format-specific parameters. Choose based on access pattern: * **Parquet**: Best for analytical reads. Column pruning and predicate pushdown apply. * **CSV**: Text-scan workloads only; set `has_header` and `delimiter` explicitly. * **JSON (newline-delimited)**: Good for ad-hoc reads; schema inference cost is linear in sampled records. * **Arrow IPC**: Fastest for Spice-to-Spice data exchange. ## Metrics[​](#metrics "Direct link to Metrics") The File connector does not register connector-specific instruments. Monitor via Spice's query execution metrics (`query_duration_ms`, `query_returned_rows`). See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. For filesystem-level issues (disk utilization, IOPS), use the underlying OS metrics (Prometheus `node_exporter`, CloudWatch agent, etc.). ## Task History[​](#task-history "Direct link to Task History") File reads participate in [task history](/docs/next/reference/task_history) through DataFusion's execution-plan spans. Listings, opens, and reads are attributed to the enclosing `sql_query` or `accelerated_table_refresh` task. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Read-only**: The File connector cannot write. * **No file watching**: File updates are not detected automatically; use `refresh_interval` on an accelerated dataset to pick up changes. * **Container portability**: Hard-coded `file://` paths in a spicepod are non-portable across environments; parameterize via env vars or use network-mounted paths with consistent mount points. * **Large CSVs**: CSV reads are single-threaded; prefer Parquet for datasets larger than a few GB. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ---------------------------------------------- | ----------------------------------------------------- | ---------------------------------------------------------------------------------- | | `No such file or directory` | Path typo, wrong working directory, or missing mount. | Verify the file exists from the Spice process context (`ls` inside the container). | | `Permission denied` | Spice process user lacks read permission. | Adjust file ACLs or mount with appropriate UID/GID. | | Schema inference is slow for JSON | Large file with sparse fields sampled. | Provide an explicit `schema`, or sample fewer records. | | Planning time dominates for glob patterns | Very large directory listings. | Prune with Hive partitioning or break the dataset into narrower prefixes. | | Query returns old data after file was replaced | No file watch; Spice sees cached schema. | Set `refresh_interval` on an accelerated dataset, or restart the runtime. | --- # Flight SQL Data Connector Connect to any Flight SQL compatible server (e.g. Influx 3.0, CnosDB, other Spice runtimes!) as a connector for federated SQL queries. ``` - from: flightsql:my_catalog.good_schemas.cool_dataset name: cool_dataset params: flightsql_endpoint: http://127.0.0.1:50051 flightsql_username: spicy flightsql_password: ${secrets:my_flightsql_pass} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `flightsql:dataset` where `dataset` is the fully qualified name of the dataset to read from. info Unquoted identifiers are normalized to lowercase. To reference a dataset with mixed-case characters, wrap each case-sensitive part in double quotes: `flightsql:my_catalog."MySchema"."MyTable"`. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter name | Description | | --------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `flightsql_endpoint` | Required. The Apache Flight endpoint used to connect to the Flight SQL server. | | `flightsql_username` | Optional. The username to use in the underlying Apache flight Handshake Request to authenticate to the server (see [reference](https://arrow.apache.org/docs/format/Flight.html#authentication)). | | `flightsql_password` | Optional. The password to use in the underlying Apache flight Handshake Request to authenticate to the server. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load the password from a secret store, e.g. `${secrets:my_flightsql_pass}`. | | `flightsql_tls_ca_certificate_file` | Optional. Path to a CA certificate file (PEM format) to use for TLS verification instead of system certificates. | | `flightsql_tls_client_certificate_file` | Optional. Path to a PEM client certificate chain for mutual TLS (mTLS). Must be set together with `flightsql_tls_client_key_file`. Mutually exclusive with `flightsql_tls_client_certificate`. | | `flightsql_tls_client_key_file` | Optional. Path to the PEM private key matching `flightsql_tls_client_certificate_file`. Must be set together with `flightsql_tls_client_certificate_file`. Mutually exclusive with `flightsql_tls_client_key`. | | `flightsql_tls_client_certificate` | Optional. Inline PEM client certificate chain for mutual TLS (mTLS). Use the [secret replacement syntax](/docs/next/components/secret-stores) to load from a secret store, e.g. `${secrets:my_cert}`. Must be set together with `flightsql_tls_client_key`. Mutually exclusive with `flightsql_tls_client_certificate_file`. | | `flightsql_tls_client_key` | Optional. Inline PEM private key for mutual TLS (mTLS). Use the [secret replacement syntax](/docs/next/components/secret-stores) to load from a secret store, e.g. `${secrets:my_key}`. Must be set together with `flightsql_tls_client_certificate`. Mutually exclusive with `flightsql_tls_client_key_file`. | ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). --- # FTP/SFTP Data Connector FTP (File Transfer Protocol) and SFTP (SSH File Transfer Protocol) are network protocols for transferring files between a client and server. FTP transmits data in plain text, while SFTP provides encrypted file transfer over SSH, making it the preferred choice for secure environments. The FTP/SFTP Data Connector enables federated SQL query across [supported file formats](/docs/next/components/data-connectors/#file-formats) stored on FTP/SFTP servers. ## Quickstart[​](#quickstart "Direct link to Quickstart") Connect to an SFTP server and query CSV files: ``` datasets: - from: sftp://files.example.com/data/sales/ name: sales params: file_format: csv sftp_user: ${secrets:sftp_user} sftp_pass: ${secrets:sftp_pass} ``` Query the data using SQL: ``` SELECT * FROM sales LIMIT 10; ``` ## FTP vs SFTP[​](#ftp-vs-sftp "Direct link to FTP vs SFTP") | Feature | FTP | SFTP | | --------------- | ------------------------- | ------------------------------ | | Default Port | 21 | 22 | | Encryption | None (plain text) | SSH encryption | | Authentication | Username/password | Username/password | | Recommended Use | Internal/trusted networks | Production and public networks | Security Recommendation Use SFTP instead of FTP whenever possible. FTP transmits credentials and data in plain text, making it vulnerable to interception. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Specifies the FTP or SFTP server and path to connect to. **Format:** `ftp:///` or `sftp:///` * ``: The server hostname or IP address * ``: Path to a file or directory on the server When pointing to a directory, Spice loads all files within that directory recursively. **Examples:** ``` # Connect to a specific file from: sftp://files.example.com/data/customers.parquet # Connect to a directory (loads all files) from: sftp://files.example.com/data/sales/ # FTP connection from: ftp://ftp.example.com/exports/reports/ ``` ### `name`[​](#name "Direct link to name") The dataset name used as the table name in SQL queries. Cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") #### FTP Parameters[​](#ftp-parameters "Direct link to FTP Parameters") | Parameter Name | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------- | | `file_format` | Required when connecting to a directory. See [File Formats](/docs/next/components/data-connectors/#file-formats). | | `ftp_user` | Required. Username for FTP authentication. | | `ftp_pass` | Required. Password for FTP authentication. Use [secrets](/docs/next/components/secret-stores) syntax: `${secrets:my_ftp_pass}`. | | `ftp_port` | FTP server port. Default: `21`. | | `client_timeout` | Wall-clock bound for one connection attempt — TCP connect, server greeting and login together. E.g. `30s`, `1m`. Default: `20s`. | | `hive_partitioning_enabled` | Enable [Hive-style partitioning](#hive-partitioning) from folder structure. Default: `false`. | #### SFTP Parameters[​](#sftp-parameters "Direct link to SFTP Parameters") | Parameter Name | Description | | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `file_format` | Required when connecting to a directory. See [File Formats](/docs/next/components/data-connectors/#file-formats). | | `sftp_user` | Required. Username for SFTP authentication. | | `sftp_pass` | Required. Password for SFTP authentication. Use [secrets](/docs/next/components/secret-stores) syntax: `${secrets:my_sftp_pass}`. | | `sftp_port` | SFTP server port. Default: `22`. | | `client_timeout` | Wall-clock bound for one connection attempt — name resolution, TCP connect, SSH handshake and password authentication together. E.g. `30s`, `1m`. Default: `20s`. | | `hive_partitioning_enabled` | Enable [Hive-style partitioning](#hive-partitioning) from folder structure. Default: `false`. | ## Examples[​](#examples "Direct link to Examples") ### Basic SFTP Connection[​](#basic-sftp-connection "Direct link to Basic SFTP Connection") Connect to an SFTP server with username and password authentication: ``` datasets: - from: sftp://sftp.example.com/data/transactions/ name: transactions params: file_format: parquet sftp_user: datauser sftp_pass: ${secrets:sftp_password} ``` ### Basic FTP Connection[​](#basic-ftp-connection "Direct link to Basic FTP Connection") Connect to an FTP server for internal file access: ``` datasets: - from: ftp://ftp.internal.local/exports/daily/ name: daily_exports params: file_format: csv ftp_user: ftpuser ftp_pass: ${secrets:ftp_password} ``` ### Reading a Single File[​](#reading-a-single-file "Direct link to Reading a Single File") When pointing to a specific file, the format is inferred from the file extension: ``` datasets: - from: sftp://files.example.com/reports/quarterly_summary.parquet name: quarterly_summary params: sftp_user: ${secrets:sftp_user} sftp_pass: ${secrets:sftp_pass} ``` ### Connection with Timeout[​](#connection-with-timeout "Direct link to Connection with Timeout") Configure a timeout for slow or unreliable connections: ``` datasets: - from: sftp://remote-server.example.com/large-datasets/ name: large_dataset params: file_format: parquet sftp_user: ${secrets:sftp_user} sftp_pass: ${secrets:sftp_pass} client_timeout: 120s ``` `client_timeout` bounds a connection attempt as a whole, not each stage of it, so a server that accepts the TCP connection and then stops responding — during the FTP greeting or login, or the SSH handshake or authentication — is abandoned once the bound expires rather than waiting indefinitely. When `client_timeout` is not set, a `20s` bound applies. ### Custom Port Configuration[​](#custom-port-configuration "Direct link to Custom Port Configuration") Connect to servers running on non-standard ports: ``` datasets: - from: sftp://secure.example.com/data/ name: secure_data params: file_format: parquet sftp_port: 2222 sftp_user: ${secrets:sftp_user} sftp_pass: ${secrets:sftp_pass} ``` ### Hive Partitioning[​](#hive-partitioning "Direct link to Hive Partitioning") Enable Hive-style partitioning to automatically extract partition columns from the folder structure: ``` datasets: - from: sftp://datalake.example.com/events/ name: events params: file_format: parquet sftp_user: ${secrets:sftp_user} sftp_pass: ${secrets:sftp_pass} hive_partitioning_enabled: true ``` Given a folder structure like: ``` /events/ year=2024/ month=01/ data.parquet month=02/ data.parquet year=2025/ month=01/ data.parquet ``` Queries can filter on partition columns: ``` SELECT * FROM events WHERE year = '2024' AND month = '01'; ``` ### Multiple Datasets from One Server[​](#multiple-datasets-from-one-server "Direct link to Multiple Datasets from One Server") Load different datasets from the same SFTP server: ``` datasets: - from: sftp://data.example.com/sales/ name: sales params: file_format: parquet sftp_user: ${secrets:sftp_user} sftp_pass: ${secrets:sftp_pass} - from: sftp://data.example.com/inventory/ name: inventory params: file_format: csv sftp_user: ${secrets:sftp_user} sftp_pass: ${secrets:sftp_pass} ``` ### Accelerated Dataset[​](#accelerated-dataset "Direct link to Accelerated Dataset") Enable local acceleration for faster repeated queries: ``` datasets: - from: sftp://archive.example.com/historical/ name: historical_data params: file_format: parquet sftp_user: ${secrets:sftp_user} sftp_pass: ${secrets:sftp_pass} acceleration: enabled: true refresh_check_interval: 1h ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores for secure credential management. Store FTP/SFTP credentials in a secret store and reference them using the `${secrets:key}` syntax. ``` datasets: - from: sftp://files.example.com/data/ name: secure_data params: file_format: parquet sftp_user: ${secrets:sftp_username} sftp_pass: ${secrets:sftp_password} ``` For detailed information, refer to the [secret stores documentation](/docs/next/components/secret-stores). ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") ### Connection Timeouts[​](#connection-timeouts "Direct link to Connection Timeouts") Connection attempts are bounded at `20s` by default. If a server is reachable but slow to complete the greeting, login, or SSH handshake, raise `client_timeout`: ``` params: client_timeout: 120s ``` ### Authentication Failures[​](#authentication-failures "Direct link to Authentication Failures") Verify credentials are correctly stored in your secret store and that the user has read access to the specified path on the server. ### File Format Errors[​](#file-format-errors "Direct link to File Format Errors") When connecting to a directory, ensure `file_format` is specified and matches the actual file types in the directory. Spice expects all files in a directory to have the same format. ## Cookbook[​](#cookbook "Direct link to Cookbook") Refer to the [FTP cookbook recipe](https://github.com/spiceai/cookbook/tree/trunk/ftp) for a complete working example. --- # GCS Data Connector The GCS Data Connector enables federated SQL queries on files stored in Google Cloud Storage. Both `gcs://` and `gs://` URI schemes are accepted. When a folder path is provided, all the contained files will be loaded. File formats are specified using the `file_format` parameter, as described in [File Formats](/docs/next/components/data-connectors/#file-formats). ``` datasets: - from: gs://my-bucket/taxi_sample.csv name: gcs_test params: gcs_service_account_path: /etc/spice/gcs-key.json file_format: csv ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Defines the GCS URI to a folder or object. Both schemes are supported and equivalent: * `from: gs:///` * `from: gcs:///` Example: `from: gs://my-bucket/path/to/file.parquet` ### `name`[​](#name "Direct link to name") Defines the dataset name, which is used as the table name within Spice. Example: ``` datasets: - from: gs://my-bucket/taxi_sample.csv name: cool_dataset params: file_format: csv ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") #### Basic parameters[​](#basic-parameters "Direct link to Basic parameters") | Parameter name | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `file_format` | Specifies the data format. Required if it cannot be inferred from the object URI. Options: `parquet`, `csv`, `json`. Refer to [File Formats](/docs/next/components/data-connectors/#file-formats) for details. | | `allow_http` | Allow insecure HTTP connections. Defaults to `false`. | | `client_timeout` | Optional. Timeout for GCS client operations. | | `hive_partitioning_enabled` | Enable partitioning using hive-style partitioning from the folder structure. Defaults to `false`. | | `schema_source_path` | Specifies the URL used to infer the dataset schema. Defaults to the most recently modified file. | #### Authentication parameters[​](#authentication-parameters "Direct link to Authentication parameters") The following authentication methods are mutually exclusive — only one can be set at a time. The runtime will fail to start if more than one is specified. * `gcs_service_account_path` * `gcs_service_account_key` * `gcs_application_default_credentials` * `gcs_skip_signature` If none of these are set, the connector accesses the bucket without explicit credentials. For public buckets, set `gcs_skip_signature: true` to skip request signing. | Parameter name | Description | | ------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `gcs_service_account_path` | Path to a GCS service account JSON key file. | | `gcs_service_account_key` | GCS service account JSON key as a string. | | `gcs_application_default_credentials` | Set to `true` to use Google [Application Default Credentials](https://cloud.google.com/docs/authentication/application-default-credentials). If `GOOGLE_APPLICATION_CREDENTIALS` is set, that path is used. Defaults to `false`. | | `gcs_skip_signature` | Set to `true` to skip signing requests. Use for public buckets. | #### Retry parameters[​](#retry-parameters "Direct link to Retry parameters") | Parameter name | Description | | ------------------------------ | --------------------------------------------- | | `gcs_max_retries` | Maximum number of retries. Defaults to `3`. | | `gcs_retry_timeout` | Total timeout for retries (e.g., `5s`, `1m`). | | `gcs_backoff_initial_duration` | Initial retry delay (e.g., `5s`). | | `gcs_backoff_max_duration` | Maximum retry delay (e.g., `1m`). | | `gcs_backoff_base` | Exponential backoff base (e.g., `0.1`). | ## Authentication[​](#authentication "Direct link to Authentication") GCS connector supports four mutually-exclusive authentication modes, as detailed in the [authentication parameters](#authentication-parameters). ### Service account JSON file[​](#service-account-json-file "Direct link to Service account JSON file") Configure a service account by setting `gcs_service_account_path` to the file path of a downloaded service account JSON key: ``` datasets: - from: gs://my-bucket/data/ name: my_data params: gcs_service_account_path: /etc/spice/gcs-key.json file_format: parquet ``` To create the key file, follow the [Google Cloud documentation for service account keys](https://cloud.google.com/iam/docs/keys-create-delete) and grant the service account `roles/storage.objectViewer` (or higher) on the bucket via the [Cloud Storage IAM](https://cloud.google.com/storage/docs/access-control/iam) settings. ### Service account JSON content[​](#service-account-json-content "Direct link to Service account JSON content") When mounting a key file is not practical (e.g., when keying off a secret store), pass the JSON contents directly via `gcs_service_account_key`: ``` datasets: - from: gs://my-bucket/data/ name: my_data params: gcs_service_account_key: ${secrets:GCS_SERVICE_ACCOUNT_JSON} file_format: parquet ``` The value should be the full JSON key as a single string, ideally provided through a [supported secret store](/docs/next/components/secret-stores). ### Application Default Credentials (ADC)[​](#application-default-credentials-adc "Direct link to Application Default Credentials (ADC)") To use [Application Default Credentials](https://cloud.google.com/docs/authentication/application-default-credentials) — for example, when running inside Google Cloud with attached service accounts (GKE Workload Identity, Compute Engine metadata, etc.) or when using `gcloud auth application-default login` locally — set `gcs_application_default_credentials: true`: ``` datasets: - from: gs://my-bucket/data/ name: my_data params: gcs_application_default_credentials: true file_format: parquet ``` If the `GOOGLE_APPLICATION_CREDENTIALS` environment variable is set to a service account JSON key path, that file is used. Otherwise, the ADC chain searches the well-known locations described in the Google Cloud documentation. ### Public buckets[​](#public-buckets "Direct link to Public buckets") For unauthenticated access to a public bucket, set `gcs_skip_signature: true`: ``` datasets: - from: gs://public-bucket/data/ name: public_data params: gcs_skip_signature: true file_format: parquet ``` ## Supported file formats[​](#supported-file-formats "Direct link to Supported file formats") Specify the file format using the `file_format` parameter. More details in [File Formats](/docs/next/components/data-connectors/#file-formats). ## Examples[​](#examples "Direct link to Examples") ### Reading a Parquet folder with a service account key file[​](#reading-a-parquet-folder-with-a-service-account-key-file "Direct link to Reading a Parquet folder with a service account key file") ``` datasets: - from: gs://my-bucket/trips/2024/ name: taxi_trips params: gcs_service_account_path: /etc/spice/gcs-key.json file_format: parquet ``` ### Reading a CSV file with the service account JSON inlined from a secret[​](#reading-a-csv-file-with-the-service-account-json-inlined-from-a-secret "Direct link to Reading a CSV file with the service account JSON inlined from a secret") ``` datasets: - from: gs://my-bucket/taxi_sample.csv name: taxi_sample params: gcs_service_account_key: ${secrets:GCS_SERVICE_ACCOUNT_JSON} file_format: csv ``` ### Reading from a public bucket[​](#reading-from-a-public-bucket "Direct link to Reading from a public bucket") ``` datasets: - from: gs://public-bucket/sample.parquet name: sample params: gcs_skip_signature: true file_format: parquet ``` ### Hive-partitioned dataset[​](#hive-partitioned-dataset "Direct link to Hive-partitioned dataset") ``` datasets: - from: gs://my-bucket/events/ name: events params: gcs_application_default_credentials: true file_format: parquet hive_partitioning_enabled: true ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple [secret stores](/docs/next/components/secret-stores) to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). `gcs_service_account_path` and `gcs_service_account_key` are marked as secrets and can be supplied through any supported secret store using the `${secrets:KEY}` replacement syntax. --- # GitHub Data Connector The GitHub Data Connector enables federated SQL queries on various GitHub resources such as files, issues, pull requests, and commits by specifying `github` as the selector in the `from` value for the dataset. ## Common Configuration[​](#common-configuration "Direct link to Common Configuration") ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field specifies the GitHub resource to query. The owner and repository name are extracted from the path (e.g., `github:github.com/spiceai/spiceai/issues` targets the `spiceai/spiceai` repository). It supports the following formats: | Format | Description | | ---------------------------------------------- | --------------------------------------------------------- | | `github:github.com///files/` | Query files from a repository at a specific branch or tag | | `github:github.com///issues` | Query issues from a repository | | `github:github.com///pulls` | Query pull requests from a repository | | `github:github.com///commits` | Query commits from a repository | | `github:github.com///stargazers` | Query stargazers from a repository | | `github:github.com//members` | Query members from an organization | ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") #### Personal Access Token[​](#personal-access-token "Direct link to Personal Access Token") | Parameter Name | Description | | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `github_token` | Required. GitHub personal access token to use to connect to the GitHub API. [Learn more](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens). | #### GitHub App Installation[​](#github-app-installation "Direct link to GitHub App Installation") GitHub Apps provide a secure and scalable way to integrate with GitHub's API, and works well when interacting with one or more GitHub organizations. [Learn more](https://docs.github.com/en/apps). | Parameter Name | Description | | ------------------------ | ------------------------------------------------------------------------------ | | `github_client_id` | Required. Specifies the client ID for GitHub App Installation auth mode. | | `github_private_key` | Required. Specifies the private key for GitHub App Installation auth mode. | | `github_installation_id` | Required. Specifies the installation ID for GitHub App Installation auth mode. | The client ID and private key are generated when creating the GitHub app. **Getting the Installation ID** If the app is installed on a GitHub organization: * Visit the settings page for the organization (`https://github.com/organizations//settings/installations`) * Click "Configure" on the app * The URL of the page will be of the form `https://github.com/organizations//settings/installations/` If the app is installed on a GitHub user: * Visit [the settings page](https://github.com/settings/installations) * Click "Configure" on the app * The URL of the page will be of the form `https://github.com/settings/installations/` Limitations With GitHub App Installation authentication, the connector's functionality depends on the permissions and scope of the GitHub App. Ensure that the app is installed on the repositories and configured with content, commits, issues and pull permissions to allow the corresponding datasets to work. #### Common Parameters[​](#common-parameters "Direct link to Common Parameters") | Parameter Name | Description | | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `github_query_mode` | Optional. Specifies whether the connector should use the GitHub [search API](https://docs.github.com/en/graphql/reference/queries#search) for improved filter performance. Defaults to `auto`, possible values of `auto` or `search`. | | `github_endpoint` | Optional. Base URL of the GitHub API. Defaults to `https://api.github.com`. Override to target a GitHub Enterprise Server instance (e.g., `https://github.example.com/api/v3`). | | `github_include_comments` | Optional. Pull-request connector only. Specifies the types of comments to fetch: `all`, `review`, `discussion`, or `none`. Defaults to `none`. See [Comments Example](#comments-example). | | `github_max_comments_fetched` | Optional. Pull-request connector only. Maximum number of comments to fetch per review thread (when `github_include_comments` is set to `review` or `all`) or per pull-request discussion (when set to `discussion` or `all`). Defaults to `25`, and is capped at `75` to protect against GitHub secondary rate limits. | | `github_include_commits` | Optional. Files connector only. Whether to fetch commit metadata (adds the `created_at` and `updated_at` timestamp columns) for each file. Set to `true` to enable. Defaults to `false`. | | `github_workflow_logs` | Optional. Workflow-runs connector only (`github.com///workflows//runs`). Set to `enabled` to download and include the workflow run logs for each row. Defaults to `disabled`. | ## Advanced Configuration[​](#advanced-configuration "Direct link to Advanced Configuration") ### Rate Limiting[​](#rate-limiting "Direct link to Rate Limiting") When using multiple GitHub datasets sharing the same GitHub token or GitHub app credentials, it is possible to exceed GitHub's primary and secondary rate limits. To mitigate this, use the `github_concurrent_connections_limit` setting under [`runtime.source_rate_control`](/docs/next/reference/spicepod/runtime#runtimesource_rate_control). This connections limit applies per GitHub token and per GitHub app installation, following GitHub's rate limit policy. Deprecated `runtime.params.github_max_concurrent_connections` is deprecated. Use `runtime.source_rate_control.github_concurrent_connections_limit` instead. Example Configuration: ``` # ... other configuration ... runtime: source_rate_control: github_concurrent_connections_limit: 5 # Defaults to 10 datasets: - from: github:github.com/spiceai/spiceai/files/v0.17.2-beta name: spiceai.files params: github_token: ${secrets:GITHUB_TOKEN} include: '**/*.txt' acceleration: enabled: true - from: github:github.com///issues name: spiceai.issues params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true # ... other configuration ... ``` The GitHub connector supports the following HTTP concurrency parameter: | Parameter Name | Description | | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `max_concurrent_requests` | Maximum number of concurrent HTTP requests to the same upstream origin. Overrides `runtime.params.http_max_concurrent_requests`. If both are unset, concurrency limiting is disabled. | The GitHub connector uses its own rate limiter based on GitHub API `X-RateLimit-*` response headers. Multiple datasets targeting the same GitHub endpoint share this rate limiter. ## Filter Push Down[​](#filter-push-down "Direct link to Filter Push Down") GitHub queries support a `github_query_mode` parameter, which can be set to either `auto` or `search` for the following types: * **Issues**: Defaults to `auto`. Query filters are only pushed down to the GitHub API in `search` mode. * **Pull Requests**: Defaults to `auto`. Query filters are only pushed down to the GitHub API in `search` mode. Commits only supports `auto` mode. Query with filter push down is only enabled for the `committed_date` column. `committed_date` supports exact matches, or greater/less than matches for dates provided in [ISO8601](https://www.iso.org/iso-8601-date-and-time-format.html) format, like `WHERE committed_date > '2024-09-24'`. When set to `search`, Issues and Pull Requests will use the GitHub [Search API](https://docs.github.com/en/search-github/searching-on-github/searching-issues-and-pull-requests) for improved filter performance when querying against the columns: * `author` and `state`; supports exact matches, or NOT matches. For example, `WHERE author = 'peasee'` or `WHERE author <> 'peasee'`. * `body` and `title`; supports exact matches, or LIKE matches. For example, `WHERE body LIKE '%duckdb%'`. * `updated_at`, `created_at`, `merged_at` and `closed_at`; supports exact matches, or greater/less than matches with dates provided in [ISO8601](https://www.iso.org/iso-8601-date-and-time-format.html) format. For example, `WHERE created_at > '2024-09-24'`. All other filters are supported when `github_query_mode` is set to `search`, but cannot be pushed down to the GitHub API for improved performance. Limitations * GitHub has a limitation in the Search API where it may return more stale data than the standard API used in the default query mode. * GitHub has a limitation in the Search API where it only returns a maximum of 1000 results for a query. Use [append mode acceleration](/docs/next/features/data-acceleration/data-refresh) to retrieve more results over time. See the [append example](#append-example) for pull requests. ## Examples[​](#examples "Direct link to Examples") ### Querying GitHub Files[​](#querying-github-files "Direct link to Querying GitHub Files") Limitations * `content` column is fetched only when acceleration is enabled. * Querying GitHub files does not support filter push down, which may result in long query times when acceleration is disabled. * Setting `github_query_mode` to `search` is not supported. - `ref` - Required. Specifies the GitHub branch or tag to fetch files from. - `include` - Optional. Specifies a pattern to include specific files. Supports glob patterns. If not specified, all files are included by default. ``` datasets: - from: github:github.com///files/ name: spiceai.files params: github_token: ${secrets:GITHUB_TOKEN} include: '**/*.json; **/*.yaml' acceleration: enabled: true ``` #### Schema[​](#schema "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ------------- | --------- | ----------- | | name | Utf8 | YES | | path | Utf8 | YES | | ref | Utf8 | NO | | size | Int64 | YES | | sha | Utf8 | YES | | mode | Utf8 | YES | | url | Utf8 | YES | | download\_url | Utf8 | YES | | created\_at | Timestamp | YES | | updated\_at | Timestamp | YES | | content | Utf8 | YES | `created_at` and `updated_at` are present only when `github_include_commits` is set to `true`. #### Example[​](#example "Direct link to Example") ``` datasets: - from: github:github.com/spiceai/spiceai/files/v0.17.2-beta name: spiceai.files params: github_token: ${secrets:GITHUB_TOKEN} include: '**/*.txt' # include txt files only acceleration: enabled: true ``` ``` sql> select * from spiceai.files +-------------+-------------+------+------------------------------------------+--------+-------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------+-------------+ | name | path | size | sha | mode | url | download_url | content | +-------------+-------------+------+------------------------------------------+--------+-------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------+-------------+ | version.txt | version.txt | 12 | ee80f747038c30e776eecb2c2ae155dec9a68187 | 100644 | https://api.github.com/repos/spiceai/spiceai/git/blobs/ee80f747038c30e776eecb2c2ae155dec9a68187 | https://raw.githubusercontent.com/spiceai/spiceai/v0.17.2-beta/version.txt | 0.17.2-beta | | | | | | | | | | +-------------+-------------+------+------------------------------------------+--------+-------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------+-------------+ Time: 0.005067 seconds. 1 rows. ``` ### Querying GitHub Issues[​](#querying-github-issues "Direct link to Querying GitHub Issues") Limitations * Querying with filters using date columns requires the use of [ISO8601 formatted dates](https://www.iso.org/iso-8601-date-and-time-format.html). For example, `WHERE created_at > '2024-09-24'`. ``` datasets: - from: github:github.com///issues name: spiceai.issues params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true ``` #### Schema[​](#schema-1 "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ---------------- | ------------------------- | ----------- | | assignees | List(Struct(login: Utf8)) | YES | | author | Utf8 | YES | | body | Utf8 | YES | | closed\_at | Timestamp | YES | | comments | List(Struct) | YES | | created\_at | Timestamp | YES | | id | Utf8 | YES | | labels | List(Struct(name: Utf8)) | YES | | milestone\_id | Utf8 | YES | | milestone\_title | Utf8 | YES | | comments\_count | Int64 | YES | | number | Int64 | YES | | state | Utf8 | YES | | title | Utf8 | YES | | updated\_at | Timestamp | YES | | url | Utf8 | YES | #### Example[​](#example-1 "Direct link to Example") ``` datasets: - from: github:github.com/spiceai/spiceai/issues name: spiceai.issues params: github_token: ${secrets:GITHUB_TOKEN} ``` ``` sql> select title, state, labels from spiceai.issues where title like '%duckdb%' +-----------------------------------------------------------------------------------------------------------+--------+----------------------+ | title | state | labels | +-----------------------------------------------------------------------------------------------------------+--------+----------------------+ | Limitation documentation duckdb accelerator about nested struct and decimal256 | CLOSED | [kind/documentation] | | Inconsistent duckdb connector params: `params.open` and `params.duckdb_file` | CLOSED | [kind/bug] | | federation across multiple duckdb acceleration tables. | CLOSED | [] | | Integration tests to cover "On Conflict" behaviors for duckdb accelerator | CLOSED | [kind/task] | | Permission denied issue while using duckdb data connector with spice using HELM for Kubernetes deployment | CLOSED | [kind/bug] | +-----------------------------------------------------------------------------------------------------------+--------+----------------------+ Time: 0.011877542 seconds. 5 rows. ``` ### Querying GitHub Pull Requests[​](#querying-github-pull-requests "Direct link to Querying GitHub Pull Requests") Limitations * Querying with filters using date columns requires the use of [ISO8601 formatted dates](https://www.iso.org/iso-8601-date-and-time-format.html). For example, `WHERE created_at > '2024-09-24'`. ``` datasets: - from: github:github.com///pulls name: spiceai.pulls params: github_token: ${secrets:GITHUB_TOKEN} # Specifies the types of comments to fetch: 'all', 'review', 'discussion', or 'none'. Defaults to 'none'. github_include_comments: none # Number of comments to fetch per discussion or review thread. # Defaults to 25, and is capped at 75 github_max_comments_fetched: 50 ``` #### Schema[​](#schema-2 "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ---------------- | -------------------------------------------------------------- | ----------- | | additions | Int64 | YES | | assignees | List(Struct(login: Utf8)) | YES | | author | Utf8 | YES | | body | Utf8 | YES | | changed\_files | Int64 | YES | | closed\_at | Timestamp | YES | | comments\_count | Int64 | YES | | commits\_count | Int64 | YES | | created\_at | Timestamp | YES | | deletions | Int64 | YES | | discussion | List(Struct(body: Utf8, author: Utf8, created\_at: Timestamp)) | YES | | hashes | List(Struct(id: Utf8)) | YES | | id | Utf8 | YES | | labels | List(Struct(name: Utf8)) | YES | | merged\_at | Timestamp | YES | | number | Int64 | YES | | review\_comments | List(Struct(body: Utf8, author: Utf8, created\_at: Timestamp)) | YES | | reviews\_count | Int64 | YES | | state | Utf8 | YES | | title | Utf8 | YES | | updated\_at | Timestamp | YES | | url | Utf8 | YES | **Note**: The `discussion` and `review_comments` columns are only included in the schema when the `github_include_comments` parameter is set accordingly. #### Example[​](#example-2 "Direct link to Example") ``` datasets: - from: github:github.com/spiceai/spiceai/pulls name: spiceai.pulls params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true ``` ``` sql> select title, url, state from spiceai.pulls where title like '%GitHub connector%' +---------------------------------------------------------------------+----------------------------------------------+--------+ | title | url | state | +---------------------------------------------------------------------+----------------------------------------------+--------+ | GitHub connector: convert `labels` and `hashes` to primitive arrays | https://github.com/spiceai/spiceai/pull/2452 | MERGED | +---------------------------------------------------------------------+----------------------------------------------+--------+ Time: 0.034996667 seconds. 1 rows. ``` #### Append Example[​](#append-example "Direct link to Append Example") ``` datasets: - from: github:github.com/spiceai/spiceai/pulls name: spiceai.pulls params: github_token: ${secrets:GITHUB_TOKEN} github_query_mode: search time_column: created_at acceleration: enabled: true refresh_mode: append refresh_check_interval: 6h # check for new results every 6 hours refresh_data_window: 90d # at initial load, load the last 90 days of pulls ``` #### Comments Example[​](#comments-example "Direct link to Comments Example") ``` datasets: - from: github:github.com/spiceai/spiceai/pulls name: spiceai.pulls params: github_token: ${secrets:GITHUB_TOKEN} github_include_comments: all github_max_comments_fetched: 75 acceleration: enabled: true ``` ``` sql> select unnest(unnest(review_comments)) from spiceai.pulls where number = 6 limit 1; +------------------------------------------------------------------+------------------------------------------------------------------------+--------------------------------------------------------------------+ | __unnest_placeholder(UNNEST(spiceai.pulls.review_comments)).body | __unnest_placeholder(UNNEST(spiceai.pulls.review_comments)).created_at | __unnest_placeholder(UNNEST(spiceai.pulls.review_comments)).author | +------------------------------------------------------------------+------------------------------------------------------------------------+--------------------------------------------------------------------+ | Nitpick - extra space. | 2021-08-11T17:36:23 | haardvark | +------------------------------------------------------------------+------------------------------------------------------------------------+--------------------------------------------------------------------+ Time: 0.034283334 seconds. 1 rows. ``` ``` sql> select unnest(unnest(discussion)) from spiceai.pulls where number = 148 limit 1; +-------------------------------------------------------------+-------------------------------------------------------------------+---------------------------------------------------------------+ | __unnest_placeholder(UNNEST(spiceai.pulls.discussion)).body | __unnest_placeholder(UNNEST(spiceai.pulls.discussion)).created_at | __unnest_placeholder(UNNEST(spiceai.pulls.discussion)).author | +-------------------------------------------------------------+-------------------------------------------------------------------+---------------------------------------------------------------+ | Do not merge until after repo goes public. | 2021-09-06T08:00:45 | lukekim | +-------------------------------------------------------------+-------------------------------------------------------------------+---------------------------------------------------------------+ Time: 0.036530584 seconds. 1 rows. ``` ### Querying GitHub Commits[​](#querying-github-commits "Direct link to Querying GitHub Commits") Limitations * Querying with filters using date columns requires the use of [ISO8601 formatted dates](https://www.iso.org/iso-8601-date-and-time-format.html). For example, `WHERE committed_date > '2024-09-24'`. * Setting `github_query_mode` to `search` is not supported. ``` datasets: - from: github:github.com///commits name: spiceai.commits params: github_token: ${secrets:GITHUB_TOKEN} ``` #### Schema[​](#schema-3 "Direct link to Schema") | Column Name | Data Type | Is Nullable | | --------------------------------- | --------- | ----------- | | additions | Int64 | YES | | associated\_pull\_request\_number | Int64 | YES | | author\_email | Utf8 | YES | | author\_name | Utf8 | YES | | changed\_files | Int64 | YES | | committed\_date | Timestamp | YES | | committer\_date | Timestamp | YES | | committer\_email | Utf8 | YES | | committer\_name | Utf8 | YES | | deletions | Int64 | YES | | id | Utf8 | YES | | message | Utf8 | YES | | message\_body | Utf8 | YES | | message\_head\_line | Utf8 | YES | | ref | Utf8 | YES | | sha | Utf8 | YES | | status | Utf8 | YES | #### Example[​](#example-3 "Direct link to Example") ``` datasets: - from: github:github.com/spiceai/spiceai/commits name: spiceai.commits params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true ``` ``` sql> select sha, message_head_line from spiceai.commits limit 10 +------------------------------------------+------------------------------------------------------------------------+ | sha | message_head_line | +------------------------------------------+------------------------------------------------------------------------+ | 2a9fab7905737e1af182e17f40aecc5c4b5dd236 | wait 2 seconds for the status to turn ready in refreshing status tes… | | b9c210a818abeaf14d2493fde5227781f47faed8 | Update README.md - Remove bigquery from tablet of connectors (#1434) | | d61e1af61ebf826f83703b8dd939f19e8b2ba426 | Add databricks_use_ssl parameter (#1406) | | f1ec55c5986e3e5d57eff94197182ffebbae1045 | wording and logs change reflected on readme (#1435) | | bfc74185584d1e048ef66c72ce3572a0b652bfd9 | Update acknowledgements (#1433) | | 0d870f1791d456e7924b4ecbbda5f3b762db1e32 | Update helm version and use v0.13.0-alpha (#1436) | | 12f930cbad69833077bd97ea43599a75cff985fc | Enable push-down federation by default (#1429) | | 6e4521090aaf39664bd61d245581d34398ce77db | Add functional tests for federation push-down (#1428) | | fa3279b7d9fcaa5e8baaa2425f69b556bb30e309 | Add LRU cache support for http-based sql queries (#1410) | | a3f93dde9d1312bfbf14f7ae3b75bdc468289212 | Add guides and examples about error handling (#1427) | +------------------------------------------+------------------------------------------------------------------------+ Time: 0.0065395 seconds. 10 rows. ``` ### Querying GitHub stars (Stargazers)[​](#querying-github-stars-stargazers "Direct link to Querying GitHub stars (Stargazers)") Limitations * Querying with filters using date columns requires the use of [ISO8601 formatted dates](https://www.iso.org/iso-8601-date-and-time-format.html). For example, `WHERE starred_at > '2024-09-24'`. * Setting `github_query_mode` to `search` is not supported. ``` datasets: - from: github:github.com///stargazers name: spiceai.stargazers params: github_token: ${secrets:GITHUB_TOKEN} ``` #### Schema[​](#schema-4 "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ----------- | --------- | ----------- | | starred\_at | Timestamp | YES | | login | Utf8 | YES | | email | Utf8 | YES | | name | Utf8 | YES | | company | Utf8 | YES | | x\_username | Utf8 | YES | | location | Utf8 | YES | | avatar\_url | Utf8 | YES | | bio | Utf8 | YES | #### Example[​](#example-4 "Direct link to Example") ``` datasets: - from: github:github.com/spiceai/spiceai/stargazers name: spiceai.stargazers params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true ``` ``` sql> select starred_at, login from spiceai.stargazers order by starred_at DESC limit 10 +----------------------+----------------------+ | starred_at | login | +----------------------+----------------------+ | 2024-09-15T13:22:09Z | cisen | | 2024-09-14T18:04:22Z | tyan-boot | | 2024-09-13T10:38:01Z | yofriadi | | 2024-09-13T10:01:33Z | FourSpaces | | 2024-09-13T04:02:11Z | d4x1 | | 2024-09-11T18:10:28Z | stephenakearns-insta | | 2024-09-09T22:17:42Z | Lrs121 | | 2024-09-09T19:56:26Z | jonathanfinley | | 2024-09-09T07:02:10Z | leookun | | 2024-09-09T03:04:27Z | royswale | +----------------------+----------------------+ Time: 0.0088075 seconds. 10 rows. ``` ### Querying Members of a GitHub Organization[​](#querying-members-of-a-github-organization "Direct link to Querying Members of a GitHub Organization") Limitations * Querying with filters using date columns requires the use of [ISO8601 formatted dates](https://www.iso.org/iso-8601-date-and-time-format.html). For example, `WHERE created_at > '2024-09-24'`. * Setting `github_query_mode` to `search` is not supported. ``` datasets: - from: github:github.com//members name: members params: github_token: ${secrets:GITHUB_TOKEN} ``` #### Schema[​](#schema-5 "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ----------- | --------- | ----------- | | username | Utf8 | YES | | name | Utf8 | YES | | avatar\_url | Utf8 | YES | | url | Utf8 | YES | | email | Utf8 | YES | | location | Utf8 | YES | | company | Utf8 | YES | | created\_at | Timestamp | YES | | bio | Utf8 | YES | #### Example[​](#example-5 "Direct link to Example") ``` datasets: - from: github:github.com/apache/members name: apache.members params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true ``` ``` sql> select created_at, username from apache.members order by created_at desc limit 10; +---------------------+-------------------+ | created_at | username | +---------------------+-------------------+ | 2023-10-09T13:14:13 | heliang666s | | 2023-04-14T11:26:44 | cortlepp | | 2023-02-16T08:28:58 | ChengJie1053 | | 2023-02-11T03:51:52 | FinalT | | 2022-11-20T12:12:56 | Yanshuming1 | | 2022-10-10T23:29:29 | bernardodemarco | | 2022-10-07T05:06:37 | coldgust | | 2022-09-06T14:38:44 | No-SilverBullet | | 2022-08-18T13:31:44 | harshithasudhakar | | 2022-07-05T10:44:08 | bearslyricattack | +---------------------+-------------------+ Time: 0.054390375 seconds. 10 rows. ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Github as a data connector in Spice. [GitHub Data Connector](https://github.com/spiceai/cookbook/tree/trunk/github#readme) --- # GitHub Data Connector Deployment Guide Production operating guide for the GitHub data connector covering authentication, GitHub API rate limits, and operational tuning. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The GitHub connector uses the GitHub REST and GraphQL APIs with a personal access token (PAT) or GitHub App installation token. | Parameter | Description | | -------------- | ------------------------------------------------------------------------------- | | `github_token` | PAT or installation token. Use `${secrets:...}` to resolve from a secret store. | Tokens must be sourced from a [secret store](/docs/next/components/secret-stores) in production. Scope the PAT to the minimum required permissions: * Public repo data only: no token required, but see the rate-limit note below. * Private repos: `repo` scope. * Issues/PRs: `repo` (private) or `public_repo` (public). * Org-level data: `read:org`. For long-running deployments, prefer GitHub App tokens (installation tokens) over user PATs — they have higher rate limits (15,000/hr vs 5,000/hr per authenticated user) and are not tied to a specific user account. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Rate Limiting[​](#rate-limiting "Direct link to Rate Limiting") GitHub's REST API rate limits: | Auth mode | Limit | | --------------------------- | --------------------- | | Unauthenticated | 60 requests/hr per IP | | Authenticated (PAT) | 5,000 requests/hr | | GitHub App installation | 15,000 requests/hr | | Enterprise Server (typical) | Configurable | The connector respects GitHub's `Retry-After` and `X-RateLimit-Reset` headers and backs off accordingly. When the remaining budget falls below a small threshold, requests pause until the next reset window. ### Pagination[​](#pagination "Direct link to Pagination") GitHub paginates at 100 items per page. Datasets backed by high-volume endpoints (e.g., `repos.commits` on a monorepo) may require many hours to initially hydrate. Use incremental acceleration with a `since` filter where possible. ### Retry Behavior[​](#retry-behavior "Direct link to Retry Behavior") Transient 5xx responses are retried with exponential backoff up to a bounded retry count. Permanent errors (401 Unauthorized, 404 Not Found, 422 Validation Failed) surface immediately. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Throughput**: Bounded by the rate limit, not network or CPU. Plan dataset refresh intervals to stay within the hourly budget. * **Latency**: Expect \~100-500ms per paginated request against `github.com`; lower for GitHub Enterprise Server on the same network. * **Initial bootstrap**: For high-volume datasets (e.g., all commits in a busy monorepo), the first materialization may exhaust the hourly budget across several runs. Plan staged ingestion if needed. ## Metrics[​](#metrics "Direct link to Metrics") The GitHub connector does not register connector-specific dataset-level instruments in the current release. Monitor via: * Spice query execution metrics (`query_duration_ms`, `query_returned_rows`, `query_failures`) from `runtime.metrics`. * GitHub's own rate-limit UI at `/settings/tokens` for token-level quota tracking. See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. ## Task History[​](#task-history "Direct link to Task History") GitHub API calls participate in [task history](/docs/next/reference/task_history) through the HTTP client's span. Each page fetch is a child of the enclosing `sql_query` or `accelerated_table_refresh` task. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Read-only**: The connector is read-only; writes (issue creation, PR comments) are not supported. * **GraphQL-only endpoints**: Some GitHub data (e.g., discussions, project v2) requires GraphQL; check the connector's documented supported endpoints. * **GitHub Enterprise Cloud with IP allowlisting**: The Spice runtime's outbound IP must be allow-listed. * **Secondary rate limits**: GitHub enforces abuse-detection "secondary" rate limits on concentrated bursts, independent of the hourly primary limit. If hit, the connector backs off. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | --------------------------------- | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | | `401 Bad credentials` | PAT expired / revoked / wrong value. | Rotate the PAT; update the secret store. | | `403 rate limit exceeded` | Primary hourly rate limit hit. | Increase refresh interval; switch to GitHub App auth for higher quota; use incremental refresh with `since`. | | `403 Secondary rate limit` | Burst of concurrent requests tripped abuse detection. | Reduce concurrent refresh; connector will back off automatically. | | `404 Not Found` on a private repo | Token lacks `repo` scope. | Regenerate PAT with `repo` scope. | | Very slow initial hydration | Large dataset + strict rate limit. | Run first refresh off-peak; use `since`/`updated_since` for incremental refreshes. | --- # Glue Data Connector The Glue Data Connector enables federated SQL querying on tables in an AWS Glue Data Catalog. ``` datasets: - from: glue:tpch.lineitem name: lineitem params: glue_region: us-east-1 glue_key: ${env:SPICE_AWS_KEY} # Optional. glue_secret: ${env:SPICE_AWS_SECRET} # Optional. ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Specify a table using the format, `glue:.` by replacing `` with the name of the Glue database and `
`with the name of the table inside of the ``. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` SELECT COUNT(*) FROM lineitem; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The following parameters are supported for configuring the connection to the Glue Data Catalog: | Parameter Name | Definition | | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `glue_region` | The AWS region for the Glue Data Catalog. E.g. `us-west-2`. | | `glue_catalog_id` | The Glue catalog ID. For Amazon S3 Tables, use the format `:s3tablescatalog/`. If not provided, the default catalog for the account is used. | | `glue_key` | Access key (e.g. AWS\_ACCESS\_KEY\_ID for AWS). If not provided, credentials will be loaded from environment variables or IAM roles. | | `glue_secret` | Secret key (e.g. AWS\_SECRET\_ACCESS\_KEY for AWS). If not provided, credentials will be loaded from environment variables or IAM roles. | | `glue_session_token` | Session token (e.g. AWS\_SESSION\_TOKEN for AWS) for temporary credentials | | `glue_iam_role_source` | Optional. IAM role credential source. `auto` (default) uses the default AWS credential chain, `metadata` uses only instance/container metadata (IMDS, ECS, EKS/IRSA), `env` uses only environment variables. | The following parameters control how the embedded S3 reader fetches Parquet/CSV data files referenced by Glue table metadata. They are inherited from the [S3 data connector](/docs/next/components/data-connectors/s3) and do not apply to Iceberg-format tables, whose object I/O is handled by the Iceberg client. | Parameter Name | Definition | | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `glue_endpoint` | Optional. Custom S3-compatible endpoint URL used when reading Parquet/CSV data files (e.g. `https://s3.us-east-1.amazonaws.com`, `http://minio.local:9000`). Leave unset for AWS S3. | | `glue_url_style` | Optional. S3 URL addressing style for Parquet/CSV data files. One of `vhost` or `path`. Auto-detected from the endpoint when unset. | | `glue_versioning` | Optional. Enables S3 object versioning support for Parquet/CSV data files when set to `enabled`. Defaults to `enabled`. | | `client_timeout` | Optional. Timeout for the underlying S3 client used to fetch Parquet/CSV data files. E.g. `30s`. | | `allow_http` | Optional. Set to `true` to allow insecure HTTP for the S3 endpoint used to read Parquet/CSV data files. Defaults to `false`. Required when `glue_endpoint` uses an `http://` scheme. | ## Examples[​](#examples "Direct link to Examples") ### Basic Glue Table[​](#basic-glue-table "Direct link to Basic Glue Table") ``` datasets: - from: glue:tpch.lineitem name: lineitem params: glue_region: us-east-1 glue_key: ${env:AWS_ACCESS_KEY_ID} glue_secret: ${env:AWS_SECRET_ACCESS_KEY} ``` ### Amazon S3 Tables[​](#amazon-s3-tables "Direct link to Amazon S3 Tables") Connect to tables in [Amazon S3 Tables](https://aws.amazon.com/s3/features/tables/) using the `glue_catalog_id` parameter with the S3 Tables catalog format: ``` datasets: - from: glue:my_namespace.orders name: orders params: glue_catalog_id: 123635965758:s3tablescatalog/my-table-bucket glue_region: us-east-2 glue_key: ${env:AWS_ACCESS_KEY_ID} glue_secret: ${env:AWS_SECRET_ACCESS_KEY} ``` ## Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. These credentials will be used to connect to the S3 bucket as well as the Glue catalog. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3/Glue permissions (e.g., `s3:ListBucket`, `glue:GetTable`) to access the tables. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ### Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access Iceberg tables in S3/Glue: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::company-bucketname-datasets" }, { "Effect": "Allow", "Action": ["s3:GetObject", "s3:PutObject"], "Resource": "arn:aws:s3:::company-bucketname-datasets/*" }, { "Effect": "Allow", "Action": [ "glue:GetCatalog", "glue:GetDatabases", "glue:GetDatabase", "glue:GetTable", "glue:GetTables" ], "Resource": "*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------- | -------------------------------------------------------------- | | `s3:ListBucket` | Required. Allows scanning all objects from the bucket | | `s3:GetObject` | Required. Allows fetching objects | | `s3:PutObject` | Required for write operations. Allows writing objects | | `glue:GetCatalog` | Required. Retrieve metadata about the specified catalog. | | `glue:GetDatabases` | Required. List the databases available in the current catalog. | | `glue:GetDatabase` | Required. Retrieve metadata about the specified database. | | `glue:GetTable` | Required. Retrieve metadata about the specified table. | | `glue:GetTables` | Required. List the tables available in the current database. | ## Write Support[​](#write-support "Direct link to Write Support") This connector supports writing data to Glue-managed Iceberg tables using SQL [`INSERT INTO`](/docs/next/reference/sql/dml#insert) statements. Writes are currently append-only — inserted data is added as new data files and registered through a new Iceberg table snapshot. Schema validation ensures inserted data matches the target table schema. To enable writes, set `access: read_write` on the dataset: ``` datasets: - from: glue:tpch.lineitem name: lineitem access: read_write params: glue_region: us-east-1 ``` ``` -- Insert with values INSERT INTO lineitem (l_orderkey, l_partkey, l_quantity) VALUES (1, 100, 5.0); -- Insert from another table INSERT INTO lineitem SELECT * FROM staging_lineitem; ``` Inserting into partitioned Iceberg tables is supported. `UPDATE` and `DELETE` operations are not currently supported. Write operations require `s3:PutObject` permission on the target S3 bucket in addition to the read permissions listed above. For more details, see [Data Ingestion](/docs/next/features/data-ingestion). ## Limitations[​](#limitations "Direct link to Limitations") Data Source/Data Format Restrictions This catalog connector is limited to tables that use the S3 data source. Kinesis and Kafka data sources are not currently supported. Additionally, this catalog connector is currently limited to Iceberg tables, tables with parquet or CSV data format only. Performance Considerations When using the Glue Data connector without acceleration, data is loaded into memory during query execution. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. Memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/next/components/data-accelerators/duckdb) and [`sqlite`](/docs/next/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. Each query retrieves data from the S3 source, which might result in significant network requests and bandwidth consumption. This can affect network performance and incur costs related to data transfer from S3. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Glue as a data connector in Spice. [Glue Data Connector](https://github.com/spiceai/cookbook/tree/trunk/glue#readme) --- # GraphQL Data Connector The [GraphQL](https://graphql.org/) Data Connector enables federated SQL queries on any GraphQL endpoint by specifying `graphql` as the selector in the `from` value for the dataset. ``` datasets: - from: graphql:your-graphql-endpoint name: my_dataset params: json_pointer: /data/some/nodes graphql_query: | { some { nodes { field1 field2 } } } ``` Limitations * The GraphQL data connector does not support variables in the query. * Filter pushdown, with the exclusion of `LIMIT`, is not currently supported. Using a `LIMIT` will reduce the amount of data requested from the GraphQL server. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form of `graphql:your-graphql-endpoint`. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The GraphQL data connector can be configured by providing the following `params`. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load the password from a secret store, e.g. `${secrets:my_graphql_auth_token}`. | Parameter Name | Description | Required | Default | | --------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ------- | | `graphql_query` | The GraphQL query to execute. See [examples](#examples) for a sample GraphQL query. | Yes | - | | `json_pointer` | The [JSON pointer](https://datatracker.ietf.org/doc/html/rfc6901) into the response body. When `graphql_query` is [paginated](#pagination), the `json_pointer` can be inferred. | No | - | | `unnest_depth` | Depth level to automatically unnest objects to. Disabled if unspecified or `0`. Maximum value is `50`. | No | `0` | | `graphql_auth_header` | A custom header name to use for authentication instead of the default `Authorization: Bearer` header. When set, the value of `graphql_auth_token` is sent as the value of this header. Useful for APIs that require authentication via a custom header (e.g. `X-Shopify-Access-Token`). | No | - | | `graphql_auth_token` | The authentication token to use to connect to the GraphQL server. Uses bearer authentication by default, or sent via the custom header specified by `graphql_auth_header`. | No | - | | `graphql_auth_user` | The username to use for basic auth. E.g. `graphql_auth_user: my_user` | No | - | | `graphql_auth_pass` | The password to use for basic auth. E.g. `graphql_auth_pass: ${secrets:my_graphql_auth_pass}` | No | - | #### GraphQL Query Example[​](#graphql-query-example "Direct link to GraphQL Query Example") ``` graphql_query: | { some { nodes { field1 field2 } } } ``` ### Examples[​](#examples "Direct link to Examples") Example using the GitHub GraphQL API and Bearer Auth. The following will use `json_pointer` to retrieve all of the nodes in starredRepositories: ``` from: graphql:https://api.github.com/graphql name: stars params: graphql_auth_token: ${env:GITHUB_TOKEN} graphql_auth_user: ${env:GRAPHQL_USER} ... graphql_auth_pass: ${env:GRAPHQL_PASS} json_pointer: /data/viewer/starredRepositories/nodes graphql_query: | { viewer { starredRepositories { nodes { name stargazerCount languages (first: 10) { nodes { name } } } } } } ``` ### Custom Auth Header Example[​](#custom-auth-header-example "Direct link to Custom Auth Header Example") Some APIs require authentication via a custom header instead of the standard `Authorization: Bearer` header. Use the `graphql_auth_header` parameter to specify a custom header name: ``` datasets: - from: graphql:https://mystore.myshopify.com/admin/api/2024-01/graphql.json name: shopify_products params: graphql_auth_header: "X-Shopify-Access-Token" graphql_auth_token: ${secrets:SHOPIFY_TOKEN} graphql_query: | { products(first: 10) { edges { node { id title } } } } json_pointer: /data/products/edges ``` | `graphql_auth_header` | `graphql_auth_token` | `graphql_auth_user`/`graphql_auth_pass` | Result | | --------------------- | -------------------- | --------------------------------------- | ---------------------------------------- | | set | set | - | Custom header: `X-Custom: ` | | not set | set | - | Default: `Authorization: Bearer ` | | not set | not set | set | HTTP Basic Auth | | not set | not set | not set | No auth | ### Rate Control Parameters[​](#rate-control-parameters "Direct link to Rate Control Parameters") The GraphQL connector supports shared HTTP rate control to limit concurrency and request rate per upstream origin. These parameters can be set per-dataset (in `params`) or globally (in `runtime.params`). Dataset-level settings override the global defaults. | Parameter Name | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `max_concurrent_requests` | Maximum number of concurrent HTTP requests to the same upstream origin. Overrides `runtime.params.http_max_concurrent_requests`. If both are unset, concurrency limiting is disabled. | | `requests_per_second_limit` | Maximum number of HTTP requests per second to the same upstream origin. Overrides `runtime.params.http_requests_per_second_limit`. If both are unset, no per-second rate limit is applied. | | `requests_per_minute_limit` | Maximum number of HTTP requests per minute to the same upstream origin. Overrides `runtime.params.http_requests_per_minute_limit`. If both are unset, no per-minute rate limit is applied. | | `rate_control_jitter_min` | Minimum random delay added before HTTP requests when rate control is active. Accepts durations such as `5ms` or `0ms`. Defaults to `5ms` when a request-rate limit is configured. | | `rate_control_jitter_max` | Maximum random delay added before HTTP requests when rate control is active. Accepts durations such as `10ms` or `0ms`. Defaults to `10ms` when a request-rate limit is configured. | Multiple datasets targeting the same GraphQL endpoint share the same rate controller. ## Pagination[​](#pagination "Direct link to Pagination") The GraphQL Data Connector supports automatic pagination of the response for queries using [cursor pagination](https://graphql.org/learn/pagination/). The `graphql_query` must include the `pageInfo` field as per [spec](https://relay.dev/graphql/connections.htm#sec-undefined.PageInfo). The connector will parse the `graphql_query`, and when `pageInfo` is present, will retrieve data until pagination completes. The query must have the correct pagination arguments in the associated paginated field. ### Example[​](#example "Direct link to Example") **Forward Pagination:** ``` { something_paginated(first: 100) { nodes { foo bar } pageInfo { endCursor hasNextPage } } } ``` **Backward Pagination:** ``` { something_paginated(last: 100) { nodes { foo bar } pageInfo { startCursor hasPreviousPage } } } ``` ## Working with JSON Data[​](#working-with-json-data "Direct link to Working with JSON Data") Tips for working with JSON data. For more information see [Datafusion Docs](https://datafusion.apache.org/user-guide/sql/scalar_functions.html#array-functions). ### Accessing objects fields[​](#accessing-objects-fields "Direct link to Accessing objects fields") You can access the fields of the object using the square bracket notation. Arrays are indexed from 1. Example for the stargazers query from [pagination section](#pagination): ``` sql> select node['login'] as login, node['name'] as name from stargazers limit 5; +--------------+----------------------+ | login | name | +--------------+----------------------+ | simsieg | Simon Siegert | | davidmathers | David Mathers | | ahmedtadde | Ahmed Tadde | | lordhamlet | Shih-Fen Cheng | | thinmy | Thinmy Patrick Alves | +--------------+----------------------+ ``` ### Piping array into rows[​](#piping-array-into-rows "Direct link to Piping array into rows") You can use Datafusion `unnest` function to pipe values from array into rows. We'll be using [countries GraphQL api](https://countries.trevorblades.com) as an example. ``` from: graphql:https://countries.trevorblades.com name: countries params: json_pointer: /data/continents graphql_query: | { continents { name countries { name capital } } } description: countries acceleration: enabled: true refresh_mode: full refresh_check_interval: 30m ``` Example query: ``` sql> select continent, country['name'] as country, country['capital'] as capital from (select name as continent, unnest(countries) as country from countries) where continent = 'North America' limit 5; +---------------+---------------------+--------------+ | continent | country | capital | +---------------+---------------------+--------------+ | North America | Antigua and Barbuda | Saint John's | | North America | Anguilla | The Valley | | North America | Aruba | Oranjestad | | North America | Barbados | Bridgetown | | North America | Saint Barthélemy | Gustavia | +---------------+---------------------+--------------+ ``` ### Unnesting object properties[​](#unnesting-object-properties "Direct link to Unnesting object properties") You can also use the `unnest_depth` parameter to control automatic unnesting of objects from GraphQL responses. This examples uses the GitHub stargazers endpoint: ``` from: graphql:https://api.github.com/graphql name: stargazers params: graphql_auth_token: ${env:GITHUB_TOKEN} unnest_depth: 2 json_pointer: /data/repository/stargazers/edges graphql_query: | { repository(name: "spiceai", owner: "spiceai") { id name stargazers(first: 100) { edges { node { id name login } } pageInfo { hasNextPage endCursor } } } } ``` If `unnest_depth` is set to 0, or unspecified, object unnesting is disabled. When enabled, unnesting automatically moves nested fields to the parent level. Without unnesting, stargazers data looks like this in a query: ``` sql> select node from stargazers limit 1; +------------------------------------------------------------+ | node | +------------------------------------------------------------+ | {id: MDQ6VXNlcjcwNzIw, login: ashtom, name: Thomas Dohmke} | +------------------------------------------------------------+ ``` With unnesting, these properties are automatically placed into their own columns: ``` sql> select node from stargazers limit 1; +------------------+--------+---------------+ | id | login | name | +------------------+--------+---------------+ | MDQ6VXNlcjcwNzIw | ashtom | Thomas Dohmke | +------------------+--------+---------------+ ``` #### Unnesting Duplicate Columns[​](#unnesting-duplicate-columns "Direct link to Unnesting Duplicate Columns") By default, the Spice Runtime will error when a duplicate column is detected during unnesting. For example, this example `spicepod.yml` query would fail due to `name` fields: ``` from: graphql:https://localhost name: stargazers params: unnest_depth: 2 json_pointer: /data/users graphql_query: | query { users { name emergency_contact { name } } } ``` This example would fail with a runtime error: ``` WARN runtime: Invalid GraphQL object access: Column 'name' already exists in the object. ``` Avoid this error by [using aliases in the query](https://www.apollographql.com/docs/kotlin/advanced/using-aliases/) where possible. In the example above, a duplicate error was introduced from `emergency_contact { name }`. The example below uses a GraphQL alias to rename `emergency_contact.name` as `emergencyContactName`. ``` from: graphql:https://localhost name: stargazers params: unnest_depth: 2 json_pointer: /data/people graphql_query: | query { users { name emergency_contact { emergencyContactName: name } } } ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure GraphQL as a data connector in Spice. [GraphQL Data Connector](https://github.com/spiceai/cookbook/tree/trunk/graphql#readme) --- # GraphQL Data Connector Deployment Guide Production operating guide for the GraphQL data connector covering authentication, pagination, and operational tuning. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") Authentication is endpoint-specific. The connector supports bearer tokens, custom headers via `graphql_auth_header`, and HTTP Basic Auth: | Parameter | Description | | --------------------- | ------------------------------------------------------------------------------------------------------- | | `graphql_auth_header` | Custom authorization header name. The value of `graphql_auth_token` is sent as this header's value. | | `graphql_auth_token` | Bearer token for GraphQL requests. Typically `"${secrets:api_token}"`. | | `graphql_auth_user` | Username for HTTP Basic Auth. | | `graphql_auth_pass` | Password for HTTP Basic Auth. | | `graphql_query` | The GraphQL query to execute. | | `json_pointer` | RFC-6901 JSON pointer to the row collection inside the response (e.g. `/data/repository/issues/nodes`). | Tokens must be sourced from a [secret store](/docs/next/components/secret-stores) in production. ### TLS[​](#tls "Direct link to TLS") Use HTTPS endpoints in production. Self-signed certificates require a trusted CA bundle in the container / host OS trust store. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Retry Behavior[​](#retry-behavior "Direct link to Retry Behavior") HTTP-level retries cover 408 (request timeout) and 5xx (server errors) plus transient network errors. 429 responses are handled proactively by the built-in rate limiter rather than retried. Retries use fibonacci backoff with a maximum of 5 attempts. ### Pagination[​](#pagination "Direct link to Pagination") The connector supports cursor-based pagination. Each page is a separate HTTP request; pagination errors mid-sequence cause the entire refresh to fail. Use `json_pointer` to select the row collection and configure the pagination variables to match the upstream schema's cursor fields. ### Server Rate Limits[​](#server-rate-limits "Direct link to Server Rate Limits") GraphQL APIs (GitHub, Shopify, etc.) typically enforce query-cost-based rate limits rather than request count. When a query returns a cost/rate-limit error, the connector surfaces it immediately. Reduce refresh frequency or narrow the query to stay within budget. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Throughput**: Bounded by the upstream rate limit, typical GraphQL endpoints cap at 100s-1000s of requests per minute. * **Query cost**: Design `graphql_query` to request only the fields you need. Request fewer nested fields to reduce query cost. * **Pagination depth**: Large datasets requiring hundreds of pages extend refresh duration linearly; plan refresh intervals accordingly. ## Metrics[​](#metrics "Direct link to Metrics") When used as a dataset connector, GraphQL exposes per-origin HTTP rate-control metrics under the `graphql` component. They are registered automatically for every GraphQL dataset — no `metrics` configuration is required — and the limit gauges report `0` when the corresponding limit is not configured. Catalog components expose none: | Metric Name | Type | Description | | ----------------------------------------- | ------- | ------------------------------------------------------------------------------------------------------ | | `inflight_operations` | Gauge | Current number of HTTP requests holding a rate-control permit. | | `rate_control_max_concurrent_requests` | Gauge | Configured maximum concurrent HTTP requests for this upstream origin; `0` means disabled. | | `rate_control_requests_per_second_limit` | Gauge | Configured HTTP request-per-second limit for this upstream origin; `0` means disabled. | | `rate_control_requests_per_minute_limit` | Gauge | Configured HTTP request-per-minute limit for this upstream origin; `0` means disabled. | | `rate_control_jitter_min_ms` | Gauge | Configured minimum rate-control jitter (ms) before HTTP requests. | | `rate_control_jitter_max_ms` | Gauge | Configured maximum rate-control jitter (ms) before HTTP requests. | | `rate_control_available_permits` | Gauge | Current available permits in the HTTP request concurrency semaphore; `0` when concurrency is disabled. | | `rate_control_acquisitions_total` | Counter | Total HTTP request rate-control permits acquired. | | `rate_control_acquire_errors_total` | Counter | Total HTTP request rate-control permit acquisition errors. | | `rate_control_wait_duration_ms` | Counter | Cumulative time (ms) spent waiting for HTTP rate-control permits, quotas, and jitter. | | `rate_limit_retry_after_updates_total` | Counter | Total upstream cooldown hints accepted from `Retry-After` or `RateLimit` reset headers. | | `rate_limit_retry_after_waits_total` | Counter | Total waits caused by `Retry-After` or `RateLimit` reset headers. | | `rate_limit_retry_after_wait_duration_ms` | Counter | Cumulative time (ms) spent waiting because of `Retry-After` or `RateLimit` reset headers. | | `rate_limit_retry_after_remaining_ms` | Gauge | Current remaining `Retry-After` / `RateLimit` cooldown (ms) for this upstream origin. | These metrics are auto-registered — no configuration is required to export them. To turn one off for a dataset, set `enabled: false` in the dataset's `metrics` section: ``` datasets: - from: graphql:https://api.example.com/graphql name: api_data metrics: - name: rate_control_wait_duration_ms enabled: false ``` Instruments are exposed with the prefix `dataset_graphql_`, and each carries an `origin` attribute (`scheme://host:port`) identifying the upstream origin instead of a dataset `name`, because datasets sharing an origin share one rate controller. See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. For broader observability, also monitor: * Spice query execution metrics (`query_duration_ms`, `query_returned_rows`, `query_failures`) from `runtime.metrics`. * The upstream GraphQL provider's rate-limit dashboards. ## Task History[​](#task-history "Direct link to Task History") GraphQL requests participate in [task history](/docs/next/reference/task_history) through the HTTP client's span. Each page fetch is a child of the enclosing `sql_query` or `accelerated_table_refresh` task. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Read-only**: Only GraphQL queries (not mutations or subscriptions) are supported. * **Single query per dataset**: Each dataset is one GraphQL query. Multi-query datasets require separate dataset definitions. * **Schema inference**: The connector infers schema from the first response; schemas with deeply-nested optional fields may require an explicit dataset `schema` override. * **Batching**: GraphQL query batching (multiple operations in one HTTP request) is not exposed. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ----------------------------------------- | ----------------------------------------------- | ----------------------------------------------------------------------------------------- | | `401 Unauthorized` | Wrong or expired token in `graphql_auth_token`. | Rotate the token; verify the header format (`Bearer` prefix, etc.). | | Rows missing from the dataset | Wrong `json_pointer`. | Inspect the response payload; JSON pointer must navigate to the array of rows. | | Refresh fails mid-pagination | Rate-limit or transient network failure. | Reduce refresh frequency; the connector will retry on retriable errors. Narrow the query. | | Query cost exceeded | Query requests too many nested fields. | Simplify the query; fetch only required fields. | | Inferred schema differs between refreshes | Optional fields appear/disappear in responses. | Provide an explicit dataset `schema` to lock down types. | --- # HTTP(s) Data Connector The HTTP(s) Data Connector enables federated SQL query across [supported file formats](/docs/next/components/data-connectors/#file-formats) stored at an HTTP(s) endpoint. The connector supports dynamic query and data refresh through SQL-based filtering. ``` datasets: - from: http://static_username@localhost:3001/report.csv name: local_report params: http_password: ${env:MY_HTTP_PASS} ``` ## Examples[​](#examples "Direct link to Examples") ### Basic Example[​](#basic-example "Direct link to Basic Example") ``` datasets: - from: https://github.com/LAION-AI/audio-dataset/raw/7fd6ae3cfd7cde619f6bed817da7aa2202a5bc28/metadata/freesound/parquet/freesound_parquet.parquet name: laion_freesound ``` ### Using Basic Authentication[​](#using-basic-authentication "Direct link to Using Basic Authentication") ``` datasets: - from: http://static_username@localhost:3001/report.csv name: local_report params: http_password: ${env:MY_HTTP_PASS} ``` The username is taken from the `user-info` section of the `from` URL (`user@host`) or from the `http_username` parameter. The password comes from the `http_password` parameter. The connector then sends a standard [RFC 7617](https://datatracker.ietf.org/doc/html/rfc7617) Basic authentication header on every request: ``` Authorization: Basic ``` For example, `static_username` with password `s3cret` produces `Authorization: Basic c3RhdGljX3VzZXJuYW1lOnMzY3JldA==`. Only one of `http_password` or user info in the URL can provide the password — setting both is not supported. ### Using Custom Headers[​](#using-custom-headers "Direct link to Using Custom Headers") Custom HTTP headers can be specified for authentication, API keys, or other requirements. Headers are treated as sensitive data and will not be logged. `http_headers` applies to **dynamic JSON API endpoints only**. A structured HTTP file dataset — `csv`, `tsv`, `parquet`, `arrow`, `avro`, `jsonl`/`ndjson`/`ldjson`, `soda`, `socrata`, `vortex`, or a static `json` file — is served by the object-store listing path, which cannot carry request headers, so the headers are ignored and a warning is logged naming the dataset. To authenticate a structured file download, use [Basic authentication](#using-basic-authentication) (`http_username` / `http_password`, or user info in the URL). ``` datasets: - from: https://api.example.com name: api_data params: http_headers: 'Authorization:Bearer ${secrets:api_token},Accept:application/json' ``` Headers can also be separated by semicolons: ``` datasets: - from: https://api.example.com name: api_data params: http_headers: 'Authorization: Bearer ${secrets:api_token}; X-API-Key: ${secrets:api_key}' ``` ### Using OAuth2 Refresh-Token Authentication[​](#using-oauth2-refresh-token-authentication "Direct link to Using OAuth2 Refresh-Token Authentication") For JSON APIs protected by OAuth2, the connector can acquire short-lived access tokens from a token endpoint and keep them fresh automatically. The **refresh-token grant** (RFC 6749 §6) exchanges a long-lived refresh token; the **client-credentials grant** (RFC 6749 §4.4) authenticates with a `client_id`/`client_secret` for machine-to-machine APIs. On startup Spice hits the configured token endpoint once, then stamps `Authorization: Bearer ` on every data request (or a custom header — see [Custom Token Header](#custom-token-header)) and refreshes the token in the background before it expires. ``` datasets: - from: https://api.example.com name: secure_data params: file_format: json allowed_request_paths: '/v1/**' auth_token_url: https://auth.example.com/oauth/token http_auth_refresh_token: ${secrets:my_refresh_token} http_auth_client_id: ${secrets:my_client_id} http_auth_client_secret: ${secrets:my_client_secret} auth_scopes: 'read:data offline_access' ``` The `http_auth_refresh_token`, `http_auth_client_id`, and `http_auth_client_secret` parameters can be loaded from any [supported secret store](/docs/next/components/secret-stores) (environment variables, Kubernetes Secrets, AWS Secrets Manager, HashiCorp Vault, the OS keychain, etc.) using the `${secrets:...}` [replacement syntax](/docs/next/components/secret-stores#using-secrets). Applies to **dynamic JSON API endpoints only** (e.g. `file_format: json` with `allowed_request_paths`). A structured HTTP file dataset (csv/parquet/etc.) goes through the object-store listing path, which cannot attach an access token, so setting any OAuth2 parameter on one is **rejected at registration** — the dataset fails to load with a configuration error naming the parameters to remove, rather than quietly sending unauthenticated requests. `http_headers` is not applied on that path either; use [Basic authentication](#using-basic-authentication) to authenticate a structured file download. See [OAuth2 Authentication](#oauth2-authentication) for the full parameter reference and behavior notes. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field specifies the HTTP(s) endpoint and can be configured in two ways: 1. **Direct URL to a file**: A complete URL pointing to a specific [supported file](/docs/next/components/data-connectors/#file-formats). ``` from: https://example.com/data/report.csv ``` 2. **Base domain/path**: A base URL that will be combined with special metadata fields to construct the complete request. ``` from: https://api.example.com/v1 ``` The connector supports templated URLs with query parameters that can be dynamically populated using `refresh_sql` filters and special metadata fields. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: http://static_username@localhost:3001/report.csv name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The connector supports authentication, timeout, connection pooling, and retry configuration via `params`. | Parameter Name | Description | | ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `http_port` | Optional. Port to create HTTP(s) connection over. Default: 80 and 443 for HTTP and HTTPS respectively. | | `http_username` | Optional. Username for HTTP basic authentication. Default: None. | | `http_password` | Optional. Password for HTTP basic authentication. Default: None. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load the password from a secret store, e.g. `${secrets:my_http_pass}`. | | `http_headers` | Optional. Custom HTTP headers as a comma-separated list of `key:value` pairs. Example: `Content-Type:application/json,Accept:application/json`. Applies to dynamic JSON API endpoints only; structured HTTP file datasets ignore these headers. Default: None. | | `allowed_request_paths` | **Required** for using `request_path` filters. Comma-separated list of allowed paths. Example: `/api/users,/api/posts`. Paths must start with `/` and cannot contain `..` segments. | | `request_query_filters` | Optional. Set to `enabled` to enable `request_query` filters. Default: `disabled`. When disabled, query parameter filters will be rejected. | | `request_body_filters` | Optional. Set to `enabled` to enable `request_body` filters for POST requests. Default: `disabled`. When disabled, request body filters will be rejected. | | `client_timeout` | Optional. Maximum time to wait for a response from the HTTP server (in seconds). Default: `30`. Applied to the entire request-response cycle. | | `connect_timeout` | Optional. Timeout for establishing HTTP(s) connections (in seconds). Default: `10`. | | `pool_max_idle_per_host` | Optional. Maximum number of idle connections to keep alive per host. Default: `10`. | | `pool_idle_timeout` | Optional. Timeout for idle connections in the pool (in seconds). Default: `90`. | | `max_retries` | Optional. Maximum number of retries for failed HTTP requests. Default: `3`. | | `retry_backoff_method` | Optional. Retry backoff strategy: `fibonacci` (default), `linear`, or `exponential`. | | `retry_max_duration` | Optional. Maximum total duration for all retries (e.g., `30s`, `5m`). If not set, retries continue up to `max_retries`. | | `retry_jitter` | Optional. Randomization factor for retry delays (0.0 to 1.0). Default: `0.3` (30% randomization). Set to `0` for no jitter. | | `max_request_query_length` | Optional. Maximum length in characters for `request_query` filter values. Default: `1024`. Maximum: `4096`. | | `max_request_body_bytes` | Optional. Maximum size in bytes for `request_body` filter values. Default: `16384` (16 KiB). Maximum: `65536` (64 KiB). | | `request_header_filters` | Optional. Set to `enabled` to allow `request_headers` filters to push down dynamic HTTP request headers. Default: `disabled`. Requires `request_header_allowlist`. | | `request_header_allowlist` | Comma-separated list of HTTP header names that `request_headers` filters may set (e.g., `x-sandbox-id, x-region`). **Required** when `request_header_filters` is enabled. The `authorization` header cannot be allowlisted when HTTP authentication is configured. | | `max_request_headers_length` | Optional. Maximum size in bytes for `request_headers` filter values. Default: `16384` (16 KiB). | | `max_request_partitions` | Optional. Maximum number of HTTP request partitions created from the cross product of `request_path`, `request_query`, `request_body`, and `request_headers` filters. If unset, partition count is unlimited. | | `health_probe` | Optional. Custom health probe path for endpoint validation during initialization (e.g., `/health`, `/api/status`). The endpoint must return a 2xx status code to pass validation. If not set, a random path is used and any status (including 404) is accepted. Must start with `/`. | | `auth_token_url` | Optional. OAuth2 token endpoint URL (must be HTTPS; `http://localhost` and loopback IPs are allowed for local testing). Enables OAuth2: the connector acquires short-lived access tokens (refresh-token grant by default, or `client_credentials` via `auth_grant_type`) and attaches them to all data requests (`Authorization: Bearer ` by default, or the bare token under a custom `auth_header_name`). Applies to dynamic JSON API endpoints only; structured HTTP file datasets reject OAuth2 params. See [OAuth2 Authentication](#oauth2-authentication). | | `auth_grant_type` | Optional. OAuth2 grant type: `refresh_token` (default, RFC 6749 §6) or `client_credentials` (RFC 6749 §4.4). `client_credentials` authenticates with `http_auth_client_id`/`http_auth_client_secret` and issues no refresh token. | | `http_auth_refresh_token` | Optional. OAuth2 refresh token exchanged against `auth_token_url` to obtain access tokens. **Required** when `auth_token_url` is set with the (default) refresh-token grant; not used by the `client_credentials` grant. Use a secret store, e.g. `${secrets:my_refresh_token}`. | | `http_auth_client_id` | Optional. OAuth2 `client_id` presented to the token endpoint. Required for confidential clients; optional for public clients. Must be paired with `http_auth_client_secret` for confidential clients. | | `http_auth_client_secret` | Optional. OAuth2 `client_secret` presented to the token endpoint. Required when the client is confidential; must be set together with `http_auth_client_id`. Use a secret store, e.g. `${secrets:my_client_secret}`. | | `auth_scopes` | Optional. Space-separated OAuth2 scopes to request when refreshing (e.g. `read:data offline_access`). Omit to inherit the scopes bound to the refresh token. | | `auth_client_auth` | Optional. How client credentials are sent to the token endpoint: `basic` (HTTP Basic header, default per RFC 6749 §2.3.1) or `body` (`client_id`/`client_secret` in the form body). Default: `basic`. | | `auth_header_name` | Optional. HTTP header that carries the access token. Default: `Authorization` (sends `Bearer `). Any other name (e.g. `X-Shopify-Access-Token`) sends the bare token instead. | #### Mutual TLS (mTLS) Client Authentication[​](#mutual-tls-mtls-client-authentication "Direct link to Mutual TLS (mTLS) Client Authentication") For upstream servers that require mutual TLS, the connector can present a client certificate during the TLS handshake. Provide the certificate and key either as file paths or inline PEM — the file-path and inline forms are mutually exclusive, and the certificate and key must always be set together. mTLS client identity applies to dynamic JSON API endpoints only; structured HTTP file datasets reject these parameters. | Parameter Name | Description | | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `http_tls_client_certificate_file` | Optional. Path to a PEM client certificate chain to present during the TLS handshake. Must be set together with `http_tls_client_key_file`. Mutually exclusive with the inline forms. | | `http_tls_client_key_file` | Optional. Path to the PEM private key matching `http_tls_client_certificate_file`. Must be set together with it. Mutually exclusive with the inline forms. | | `http_tls_client_certificate` | Optional. Inline PEM client certificate chain (or `${secrets:...}` reference). Must be set together with `http_tls_client_key`. Mutually exclusive with the file-path forms. | | `http_tls_client_key` | Optional. Inline PEM private key (or `${secrets:...}` reference) matching `http_tls_client_certificate`. Must be set together with it. Mutually exclusive with the file-path forms. | #### Rate Control Parameters[​](#rate-control-parameters "Direct link to Rate Control Parameters") HTTP-based connectors share a rate control system that limits concurrency and request rate per upstream origin. These parameters can be set per-dataset (in `params`) or globally (in `runtime.params`). Dataset-level settings override the global defaults. | Parameter Name | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `max_concurrent_requests` | Maximum number of concurrent HTTP requests to the same upstream origin. Overrides `runtime.params.http_max_concurrent_requests`. If both are unset, concurrency limiting is disabled. | | `requests_per_second_limit` | Maximum number of HTTP requests per second to the same upstream origin. Overrides `runtime.params.http_requests_per_second_limit`. If both are unset, no per-second rate limit is applied. | | `requests_per_minute_limit` | Maximum number of HTTP requests per minute to the same upstream origin. Overrides `runtime.params.http_requests_per_minute_limit`. If both are unset, no per-minute rate limit is applied. | | `rate_control_jitter_min` | Minimum random delay added before HTTP requests when rate control is active. Accepts durations such as `5ms` or `0ms`. Defaults to `5ms` when a request-rate limit is configured. | | `rate_control_jitter_max` | Maximum random delay added before HTTP requests when rate control is active. Accepts durations such as `10ms` or `0ms`. Defaults to `10ms` when a request-rate limit is configured. | Multiple datasets targeting the same origin share the same rate controller, ensuring the limits apply across all datasets for that origin. ``` runtime: params: http_max_concurrent_requests: 10 http_requests_per_second_limit: 5 datasets: - from: https://api.example.com/v1 name: api_data params: file_format: json allowed_request_paths: '/data/**' max_concurrent_requests: 3 # Override: this dataset uses at most 3 concurrent requests ``` #### Pagination Parameters[​](#pagination-parameters "Direct link to Pagination Parameters") | Parameter Name | Description | | ------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `pagination` | Optional. Pagination mode: `auto` (default) auto-detects `Link` headers, `enabled` explicitly enables pagination with configuration below, `disabled` turns off pagination. | | `pagination_next_pointer` | Optional. JSON pointer ([RFC 6901](https://datatracker.ietf.org/doc/html/rfc6901)) to the next page URL or cursor in the response body (e.g., `/next`, `/pagination/cursor`, `/links/next`). | | `pagination_link_header` | Optional. Whether to follow HTTP `Link` headers with `rel="next"` for pagination. Default: `enabled`. Set to `disabled` to ignore `Link` headers. | | `pagination_token_param` | Optional. When set, the value from `pagination_next_pointer` is treated as a cursor/token and passed as this query parameter name in subsequent requests. When not set, the value is treated as a full URL. | | `pagination_data_pointer` | Optional. JSON pointer ([RFC 6901](https://datatracker.ietf.org/doc/html/rfc6901)) to the data array in each page's response (e.g., `/data`, `/results`, `/items`). When set, only the array at this path is returned as data rows. | | `pagination_max_pages` | Optional. Maximum number of pages to fetch. Default: `100`. Set to `nolimit` to disable the page cap and fetch all available pages. | | `pagination_data_map_to_array` | Optional. When `enabled`, if the data at `pagination_data_pointer` (or the top-level response) is a JSON object/map, extracts its values as rows instead of treating it as a single row. Default: `disabled`. Requires pagination to be enabled. | | `pagination_query_params` | Optional. Query parameter template for client-driven pagination. Supports `{offset}`, `{limit}`, and `{page}` variables (e.g., `offset={offset}&limit={limit}`). Requires `pagination_page_size`. Mutually exclusive with `pagination_next_pointer` and `pagination_token_param`. | | `pagination_page_size` | Optional. Number of items per page for query-parameter pagination. Must be a positive integer. Expands `{limit}` in `pagination_query_params` and detects the last page (fewer results than `page_size` means done). Requires `pagination_query_params`. | ### Caching Mode Parameters[​](#caching-mode-parameters "Direct link to Caching Mode Parameters") When using [`refresh_mode: caching`](/docs/next/features/data-acceleration/refresh-modes/caching), cache freshness is controlled by additional parameters placed under `acceleration.params` — **not** under the top-level `params` block. | Parameter | Description | Default | | ------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | | `caching_ttl` | How long cached data is considered fresh. After this period, data becomes stale and a background refresh is triggered. **Defaults to `30s`.** | `30s` | | `caching_stale_while_revalidate_ttl` | How long after `caching_ttl` expires to continue serving stale data while a background refresh runs. If omitted, queries wait for fresh data once TTL expires. | None | | `caching_stale_if_error` | When set to `enabled`, serves expired cached data if the upstream source returns an error rather than failing the query. | `disabled` | `caching_ttl` defaults to 30 seconds If you set `refresh_check_interval: 15m` but leave `caching_ttl` at its default, cached entries are considered stale after only **30 seconds** — not 15 minutes. Always set `caching_ttl` explicitly to match your intended freshness window. ``` datasets: - from: https://api.example.com/v1 name: api_cache params: file_format: json allowed_request_paths: '/data/**' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching refresh_check_interval: 15m on_zero_results: use_source params: caching_ttl: 15m # explicit — default is only 30s caching_stale_while_revalidate_ttl: 5m # serve stale data while refreshing ``` See [Caching Refresh Mode](/docs/next/features/data-acceleration/refresh-modes/caching) for full TTL semantics, stale-while-revalidate behaviour, and cache persistence options. ## HTTP Response Headers[​](#http-response-headers "Direct link to HTTP Response Headers") When querying HTTP(s) datasets, Spice respects standard HTTP caching headers in responses. The connector supports the following cache-related response headers: ### `Cache-Control`[​](#cache-control "Direct link to cache-control") The `Cache-Control` response header from the HTTP(s) endpoint is passed through to clients querying Spice. When the HTTP(s) server returns a `Cache-Control` header with the `stale-while-revalidate` directive, clients can use this value to determine appropriate caching behavior. For example, if the HTTP(s) endpoint returns: ``` Cache-Control: max-age=10, stale-while-revalidate=10 ``` Clients querying Spice will receive this header and can: 1. Serve fresh data for 10 seconds after fetching. 2. Between 10-20 seconds, serve stale data while fetching fresh data in the background. 3. After 20 seconds, fetch fresh data before serving the next request. The stale-while-revalidate behavior in Spice is controlled by the `stale_while_revalidate_ttl` parameter in the [caching configuration](/docs/next/features/caching#stale-while-revalidate). When `stale_while_revalidate_ttl` is set to `0` (default), stale data will not be served. When set to a non-zero value, Spice serves stale cache entries while revalidating in the background. ## Advanced Features[​](#advanced-features "Direct link to Advanced Features") The HTTP connector provides advanced capabilities for working with dynamic APIs and RESTful services, including built-in pagination and special metadata fields. ### Pagination[​](#pagination "Direct link to Pagination") The HTTP connector supports automatic pagination for REST APIs that return data across multiple pages. Pagination is configured via `params` and works transparently with acceleration (caching, append, and full refresh modes) — each page is streamed as a separate batch without buffering entire result sets in memory. #### Pagination Modes[​](#pagination-modes "Direct link to Pagination Modes") There are three pagination modes: **URL mode** — The next page URL is extracted from the response body (via `pagination_next_pointer`) or from the HTTP `Link` header with `rel="next"`. ``` datasets: - from: https://api.example.com/v1/items name: items params: pagination: enabled pagination_next_pointer: /links/next pagination_data_pointer: /data pagination_max_pages: 50 ``` **Token mode** — A cursor/token is extracted from the response body (via `pagination_next_pointer`) and passed as a query parameter (specified by `pagination_token_param`) in subsequent requests. ``` datasets: - from: https://api.example.com/v1/items name: items params: pagination: enabled pagination_next_pointer: /pagination/cursor pagination_token_param: cursor pagination_data_pointer: /results ``` **Query-parameter mode** — The client drives pagination by expanding a template (`pagination_query_params`) with `{offset}`, `{limit}`, and `{page}` variables. Pagination stops when a page returns fewer rows than `pagination_page_size`. ``` datasets: - from: https://api.example.com/v1/widgets name: widgets params: pagination: enabled pagination_query_params: "offset={offset}&limit={limit}" pagination_page_size: "100" pagination_max_pages: "50" ``` #### Map-to-Array Conversion[​](#map-to-array-conversion "Direct link to Map-to-Array Conversion") Some APIs return data as a JSON object/map (e.g., `{"1": {...}, "2": {...}}`) instead of an array. Set `pagination_data_map_to_array: enabled` to extract the map values as individual rows. ``` datasets: - from: https://api.example.com/v1/records name: records params: pagination: enabled pagination_data_map_to_array: enabled pagination_query_params: "offset={offset}&limit={limit}" pagination_page_size: "100" ``` #### Auto Mode[​](#auto-mode "Direct link to Auto Mode") By default, `pagination` is set to `auto`, which automatically follows HTTP `Link` headers with `rel="next"` if present in responses. Set `pagination: disabled` to turn off all pagination behavior, or `pagination: enabled` to explicitly configure pagination with the parameters above. #### Validation Rules[​](#validation-rules "Direct link to Validation Rules") * `pagination_query_params` requires `pagination_page_size` (and vice versa) * `pagination_query_params` is mutually exclusive with `pagination_next_pointer` and `pagination_token_param` * `pagination_query_params` must contain `{offset}` or `{page}` to ensure pages advance * `pagination_token_param` requires `pagination_next_pointer` * `pagination_next_pointer` and `pagination_data_pointer` must be valid JSON pointers starting with `/` #### SSRF Protection[​](#ssrf-protection "Direct link to SSRF Protection") When using URL mode, next-page URLs extracted from response bodies are validated to share the same origin as the base URL configured in `from`. Cross-origin redirects are rejected. ### Special Metadata Fields[​](#special-metadata-fields "Direct link to Special Metadata Fields") The HTTP connector supports special metadata fields that provide fine-grained control over HTTP requests. These fields can be included in your dataset schema to dynamically construct request URLs and payloads. Security Requirements For security, these metadata fields require explicit configuration to prevent unauthorized access: * `request_path` requires `allowed_request_paths` to be configured with glob patterns * `request_query` requires `request_query_filters: enabled` * `request_body` requires `request_body_filters: enabled` * `request_headers` requires `request_header_filters: enabled` and `request_header_allowlist` | Field Name | Type | Description | | ----------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `request_path` | String | Specifies the URL path to append to the base URL from the `from` field. When using a base domain/path in `from`, `request_path` constructs the complete endpoint. Example: If `from: https://api.example.com` and `request_path: /users/123`, the request will be made to `https://api.example.com/users/123`. **Requires `allowed_request_paths` parameter.** | | `request_query` | String | Defines query parameters to append to the request URL. Formatted as a query string (e.g., `key1=value1&key2=value2`). These parameters are appended to the URL after any path specified in `request_path`. **Requires `request_query_filters: enabled`.** Maximum length: configurable via `max_request_query_length` (default: 1024 characters). | | `request_body` | String | Contains the request body for POST/PUT requests. Typically used with REST APIs that require a JSON or form-encoded payload. The content type should be specified using `http_headers`. **Requires `request_body_filters: enabled`.** Maximum size: configurable via `max_request_body_bytes` (default: 16 KiB). | | `request_headers` | String | A JSON object of HTTP request headers to set on a per-request basis (e.g., `'{"x-sandbox-id":"sandbox-1"}'`). Only headers listed in `request_header_allowlist` are permitted. **Requires `request_header_filters: enabled` and `request_header_allowlist`.** Maximum size: configurable via `max_request_headers_length` (default: 16 KiB). | These metadata fields work in combination: * If `from` specifies a complete file URL, these fields are ignored * If `from` specifies a base URL, these fields construct the full request dynamically * `request_path` is appended to the base URL * `request_query` is appended as query parameters * `request_body` is sent as the request payload (requires appropriate HTTP method configuration) * `request_headers` sets per-request HTTP headers (allowlisted names only) OR filter restriction `OR` expressions across **different** filter columns (e.g., `WHERE request_query = 'a' OR request_headers = 'b'`) are rejected because the connector would issue a single combined HTTP request instead of separate ones. Use `UNION ALL` for alternative requests across different columns. `OR` within a single column (e.g., `WHERE request_path = '/a' OR request_path = '/b'`) is supported. ### Response Metadata Fields[​](#response-metadata-fields "Direct link to Response Metadata Fields") In addition to request metadata, the HTTP connector includes response metadata fields in the dataset schema. These fields capture information about the HTTP response and are available in SQL queries. | Field Name | Type | Description | | ------------------ | ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `content` | String | The response body content. | | `response_status` | UInt16 | The HTTP status code of the response (e.g., `200`, `404`, `500`). | | `response_headers` | Map(String, String) | The HTTP response headers as key-value pairs. Each header name maps to its value. Available for inspection in queries, e.g., to check `content-type` or custom headers returned by the API. | | `_fetched_at` | Timestamp (Nanosecond) | The timestamp when the data was fetched. Uses the HTTP `Date` response header when available, falling back to the current system time. Always present in the dataset schema, even when not declared explicitly — this is required for caching TTL eviction and for append-mode datasets that set `time_column: _fetched_at`. | #### Querying Response Metadata[​](#querying-response-metadata "Direct link to Querying Response Metadata") ``` -- Check the HTTP status of cached responses SELECT request_path, response_status, _fetched_at FROM my_http_dataset; -- Inspect response headers SELECT request_path, response_headers FROM my_http_dataset WHERE request_path = '/api/data'; ``` note When using [caching refresh mode](/docs/next/features/data-acceleration/refresh-modes/caching), transient HTTP error responses (5xx server errors and 429 Too Many Requests) are automatically excluded from the cache. These responses are still returned to the querying client but are not persisted, preventing temporary failures from polluting cached data. ### Metadata Columns with JSON Schema Decomposition[​](#metadata-columns-with-json-schema-decomposition "Direct link to Metadata Columns with JSON Schema Decomposition") When a dataset uses JSON schema decomposition (`metadata.json_object: "*"`), columns whose names match a reserved HTTP metadata field are populated from the HTTP request/response — with their original Arrow types — instead of being decomposed from the JSON body. This lets a single dataset expose both decomposed body columns and typed HTTP metadata. Reserved metadata field names: `request_path`, `request_query`, `request_body`, `request_headers`, `content`, `response_status`, `response_headers`, `_fetched_at`. `_fetched_at` is auto-injected into the schema even when not declared. ``` datasets: - from: https://api.tvmaze.com/shows name: tvmaze_shows columns: - name: request_path # Utf8, from HTTP request metadata - name: response_status # UInt16, from HTTP response metadata - name: _fetched_at # Timestamp(ns), from response Date header (auto-injected if omitted) - name: id # Utf8, decomposed from JSON body - name: name # Utf8, decomposed from JSON body - name: premiered # Utf8, decomposed from JSON body - name: details # catch-all for remaining JSON keys metadata: json_object: "*" ``` ``` SELECT id, name, response_status, _fetched_at FROM tvmaze_shows WHERE request_path = '/shows/1' AND response_status = 200; ``` Metadata columns retain their native types (`UInt16` for `response_status`, `Timestamp(Nanosecond)` for `_fetched_at`, `Map(String, String)` for `response_headers`), while body-derived columns are `Utf8`. `_fetched_at` is appended to the schema automatically when omitted, so caching TTL eviction and `time_column: _fetched_at` work without requiring the column to be declared. **Collision rules:** * A JSON body key that collides with a reserved metadata name is dropped — it does not shadow the real HTTP value and does not leak into the catch-all column. * Declaring the catch-all column itself (`json_object: "*"`) with a reserved metadata name is rejected at registration time. * Datasets that don't use JSON schema decomposition are unaffected. ### Endpoint Validation[​](#endpoint-validation "Direct link to Endpoint Validation") The HTTP connector validates the configured endpoint during initialization to detect issues such as DNS errors, connection problems, or invalid URLs early in the startup process. #### Default Validation Behavior[​](#default-validation-behavior "Direct link to Default Validation Behavior") By default, the connector performs a health check by requesting a randomly generated path (e.g., `/__spice_health_check_abc123def456`) that is expected to return a 404 status. Any HTTP response, including 404 Not Found, indicates that the endpoint is reachable and the dataset will initialize successfully. This default behavior works for most HTTP endpoints but may not be suitable for APIs that: * Return error responses for unknown paths without proper HTTP status codes * Have strict path validation that rejects requests to non-existent endpoints * Require authentication for all paths, including health check endpoints #### Custom Health Probe[​](#custom-health-probe "Direct link to Custom Health Probe") For endpoints that require a specific health check path, configure the `health_probe` parameter: ``` datasets: - from: https://api.example.com/v1 name: api_data params: health_probe: /health ``` When a custom health probe is configured: * The connector validates the endpoint by requesting the specified path * The health probe endpoint must return a 2xx status code (200-299) for validation to succeed * If the health probe returns a non-2xx status code, the dataset will fail to initialize with an error message This provides more reliable validation for APIs with dedicated health check endpoints. ##### Example with Authentication[​](#example-with-authentication "Direct link to Example with Authentication") ``` datasets: - from: https://api.example.com name: authenticated_api params: http_headers: 'Authorization:Bearer ${secrets:api_token}' health_probe: /api/status ``` In this configuration, the health probe request to `/api/status` will include the authentication header, ensuring that the validation succeeds for APIs that require authentication on all endpoints. ##### Health Probe Requirements[​](#health-probe-requirements "Direct link to Health Probe Requirements") The `health_probe` parameter has the following requirements: * Must start with `/` * Cannot exceed 2048 characters in length * The target endpoint must return a 2xx HTTP status code for validation to succeed ### OAuth2 Authentication[​](#oauth2-authentication "Direct link to OAuth2 Authentication") The HTTP connector supports two OAuth2 grants for JSON APIs — the **refresh-token grant** (RFC 6749 §6, the default) and the **client-credentials grant** (RFC 6749 §4.4). Both acquire short-lived access tokens from a token endpoint and keep them fresh in the background. For the (default) refresh-token grant, given a long-lived refresh token and a token endpoint, Spice will: 1. Exchange the refresh token for an access token at dataset startup. 2. Attach the access token to every data request — `Authorization: Bearer ` by default, or under a custom header (see [Custom Token Header](#custom-token-header)). 3. Refresh the access token in the background, 60 seconds before it expires, for the lifetime of the process. 4. Honor rotated refresh tokens — when the token endpoint returns a new `refresh_token`, Spice uses it for the next exchange. The [client-credentials grant](#client-credentials-grant) is designed for machine-to-machine APIs that authenticate with a `client_id`/`client_secret` and issue no refresh token; it re-runs the same token exchange in the background before expiry. In both cases Spice does **not** perform an interactive authorization flow (the authorization-code and device-code flows are not exposed), nor does it retry data requests on 401 — keeping the token continuously fresh in the background is the only recovery path. #### Basic Configuration[​](#basic-configuration "Direct link to Basic Configuration") ``` datasets: - from: https://api.example.com name: secure_data params: file_format: json allowed_request_paths: '/v1/**' auth_token_url: https://auth.example.com/oauth/token http_auth_refresh_token: ${secrets:my_refresh_token} http_auth_client_id: ${secrets:my_client_id} http_auth_client_secret: ${secrets:my_client_secret} ``` #### Parameter Reference[​](#parameter-reference "Direct link to Parameter Reference") | Parameter | Kind | Required | Description | | ------------------------- | ----------------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `auth_token_url` | runtime | yes (for OAuth) | OAuth2 token endpoint URL. Must be HTTPS; `http://localhost`, `http://127.0.0.1`, and `http://[::1]` are accepted for local testing. | | `auth_grant_type` | runtime | no | OAuth2 grant type: `refresh_token` (default, RFC 6749 §6) or `client_credentials` (RFC 6749 §4.4). See [Client-Credentials Grant](#client-credentials-grant). | | `http_auth_refresh_token` | component, secret | refresh-token | Long-lived refresh token. Exchanged on startup for the first access token. **Required** for the (default) refresh-token grant; must not be set for `client_credentials`. Can be loaded from any [supported secret store](/docs/next/components/secret-stores) via `${secrets:...}`. | | `http_auth_client_id` | component, secret | confidential | `client_id`. Required for confidential clients and for the `client_credentials` grant, optional for public clients. When set together with `http_auth_client_secret`, both are sent to the token endpoint. | | `http_auth_client_secret` | component, secret | confidential | `client_secret`. Must be paired with `http_auth_client_id`; required for the `client_credentials` grant. Can be loaded from any [supported secret store](/docs/next/components/secret-stores) via `${secrets:...}`. | | `auth_scopes` | runtime | no | Space-separated OAuth2 scopes (e.g. `read:data offline_access`). Omit to inherit the scopes bound to the refresh token. | | `auth_client_auth` | runtime | no | How client credentials are sent to the token endpoint: `basic` (default, HTTP Basic header per RFC 6749 §2.3.1) or `body` (as `client_id`/`client_secret` form fields). | | `auth_header_name` | runtime | no | HTTP header that carries the access token. Default `Authorization` (sends `Bearer `); any other name sends the bare token (e.g. `X-Shopify-Access-Token`). See [Custom Token Header](#custom-token-header). | Parameter naming convention Component/secret parameters carry the `http_` prefix when set in a dataset (`http_auth_refresh_token`, `http_auth_client_id`, `http_auth_client_secret`). Runtime parameters do not (`auth_token_url`, `auth_grant_type`, `auth_scopes`, `auth_client_auth`, `auth_header_name`). This follows the same convention as `http_password` vs `client_timeout`. Loading secrets from a secret store The refresh token and client secret should never be committed to source. Reference them from any [supported secret store](/docs/next/components/secret-stores) — environment variables, Kubernetes Secrets, AWS Secrets Manager, HashiCorp Vault, or the OS keychain — using the `${secrets:...}` [replacement syntax](/docs/next/components/secret-stores#using-secrets). For example, with Kubernetes Secrets enabled: ``` params: auth_token_url: https://auth.example.com/oauth/token http_auth_refresh_token: ${secrets:my_refresh_token} http_auth_client_id: ${secrets:my_client_id} http_auth_client_secret: ${secrets:my_client_secret} ``` #### Public Clients (No Client Secret)[​](#public-clients-no-client-secret "Direct link to Public Clients (No Client Secret)") For public clients the `client_secret` is omitted. If you still want to send a `client_id` for correlation, set `http_auth_client_id` without `http_auth_client_secret`: ``` params: auth_token_url: https://auth.example.com/oauth/token http_auth_refresh_token: ${secrets:my_refresh_token} http_auth_client_id: ${secrets:my_public_client_id} ``` #### Sending Credentials in the Body Instead of Basic Auth[​](#sending-credentials-in-the-body-instead-of-basic-auth "Direct link to Sending Credentials in the Body Instead of Basic Auth") Some token endpoints require `client_id`/`client_secret` in the form body rather than via the HTTP Basic header. Set `auth_client_auth: body`: ``` params: auth_token_url: https://auth.example.com/oauth/token http_auth_refresh_token: ${secrets:my_refresh_token} http_auth_client_id: ${secrets:my_client_id} http_auth_client_secret: ${secrets:my_client_secret} auth_client_auth: body ``` #### Client-Credentials Grant[​](#client-credentials-grant "Direct link to Client-Credentials Grant") For machine-to-machine APIs that authenticate with a `client_id`/`client_secret` and issue no refresh token, set `auth_grant_type: client_credentials`. Both `http_auth_client_id` and `http_auth_client_secret` are required, and `http_auth_refresh_token` must **not** be set: ``` params: auth_token_url: https://auth.example.com/oauth/token auth_grant_type: client_credentials http_auth_client_id: ${secrets:my_client_id} http_auth_client_secret: ${secrets:my_client_secret} auth_scopes: 'read:data' ``` Spice re-runs the client-credentials exchange in the background before the access token expires, reusing the same refresh machinery as the refresh-token grant. `auth_client_auth` (`basic`/`body`) applies to this grant as well. #### Custom Token Header[​](#custom-token-header "Direct link to Custom Token Header") By default the access token is attached as `Authorization: Bearer `. Some APIs expect the token in a non-standard header and without the `Bearer` prefix — for example the Shopify Admin API uses `X-Shopify-Access-Token`. Set `auth_header_name` to send the bare token under that header instead: ``` params: auth_token_url: https://auth.example.com/oauth/token auth_grant_type: client_credentials auth_header_name: X-Shopify-Access-Token http_auth_client_id: ${secrets:my_client_id} http_auth_client_secret: ${secrets:my_client_secret} ``` `auth_header_name` is independent of the grant type — it works with the refresh-token grant too. #### Local Testing[​](#local-testing "Direct link to Local Testing") The connector rejects `http://` token URLs by default, but allows `http://localhost`, `http://127.0.0.1`, and `http://[::1]` so you can run a mock OAuth server for development: ``` params: auth_token_url: http://localhost:8080/oauth/token http_auth_refresh_token: local-dev-token ``` #### Error Behavior[​](#error-behavior "Direct link to Error Behavior") The connector classifies token-endpoint errors to make remediation easy: * **Configuration errors** (fail-fast at dataset init, surfaces as `InvalidConfiguration`): * Malformed or insecure `auth_token_url` * Token endpoint returns `400`, `401`, or `403` (typically an invalid refresh token, client credentials, or scope) * Token endpoint returns a non-`Bearer` `token_type` * Incomplete config (e.g. `http_auth_refresh_token` without `auth_token_url`, or `http_auth_client_secret` without `http_auth_client_id`) * `auth_grant_type: client_credentials` without both `http_auth_client_id` and `http_auth_client_secret`, or with `http_auth_refresh_token` set (the client-credentials grant issues no refresh token) * Both OAuth2 auth *and* an `Authorization` header in `http_headers` — remove one * **Transient / connection errors** (surfaces as `UnableToConnect`, retried in the background): * Network / TLS failures * `5xx`, `408`, or `429` from the token endpoint * Parse failures on the token response Error bodies returned by the token endpoint are truncated to 512 bytes and whitespace-collapsed before being surfaced in errors or logs, so hostile or misbehaving endpoints cannot force unbounded buffering or leak multi-line payloads into logs. #### Limitations[​](#limitations "Direct link to Limitations") * **Dynamic JSON APIs only.** A structured HTTP file dataset (csv, parquet, etc.) is served by the object-store listing path, which cannot attach an access token. Setting any OAuth2 parameter on one fails registration with a configuration error naming the parameters to remove, rather than sending unauthenticated requests. `http_headers` is not applied on that path either — use [Basic authentication](#using-basic-authentication) to authenticate a structured file download. * **No interactive auth flows.** The refresh-token and client-credentials grants are supported; the authorization-code and device-code flows are not. For the refresh-token grant, obtain the initial refresh token out-of-band. * **No 401→refresh-and-retry.** Background refresh keeps the token fresh; if a data request 401s, it propagates to the caller. * **One authenticator per dataset.** Configure either OAuth2 or an `Authorization` header in `http_headers`, not both — the connector rejects the combination at registration time. ## Advanced Usage[​](#advanced-usage "Direct link to Advanced Usage") ### Using Special Metadata Fields with Base URL[​](#using-special-metadata-fields-with-base-url "Direct link to Using Special Metadata Fields with Base URL") When using a base URL with special metadata fields, you can dynamically construct different API endpoints: ``` datasets: - from: https://api.example.com/v1 name: api_requests params: http_headers: 'Content-Type:application/json' allowed_request_paths: '/users,/data/upload,/api/**' request_query_filters: enabled request_body_filters: enabled ``` With the above configuration, you can query different endpoints by providing values for the special metadata fields: ``` -- Query a specific user endpoint SELECT * FROM api_requests WHERE request_path = '/users/123' AND request_query = 'include=profile,settings'; -- Make a POST request with a body SELECT * FROM api_requests WHERE request_path = '/data/upload' AND request_body = '{"name":"example","value":42}'; ``` The connector will construct requests like: * `https://api.example.com/v1/users/123?include=profile,settings` * `https://api.example.com/v1/data/upload` with the JSON body #### Securing Paths with Glob Patterns[​](#securing-paths-with-glob-patterns "Direct link to Securing Paths with Glob Patterns") The `allowed_request_paths` parameter supports glob patterns to flexibly and securely match request paths. This provides a flexible way to configure path filtering without listing every possible endpoint. **Pattern Types:** * **Single wildcard (`*`)**: Matches any characters within a single path segment * Example: `/shows/*` matches `/shows/123` and `/shows/breaking-bad` * Does not match across path separators: `/shows/*` does not match `/shows/123/episodes` * \*\*Recursive wildcard (`**`)\*\*: Matches any number of path segments * Example: `/api/**` matches `/api/users`, `/api/v1/users`, and `/api/v2/posts/123` * Use for flexible API version matching or deep hierarchies * **Character classes (`[...]`)**: Matches one character from a set * Example: `/api/v[0-9]/*` matches `/api/v1/users` and `/api/v2/posts` * Example: `/api/v[1-3]/*` matches `/api/v1/users`, `/api/v2/posts`, and `/api/v3/data` **Examples:** ``` datasets: - from: https://api.tvmaze.com name: tv_api params: # Match any show ID allowed_request_paths: '/shows/*' ``` ``` -- Matches because /shows/82 matches the pattern /shows/* SELECT * FROM tv_api WHERE request_path = '/shows/82'; ``` ``` datasets: - from: https://api.example.com name: versioned_api params: # Match all endpoints under any API version allowed_request_paths: '/api/**' ``` ``` -- All of these match the pattern /api/** SELECT * FROM versioned_api WHERE request_path = '/api/users'; SELECT * FROM versioned_api WHERE request_path = '/api/v1/users'; SELECT * FROM versioned_api WHERE request_path = '/api/v2/products/electronics'; ``` ``` datasets: - from: https://api.example.com name: specific_versions params: # Match only API versions 1-9 allowed_request_paths: '/api/v[0-9]/*' ``` ``` -- Matches because /api/v1/users matches /api/v[0-9]/* SELECT * FROM specific_versions WHERE request_path = '/api/v1/users'; -- Does NOT match because v10 has two digits SELECT * FROM specific_versions WHERE request_path = '/api/v10/users'; ``` ### Dynamic Filters with Metadata Fields[​](#dynamic-filters-with-metadata-fields "Direct link to Dynamic Filters with Metadata Fields") The special metadata fields can be combined with dynamic filters to create sophisticated data refresh patterns. #### Dynamic API Queries with SQL[​](#dynamic-api-queries-with-sql "Direct link to Dynamic API Queries with SQL") ``` datasets: - from: https://api.tvmaze.com name: tv_shows params: http_headers: 'Accept:application/json' allowed_request_paths: '/search/shows,/shows/*,/shows/*/episodes' request_query_filters: enabled ``` Query specific API endpoints dynamically: ``` -- Search for shows by name SELECT * FROM tv_shows WHERE request_path = '/search/shows' AND request_query = 'q=game+of+thrones'; -- Get a specific show by ID (matches /shows/* pattern) SELECT * FROM tv_shows WHERE request_path = '/shows/82'; -- Get episodes for a show with filters (matches /shows/*/episodes pattern) SELECT * FROM tv_shows WHERE request_path = '/shows/82/episodes' AND request_query = 'season=1'; ``` #### Incremental Loading with Metadata Fields[​](#incremental-loading-with-metadata-fields "Direct link to Incremental Loading with Metadata Fields") ``` datasets: - from: https://api.example.com name: events params: allowed_request_paths: '/events,/events/*' request_query_filters: enabled acceleration: enabled: true refresh_mode: append refresh_sql: | SELECT * FROM events WHERE request_path = '/events' AND request_query = CONCAT('since=', (SELECT MAX(created_at) FROM events)) ``` This configuration: * Uses `request_path` to specify the `/events` endpoint * Dynamically constructs the `request_query` parameter using the latest timestamp from existing data * On each refresh, only fetches events created after the last refresh #### Paginated Data Loading[​](#paginated-data-loading "Direct link to Paginated Data Loading") tip For APIs with standard pagination patterns, consider using the built-in [pagination](#pagination) feature instead of manual `refresh_sql` pagination. Built-in pagination handles page traversal automatically with streaming execution. ``` datasets: - from: https://api.example.com/v2 name: paginated_data params: http_headers: 'Content-Type:application/json' allowed_request_paths: '/data' request_query_filters: enabled acceleration: enabled: true refresh_mode: append refresh_sql: | SELECT * FROM paginated_data WHERE request_path = '/data' AND request_query = CONCAT('page=', COALESCE((SELECT MAX(page_number) FROM paginated_data) + 1, 1), '&limit=100') ``` This incrementally loads pages of data by: * Tracking the last loaded page number * Constructing the next page query parameter * Fetching 100 records per page #### POST Request with Dynamic Body[​](#post-request-with-dynamic-body "Direct link to POST Request with Dynamic Body") ``` datasets: - from: https://api.example.com name: search_results params: http_headers: 'Content-Type:application/json' allowed_request_paths: '/search' request_body_filters: enabled acceleration: enabled: true refresh_mode: full refresh_sql: | SELECT * FROM search_results WHERE request_path = '/search' AND request_body = '{"query": {"match": {"status": "active"}}, "from": 0, "size": 1000}' ``` This example demonstrates: * Using `_body` to send a JSON payload for a POST request * Executing complex search queries against REST APIs * Fetching results based on structured query syntax #### Subquery-Driven HTTP Requests[​](#subquery-driven-http-requests "Direct link to Subquery-Driven HTTP Requests") The HTTP connector supports `IN (SELECT ...)` subqueries on filter columns (`request_path`, `request_query`, `request_body`, `request_headers`). Instead of fetching the entire HTTP dataset and joining in memory, the optimizer produces one HTTP request per unique subquery value. ``` datasets: - from: s3://my-bucket/org_list.csv name: orgs - from: https://api.example.com name: org_api params: file_format: json allowed_request_paths: '/headers' http_headers: 'x-static-header: static-value' request_header_filters: enabled request_header_allowlist: x-org-id max_request_partitions: 100 ``` ``` WITH org_headers AS ( SELECT '{"x-org-id":"' || org_id || '"}' AS hdr FROM orgs ) SELECT request_headers, content FROM org_api WHERE request_path = '/headers' AND request_headers IN (SELECT hdr FROM org_headers); ``` Each unique `hdr` value from the subquery triggers a separate HTTP request with the corresponding `x-org-id` header. The connector deduplicates values and caps the build side at 20,000 unique values. Use `max_request_partitions` to limit the total number of HTTP requests. JOIN is not supported for HTTP filter columns `JOIN ... ON` queries where the join key is an HTTP filter column (e.g., `request_headers`, `request_path`) are not supported and will return an error. Use `IN (SELECT ...)` instead: ``` -- This will error: SELECT h.content, p.path FROM http_api h JOIN params p ON h.request_path = p.path; -- Use this instead: SELECT content FROM http_api WHERE request_path IN (SELECT path FROM params); ``` ### Processing JSON Responses[​](#processing-json-responses "Direct link to Processing JSON Responses") APIs often return JSON data that requires parsing to extract specific fields. Spice provides [JSON functions](/docs/next/reference/sql/json) to process and transform JSON responses directly in SQL queries. #### Extracting Fields from JSON[​](#extracting-fields-from-json "Direct link to Extracting Fields from JSON") ``` datasets: - from: https://api.tvmaze.com name: tvmaze params: file_format: json allowed_request_paths: '/shows/*' ``` Extract specific fields from JSON responses: ``` -- Extract the show name from a JSON response SELECT json_get_str(content, 'name') as name FROM tvmaze WHERE request_path = '/shows/169'; ``` #### Working with Nested JSON[​](#working-with-nested-json "Direct link to Working with Nested JSON") APIs often return deeply nested JSON structures that require parsing to extract specific fields. Use chained JSON functions to navigate nested objects: ``` -- Extract nested fields from a show's network information SELECT json_get_str(content, 'name') as show_name, json_get_str(json_get(content, 'network'), 'name') as network_name, json_get_str(json_get(json_get(content, 'network'), 'country'), 'name') as country, json_get_str(json_get(json_get(content, 'network'), 'country'), 'code') as country_code FROM tvmaze WHERE request_path = '/shows/82'; ``` This demonstrates extracting nested objects step by step: * `json_get(content, 'network')` extracts the network object * `json_get_str(json_get(content, 'network'), 'name')` gets the network name from the nested object * Multiple `json_get` calls can be chained to navigate deeper levels #### Extracting Multiple Fields[​](#extracting-multiple-fields "Direct link to Extracting Multiple Fields") ``` -- Parse multiple fields from a TV show API response SELECT json_get_str(content, 'name') as show_name, json_get_str(content, 'type') as show_type, json_get_str(content, 'language') as language, json_get_int(content, 'runtime') as runtime_minutes, json_get_str(content, 'premiered') as premiere_date, json_get_str(content, 'status') as status FROM tvmaze WHERE request_path = '/shows/169'; ``` #### Processing JSON Arrays[​](#processing-json-arrays "Direct link to Processing JSON Arrays") ``` -- Extract genres from a JSON array SELECT json_get_str(content, 'name') as show_name, json_get_array(content, 'genres') as genres_array FROM tvmaze WHERE request_path = '/shows/82'; ``` For more details on available JSON functions including `json_get`, `json_get_str`, `json_get_int`, `json_get_bool`, and others, refer to the [JSON functions reference](/docs/next/reference/sql/json). ### Refresh SQL with Dynamic Filters[​](#refresh-sql-with-dynamic-filters "Direct link to Refresh SQL with Dynamic Filters") The HTTP connector supports dynamic URL construction through `refresh_sql` with templated query parameters. This enables incremental data loading by appending filter conditions from the SQL query to the HTTP request URL. #### How It Works[​](#how-it-works "Direct link to How It Works") When `refresh_sql` is specified with filters, the connector extracts filter conditions and appends them as query parameters to the URL. This is particularly useful for APIs that support filtering via query parameters. #### Time-Based Incremental Loading[​](#time-based-incremental-loading "Direct link to Time-Based Incremental Loading") ``` datasets: - from: https://api.example.com/data.csv?start_time={start_time}&end_time={end_time} name: incremental_data acceleration: enabled: true refresh_mode: append refresh_sql: | SELECT * FROM incremental_data WHERE timestamp > (SELECT MAX(timestamp) FROM incremental_data) ``` In this example: * The `{start_time}` and `{end_time}` placeholders in the URL are replaced with values extracted from the `WHERE` clause in `refresh_sql` * Each refresh appends only new data since the last refresh * The connector automatically maps SQL filter conditions to URL query parameters #### Supported Filter Operations[​](#supported-filter-operations "Direct link to Supported Filter Operations") The dynamic filter feature supports the following SQL operations: * Equality comparisons (`=`) * Greater than (`>`) * Less than (`<`) * Greater than or equal (`>=`) * Less than or equal (`<=`) * Range queries with `BETWEEN` #### Notes[​](#notes "Direct link to Notes") * URL parameters must match filter column names in the `refresh_sql` * Only filters that can be pushed down to the HTTP source will be applied to the URL * Complex filters may not be supported for URL templating ## Limitations[​](#limitations-1 "Direct link to Limitations") ### Security Constraints[​](#security-constraints "Direct link to Security Constraints") For security and to prevent unauthorized access, the HTTP connector enforces the following constraints on special metadata fields: #### Request Path Limitations[​](#request-path-limitations "Direct link to Request Path Limitations") * **Explicit Allow-List Required**: The `request_path` field cannot be used without configuring `allowed_request_paths` * **Path Pattern Format**: All patterns in `allowed_request_paths` must: * Start with `/` * Not contain `..` path traversal segments * Not exceed 2048 characters in length * **Glob Pattern Matching**: Query filters are matched against glob patterns in the `allowed_request_paths` list using: * `*` matches a single path segment (e.g., `/shows/*` matches `/shows/123` but not `/shows/123/episodes`) * `**` matches multiple path segments recursively (e.g., `/api/**` matches `/api/v1/users` and `/api/v2/posts/123`) * `[...]` character classes (e.g., `/api/v[0-9]/*` matches `/api/v1/users` but not `/api/v10/users`) * **Empty Paths**: Empty `request_path` filters are rejected Example error when `allowed_request_paths` is not configured: ``` request_path filters are disabled for this dataset. Configure allowed_request_paths to enable them. ``` #### Request Query Limitations[​](#request-query-limitations "Direct link to Request Query Limitations") * **Explicit Enable Required**: The `request_query` field requires `request_query_filters: enabled` * **Length Limit**: Query strings are limited to 1024 characters by default (configurable up to 4096 via `max_request_query_length`) * **Control Characters**: Query strings cannot contain control characters * **Leading Question Mark**: The connector automatically strips leading `?` if present Example error when query filters are not enabled: ``` request_query filters are disabled for this dataset. Enable request_query_filters to use them. ``` #### Request Body Limitations[​](#request-body-limitations "Direct link to Request Body Limitations") * **Explicit Enable Required**: The `request_body` field requires `request_body_filters: enabled` * **Size Limit**: Request bodies are limited to 16 KiB (16,384 bytes) by default (configurable up to 64 KiB via `max_request_body_bytes`) * **POST Method**: When a `request_body` filter is present, the HTTP method automatically changes to POST Example error when body filters are not enabled: ``` request_body filters are disabled for this dataset. Enable request_body_filters to use them. ``` #### Request Headers Limitations[​](#request-headers-limitations "Direct link to Request Headers Limitations") * **Explicit Enable Required**: The `request_headers` field requires `request_header_filters: enabled` * **Allowlist Required**: Every header name in the JSON object must be listed in `request_header_allowlist` * **Size Limit**: Header filter values are limited to 16 KiB (16,384 bytes) by default (configurable via `max_request_headers_length`) * **Authorization Blocked**: The `authorization` header cannot be allowlisted when HTTP authentication (Basic or OAuth2) is configured * **OR Across Columns Not Supported**: `OR` expressions that span different filter columns (e.g., `request_headers OR request_query`) are rejected. Use `UNION ALL` for cross-column alternatives. #### Partition Limits[​](#partition-limits "Direct link to Partition Limits") When multiple filter columns are used together with `AND`, the connector creates a cross product of all filter values. For example, 3 `request_path` values × 2 `request_headers` values = 6 HTTP requests. Use the `max_request_partitions` parameter to cap this cross product and prevent runaway request counts. #### Subquery Limitations[​](#subquery-limitations "Direct link to Subquery Limitations") * **`IN (SELECT ...)` only**: Subqueries against HTTP filter columns must use `IN (SELECT ...)`. `JOIN ... ON` with HTTP filter columns is not supported and returns an error. * **Build-side value cap**: The subquery (build side) is capped at 20,000 unique values. Values are deduplicated before creating HTTP requests. * **Partition limit applies**: The expanded partitions from subquery values are subject to `max_request_partitions`. If the cross product of existing partitions and subquery values exceeds the limit, the query fails with an error. ### Configuration Requirements[​](#configuration-requirements "Direct link to Configuration Requirements") To use the special metadata fields (`request_path`, `request_query`, `request_body`, `request_headers`), you must: 1. **For `request_path`**: Configure `allowed_request_paths` with a comma-separated list of allowed path patterns (supports glob patterns) 2. **For `request_query`**: Set `request_query_filters: enabled` in params 3. **For `request_body`**: Set `request_body_filters: enabled` in params 4. **For `request_headers`**: Set `request_header_filters: enabled` and `request_header_allowlist` in params Example minimal configuration for all four fields: ``` datasets: - from: https://api.example.com name: my_api params: allowed_request_paths: '/users,/posts,/comments,/api/**' request_query_filters: enabled request_body_filters: enabled request_header_filters: enabled request_header_allowlist: x-sandbox-id, x-region max_request_partitions: 10000 ``` ### Performance Considerations[​](#performance-considerations "Direct link to Performance Considerations") * **Connection Pooling**: The connector maintains up to 10 idle connections per host by default * **Retry Overhead**: With the default 3 retries and Fibonacci backoff, failed requests may take several seconds before returning an error * **Cache Behavior**: HTTP responses are cached based on the combination of path, query, body, and headers parameters * **Partition Limits**: Use `max_request_partitions` to cap the number of HTTP requests created from cross-product filters ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure an HTTP/HTTPS endpoint as a data connector in Spice. [HTTP Data Connector](https://github.com/spiceai/cookbook/tree/trunk/http#readme) --- # HTTP(s) Data Connector Deployment Guide Production operating guide for the HTTP(s) data connector covering authentication, rate control, retry tuning, and observability. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The connector supports HTTP Basic, custom-header, and OAuth2 (refresh-token and client-credentials grants) authentication. Secrets must be sourced from a [secret store](/docs/next/components/secret-stores) in production. | Parameter | Description | | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `http_username` | Username for HTTP Basic authentication. | | `http_password` | Password for HTTP Basic authentication. Use `${secrets:...}` to resolve from a secret store. | | `http_headers` | Custom headers (e.g. `Authorization:Bearer ${secrets:api_token}`). Treated as sensitive — not logged. Dynamic JSON API endpoints only; structured HTTP file datasets ignore these headers. | | `auth_token_url` | OAuth2 token endpoint URL (must be HTTPS in production). | | `auth_grant_type` | OAuth2 grant: `refresh_token` (default) or `client_credentials`. | | `http_auth_refresh_token` | OAuth2 refresh token. Required for the (default) refresh-token grant; unused by `client_credentials`. | | `http_auth_client_id` | OAuth2 client ID (required for confidential clients and for `client_credentials`). | | `http_auth_client_secret` | OAuth2 client secret (required for confidential clients and for `client_credentials`). Use `${secrets:...}`. | | `auth_header_name` | Header carrying the access token. Default `Authorization` (`Bearer `); any other name sends the bare token. | For OAuth2-protected APIs, prefer refresh-token flow over storing long-lived bearer tokens. The connector exchanges the refresh token for short-lived access tokens at startup and refreshes them before expiry. ### TLS[​](#tls "Direct link to TLS") Use HTTPS endpoints in production. `auth_token_url` must use HTTPS (loopback addresses are allowed for local testing only). Self-signed certificates require a trusted CA bundle in the container or host OS trust store. For upstream servers that require mutual TLS (mTLS), the connector can present a client certificate during the TLS handshake. Supply the certificate and key as file paths or inline PEM — the two forms are mutually exclusive, and the certificate and key must be set together. mTLS client identity applies to dynamic JSON API endpoints only. | Parameter | Description | | ---------------------------------- | ------------------------------------------------------------------------------------------- | | `http_tls_client_certificate_file` | Path to a PEM client certificate chain. Pair with `http_tls_client_key_file`. | | `http_tls_client_key_file` | Path to the PEM private key matching the client certificate file. | | `http_tls_client_certificate` | Inline PEM client certificate chain. Use `${secrets:...}`. Pair with `http_tls_client_key`. | | `http_tls_client_key` | Inline PEM private key matching the inline certificate. Use `${secrets:...}`. | ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Rate Control[​](#rate-control "Direct link to Rate Control") The HTTP connector participates in the shared HTTP rate control system. Concurrency and per-second/per-minute request limits can be configured per-dataset (in `params`) or globally (in `runtime.params`). Dataset-level settings override the global defaults. Multiple datasets targeting the same upstream origin share a single rate controller. | Parameter | Description | | --------------------------- | ------------------------------------------------------------------------------------- | | `max_concurrent_requests` | Maximum concurrent HTTP requests to the same origin. Disabled when unset. | | `requests_per_second_limit` | Maximum HTTP requests per second to the same origin. Disabled when unset. | | `requests_per_minute_limit` | Maximum HTTP requests per minute to the same origin. Disabled when unset. | | `rate_control_jitter_min` | Minimum random delay before requests when rate control is active. Defaults to `5ms`. | | `rate_control_jitter_max` | Maximum random delay before requests when rate control is active. Defaults to `10ms`. | The runtime equivalents (`http_max_concurrent_requests`, `http_requests_per_second_limit`, `http_requests_per_minute_limit`, `http_rate_control_jitter_min`, `http_rate_control_jitter_max`) set defaults that apply to every HTTP-based connector unless overridden per dataset. ``` runtime: params: http_max_concurrent_requests: 10 http_requests_per_second_limit: 5 datasets: - from: https://api.example.com/v1 name: api_data params: file_format: json allowed_request_paths: '/data/**' max_concurrent_requests: 3 # Override for this dataset requests_per_minute_limit: 60 ``` Use rate control when the upstream API enforces request quotas, when many datasets share a single origin, or when running large `IN`-list refreshes that would otherwise burst hundreds of concurrent requests. ### Retry Behavior[​](#retry-behavior "Direct link to Retry Behavior") HTTP-level retries follow the shared `resilient_http` policy: 408, 429, and 5xx responses plus transient network errors are retried. The connector respects `Retry-After`, `retry-after-ms`, and `x-retry-after-ms` headers. | Parameter | Default | Description | | ---------------------- | ----------- | ------------------------------------------------------------------------------------------------------------- | | `max_retries` | `3` | Maximum retry attempts per request. | | `retry_backoff_method` | `fibonacci` | Backoff strategy. Options: `fibonacci`, `linear`, `exponential`. | | `retry_max_duration` | unset | Maximum total duration across all retries (e.g. `30s`, `5m`). When set, retries stop after this elapsed time. | | `retry_jitter` | `0.3` | Randomization factor (`0.0`–`1.0`) applied to retry delays. Set to `0` to disable jitter. | Retries are independent of rate control. If a retry would exceed the configured per-second or per-minute rate, it waits for the rate window to open before issuing the request. ### Timeouts and Connection Pool[​](#timeouts-and-connection-pool "Direct link to Timeouts and Connection Pool") | Parameter | Default | Description | | ------------------------ | ------- | --------------------------------------------------------------------- | | `client_timeout` | `30` | Maximum time (seconds) to wait for the entire request-response cycle. | | `connect_timeout` | `10` | Maximum time (seconds) to establish a TCP/TLS connection. | | `pool_max_idle_per_host` | `10` | Maximum idle connections held per upstream host. | | `pool_idle_timeout` | `90` | Idle connection lifetime (seconds) before the pool closes them. | Increase `client_timeout` for endpoints with large response bodies or expensive server-side computation. Reduce `pool_max_idle_per_host` when running many small datasets against the same host to keep the runtime's open file descriptors bounded. ### Caching Mode[​](#caching-mode "Direct link to Caching Mode") When using `refresh_mode: caching`, transient HTTP errors (5xx, 429) are excluded from the cache and propagated to clients. Set `caching_stale_if_error: enabled` to serve expired cached data on upstream failure. Always set `caching_ttl` explicitly — the default of `30s` is rarely the desired window. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Throughput**: Bounded by the upstream rate limit, then by `max_concurrent_requests` and `connect_timeout`. Plan limits to stay within the API quota. * **Memory**: Response bodies are streamed; memory footprint is bounded by `max_request_body_bytes` (filter inputs) and DataFusion's record-batch size for response rows. * **Connection setup**: TLS handshake adds latency. The connection pool keeps `pool_max_idle_per_host` warm connections to absorb burst traffic. * **Partitioned refreshes**: When using `IN`-list filters or cross-product partitioning, the runtime issues one HTTP request per partition. Use `max_request_partitions` to cap the request count for unbounded filter combinations, and `max_concurrent_requests` to throttle their fan-out. ## Metrics[​](#metrics "Direct link to Metrics") The connector exposes per-origin HTTP rate-control metrics for every dynamic JSON API dataset. They are registered automatically — no `metrics` configuration is required — and the limit gauges report `0` when the corresponding limit is not configured. Structured file-format datasets (`parquet`, `csv`, and the other listing-table formats) do not expose them: | Metric Name | Type | Description | | ----------------------------------------- | ------- | ------------------------------------------------------------------------------------------------------ | | `inflight_operations` | Gauge | Current number of HTTP requests holding a rate-control permit. | | `rate_control_max_concurrent_requests` | Gauge | Configured maximum concurrent HTTP requests for this upstream origin; `0` means disabled. | | `rate_control_requests_per_second_limit` | Gauge | Configured HTTP request-per-second limit for this upstream origin; `0` means disabled. | | `rate_control_requests_per_minute_limit` | Gauge | Configured HTTP request-per-minute limit for this upstream origin; `0` means disabled. | | `rate_control_jitter_min_ms` | Gauge | Configured minimum rate-control jitter (ms) before HTTP requests. | | `rate_control_jitter_max_ms` | Gauge | Configured maximum rate-control jitter (ms) before HTTP requests. | | `rate_control_available_permits` | Gauge | Current available permits in the HTTP request concurrency semaphore; `0` when concurrency is disabled. | | `rate_control_acquisitions_total` | Counter | Total HTTP request rate-control permits acquired. | | `rate_control_acquire_errors_total` | Counter | Total HTTP request rate-control permit acquisition errors. | | `rate_control_wait_duration_ms` | Counter | Cumulative time (ms) spent waiting for HTTP rate-control permits, quotas, and jitter. | | `rate_limit_retry_after_updates_total` | Counter | Total upstream cooldown hints accepted from `Retry-After` or `RateLimit` reset headers. | | `rate_limit_retry_after_waits_total` | Counter | Total waits caused by `Retry-After` or `RateLimit` reset headers. | | `rate_limit_retry_after_wait_duration_ms` | Counter | Cumulative time (ms) spent waiting because of `Retry-After` or `RateLimit` reset headers. | | `rate_limit_retry_after_remaining_ms` | Gauge | Current remaining `Retry-After` / `RateLimit` cooldown (ms) for this upstream origin. | These metrics are auto-registered — no configuration is required to export them. To turn one off for a dataset, set `enabled: false` in the dataset's `metrics` section: ``` datasets: - from: https://api.example.com/v1 name: api_data params: file_format: json metrics: - name: rate_control_wait_duration_ms enabled: false ``` Instruments are exposed with the prefix `dataset_http_` — the HTTP connector's component name is `http`, not `https` — and each carries an `origin` attribute (`scheme://host:port`) identifying the upstream origin instead of a dataset `name`, because datasets sharing an origin share one rate controller. See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. For broader observability, also monitor: * Spice query execution metrics (`query_duration_ms`, `query_returned_rows`, `query_failures`) from `runtime.metrics`. ## Task History[​](#task-history "Direct link to Task History") HTTP requests participate in [task history](/docs/next/reference/task_history) through the HTTP client's span. Each partitioned request and each pagination page is a child of the enclosing `sql_query` or `accelerated_table_refresh` task. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Read-only**: The connector is read-only. Only `GET` and `POST` (via `request_body` filters) are supported. * **Filter pushdown is opt-in**: `request_path`, `request_query`, `request_body`, and `request_headers` filters require explicit allowlists or `_filters: enabled` parameters. * **OAuth2 OOS scope**: The refresh-token and client-credentials grants are supported. The authorization-code and device-code flows are not exposed. * **OR across virtual filter columns**: `WHERE request_path = '/a' OR request_query = 'b=1'` is rejected. Use separate datasets or `UNION ALL` for cross-column alternatives. Single-column `OR` (and `IN`-lists) is supported. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ------------------------------------------------- | --------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | | `401 Unauthorized` | Wrong/expired token or password. | Rotate the credential in the secret store. | | `429 Too Many Requests` (frequent) | Upstream rate limit hit; concurrency too high. | Set `requests_per_second_limit` / `requests_per_minute_limit`; reduce `max_concurrent_requests`. | | Refresh blocked / queue building up | `max_concurrent_requests` set too low for the workload. | Raise the dataset-level limit or move heavy datasets to their own origin. | | OAuth2 token refresh fails | `auth_token_url` not HTTPS, or wrong client credentials. | Verify the token endpoint URL; check `http_auth_client_id`/`secret` and required scopes. | | Request rejected: "OR across HTTP filter columns" | `WHERE request_path = '...' OR request_query = '...'`. | Split into separate refreshes or `UNION ALL`. | | Many partitions created from cross-product | Multiple `IN`-list filters multiplied into many requests. | Set `max_request_partitions` to cap; tighten filters. | | Slow first refresh | Cold connection pool + TLS handshake per request. | Raise `pool_max_idle_per_host`; ensure `pool_idle_timeout` is long enough to keep connections warm. | --- # Iceberg Data Connector The Iceberg Data Connector helps query [Apache Iceberg](https://iceberg.apache.org/) tables using federated SQL. Iceberg table format versions V1, V2, and V3 are supported. Every Iceberg dataset requires an Iceberg catalog to provide table metadata and manage access. When working with multiple datasets, it is recommended to use a catalog connector (instead of a data connector), such as the [Iceberg Catalog Connector](/docs/next/components/catalogs/iceberg) or [AWS Glue Catalog Connector](/docs/next/components/catalogs/glue) instead of configuring individual datasets. Iceberg catalogs can be of several types: * **Iceberg REST Catalog**: The most common and recommended approach. REST Catalogs expose Iceberg tables over HTTP(S) endpoints and are compatible with most managed Iceberg services and cloud providers. * **AWS Glue Catalog**: Integrates with AWS Glue as a catalog provider, supporting Iceberg tables stored in S3. This is the preferred method for AWS environments. * **Hadoop-style Catalogs**: Use file-based storage (e.g., `file://`, `s3://`, `s3a://`) to manage table metadata. This approach is typically used for local development or legacy deployments. Hadoop-style Catalogs For production and cloud environments, REST and AWS Glue catalogs are recommended. Hadoop-style catalogs are supported but less common and not recommended for most new deployments. ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: my_table ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field specifies the Iceberg table to connect to, in the format `iceberg:`. The `table_path` is the URL to the Iceberg table in the catalog provider. For REST Catalogs, use the format `http[s]:///v1/{prefix}/namespaces//tables/`. For AWS Glue catalogs, the URL format is `https://glue..amazonaws.com/iceberg/v1/catalogs//namespaces`, where `` is the AWS account ID. While possible to connect to Iceberg tables hosted by Glue using this generic connector, it is recommended to instead use the [AWS Glue Data Connector](/docs/next/components/data-connectors/glue) for connecting to Iceberg tables managed by Glue for a better experience. Example (REST Catalog): ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: my_table ``` Example (AWS Glue Catalog): ``` datasets: - from: iceberg:https://glue.us-east-1.amazonaws.com/iceberg/v1/catalogs/123456789012/namespaces/my_namespace/tables/my_table name: glue_table ``` Hadoop-style catalogs use file-based paths such as `file://`, `s3://`, or `s3a://`. For these, specify the warehouse path as the table location. This is typically only used for local development or legacy setups. Example (Hadoop Catalog, local): ``` datasets: - from: iceberg:file:///tmp/hadoop_warehouse/test/my_table_1 name: local_hadoop ``` Example (Hadoop Catalog, S3): ``` datasets: - from: iceberg:s3a://my-bucket/hadoop_warehouse/test/my_table_2 name: s3_hadoop ``` ### `name`[​](#name "Direct link to name") The `name` field sets the table name within Spice. This name is used to reference the dataset in SQL queries. The name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). Example: ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: transactions params: iceberg_token: ${secrets:iceberg_token} ``` ``` SELECT COUNT(*) FROM transactions; ``` ``` +----------+ | count(*) | +----------+ | 1234567 | +----------+ ``` ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `iceberg_token` | Bearer token value to use for Authorization header. | | `iceberg_oauth2_credential` | Credential to use for OAuth2 client credential flow when connecting to the table. Format: `:` | | `iceberg_oauth2_scope` | Scope to use for OAuth2 client credential flow when connecting to the table. Default: `catalog` | | `iceberg_oauth2_server_url` | URL of the OAuth2 server tokens endpoint for the client credential flow. | | `iceberg_s3_endpoint` | S3-compatible endpoint where the Iceberg table data is stored. | | `iceberg_s3_region` | Region of the S3-compatible endpoint. | | `iceberg_s3_access_key_id` | The AWS access key ID to use for S3 storage. If not provided, credentials will be loaded from environment variables or IAM roles. | | `iceberg_s3_secret_access_key` | The AWS secret access key to use for S3 storage. If not provided, credentials will be loaded from environment variables or IAM roles. | | `iceberg_s3_session_token` | Session token for the S3-compatible endpoint. | | `iceberg_s3_role_arn` | ARN of the IAM role to assume when accessing the S3-compatible endpoint. | | `iceberg_s3_role_session_name` | Session name to use when assuming the IAM role. | | `iceberg_s3_iam_role_source` | Optional. IAM role credential source. `auto` (default) uses the default AWS credential chain, `metadata` uses only instance/container metadata (IMDS, ECS, EKS/IRSA), `env` uses only environment variables. | | `iceberg_s3_connect_timeout` | Connection timeout in seconds for the S3-compatible endpoint. Default: `60`. **Note:** This parameter is currently accepted but has no effect — it is not consumed by any code path. | | `iceberg_s3_path_style_access` | Controls S3 addressing style. `true` (default) uses path-style (`endpoint/bucket`), required for object stores such as MinIO; `false` uses virtual-hosted-style (`bucket.endpoint`). | | `iceberg_sigv4_enabled` | Enable SigV4 (AWS Signature Version 4) authentication when connecting to the catalog. Automatically enabled if the URL in `from` is an AWS Glue catalog. Default: `false` | | `iceberg_signing_region` | Region to use for SigV4 authentication. Extracted from the URL in `from` if not specified. | | `iceberg_signing_name` | Service name to use for SigV4 authentication. Default: `glue`. | | `iceberg_warehouse` | Name of the Iceberg warehouse. Used for catalog types that support it (e.g. Lakekeeper). | | `iceberg_gcs_project_id` | The Google Cloud project ID for GCS storage. | | `iceberg_gcs_credentials` | Base64-encoded Google Cloud service account credentials JSON for GCS storage. | | `iceberg_gcs_token` | OAuth2 token to use for GCS authentication. | | `iceberg_gcs_service_path` | Custom endpoint URL for GCS (for emulators or custom endpoints). | | `iceberg_gcs_no_auth` | Set to `true` to allow anonymous access to GCS (for public buckets). | | `metadata_path` | The path including scheme to the metadata file for the Hadoop table. Must specify a path to a `.json` file. For example, `s3a://my-bucket/warehouse/namespace/table/metadata/v1.metadata.json` | ## Authentication[​](#authentication "Direct link to Authentication") Authentication to the Iceberg catalog. Supported methods include: * **Bearer Token**: Use `iceberg_token` for Authorization header. * **OAuth2 Client Credentials**: Use `iceberg_oauth2_credential`, `iceberg_oauth2_scope`, and `iceberg_oauth2_server_url`. * **AWS SigV4**: For AWS Glue, set `iceberg_sigv4_enabled: true` (or use a Glue URL). * **S3 Authentication**: Use `iceberg_s3_*` parameters for S3 data access. ### AWS Authentication[​](#aws-authentication "Direct link to AWS Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. These credentials will be used to connect to the S3 bucket as well as the Glue catalog (if configured). 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3/Glue permissions (e.g., `s3:ListBucket`, `s3:GetObject`) to access the tables. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ### Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access Iceberg tables in S3/Glue: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::company-bucketname-datasets" }, { "Effect": "Allow", "Action": ["s3:GetObject", "s3:PutObject"], "Resource": "arn:aws:s3:::company-bucketname-datasets/*" }, { "Effect": "Allow", "Action": [ "glue:GetCatalog", "glue:GetDatabases", "glue:GetDatabase", "glue:GetTable", "glue:GetTables" ], "Resource": "*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------- | -------------------------------------------------------------- | | `s3:ListBucket` | Required. Allows scanning all objects from the bucket | | `s3:GetObject` | Required. Allows fetching objects | | `s3:PutObject` | Required for write operations. Allows writing objects | | `glue:GetCatalog` | Required. Retrieve metadata about the specified catalog. | | `glue:GetDatabases` | Required. List the databases available in the current catalog. | | `glue:GetDatabase` | Required. Retrieve metadata about the specified database. | | `glue:GetTable` | Required. Retrieve metadata about the specified table. | | `glue:GetTables` | Required. List the tables available in the current database. | ## Write Support[​](#write-support "Direct link to Write Support") This connector supports writing data to Iceberg tables using SQL [`INSERT INTO`](/docs/next/reference/sql/dml#insert) statements. Writes are currently append-only — inserted data is added as new data files and registered through a new Iceberg table snapshot. Schema validation ensures inserted data matches the target table schema. To enable writes, set `access: read_write` on the dataset: ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: my_table access: read_write params: iceberg_token: ${secrets:iceberg_token} ``` ``` -- Insert with values INSERT INTO my_table (id, name, amount) VALUES (1, 'Alice', 100.0), (2, 'Bob', 200.0); -- Insert from another table INSERT INTO my_table SELECT * FROM source_table; ``` Inserting into partitioned Iceberg tables is supported. `DELETE FROM` is supported via equality delete files (Iceberg v2+ tables only). `UPDATE` operations are not currently supported. Write operations require `s3:PutObject` permission on the target S3 bucket in addition to the read permissions listed above. For more details, see [Data Ingestion](/docs/next/features/data-ingestion). ## Examples[​](#examples "Direct link to Examples") ### Basic Example (REST Catalog)[​](#basic-example-rest-catalog "Direct link to Basic Example (REST Catalog)") Connect to an Iceberg table with token authentication: ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: my_table params: iceberg_token: ${secrets:iceberg_token} ``` ### AWS Glue Catalog Example[​](#aws-glue-catalog-example "Direct link to AWS Glue Catalog Example") Connect to an Iceberg table in AWS Glue catalog: ``` datasets: - from: iceberg:https://glue.us-east-1.amazonaws.com/iceberg/v1/catalogs/123456789012/namespaces/my_namespace/tables/my_table name: glue_table params: iceberg_sigv4_enabled: true ``` ### OAuth2 Authentication Example[​](#oauth2-authentication-example "Direct link to OAuth2 Authentication Example") Connect to an Iceberg table using OAuth2 authentication: ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: oauth_table params: iceberg_oauth2_credential: ${secrets:client_id}:${secrets:client_secret} iceberg_oauth2_scope: catalog iceberg_oauth2_server_url: https://iceberg-catalog-host.com/oauth2/token ``` ### S3 Storage Example[​](#s3-storage-example "Direct link to S3 Storage Example") Connect to an Iceberg table with custom S3 storage configuration: ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: s3_table params: iceberg_token: ${secrets:iceberg_token} iceberg_s3_endpoint: http://localhost:9000 iceberg_s3_region: us-west-2 iceberg_s3_access_key_id: ${secrets:aws_access_key_id} iceberg_s3_secret_access_key: ${secrets:aws_secret_access_key} ``` ### Hadoop Catalog Example[​](#hadoop-catalog-example "Direct link to Hadoop Catalog Example") Connect to an Iceberg table using Hadoop Catalog with a local warehouse: ``` datasets: - from: iceberg:file:///tmp/hadoop_warehouse/test/my_table_1 name: local_hadoop params: metadata_path: file:///tmp/hadoop_warehouse/test/my_table_1/metadata/v1.metadata.json ``` Connect to an Iceberg table using Hadoop Catalog with S3: ``` datasets: - from: iceberg:s3a://my-bucket/hadoop_warehouse/test/my_table_2 name: s3_hadoop params: metadata_path: s3a://my-bucket/hadoop_warehouse/test/my_table_2/metadata/v1.metadata.json ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") A cookbook recipe to configure Iceberg as a catalog connector in Spice, including write operations: [Iceberg Catalog Connector](https://github.com/spiceai/cookbook/tree/trunk/catalogs/iceberg#readme). ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Limitations[​](#limitations "Direct link to Limitations") Performance Considerations When querying Iceberg tables, performance depends on the size of the table, the complexity of the query, and the underlying storage system. For large tables, consider using appropriate filtering to limit the amount of data scanned. The connector needs to access both the Iceberg catalog metadata and the underlying data files (typically stored in S3 or a compatible object store). Ensure proper network connectivity and authentication for both systems. --- # IMAP Data Connector The IMAP Data Connector enables federated SQL query across emails stored in an IMAP email server. ``` datasets: - from: imap:myawesomeemail@example.com name: emails params: imap_password: ${secrets:IMAP_PASSWORD} ``` ## Schema[​](#schema "Direct link to Schema") | Field Name | Data Type | Nullable | Description | | ------------- | ------------------------------ | -------- | -------------------------------------------------------------------------------- | | `date` | `Timestamp(Millisecond, None)` | No | The date and time when the email was sent, as milliseconds since the Unix epoch. | | `subject` | `Utf8` | Yes | The subject line of the email. | | `from` | `List` | Yes | The sender(s) of the email. | | `to` | `List` | Yes | The primary recipient(s) of the email. | | `cc` | `List` | Yes | The carbon copy recipient(s) of the email. | | `bcc` | `List` | Yes | The blind carbon copy recipient(s) of the email. | | `reply_to` | `List` | Yes | The email address(es) to which replies should be sent. | | `message_id` | `Utf8` | Yes | A unique identifier for the email message. | | `in_reply_to` | `Utf8` | Yes | The `message_id` of the email this message is replying to, if applicable. | | `content` | `Utf8` | Yes | The raw email body of this message. Not retrieved when acceleration is disabled. | If a MIME-encoded value is retrieved for a field, it is not decoded and the MIME-encoded value is returned in SQL queries. Most fields are optional, and depend on the implementation of the specific IMAP server being connected to. For example, the IMAP RFC specifies the `message_id` field SHOULD be supplied but the field is an optional field. For more information, refer to the [IMAP RFC 2822 - Section 3.6](https://www.rfc-editor.org/rfc/rfc2822#section-3.6). ## Retrieving email body contents[​](#retrieving-email-body-contents "Direct link to Retrieving email body contents") When the IMAP Data Connector is used without acceleration, the email body will not be retrieved - only header/subject values. To load the email body contents, specify an acceleration: ``` datasets: - from: imap:myawesomeemail@example.com name: emails params: imap_password: ${secrets:IMAP_PASSWORD} acceleration: enabled: true ``` With an acceleration enabled, the `content` field will be populated with the complete email body including headers, without any decoding applied. This field could be used for post-processing the email, like retrieving custom header values or decoding MIME-encoded content. Limitations * Email attachments are currently not parsed from the email body into separate dataset fields. To read email attachments, parse the multipart encodings from the `content` field. ## Query performance[​](#query-performance "Direct link to Query performance") Filters on the `date` column are pushed down to the IMAP server as a `SEARCH` command, so a query or refresh transfers only the matching messages instead of the whole mailbox. * **Supported filter shapes.** A comparison of `date` against a timestamp or date literal (`>`, `>=`, `<`, `<=`, `=`, with the column on either side) and conjunctions (`AND`) of those. Lower bounds become `SENTSINCE`, upper bounds `SENTBEFORE`. * **Everything else scans the full mailbox.** A filter on any other column, a disjunction (`OR`), `!=`, or a non-temporal literal contributes no `SEARCH` criteria, and the connector fetches every message and filters locally, exactly as before. * **Server-side matching is a superset.** IMAP `SEARCH` compares the calendar day of the `Date:` header, disregarding time and timezone ([RFC 3501 §6.4.4](https://www.rfc-editor.org/rfc/rfc3501#section-6.4.4)), so each bound is widened by a day and Spice re-applies the predicate exactly on the returned rows. Results are identical to an unpushed scan; only the bytes transferred change. * **Very large match sets fall back to a full fetch.** When the set of matching messages is too large to express compactly as an IMAP identifier set, the connector fetches the whole mailbox instead. The filters are applied above the scan either way, so results are unaffected. To make an [acceleration](/docs/next/features/data-acceleration) refresh incremental rather than re-reading the mailbox each cycle, set `time_column: date` so that `refresh_mode: append` and [`refresh_data_window`](/docs/next/reference/spicepod/datasets#accelerationrefresh_data_window) generate a `date` predicate for the connector to push down: ``` datasets: - from: imap:jsmith@example.com name: emails time_column: date params: imap_password: ${secrets:IMAP_PASSWORD} acceleration: enabled: true refresh_mode: append refresh_check_interval: 10m ``` Scans also only request the raw message body when it is needed to populate `content` — that is, when acceleration is enabled (see [Retrieving email body contents](#retrieving-email-body-contents)). A dataset without acceleration transfers envelopes only, and never the MIME parts and attachments of every message. Every body section is requested with `.PEEK`, so querying a dataset never marks messages as read (`\Seen`) on the server. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field must contain the email address for the mailbox to connect to. For example, `me@outlook.com`, or `jsmith@example.com`. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: imap:jsmith@example.com name: emails params: ... ``` ``` SELECT COUNT(*) FROM emails; ``` ``` +----------+ | count(*) | +----------+ | 1234 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The IMAP connector supports the following connection and authentication parameters: | Parameter Name | Description | | --------------- | ---------------------------------------------------------------------------------------------------------------------------- | | `imap_username` | Optional. The username to use for the IMAP connection. Defaults to the value of the `from:` mailbox field. | | `imap_password` | Required. The password to use for the IMAP connection, in plaintext authentication mode. | | `imap_host` | Optional. The host or IP address of the IMAP server to connect to. Not required for known connections like Outlook or Gmail. | | `imap_port` | Optional. The port of the IMAP server to connect to. Defaults to `993`. | | `imap_mailbox` | Optional. The mailbox to read mail from. Defaults to `INBOX`, the standard email inbox. | | `imap_ssl_mode` | Optional. The IMAP SSL mode to use. Defaults to `auto`, permitted values of `tls`, `starttls`, `disabled` or `auto`. | ## Examples[​](#examples "Direct link to Examples") ### Basic example[​](#basic-example "Direct link to Basic example") ``` datasets: - from: imap:jsmith@example.com name: emails params: imap_host: mail.example.com imap_password: ${ secrets:IMAP_PASSWORD } ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure IMAP as a data connector in Spice. [IMAP Data Connector](https://github.com/spiceai/cookbook/tree/trunk/imap/#readme) --- # Kafka Data Connector The Kafka Data Connector enables direct acceleration of data from [Apache Kafka](https://kafka.apache.org/) topics using `refresh_mode: append` [acceleration](/docs/next/components/data-accelerators). This provides direct integration with existing Kafka-based event streaming infrastructure for real-time data acceleration and analytics. ``` datasets: - from: kafka:my_kafka_topic name: my_dataset params: kafka_bootstrap_servers: broker1:9092,broker2:9092,broker3:9092 # Required. A comma separated list of Kafka broker servers. kafka_security_protocol: sasl_ssl # Default is `sasl_ssl`. Valid values are `plaintext`, `ssl`, `sasl_plaintext`, `sasl_ssl`. kafka_sasl_mechanism: SCRAM-SHA-512 # Default is `SCRAM-SHA-512`. Valid values are `PLAIN`, `SCRAM-SHA-256`, `SCRAM-SHA-512`. kafka_sasl_username: kafka # Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. kafka_sasl_password: ${secrets:kafka_sasl_password} # Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. kafka_ssl_ca_location: ./certs/kafka_ca_cert.pem # Optional. Used to verify the SSL/TLS certificate of the Kafka broker. kafka_ssl_certificate_location: ./certs/client_cert.pem # Optional. Client SSL/TLS certificate for mTLS authentication. kafka_ssl_key_location: ./certs/client_key.pem # Optional. Client SSL/TLS private key for mTLS authentication. kafka_ssl_key_password: ${secrets:kafka_ssl_key_password} # Optional. Password for the client SSL/TLS private key, if encrypted. kafka_enable_ssl_certificate_verification: true # Default is `true`. Set to `false` to disable SSL/TLS certificate verification. kafka_ssl_endpoint_identification_algorithm: https # Default is `https`. Valid values are `none` and `https`. batch_max_size: 100000 # Default is `10000`. Maximum number of change events to batch together before processing. batch_max_duration: 1s # Default is `1s`. Maximum time to wait for a batch to fill before processing. acceleration: enabled: true # Acceleration is required for the kafka connector. engine: duckdb # `duckdb`, `sqlite` and `postgres` are supported acceleration engines for Kafka. refresh_mode: append # Required. Must be set to `append` for the Kafka connector. mode: file # Persistence is recommended to not have to fully rebuild the table each time Spice starts. ``` ## Overview[​](#overview "Direct link to Overview") Upon startup, Spice subscribes to the specified topic using either a uniquely generated consumer group or a custom one specified via `kafka_consumer_group_id`. If a persistent acceleration engine is used (with `mode: file`), data is fetched starting from the last processed record, so Spice can resume without reprocessing all historical data. Schema is automatically inferred from the first available topic message in JSON format. The connector creates the appropriate table schema for acceleration based on the detected data structure. ## Consumer Group Management[​](#consumer-group-management "Direct link to Consumer Group Management") The Kafka connector manages consumer groups to ensure data consistency across restarts. Offsets are committed to Kafka, so Spice can track consumption progress. **Default behavior:** When no `kafka_consumer_group_id` is specified, Spice automatically generates a unique consumer group ID and stores it in the acceleration metadata. On subsequent restarts, Spice retrieves and reuses this stored consumer group ID to maintain offset tracking and resume consumption from where it left off. **Custom consumer group:** If you specify a custom `kafka_consumer_group_id`, Spice stores this ID in the acceleration metadata. The same consumer group must be used on subsequent restarts. If no acceleration data exists and a custom consumer group is provided, Spice will reset its position to the oldest available offset and begin consuming from the start of the topic. **Consumer group mismatch error:** Spice will return an error if a restart is attempted with a different consumer group than what is stored in the acceleration metadata. This applies to both auto-generated and custom consumer group IDs. This safeguard prevents data inconsistency that could occur from mixing offsets between different consumer groups. To resolve a consumer group mismatch, either: * Use the same consumer group ID as stored in the acceleration * Reset the acceleration data to start fresh with a new consumer group ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form of `kafka:kafka_topic` where `kafka_topic` is the name of the Kafka topic to consume from. ``` datasets: - from: kafka:user_events name: events ... ``` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. ``` datasets: - from: kafka:orders_events name: orders ... ``` ``` SELECT COUNT(*) FROM orders; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `kafka_bootstrap_servers` | **Required**. A list of host/port pairs for establishing the initial Kafka cluster connection. The client will use all servers, regardless of the bootstrapping servers specified here. This list only affects the initial hosts used to discover the full server set and should be formatted as `host1:port1,host2:port2,...`. | | `kafka_security_protocol` | Security protocol for Kafka connections. Default: `sasl_ssl`. Options: - `plaintext`
- `ssl`
- `sasl_plaintext`
- `sasl_ssl` | | `kafka_sasl_mechanism` | SASL (Simple Authentication and Security Layer) authentication mechanism. Default: `SCRAM-SHA-512`. Options: - `PLAIN`
- `SCRAM-SHA-256`
- `SCRAM-SHA-512` | | `kafka_sasl_username` | SASL username. Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. | | `kafka_sasl_password` | SASL password. Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. | | `kafka_ssl_ca_location` | Path to the SSL/TLS CA certificate file for server verification. | | `kafka_ssl_certificate_location` | Path to the client SSL/TLS certificate file for mTLS authentication. | | `kafka_ssl_key_location` | Path to the client SSL/TLS private key file for mTLS authentication. | | `kafka_ssl_key_password` | Password for the client SSL/TLS private key, if encrypted. | | `kafka_enable_ssl_certificate_verification` | Enable SSL/TLS certificate verification. Default: `true`. | | `kafka_ssl_endpoint_identification_algorithm` | SSL/TLS endpoint identification algorithm. Default: `https`. Options: - `none`
- `https` | | `kafka_consumer_group_id` | Kafka consumer group id to use. If not set, a unique id will be generated. | | `schema_infer_max_records` | Number of Kafka messages to sample for schema inference. Default: `1`. Increase if your data has optional fields or varying structure. | | `flatten_json` | Set `true` to flatten nested structs in JSON as separate columns. | | `batch_max_size` | Maximum number of change events to batch together before processing. Default: `10000`. | | `batch_max_duration` | Maximum time to wait for a batch to fill before processing. Default: `1s`. | ### `metrics`[​](#metrics "Direct link to metrics") The connector supports the following optional [component metrics](/docs/next/features/observability/component_metrics): | Metric Name | Type | Description | | ------------------------ | ------- | ------------------------------------------------------------------------------------ | | `bytes_consumed_total` | Counter | Total number of bytes consumed from the Kafka topic | | `records_consumed_total` | Counter | Total number of records (messages) consumed from Kafka topics | | `records_lag` | Gauge | Total consumer lag across all topic partitions (number of messages not yet consumed) | These metrics are not enabled by default, enable them by setting the `metrics` parameter: ``` datasets: - from: kafka:user_events name: events metrics: - name: records_lag - name: records_consumed_total - name: bytes_consumed_total params: ... ``` ### Acceleration Settings[​](#acceleration-settings "Direct link to Acceleration Settings") warning Using the Kafka connector **requires** [acceleration](/docs/next/components/data-accelerators) with `refresh_mode: append` enabled. The following settings are required: | Parameter Name | Description | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `enabled` | Required. Must be set to `true` to enable acceleration. | | `engine` | Required. The acceleration engine to use. Possible valid values: - `duckdb`: Use [DuckDB](/docs/next/components/data-accelerators/duckdb) as the acceleration engine.
- `sqlite`: Use [SQLite](/docs/next/components/data-accelerators/sqlite) as the acceleration engine.
- `postgres`: Use [PostgreSQL](/docs/next/components/data-accelerators/postgres) as the acceleration engine. | | `refresh_mode` | Required. The refresh mode to use. Must be set to `append` for the Kafka connector. | | `mode` | Optional. The persistence mode to use. When using the `duckdb` and `sqlite` engines, it is recommended to set this to `file` to persist the data across restarts. Spice persists metadata about the dataset, so it can resume from the last known state instead of re-processing all messages. | ## Data Format Support[​](#data-format-support "Direct link to Data Format Support") The Kafka connector currently supports JSON-formatted messages. Schema is automatically inferred from the first available message in the topic, and all subsequent messages are expected to follow a compatible structure. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * See how to query Kafka real-time data with other datasets using federated queries in [Live Orders Analytics example](https://github.com/spiceai/cookbook/blob/trunk/kafka/README.md). --- # Localpod Data Connector The Localpod Data Connector enables setting up a parent/child relationship between datasets in the current Spicepod. This can be used for configuring multiple/tiered accelerations for a single dataset, and ensuring that the data is only downloaded once from the remote source. For example, you can use the `localpod` connector to create a child dataset that is accelerated in-memory, while the parent dataset is accelerated to a file. The dataset created by the `localpod` connector will logically have the same data as the parent dataset. ## Synchronized Refreshes[​](#synchronized-refreshes "Direct link to Synchronized Refreshes") The `localpod` connector supports synchronized refreshes, which ensures that the child dataset is refreshed from the same data as the parent dataset. Synchronized refreshes require that both the parent and child datasets are accelerated with `refresh_mode: full` (which is the default) or `refresh_mode: caching`. When synchronization is enabled, the following logs will be emitted: ``` 2024-10-28T15:45:24.220665Z INFO runtime::datafusion: Localpod dataset test_local synchronizing refreshes with parent table test ``` ### Examples[​](#examples "Direct link to Examples") ``` datasets: - from: postgres:cleaned_sales_data name: test params: ... acceleration: enabled: true # This dataset will be accelerated into a DuckDB file engine: duckdb mode: file refresh_check_interval: 10s - from: localpod:test name: test_local acceleration: enabled: true # This dataset accelerates the parent `test` dataset into in-memory Arrow records and is synchronized with the parent ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Localpod as a data connector in Spice. [Local dataset replication (Localpod)](https://github.com/spiceai/cookbook/tree/trunk/localpod#readme) --- # Memory Data Connector The Memory Data Connector enables configuring an in-memory dataset for tables used, or produced by the Spice runtime. Only certain tables, with predefined schemas, can be defined by the connector. These are: * `store`: Defines a table that LLMs, with [memory tooling](/docs/next/features/large-language-models/memory), can store data in. Requires `access: read_write`. ### Examples[​](#examples "Direct link to Examples") ``` datasets: - from: memory:store name: llm_memory access: read_write columns: - name: value embeddings: # Easily make your LLM learnings searchable. - from: all-MiniLM-L6-v2 embeddings: - name: all-MiniLM-L6-v2 from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to provide persistent memory capabilities for language models in Spice. [LLM Memory](https://github.com/spiceai/cookbook/tree/trunk/llm-memory#readme) --- # MongoDB Data Connector MongoDB is an open-source NoSQL database that stores data in flexible, JSON-like documents, supporting dynamic schemas and easy scalability. The MongoDB Data Connector enables federated/accelerated SQL queries on data stored in MongoDB databases. ``` datasets: - from: mongodb:mytable name: my_dataset params: mongodb_host: localhost mongodb_port: 27017 mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} mongodb_pool_min: 1 mongodb_pool_max: 10 ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `mongodb:{table_name}` where `table_name` is the table identifer in the MongoDB server to read from. info Unquoted identifiers are normalized to lowercase. To reference a collection with mixed-case characters, wrap it in double quotes: `mongodb:"MixedCaseCollection"`. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). ``` datasets: - from: mongodb:mytable name: my_dataset params: mongodb_db: my_database ... ``` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: mongodb:my_dataset name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords) ### `params`[​](#params "Direct link to params") The MongoDB data connector can be configured by providing the following `params`. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:my_mongodb_conn_string}`. | Parameter Name | Description | | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mongodb_connection_string` | The connection string to use to connect to the MongoDB server. This can be used instead of providing individual connection parameters. | | `mongodb_user` | The MongoDB username. | | `mongodb_pass` | The password to connect with. | | `mongodb_host` | The hostname of the MongoDB server. Defaults to `localhost`. | | `mongodb_port` | The port of the MongoDB server. Defaults to `27017`. | | `mongodb_db` | The name of the database to connect to. Defaults to `default`. | | `mongodb_sslmode` | Optional. Specifies the SSL/TLS behavior for the connection, supported values:
- `required`: (default) This mode requires an SSL connection. If a secure connection cannot be established, server will not connect.
- `preferred`: Establishes an encrypted TLS/SSL connection but does not validate the server certificate or hostname (accepts invalid or self-signed certificates). If the server does not support TLS, the connection fails; it is never downgraded to plaintext.
- `disabled`: This mode will not attempt to use an SSL connection, even if the server supports it. | | `mongodb_sslrootcert` | Optional parameter specifying the path to a custom PEM certificate that the connector will trust. | | `mongodb_time_zone` | Optional. Specifies connection time zone. Default is `UTC`. Accepts:
- Fixed offsets (e.g., `+02:00`).
- IANA time zone names (e.g., `America/Los_Angeles`) | | `mongodb_auth_source` | Optional. Authentication source database. Overrides the default auth source in the connection string. | | `mongodb_direct_connection` | Optional. Whether to connect directly to a single MongoDB host instead of discovering the topology. Accepts `true` or `false`. | | `mongodb_srv` | Optional. Use the `mongodb+srv://` connection scheme for DNS SRV record discovery (MongoDB Atlas). Auto-detected (and `mongodb_port` is ignored) when `mongodb_host` ends with `.mongodb.net`. Accepts `true` or `false`. Default: `false`. | | `mongodb_unnest_depth` | Optional. Maximum nesting depth for unnesting embedded documents into a flattened structure. Higher values expand deeper nested fields. Default: `0` | | `mongodb_schema_infer_max_records` | Optional. Number of documents to use to infer the schema. Defaults to 400. | | `mongodb_num_docs_to_infer_schema` | Optional. **Deprecated** — use `mongodb_schema_infer_max_records` instead. Number of documents to use to infer the schema. Defaults to 400. If both are set, `mongodb_schema_infer_max_records` takes precedence. | | `mongodb_pool_min` | The minimum number of connections to keep open in the pool, lazily created when requested. Default: `1` | | `mongodb_pool_max` | The maximum number of connections to allow in the pool. Default: `5` | | `mongodb_resume_token_invalid_behavior` | Optional. Used with `refresh_mode: changes`. Behavior when a persisted Change Stream resume token is rejected by the server (e.g. it is past the oplog window). `error` (default) surfaces a clear error; `rebootstrap` drops the persisted token and re-snapshots the collection. See [Using MongoDB Change Streams](#using-mongodb-change-streams). | | `change_stream_batch_max_size` | Optional. Used with `refresh_mode: changes`. Maximum number of Change Stream events grouped into one CDC batch before applying it. Default: `1000`. | | `change_stream_batch_max_duration` | Optional. Used with `refresh_mode: changes`. Maximum time to wait for a Change Stream batch to fill before applying it. Accepts [fundu](https://docs.rs/fundu) duration strings. Default: `1s`. | | `change_stream_max_await_time` | Optional. Used with `refresh_mode: changes`. Maximum time MongoDB waits for new Change Stream events before returning an empty server batch. Accepts [fundu](https://docs.rs/fundu) duration strings. Default: `1s`. | | `change_stream_batch_size` | Optional. Used with `refresh_mode: changes`. Number of Change Stream events MongoDB should request from the server per batch. Default: `1000`. | ## Types[​](#types "Direct link to Types") The table below shows the MongoDB data types supported, along with the type mapping to Apache Arrow types in Spice. | MongoDB Type | Arrow Type | | ------------------------- | ------------------------------------ | | `String` | `Utf8` | | `Boolean` | `Boolean` | | `Int32` | `Int32` | | `Int64` | `Int64` | | `Double` | `Float64` | | `Decimal128` | `Decimal128` | | `Binary` | `Binary` | | `Datetime` without time | `Date32` | | `Datetime` with time | `Timestamp(Millisecond, )` | | `Timestamp` | `Timestamp(Millisecond, None)` | | `Array` | `List` | | `Null` | `Null` | | `Undefined` | `Null` | | `RegularExpression` | `Utf8` | | `JavaScriptCode` | `Utf8` | | `JavaScriptCodeWithScope` | `Utf8` | | `Symbol` | `Utf8` | | `MaxKey` | `Utf8` | | `MinKey` | `Utf8` | | `DbPointer` | `Utf8` | | `ObjectId` | `Utf8` | | `Document` | See unnesting section | note * The MongoDB `Datetime` value is [retrieved as a UTC time value](https://www.mongodb.com/docs/manual/reference/method/Date/) by default. Use the `mongodb_time_zone` configuration parameter to specify the desired time zone for interpreting `TIMESTAMP` values during data retrieval. ## Unnesting[​](#unnesting "Direct link to Unnesting") Consider the following document: ``` { "a": 1, "b": { "x": 2, "y": { "z": 3 } } } ``` Using `mongodb_unnest_depth` you can control the unnesting behavior. Here are the examples: ### mongodb\_unnest\_depth: 0[​](#mongodb_unnest_depth-0 "Direct link to mongodb_unnest_depth: 0") ``` sql> select * from test_table; +-----------+---------------------+ | a (Int32) | b (Utf8) | +-----------+---------------------+ | 1 | {"x":2,"y":{"z":3}} | +---+-----------------------------+ ``` ### mongodb\_unnest\_depth: 1[​](#mongodb_unnest_depth-1 "Direct link to mongodb_unnest_depth: 1") ``` sql> select * from test_table; +-----------+-------------+------------+ | a (Int32) | b.x (Int32) | b.y (Utf8) | +-----------+-------------+------------+ | 1 | 2 | {"z":3} | +-----------+-------------+------------+ ``` ### mongodb\_unnest\_depth: 2[​](#mongodb_unnest_depth-2 "Direct link to mongodb_unnest_depth: 2") ``` sql> select * from test_table; +-----------+-------------+---------------+ | a (Int32) | b.x (Int32) | b.y.z (Int32) | +-----------+-------------+---------------+ | 1 | 2 | 3 | +-----------+-------------+---------------+ ``` ## JSON Nesting[​](#json-nesting "Direct link to JSON Nesting") When a MongoDB collection has many fields but you only need a few as discrete columns, you can consolidate the rest into a single JSON column using the `json_object` metadata option. Declare the fields you want as top-level columns explicitly in the `columns` list, then add a "catch-all" column with `json_object: "*"` metadata — every field not otherwise listed is nested into it as a JSON object. This applies to both the query (scan) path and the [Change Streams](#using-mongodb-change-streams) (CDC) path. ``` datasets: - from: mongodb:orders name: orders params: mongodb_host: localhost mongodb_db: shop columns: - name: _id - name: status - name: data_json metadata: json_object: '*' ``` Any field other than `_id`, `status`, and `data_json` is folded into `data_json` as a JSON object. The `_id` field must be declared explicitly and cannot be folded into the catch-all column. Limitations * The `json_object` metadata only accepts `"*"` as its value, which captures all unspecified fields. * Only one column can carry the `json_object` metadata; declaring more than one is an error. ## Examples[​](#examples "Direct link to Examples") ### Connecting using username and password and custom auth table[​](#connecting-using-username-and-password-and-custom-auth-table "Direct link to Connecting using username and password and custom auth table") ``` datasets: - from: mongodb:my_dataset name: my_dataset params: mongodb_host: localhost mongodb_port: 27017 mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} mongodb_auth_source: admin ``` ### Connecting using SSL[​](#connecting-using-ssl "Direct link to Connecting using SSL") ``` datasets: - from: mongodb:my_dataset name: my_dataset params: mongodb_host: localhost mongodb_port: 27017 mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} mongodb_sslmode: preferred mongodb_sslrootcert: ./custom_cert.pem ``` ### Connecting using a Connection String[​](#connecting-using-a-connection-string "Direct link to Connecting using a Connection String") ``` datasets: - from: mongodb:my_dataset name: my_dataset params: mongodb_connection_string: mongodb://${secrets:my_user}:${secrets:my_password}@localhost:27017/my_db?authSource=admin ``` ### Connecting to MongoDB Atlas (SRV)[​](#connecting-to-mongodb-atlas-srv "Direct link to Connecting to MongoDB Atlas (SRV)") When `mongodb_host` ends with `.mongodb.net`, `mongodb_srv` is automatically enabled and `mongodb_port` is ignored — `mongodb+srv://` discovers host/port via DNS SRV records. ``` datasets: - from: mongodb:my_collection name: my_dataset params: mongodb_host: cluster0.abc123.mongodb.net mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} ``` Set `mongodb_srv: true` explicitly for non-Atlas hosts that are configured with SRV records: ``` datasets: - from: mongodb:my_collection name: my_dataset params: mongodb_srv: true mongodb_host: mongo.example.com mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} ``` ### With custom connection pool settings[​](#with-custom-connection-pool-settings "Direct link to With custom connection pool settings") ``` datasets: - from: mongodb:my_dataset name: my_dataset params: mongodb_host: localhost mongodb_port: 27017 mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} mongodb_pool_min: 5 mongodb_pool_max: 10 ``` ### Using MongoDB Change Streams[​](#using-mongodb-change-streams "Direct link to Using MongoDB Change Streams") Spice supports real-time Change Data Capture (CDC) from MongoDB using native [MongoDB Change Streams](https://www.mongodb.com/docs/manual/changeStreams/). This streams inserts, updates, replacements, deletes, and collection-level invalidation events from MongoDB collections directly into Spice accelerators. #### How it works[​](#how-it-works "Direct link to How it works") On startup, Spice opens a Change Stream on the source collection (`fullDocument=updateLookup`), emits a CDC `TRUNCATE`, applies a full snapshot of the collection as upsert rows, signals readiness, then processes Change Stream events in batches. Opening the Change Stream before the snapshot prevents gaps between the snapshot and the live stream. File-accelerated datasets persist resume tokens and resume from the last committed token on restart. In-memory accelerators re-bootstrap from a fresh snapshot. #### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * MongoDB 4.0+ with Change Streams enabled. MongoDB requires a replica set or sharded cluster for Change Streams. * The MongoDB user must have `changeStream` privileges. * The accelerator must support upsert behavior. Use `duckdb`, `sqlite`, `postgres`, `turso`, or `cayenne`. * `acceleration.primary_key: _id` is required. Delete events only include the document key, so Spice needs `_id` to route deletes. * `acceleration.on_conflict` must specify `upsert` on `_id` so update and replace events overwrite existing rows. #### Minimal configuration[​](#minimal-configuration "Direct link to Minimal configuration") ``` datasets: - from: mongodb:users name: users params: mongodb_host: localhost mongodb_port: '27017' mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} acceleration: enabled: true engine: duckdb refresh_mode: changes primary_key: _id on_conflict: _id: upsert ``` #### Change Stream parameters[​](#change-stream-parameters "Direct link to Change Stream parameters") These optional runtime parameters live under dataset `params:`. The first four are not prefixed with `mongodb_`. | Parameter Name | Default | Description | | --------------------------------------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `change_stream_batch_max_size` | `1000` | Maximum number of Change Stream events to group into one CDC batch before applying it. Must be greater than 0. | | `change_stream_batch_max_duration` | `1s` | Maximum time to wait for a Change Stream batch to fill before applying it. Accepts [fundu](https://docs.rs/fundu) duration strings; must be greater than 0. | | `change_stream_max_await_time` | `1s` | Maximum time MongoDB waits for new Change Stream events before returning an empty server batch. Accepts [fundu](https://docs.rs/fundu) duration strings; must be greater than 0. | | `change_stream_batch_size` | `1000` | Number of Change Stream events MongoDB should request from the server per batch. Must fit in a `u32` and be greater than 0. | | `mongodb_resume_token_invalid_behavior` | `error` | Behavior when a persisted Change Stream resume token cannot be honored by the server (e.g. the token is past the oplog retention window). `error` surfaces a clear error so the operator can decide; `rebootstrap` drops the persisted token and re-snapshots the collection. | The existing `mongodb_unnest_depth` parameter also applies to Change Stream documents, so nested BSON is flattened the same way as normal MongoDB reads. #### Event mapping[​](#event-mapping "Direct link to Event mapping") * `insert`: create/upsert, using `fullDocument`. * `update`: update/upsert, using `fullDocument` from `fullDocument=updateLookup`. * `replace`: update/upsert, using `fullDocument`. * `delete`: delete, using `documentKey`; non-key columns are `null`. * `drop`, `rename`, `dropDatabase`, `invalidate`: truncate, because collection continuity is no longer guaranteed. If MongoDB omits `fullDocument` for an `update` or `replace` event (for example, the document was deleted before the `fullDocument=updateLookup` post-image could be read), Spice substitutes a synthetic delete keyed on `documentKey` and logs a warning; if `documentKey` is also unavailable, the event is skipped with a warning. An `insert` event missing `fullDocument`, or a `delete` event missing `documentKey`, still fails the stream. #### Resumability across restarts[​](#resumability-across-restarts "Direct link to Resumability across restarts") For file-accelerated datasets (acceleration `mode: file` / `file_create` / `file_update`, or `engine: postgres`), Spice persists the most recent Change Stream resume token in a sidecar table named `spice_sys_mongodb`, stored alongside the accelerator data. The token is committed only after the downstream accelerator write succeeds (at-least-once semantics). On restart with a persisted token, Spice resumes the Change Stream from that token and skips the collection snapshot. If MongoDB rejects the token (typical codes `ChangeStreamHistoryLost` 286 or `ChangeStreamFatalError` 280, e.g. when the oplog window has rolled past the token's position), the behavior is governed by `mongodb_resume_token_invalid_behavior` above. Re-snapshotting a large collection is opt-in by default. Datasets that are not file-accelerated (in-memory Arrow, etc.) do not get a sidecar row; restarts re-bootstrap from a fresh snapshot. *** ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure MongoDB as a data connector in Spice. [MongoDB Data Connector](https://github.com/spiceai/cookbook/tree/trunk/mongodb/connector#readme) --- # Microsoft SQL Server Data Connector [Microsoft SQL Server](https://www.microsoft.com/en-us/sql-server) is a relational database management system developed by Microsoft. The Microsoft SQL Server Data Connector enables federated/accelerated SQL queries on data stored in MSSQL databases. Limitations 1. The connector supports SQL Server authentication (SQL Login and Password) only. 2. Spatial types (`geography`) are not supported, and columns with these types will be ignored. 3. `DATETIME2` and `DATETIMEOFFSET` columns are mapped to Arrow `Timestamp(Nanosecond)`. Timestamps outside the nanosecond range (approximately years 1677–2262) will return an error. This is an inherent limitation of Arrow's nanosecond timestamp representation. ``` datasets: - from: mssql:path.to.my_dataset name: my_dataset params: mssql_connection_string: ${secrets:mssql_connection_string} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `mssql:database.schema.table` where `database.schema.table` is the fully-qualified table name in the SQL server. info Unquoted identifiers are normalized to lowercase. To reference a table, schema, or database with mixed-case characters, wrap each case-sensitive part in double quotes: `mssql:my_database."MySchema"."MyTable"`. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: mssql:path.to.my_dataset name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords) or any of the following keywords that are reserved by Microsoft SQL Server: * `OUTER` * `SET` * `QUALIFY` * `WINDOW` * `END` * `FOR` ### `params`[​](#params "Direct link to params") The data connector supports the following `params`. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:my_mssql_conn_string}`. | Parameter Name | Description | | -------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mssql_connection_string` | The ADO connection string to use to connect to the server. This can be used instead of providing individual connection parameters, and is the only way to set `ApplicationIntent` — see [Availability groups and read-only routing](#availability-groups-and-read-only-routing). | | `mssql_host` | The hostname or IP address of the Microsoft SQL Server instance. | | `mssql_port` | (Optional) The port of the Microsoft SQL Server instance. Default value is 1433. | | `mssql_database` | (Optional) The name of the database to connect to. The default database (`master`) will be used if not specified. | | `mssql_username` | The username for the SQL Server authentication. | | `mssql_password` | The password for the SQL Server authentication. | | `mssql_encrypt` | (Optional) Specifies whether encryption is required for the connection.
- `true` or `require`: (default) This mode requires an SSL connection. If a secure connection cannot be established, server will not connect.
- `false` or `disable`: This mode will not attempt to use an SSL connection, even if the server supports it. Only the login procedure is encrypted. | | `mssql_trust_server_certificate` | (Optional) Specifies whether the server certificate should be trusted without validation when encryption is enabled.
- `true`: The server certificate will not be validated and it is accepted as-is.
- `false`: (default) Server certificate will be validated against system's certificate storage. | ### Example[​](#example "Direct link to Example") ``` datasets: - from: mssql:SalesLT.Customer name: customer params: mssql_host: mssql-host.database.windows.net mssql_database: my_catalog mssql_username: my_user mssql_password: ${secrets:mssql_pass} mssql_encrypt: true mssql_trust_server_certificate: true ``` ## Availability groups and read-only routing[​](#availability-groups-and-read-only-routing "Direct link to Availability groups and read-only routing") When a dataset connects to an [Always On availability group](https://learn.microsoft.com/en-us/sql/database-engine/availability-groups/windows/always-on-availability-groups-sql-server) listener with read intent, the listener answers the login by naming the secondary replica the session belongs on instead of completing it. Spice follows that redirect automatically and re-dials the named replica, so a read-intent dataset lands on a readable secondary rather than the primary. Read intent is requested through the connection string, so it requires `mssql_connection_string` — the individual connection parameters (`mssql_host`, `mssql_username`, and so on) have no equivalent: ``` datasets: - from: mssql:SalesLT.Customer name: customer params: mssql_connection_string: ${secrets:mssql_connection_string} ``` With a connection string of the form: ``` Server=tcp:ag-listener.example.com,1433;Database=Sales;User ID=my_user;Password=my_password;ApplicationIntent=ReadOnly;Encrypt=true ``` `ApplicationIntent` is the only value that enables routing, and `ReadOnly` is matched exactly — other spellings are treated as read-write intent and are not routed. The redirect carries the settings the dataset supplied: only the host and port are replaced, so the routed replica is reached with the same credentials, database, encryption level and certificate-trust configuration. info Spice follows up to **3** redirects (4 connection attempts) before reporting the chain as broken. A routing list that points back at the listener redirects indefinitely, so exceeding the limit fails the connection with an error naming the last address it was routed to — check the availability group's read-only routing list when this happens. ## Performance[​](#performance "Direct link to Performance") See the dedicated **[Performance](/docs/next/components/data-connectors/mssql/performance)** page for query-pushdown tuning — TopK / `ORDER BY ... LIMIT` pushdown and its NULL-ordering rules. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Microsoft SQL Server as a data connector in Spice. [MSSQL (Microsoft SQL Server) Connector](https://github.com/spiceai/cookbook/tree/trunk/mssql#readme) --- # Microsoft SQL Server Performance Performance tuning for the [Microsoft SQL Server connector](/docs/next/components/data-connectors/mssql). ## TopK / ORDER BY ... LIMIT pushdown[​](#topk--order-by--limit-pushdown "Direct link to TopK / ORDER BY ... LIMIT pushdown") Spice pushes `ORDER BY ... LIMIT N` queries down to SQL Server as `SELECT TOP N ... ORDER BY ...`, avoiding transferring unnecessary rows over the network. This pushdown is applied when the sort can be satisfied exactly by SQL Server — which depends on NULL ordering. SQL Server treats `NULL` as the smallest possible value, so its native ordering is: | Direction | NULLs position | | --------- | -------------- | | `ASC` | First | | `DESC` | Last | Most SQL clients and tools (including Spice's default planner) use the opposite convention (`ASC NULLS LAST`, `DESC NULLS FIRST`). When the requested NULL ordering doesn't match SQL Server's native behavior, Spice falls back to fetching all matching rows and applying the limit locally. **To guarantee TopK pushdown on nullable columns**, explicitly specify the NULL ordering that matches SQL Server's native behavior: ``` -- Pushed down: DESC NULLS LAST matches SQL Server native ordering SELECT id, value FROM my_dataset ORDER BY value DESC NULLS LAST LIMIT 10; -- Pushed down: ASC NULLS FIRST matches SQL Server native ordering SELECT id, value FROM my_dataset ORDER BY value ASC NULLS FIRST LIMIT 10; ``` tip Sorting on `NOT NULL` columns (e.g. primary keys) always pushes the limit down regardless of the `NULLS` clause, since there are no NULLs to order. --- # MySQL Data Connector MySQL is an open-source relational database management system that uses structured query language (SQL) for managing and manipulating databases. The MySQL Data Connector enables federated/accelerated SQL queries on data stored in MySQL databases. ``` datasets: - from: mysql:mytable name: my_dataset params: mysql_host: localhost mysql_tcp_port: 3306 mysql_db: my_database mysql_user: my_user mysql_pass: ${secrets:mysql_pass} mysql_pool_min: 1 mysql_pool_max: 5 ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `mysql:database_name.table_name` where `database_name` is the fully-qualified table name in the SQL server. If the `database_name` is omitted in the `from` field, the connector will use the database specified in the `mysql_db` parameter. If the `mysql_db` parameter is not provided, it will default to the user's default database. info Unquoted identifiers are normalized to lowercase. To reference a table or database with mixed-case characters, wrap each case-sensitive part in double quotes: `mysql:my_database."MixedCaseTable"`. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). These two examples are identical: ``` datasets: - from: mysql:mytable name: my_dataset params: mysql_db: my_database ... ``` ``` datasets: - from: mysql:my_database.mytable name: my_dataset params: ... ``` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: mysql:path.to.my_dataset name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords) or any of the following keywords that are reserved by MySQL: * `PARTITION` ### `params`[​](#params "Direct link to params") The MySQL data connector can be configured by providing the following `params`. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:my_mysql_conn_string}`. | Parameter Name | Description | | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `mysql_connection_string` | The connection string to use to connect to the MySQL server. This can be used instead of providing individual connection parameters. | | `mysql_host` | The hostname of the MySQL server. | | `mysql_tcp_port` | The port of the MySQL server. | | `mysql_db` | The name of the database to connect to. | | `mysql_user` | The MySQL username. | | `mysql_pass` | The password to connect with. | | `mysql_sslmode` | Optional. Specifies the SSL/TLS behavior for the connection, supported values:
- `required`: (default) Require TLS and verify the server certificate against system root CAs and domain name (equivalent to `verify_identity`). Set `mysql_sslrootcert` to verify against a specific CA bundle instead of the system trust store.
- `preferred`: Attempt a TLS connection but skip certificate and hostname verification (invalid or self-signed certificates are accepted). The connection is still encrypted and does **not** fall back to plaintext — if the server does not support TLS, the connection fails. Not recommended for production, as it does not protect against man-in-the-middle attacks.
- `disabled`: Do not attempt to use an SSL connection, even if the server supports it. | | `mysql_sslrootcert` | Optional parameter specifying the path to a custom PEM certificate that the connector will trust. | | `mysql_time_zone` | Optional. Specifies connection time zone. Default is `+00:00` (UTC). Accepts:
- Fixed offsets (e.g., `+02:00`).
- IANA time zone names (e.g., `America/Los_Angeles`), if supported by the MySQL server.
- `system`: The MySQL server host’s OS time zone.
- `local_system`: The local runtime OS time zone. | | `mysql_pool_min` | The minimum number of connections to keep open in the pool, lazily created when requested. Default: `1` | | `mysql_pool_max` | The maximum number of connections created in the connection pool. Default: `5` | | `mysql_zero_date_behavior` | Optional. How to handle the MySQL `0000-00-00` / `0000-00-00 00:00:00` zero-date sentinel for DATE/DATETIME/TIMESTAMP columns. Supported values:
- `null`: (default) Coerces zero dates to NULL and reports such columns as nullable in the Arrow schema.
- `error`: Fails the scan when a zero date is encountered and honors the source NOT NULL constraint exactly. | #### Replication parameters[​](#replication-parameters "Direct link to Replication parameters") The following parameters configure MySQL [binlog replication](/docs/next/features/cdc/mysql-replication) (native CDC) when using `refresh_mode: changes`. `primary_key` + `on_conflict: upsert` are required on the accelerator (except on the append-only `arrow` engine). See [MySQL Binlog Replication](/docs/next/features/cdc/mysql-replication) for source prerequisites, semantics, and metrics. | Parameter Name | Description | | ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `mysql_replication_server_id` | Optional. The `server_id` this replica registers with. Must be unique among all replicas attached to the source. Default: derived from the dataset name and process. | | `mysql_replication_initial_snapshot` | Optional. When existing rows load: `auto` (default) snapshots when no resumable position exists and resumes without a snapshot when one does; `disabled` streams changes only; `always` re-snapshots on every start. | | `mysql_replication_checkpoint_interval` | Optional. How often the committed binlog position persists to the accelerator sidecar (e.g. `10s`). Default: `10s`. | | `mysql_replication_bootstrap_batch_size` | Optional. Number of rows per emitted batch during the initial snapshot. Default: `8192`. Maximum: `1048576`. | | `mysql_replication_invalid_checkpoint_behavior` | Optional. What to do when the persisted position cannot be resumed losslessly — either it was purged from the source, or the source table's column layout drifted incompatibly with the recorded position: `error` (default) or `restart` (drop the position and re-snapshot). Default: `error`. | | `mysql_replication_ready_lag` | Optional. For `refresh_mode: changes`, the dataset is marked Ready once its replication lag (now minus the newest applied commit's binlog-header timestamp) falls below this. Default: `2s`. | ### `metrics`[​](#metrics "Direct link to metrics") The MySQL data connector supports the following optional [component metrics](/docs/next/features/observability/component_metrics): | Metric Name | Type | Description | | ------------------------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ | | `connection_count` | Gauge | Gauge of active connections to the database server | | `connections_in_pool` | Gauge | Gauge of active connections that are idling in the pool | | `active_wait_requests` | Gauge | Gauge of requests that are waiting for a connection to be returned to the pool | | `create_failed` | Counter | Counter of connections that failed to be created | | `discarded_superfluous_connection` | Counter | Counter of connections that were closed because there were already enough idle connections in the pool | | `discarded_unestablished_connection` | Counter | Counter of connections that were closed because they could not be established | | `dirty_connection_return` | Counter | Counter of connections that were returned to the pool but were dirty (ie. open transactions, pending queries, etc) | | `discarded_expired_connection` | Counter | Counter of connections that were discarded because they were expired by the pool constraints (i.e. TTL expired) | | `resetting_connection` | Counter | Counter of connections that were reset | | `discarded_error_during_cleanup` | Counter | Counter of connections that were discarded because they returned an error during cleanup | | `connection_returned_to_pool` | Counter | Counter of connections that were returned to the pool | These metrics are not enabled by default, enable them by setting the `metrics` parameter: ``` datasets: - from: mysql:mytable name: my_dataset metrics: - name: connection_count - name: connections_in_pool - name: active_wait_requests - name: create_failed - name: discarded_superfluous_connection - name: discarded_unestablished_connection - name: dirty_connection_return - name: discarded_expired_connection - name: resetting_connection - name: discarded_error_during_cleanup - name: connection_returned_to_pool params: ¶ms mysql_host: localhost mysql_tcp_port: 3306 mysql_user: my_user mysql_pass: ${secrets:mysql_pass} ``` ## Types[​](#types "Direct link to Types") The table below shows the MySQL data types supported, along with the type mapping to Apache Arrow types in Spice. | MySQL Type | Arrow Type | | ------------ | ------------------------------ | | `TINYINT` | `Int8` | | `SMALLINT` | `Int16` | | `INT` | `Int32` | | `MEDIUMINT` | `Int32` | | `BIGINT` | `Int64` | | `DECIMAL` | `Decimal128` / `Decimal256` | | `FLOAT` | `Float32` | | `DOUBLE` | `Float64` | | `DATETIME` | `Timestamp(Microsecond, None)` | | `TIMESTAMP` | `Timestamp(Microsecond, None)` | | `YEAR` | `Int16` | | `TIME` | `Time64(Nanosecond)` | | `DATE` | `Date32` | | `CHAR` | `Utf8` | | `BINARY` | `Binary` | | `VARCHAR` | `Utf8` | | `VARBINARY` | `Binary` | | `TINYBLOB` | `Binary` | | `TINYTEXT` | `Utf8` | | `BLOB` | `Binary` | | `TEXT` | `Utf8` | | `MEDIUMBLOB` | `Binary` | | `MEDIUMTEXT` | `Utf8` | | `LONGBLOB` | `LargeBinary` | | `LONGTEXT` | `LargeUtf8` | | `JSON` | `LargeUtf8` | | `SET` | `Utf8` | | `ENUM` | `Dictionary(UInt16, Utf8)` | | `BIT` | `UInt64` | note * The MySQL `TIMESTAMP` value is [retrieved as a UTC time value](https://dev.mysql.com/doc/refman/8.4/en/datetime.html) by default. Use the `mysql_time_zone` configuration parameter to specify the desired time zone for interpreting `TIMESTAMP` values during data retrieval. ## Limitations[​](#limitations "Direct link to Limitations") * MySQL has no native nested or array column types (see the [MySQL data types reference](https://dev.mysql.com/doc/refman/8.4/en/data-types.html)), so columns containing Arrow `Struct`, `List`, or `LargeList` values are not supported by this connector. * [MySQL spatial types](https://dev.mysql.com/doc/refman/8.4/en/spatial-types.html) (`GEOMETRY`, `POINT`, `LINESTRING`, `POLYGON`, `MULTIPOINT`, `MULTILINESTRING`, `MULTIPOLYGON`, `GEOMETRYCOLLECTION`) are not currently supported — only the types listed in the [Types](#types) table are mapped to Arrow. * Nested or array-valued data should be stored in a `JSON` column, which is read as `LargeUtf8` and can be decoded with SQL JSON functions. ## Examples[​](#examples "Direct link to Examples") ### Connecting using username and password[​](#connecting-using-username-and-password "Direct link to Connecting using username and password") ``` datasets: - from: mysql:path.to.my_dataset name: my_dataset params: mysql_host: localhost mysql_tcp_port: 3306 mysql_db: my_database mysql_user: my_user mysql_pass: ${secrets:mysql_pass} ``` ### Connecting using SSL[​](#connecting-using-ssl "Direct link to Connecting using SSL") ``` datasets: - from: mysql:path.to.my_dataset name: my_dataset params: mysql_host: localhost mysql_tcp_port: 3306 mysql_db: my_database mysql_user: my_user mysql_pass: ${secrets:mysql_pass} mysql_sslmode: preferred mysql_sslrootcert: ./custom_cert.pem ``` ### Connecting using a Connection String[​](#connecting-using-a-connection-string "Direct link to Connecting using a Connection String") ``` datasets: - from: mysql:path.to.my_dataset name: my_dataset params: mysql_connection_string: mysql://${secrets:my_user}:${secrets:my_password}@localhost:3306/my_db ``` ### Connecting to the default database[​](#connecting-to-the-default-database "Direct link to Connecting to the default database") ``` datasets: - from: mysql:mytable name: my_dataset params: mysql_host: localhost mysql_tcp_port: 3306 mysql_user: my_user mysql_pass: ${secrets:mysql_pass} ``` ### With custom connection pool settings[​](#with-custom-connection-pool-settings "Direct link to With custom connection pool settings") ``` datasets: - from: mysql:path.to.my_dataset name: my_dataset params: mysql_host: localhost mysql_tcp_port: 3306 mysql_db: my_database mysql_user: my_user mysql_pass: ${secrets:mysql_pass} mysql_pool_min: 5 mysql_pool_max: 10 ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure MySQL as a data connector in Spice. [MySQL Data Connector](https://github.com/spiceai/cookbook/tree/trunk/mysql/connector#readme) * A cookbook recipe to configure AWS RDS Aurora (MySQL Compatible) as a data connector in Spice. [AWS RDS Aurora (MySQL Data Connector)](https://github.com/spiceai/cookbook/tree/trunk/mysql/rds-aurora#readme) * A cookbook recipe to configure Planetscale as a data connector in Spice. [Planetscale (MySQL Data Connector)](https://github.com/spiceai/cookbook/tree/trunk/mysql/planetscale#readme) --- # MySQL Data Connector Deployment Guide Production operating guide for the MySQL data connector covering authentication, connection pool sizing, TLS, metrics, and observability. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The MySQL connector supports username/password authentication and connects over TCP. Credentials are supplied via the following parameters: | Parameter | Description | | ------------------------- | ------------------------------------------------------------------------- | | `mysql_host` | MySQL server hostname. | | `mysql_tcp_port` | TCP port (default `3306`). | | `mysql_db` | Database name. | | `mysql_user` | Database user. | | `mysql_pass` | Password. Use `${secrets:...}` to resolve from a configured secret store. | | `mysql_connection_string` | Alternative to the individual parameters. | Passwords must be sourced from a secret store in production. See [Secret Stores](/docs/next/components/secret-stores) for configuration options (environment variables, file, Kubernetes, AWS Secrets Manager, HashiCorp Vault). ### TLS[​](#tls "Direct link to TLS") TLS is controlled via `mysql_sslmode`: | Value | Behavior | | ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `disabled` | No TLS. | | `preferred` | Use TLS but skip certificate and hostname verification (accepts invalid or self-signed certs). The connection is still encrypted and does not fall back to plaintext — it fails if the server lacks TLS support. Not recommended for production. | | `required` | Require TLS and verify the server certificate against system root CAs. | For production, use `required` (the default). To verify against a specific CA rather than the system trust store, also set `mysql_sslrootcert` to the CA bundle path. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Connection Pool Sizing[​](#connection-pool-sizing "Direct link to Connection Pool Sizing") The connector uses a per-dataset connection pool with the following defaults: | Parameter | Default | Description | | ---------------- | ------- | ------------------------------------------ | | `mysql_pool_min` | `1` | Minimum idle connections held by the pool. | | `mysql_pool_max` | `5` | Maximum connections the pool will open. | Both values are parsed leniently, and an unusable value produces **no error and no log message**. A non-integer value is discarded and the pool falls back to the underlying `mysql_async` pool defaults — `10` for `mysql_pool_min` and `100` for `mysql_pool_max`, not the `1` and `5` above. A `mysql_pool_min` greater than `mysql_pool_max` (or a `mysql_pool_max` of `0`) is likewise **not** rejected at startup: the invalid pair is discarded and the same `10`/`100` defaults apply. Because a typo silently widens the pool instead of failing, double-check both values when tuning them. Size the pool to match the concurrent query and refresh load for the dataset; the upper bound should respect the MySQL server's `max_connections` budget shared across all Spice datasets and external clients. ### Retry Behavior[​](#retry-behavior "Direct link to Retry Behavior") Transient query failures are not automatically retried at the connector layer. Dataset refresh retries are controlled by the acceleration refresh policy (see [Data Refresh](/docs/next/features/data-acceleration/data-refresh)). Connection failures surface to the caller and to the connection pool metrics below. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Network**: MySQL traffic is TCP. Plan for the sum of `mysql_pool_max` across all Spice datasets targeting the same MySQL instance when sizing database `max_connections`. * **Memory**: Each pooled connection holds a small amount of client-side state. Result sets stream in batches; memory footprint for federated reads is bounded by DataFusion's record batch size (8192 rows default). * **Throughput**: For full-table materialization (acceleration refresh), query latency scales with the source table size and the presence of indexes on the refresh partitioning/ordering columns. ## Metrics[​](#metrics "Direct link to Metrics") The MySQL connector exposes observable metrics for its connection pool. Enable them in the dataset's `metrics` section. See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. | Metric Name | Type | Description | | ------------------------------------ | --------------- | --------------------------------------------------------------------------------------------- | | `connection_count` | ObservableGauge | Active connections to the database server. | | `connections_in_pool` | ObservableGauge | Idle connections sitting in the pool. | | `active_wait_requests` | ObservableGauge | Requests waiting for a connection (saturation signal). | | `create_failed` | Counter | Connections that failed to be created. | | `discarded_superfluous_connection` | Counter | Connections closed because the pool already had enough idle connections. | | `discarded_unestablished_connection` | Counter | Connections closed because they could not be established. | | `dirty_connection_return` | Counter | Connections returned to the pool in a dirty state (open transactions, pending queries, etc.). | | `discarded_expired_connection` | Counter | Connections discarded because they were expired by pool constraints (i.e. TTL expired). | | `resetting_connection` | Counter | Connections that were reset. | | `discarded_error_during_cleanup` | Counter | Connections discarded because they returned an error during cleanup. | | `connection_returned_to_pool` | Counter | Connections returned to the pool. | Metric instruments are exposed with the prefix `dataset_mysql_`. Each instrument carries a `name` attribute set to the dataset name. Key signals to alert on: * `active_wait_requests > 0` sustained → pool is saturated, increase `mysql_pool_max` or the server's `max_connections`. * `create_failed` increasing → credentials, network, or server availability problem. * `dirty_connection_return` increasing → a query is not cleaning up its transaction state; investigate long-running or aborted queries. ## Task History[​](#task-history "Direct link to Task History") MySQL operations participate in Spice [task history](/docs/next/reference/task_history) via the shared SQL data-connector spans. Queries executed against MySQL are captured as child spans of the enclosing `sql_query` or `accelerated_table_refresh` task. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * Only TCP connections are supported. Unix socket connections are not exposed through Spice configuration. * Only `disabled`, `preferred`, and `required` SSL modes are exposed. The `required` mode verifies the server certificate and domain name (equivalent to `verify_identity`); `preferred` encrypts the connection but skips certificate and hostname verification. * Large text/blob columns are fetched in their entirety per row; consider selecting only the columns you need when federating. * `mysql_sslmode: preferred` encrypts the connection but skips certificate and hostname verification (accepting invalid or self-signed certificates), so it does not protect against man-in-the-middle attacks and is not recommended for production. It does not fall back to plaintext — connecting to a server without TLS support fails. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ------------------------------------------------------------------------ | ------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | | `Access denied for user` | Incorrect credentials or user lacks `SELECT` on the DB. | Verify credentials; confirm the user has read access on the required tables. | | `Too many connections` | Sum of Spice pool sizes + other clients exceeds server `max_connections`. | Reduce `mysql_pool_max` or raise the server limit. | | Sustained `active_wait_requests > 0` | Pool saturation. | Increase `mysql_pool_max`; reduce concurrent dataset refreshes. | | `SSL connection error` | Certificate mismatch or TLS negotiation failure. | Verify `mysql_sslrootcert` matches the server's issuing CA. Use `openssl s_client -connect` to inspect. | | Connection accepted despite an invalid or self-signed server certificate | `mysql_sslmode: preferred` skips certificate and hostname verification. | Switch to `required` (optionally with `mysql_sslrootcert`) to enforce verification. | --- # NFS Data Connector NFS (Network File System) is a distributed file system protocol that provides transparent remote access to shared files across networks. Originally developed by Sun Microsystems, it is widely used in Unix/Linux environments for sharing files between systems. The NFS Data Connector enables federated SQL query across [supported file formats](/docs/next/components/data-connectors/#file-formats) stored on NFS exports. Enterprise edition The NFS Data Connector is available in the Spice [Enterprise edition](https://docs.spice.ai/docs/enterprise/getting-started/distributions). It is feature-gated and not included in the default open source build; open source users can build it from source with the `nfs` feature (requires the system `libnfs` library). Platform Support The NFS Data Connector is available on Linux and macOS. It is not supported on Windows. ## Quickstart[​](#quickstart "Direct link to Quickstart") Connect to an NFS export and query Parquet files: ``` datasets: - from: nfs://nfs-server.local/exports/data/ name: data params: file_format: parquet ``` Query the data using SQL: ``` SELECT * FROM data LIMIT 10; ``` ## How It Works[​](#how-it-works "Direct link to How It Works") Unlike SMB and FTP/SFTP connectors, NFS uses host-based authentication rather than username/password credentials. The NFS server grants access based on the client machine's IP address or hostname, as configured in the server's export rules (typically in `/etc/exports`). Ensure the host running Spice is authorized to mount the NFS export before configuring the connector. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Specifies the NFS server and export path to connect to. **Format:** `nfs:///` * ``: The server hostname or IP address * ``: The exported directory path on the server When pointing to a directory, Spice loads all files within that directory recursively. **Examples:** ``` # Connect to an export root from: nfs://nfs-server.local/exports/data/ # Connect to a subdirectory within an export from: nfs://nfs-server.local/exports/data/sales/2024/ # Connect to a specific file from: nfs://192.168.1.100/data/reports/quarterly.parquet ``` ### `name`[​](#name "Direct link to name") The dataset name used as the table name in SQL queries. Cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | --------------------------- | ----------------------------------------------------------------------------------------------------------------- | | `file_format` | Required when connecting to a directory. See [File Formats](/docs/next/components/data-connectors/#file-formats). | | `client_timeout` | Connection timeout duration. E.g. `30s`, `1m`. No timeout when unset. | | `hive_partitioning_enabled` | Enable [Hive-style partitioning](#hive-partitioning) from folder structure. Default: `false`. | ## Examples[​](#examples "Direct link to Examples") ### Basic Connection[​](#basic-connection "Direct link to Basic Connection") Connect to an NFS export: ``` datasets: - from: nfs://storage.local/exports/analytics/ name: analytics params: file_format: parquet ``` ### Reading a Single File[​](#reading-a-single-file "Direct link to Reading a Single File") When pointing to a specific file, the format is inferred from the file extension: ``` datasets: - from: nfs://nfs-server.local/data/reports/summary.parquet name: summary ``` ### Connection with Timeout[​](#connection-with-timeout "Direct link to Connection with Timeout") Configure a timeout for slow or unreliable network connections: ``` datasets: - from: nfs://remote-nfs.example.com/exports/large-data/ name: large_data params: file_format: parquet client_timeout: 120s ``` ### Reading CSV Files[​](#reading-csv-files "Direct link to Reading CSV Files") Connect to an NFS export containing CSV files: ``` datasets: - from: nfs://storage.local/exports/csv-data/ name: csv_data params: file_format: csv csv_has_header: true ``` ### Hive Partitioning[​](#hive-partitioning "Direct link to Hive Partitioning") Enable Hive-style partitioning to automatically extract partition columns from the folder structure: ``` datasets: - from: nfs://datalake.local/warehouse/events/ name: events params: file_format: parquet hive_partitioning_enabled: true ``` Given a folder structure like: ``` /events/ date=2024-01-01/ data.parquet date=2024-01-02/ data.parquet date=2024-01-03/ data.parquet ``` Queries can filter on partition columns: ``` SELECT * FROM events WHERE date = '2024-01-01'; ``` ### Multiple Exports from One Server[​](#multiple-exports-from-one-server "Direct link to Multiple Exports from One Server") Load different datasets from multiple exports on the same NFS server: ``` datasets: - from: nfs://storage.local/exports/sales/ name: sales params: file_format: parquet - from: nfs://storage.local/exports/inventory/ name: inventory params: file_format: csv ``` ### Accelerated Dataset[​](#accelerated-dataset "Direct link to Accelerated Dataset") Enable local acceleration for faster repeated queries: ``` datasets: - from: nfs://archive.local/historical/ name: historical_data params: file_format: parquet acceleration: enabled: true refresh_check_interval: 1h ``` ### Combined with Other Data Sources[​](#combined-with-other-data-sources "Direct link to Combined with Other Data Sources") Join NFS data with other data sources: ``` datasets: - from: nfs://datalake.local/exports/transactions/ name: transactions params: file_format: parquet - from: postgres://db.local/customers name: customers - from: s3://analytics-bucket/products/ name: products params: file_format: parquet ``` ``` SELECT c.name, p.product_name, t.amount FROM transactions t JOIN customers c ON t.customer_id = c.id JOIN products p ON t.product_id = p.id; ``` ## Authentication and Access Control[​](#authentication-and-access-control "Direct link to Authentication and Access Control") NFS relies on the server's export configuration for access control. The NFS server administrator must configure the export rules to grant read access to the host running Spice. **Example NFS server export configuration** (`/etc/exports`): ``` /exports/data 192.168.1.0/24(ro,sync,no_subtree_check) /exports/data spice-host.local(ro,sync,no_subtree_check) ``` This grants read-only access to all hosts on the `192.168.1.0/24` subnet and to `spice-host.local`. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") ### Connection Errors[​](#connection-errors "Direct link to Connection Errors") If you receive "connection refused" or "mount failed" errors: * Verify the NFS server is running and accessible from the Spice host * Check that the export path exists on the server * Ensure the Spice host is authorized in the server's export configuration * Verify firewall rules permit NFS traffic (typically TCP/UDP ports 111 and 2049) ### Permission Denied[​](#permission-denied "Direct link to Permission Denied") If you can connect but receive permission errors: * Verify the export grants read access to the Spice host * Check file permissions on the exported directory * Ensure the export uses appropriate UID/GID mapping options ### Timeout Errors[​](#timeout-errors "Direct link to Timeout Errors") For slow or unreliable network connections, increase the timeout: ``` params: client_timeout: 120s ``` ### File Format Errors[​](#file-format-errors "Direct link to File Format Errors") When connecting to a directory, ensure `file_format` is specified and matches the actual file types in the directory. Spice expects all files in a directory to have the same format. ## Platform Requirements[​](#platform-requirements "Direct link to Platform Requirements") The NFS connector requires the `libnfs` library to be installed on the system: **Ubuntu/Debian:** ``` sudo apt-get install libnfs-dev ``` **macOS:** ``` brew install libnfs ``` --- # ODBC Data Connector ODBC (Open Database Connectivity) is a standard API for connecting applications to various database management systems using a common interface. To connect to any ODBC database for federated/accelerated SQL queries, specify `odbc` as the selector in the `from` value for the dataset. The `odbc_connection_string` parameter is required. Enterprise edition The ODBC Data Connector is available in the Spice [Enterprise edition](https://docs.spice.ai/docs/enterprise/getting-started/distributions). It is feature-gated and not included in the default open source build; open source users can [build it from source](#building-spice-with-odbc). warning In the open source edition, Spice must be [built with the `odbc` feature](#building-spice-with-odbc); [Spice.ai Enterprise](https://docs.spice.ai/docs/enterprise/getting-started/distributions) distributions ship with ODBC support built in. Either way, the host/container must have a [valid ODBC configuration](https://www.unixodbc.org/odbcinst.html). The published open source `spiceai/spiceai` Docker images do **not** include ODBC support — they are built without the `odbc` feature and do not ship an ODBC Driver Manager. To run ODBC in a container, use an Enterprise distribution or [bake your own image](#baking-an-image-with-odbc-support) from an ODBC-enabled build. ``` datasets: - from: odbc:path.to.my_dataset name: my_dataset params: odbc_connection_string: Driver={Foo Driver};Host=db.foo.net;Param=Value ``` An ODBC connection requires a compatible ODBC driver and valid driver configuration. ODBC drivers are available from their respective vendors. Here are a few examples: * [PostgreSQL](https://odbc.postgresql.org/) * [MySQL](https://dev.mysql.com/downloads/connector/odbc/) * [Databricks](https://www.databricks.com/spark/odbc-drivers-download) * [AWS Athena](https://docs.aws.amazon.com/athena/latest/ug/connect-with-odbc.html) Non-Windows systems additionally require the installation of an ODBC Driver Manager like `unixodbc`. * Ubuntu: `sudo apt-get install unixodbc` * MacOS: `brew install unixodbc` info For the best `JOIN` performance, ensure all ODBC datasets from the same database are configured with the exact same `odbc_connection_string` in Spice. ## ODBC Connection String[​](#odbc-connection-string "Direct link to ODBC Connection String") The ODBC connection string requires the use of an installed and registered driver based on your system type: * Unix systems; ODBC driver installations can be managed using [unixODBC](https://www.unixodbc.org/), or directly edited through `/etc/odbc.ini` or `/etc/odbcinst.ini`. For example, in the [Databricks DSN Connection Setup Guide](https://docs.databricks.com/en/integrations/odbc/dsn.html#linux) for Linux. * Windows systems; ODBC driver installations are managed using the [ODBC Data Source Administrator](https://support.microsoft.com/en-au/office/administer-odbc-data-sources-b19f856b-5b9b-48c9-8b93-07484bfab5a7). For an example Unix system with an installed PostgreSQL driver where the contents of `/etc/odbcinst.ini` is: ``` [PostgreSQL Unicode] Description=PostgreSQL ODBC driver (Unicode version) Driver=psqlodbcw.so Setup=libodbcpsqlS.so Debug=0 CommLog=1 UsageCount=1 ``` The Spice Runtime can use this driver installation where `Driver={PostgreSQL Unicode}` is used in the connection string, like: ``` datasets: - from: odbc:my_table name: my_dataset params: odbc_connection_string: Driver={PostgreSQL Unicode};Server=localhost;Port=5432;Database=postgres;Uid=myuser;Pwd=mypass ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `odbc:path.to.my.dataset` where `path.to.my.dataset` is the table name in the ODBC-supporting server to read from. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: odbc:my.cool.table name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter | Type | Description | | ----------------------------- | -------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `sql_dialect` | string | Override what SQL dialect is used for the ODBC connection. Supports `postgresql`, `mysql`, `sqlite`, `athena` or `databricks` values. Default is unset (auto-detected). | | `odbc_max_bytes_per_batch` | number (bytes) | Maximum number of bytes transferred in each query record batch. A lower value may improve performance on low-memory systems. Default is `512_000_000`. | | `odbc_max_num_rows_per_batch` | number (rows) | Maximum number of rows transferred in each query record batch. A higher value may speed up query results, but requires more memory in conjunction with `odbc_max_bytes_per_batch`. Default is `4000`. | | `odbc_max_text_size` | number (bytes) | A limit for the maximum size of text columns transmitted between the ODBC driver and the Runtime. Default is unset (allocates driver-reported max column size). | | `odbc_max_binary_size` | number (bytes) | A limit for the maximum size of binary columns transmitted between the ODBC driver and the Runtime. Default is unset (allocates driver-reported max column size). | | `odbc_connection_string` | string | Connection string to use to connect to the ODBC server | ``` datasets: - from: odbc:path.to.my_dataset name: my_dataset params: odbc_connection_string: Driver={Foo Driver};Host=db.foo.net;Param=Value ``` ## Selecting SQL Dialect[​](#selecting-sql-dialect "Direct link to Selecting SQL Dialect") The default SQL dialect may not be supported by every ODBC connection. The `sql_dialect` parameter supports overriding the selected SQL dialect for a specified connection. The runtime will attempt to detect the dialect to use for a connection based on the contents of `Driver=` in the `odbc_connection_string`. The runtime will detect the correct SQL dialect for the following connection types, when setup with a standard driver configuration: * PostgreSQL * MySQL * SQLite * Databricks * AWS Athena These connection types are also the supported values for overriding dialect in `sql_dialect`, in lowercase format: `postgresql`, `mysql`, `sqlite`, `databricks`, `athena`. For example, overriding the dialect for your connection to a `postgresql` style dialect: ``` datasets: - from: odbc:path.to.my_dataset name: my_dataset params: sql_dialect: postgresql odbc_connection_string: Driver={Foo Driver};Host=db.foo.net;Param=Value ``` ## Building Spice with ODBC[​](#building-spice-with-odbc "Direct link to Building Spice with ODBC") ODBC support is built into [Spice.ai Enterprise](https://docs.spice.ai/docs/enterprise/getting-started/distributions) distributions. It is not included in the open source released binaries or in the published `spiceai/spiceai` Docker images — to use ODBC with the open source build, [checkout and compile the code](https://github.com/spiceai/spiceai/blob/trunk/CONTRIBUTING#building) with the `--features odbc` flag (`cargo build --release --features odbc`). To build a container image, pass the feature through to the repository `Dockerfile`, which also installs the `unixodbc` Driver Manager when the feature is enabled: ``` docker build --build-arg CARGO_FEATURES=release,models,odbc -t spiceai-odbc:local . ``` ## Baking an image with ODBC Support[​](#baking-an-image-with-odbc-support "Direct link to Baking an image with ODBC Support") There are many dozens of ODBC adapters; this recipe covers making a custom image and configuring it to work with Spice. The base image must already contain an ODBC-enabled `spiced` build — the published `spiceai/spiceai` images do not (see [Building Spice with ODBC](#building-spice-with-odbc)). ``` FROM spiceai-odbc:local RUN apt update \ && apt install --yes libsqliteodbc --no-install-recommends \ && rm -rf /var/lib/{apt,dpkg,cache,log} ``` Build the container: ``` docker build -t spice-libsqliteodbc . ``` Validate that the ODBC configuration was updated to reference the newly installed driver: Note Since `libsqliteodbc` is vendored by Debian, the package install hooks append the driver configuration to `/etc/odbcinst.ini`. When using a custom driver (e.g. [Databricks Simba](https://www.databricks.com/spark/odbc-drivers-download)), it is your responsibility to update `/etc/odbcinst.ini` to point at the location of the newly installed driver. ``` $ docker run --entrypoint /bin/bash -it spice-libsqliteodbc root@f8ceccc94d6a:/# odbcinst -j unixODBC 2.3.11 DRIVERS............: /etc/odbcinst.ini SYSTEM DATA SOURCES: /etc/odbc.ini FILE DATA SOURCES..: /etc/ODBCDataSources USER DATA SOURCES..: /root/.odbc.ini SQLULEN Size.......: 8 SQLLEN Size........: 8 SQLSETPOSIROW Size.: 8 root@f8ceccc94d6a:/# cat /etc/odbcinst.ini [SQLite] Description=SQLite ODBC Driver Driver=libsqliteodbc.so Setup=libsqliteodbc.so UsageCount=1 [SQLite3] Description=SQLite3 ODBC Driver Driver=libsqlite3odbc.so Setup=libsqlite3odbc.so UsageCount=1 ``` ### `test.db`[​](#testdb "Direct link to testdb") To fully test the image, make an example SQLite database (`test.db`) and spicepod on your host: ``` $ sqlite3 test.db SQLite version 3.43.2 2023-10-10 13:08:14 Enter ".help" for usage hints. sqlite> create table spice_test (name text not null); sqlite> insert into spice_test values ("Lala"); sqlite> insert into spice_test values ("Hopper"); sqlite> insert into spice_test values ("Linus"); ``` ### `spicepod.yaml`[​](#spicepodyaml "Direct link to spicepodyaml") Make sure that the `DRIVER` parameter matches the name of the driver section in `odbcinst.ini`. ``` version: v1 kind: Spicepod name: sqlite datasets: - from: odbc:spice_test name: spice_test mode: read acceleration: enabled: false params: odbc_connection_string: DRIVER={SQLite3};SERVER=localhost;DATABASE=test.db;Trusted_connection=yes ``` All together now: ``` $ docker run -p8090:8090 -p50051:50051 -v $(pwd)/spicepod.yaml:/spicepod.yaml -v $(pwd)/test.db:/test.db -it spice-libsqliteodbc --http=0.0.0.0:8090 --flight=0.0.0.0:50051 $ spice sql Welcome to the interactive Spice.ai SQL Query Utility! Type 'help' for help. show tables; -- list available tables sql> show tables; +------------+ | table_name | +------------+ | spice_test | +------------+ Query took: 0.059305583 seconds. 1/1 rows displayed. sql> select * from spice_test; +--------+ | name | +--------+ | Hopper | | Lala | | Linus | +--------+ Query took: 1.8504053329999999 seconds. 3/3 rows displayed. ``` ## Examples[​](#examples "Direct link to Examples") ### Connecting to an SQLite database[​](#connecting-to-an-sqlite-database "Direct link to Connecting to an SQLite database") ``` version: v1 kind: Spicepod name: sqlite datasets: - from: odbc:spice_test name: spice_test mode: read acceleration: enabled: false params: odbc_connection_string: DRIVER={SQLite3};SERVER=localhost;DATABASE=test.db;Trusted_connection=yes ``` ### Connecting to Postgres[​](#connecting-to-postgres "Direct link to Connecting to Postgres") Ensure that the Postgres ODBC driver is installed. On Unix systems, this will create an entry in `/etc/odbcinst.ini` similar to: ``` [PostgreSQL Unicode] Description=PostgreSQL ODBC driver (Unicode version) Driver=psqlodbcw.so Setup=libodbcpsqlS.so Debug=0 CommLog=1 UsageCount=1 ``` Then, in your `spicepod.yaml` the `odbc_connection_string` parameter can be used for the ODBC connection string: ``` version: v1 kind: Spicepod name: odbc-demo datasets: - from: odbc:taxi_trips name: taxi_trips params: odbc_connection_string: Driver={PostgreSQL Unicode};Server=localhost;Port=5432;Database=spice_demo;Uid=postgres ``` See the [ODBC Cookbook](https://github.com/spiceai/cookbook/blob/trunk/odbc/README.md) for more help on getting started with ODBC and Postgres. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure ODBC as a data connector in Spice. [ODBC Data Connector](https://github.com/spiceai/cookbook/tree/trunk/odbc#readme) --- # Oracle Data Connector The Oracle Data Connector enables SQL queries on data stored in Oracle databases, including on-premises instances, Oracle Cloud User-Managed Databases, and Oracle Cloud Autonomous Databases (ADB). ``` datasets: - from: oracle:"SH"."PRODUCTS" name: my_dataset params: oracle_host: localhost oracle_port: 1521 oracle_username: scott oracle_password: ${secrets:oracle_password} oracle_service_name: XEPDB1 ``` Limitations 1. Only basic filter predicates are currently pushed down to the Oracle database. Full query federation is not currently supported. Joins, subqueries, and complex query constructs are not pushed down to the Oracle database; these operations are performed in-memory after data retrieval. **Enable [Data Acceleration](/docs/next/features/data-acceleration) for full federation support**. 2. The Oracle connector does not support filter push-down optimization for datetime columns. Filtering on these columns is performed in-memory after data retrieval. 3. The following Oracle data types are not currently supported; columns with these types will be ignored: `INTERVAL YEAR TO MONTH` (Code 182), `INTERVAL DAY TO SECOND` (Code 183), `UROWID` (Code 208), `BFILE` (Code 114), `JSON` (Code 119). ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `oracle:"schema_name"."table_name"` where both schema and table names should be quoted to handle case sensitivity properly. Example: ``` datasets: - from: oracle:"SH"."PRODUCTS" name: products params: oracle_host: localhost oracle_username: scott oracle_password: ${secrets:ORACLE_PASSWORD} ``` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: oracle:"SH"."PRODUCTS" name: products params: ... ``` ``` SELECT COUNT(*) FROM products; ``` ``` +----------+ | count(*) | +----------+ | 10500 | +----------+ ``` ### `params`[​](#params "Direct link to params") The Oracle data connector can be configured by providing the following `params`. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:MY_ORACLE_PASSWORD}`. | Parameter Name | Description | | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `oracle_connection_string` | The connection string to use to connect to the Oracle server. This can be a TNS alias from `tnsnames.ora` for local mTLS/Wallet connections or an [Easy Connect](https://download.oracle.com/ocomdocs/global/Oracle-Net-Easy-Connect-Plus.pdf) string. | | `oracle_host` | The hostname or IP address of the Oracle Database instance. Required when not using `oracle_connection_string`. | | `oracle_port` | Optional. The port of the Oracle Database server. Default: `1521` | | `oracle_username` | The Oracle username. Required. | | `oracle_password` | The password to connect with. Required. | | `oracle_service_name` | The Oracle Database service name to connect to. Default: `XEPDB1` | | `oracle_wallet_sso_cert` | The base64-encoded `cwallet.sso` (wallet auto-login certificate) to use for mTLS authentication with Oracle Cloud. | | `oracle_wallet` | Specifies the Oracle wallet directory for mTLS connections — either an existing/pre-downloaded wallet, or the destination to save the decoded `oracle_wallet_sso_cert`. The `.oracle` default applies only as the save destination when `oracle_wallet_sso_cert` is set; otherwise no wallet directory is used unless this is specified. | ## Types[​](#types "Direct link to Types") The table below shows the Oracle data types supported, along with the type mapping to Apache Arrow types in Spice. | Oracle Type | Arrow Type | | -------------------------------- | -------------------------------------------------------------------------------- | | `ROWID` | `Utf8` | | `CHAR` | `Utf8` | | `NCHAR` | `Utf8` | | `VARCHAR2` | `Utf8` | | `NVARCHAR2` | `Utf8` | | `LONG` | `Utf8` | | `CLOB` | `LargeUtf8` | | `NCLOB` | `LargeUtf8` | | `NUMBER` | `Int64` for integer types (scale=0, precision≤18), otherwise `Decimal128` | | `FLOAT` | `Float32` for precision≤24, otherwise `Float64` | | `BINARY_FLOAT` | `Float32` | | `BINARY_DOUBLE` | `Float64` | | `BOOLEAN` | `Boolean` | | `DATE` | `Timestamp(Second)` | | `TIMESTAMP` | `Timestamp(Second)` for precision=0, otherwise `Timestamp(Nanosecond)` | | `TIMESTAMP WITH TIME ZONE` | `Timestamp(Second, UTC)` for precision=0, otherwise `Timestamp(Nanosecond, UTC)` | | `TIMESTAMP WITH LOCAL TIME ZONE` | `Timestamp(Second, UTC)` for precision=0, otherwise `Timestamp(Nanosecond, UTC)` | | `RAW` | `Binary` | | `LONG RAW` | `Binary` | | `BLOB` | `LargeBinary` | note * Oracle `DATE` is **not** a date-only type — it stores hour, minute, and second alongside the calendar date — so it maps to `Timestamp(Second)`, matching Oracle's 1-second resolution. Earlier versions mapped it to `Date32`, which silently truncated every value to midnight. * The Oracle `TIMESTAMP WITH LOCAL TIME ZONE` value is retrieved as a UTC time value. * `TIMESTAMP`, `TIMESTAMP WITH TIME ZONE`, and `TIMESTAMP WITH LOCAL TIME ZONE` columns with non-zero precision are mapped to `Timestamp(Nanosecond)`. Timestamps outside the nanosecond range (approximately years 1677–2262) will return an error. This is an inherent limitation of Arrow's nanosecond timestamp representation. ## Examples[​](#examples "Direct link to Examples") ### Connecting to On-Premises Oracle Database[​](#connecting-to-on-premises-oracle-database "Direct link to Connecting to On-Premises Oracle Database") ``` datasets: - from: oracle:"SH"."PRODUCTS" name: products params: oracle_host: localhost oracle_port: 1521 oracle_username: scott oracle_password: ${secrets:ORACLE_PASSWORD} oracle_service_name: XEPDB1 ``` ### Connecting to Oracle Cloud Autonomous Database with mTLS (Wallet-based)[​](#connecting-to-oracle-cloud-autonomous-database-with-mtls-wallet-based "Direct link to Connecting to Oracle Cloud Autonomous Database with mTLS (Wallet-based)") ### Wallet Folder Exists Locally[​](#wallet-folder-exists-locally "Direct link to Wallet Folder Exists Locally") If your Oracle Cloud Autonomous Database wallet folder is available locally, specify its path using the `oracle_wallet` parameter. Set the `oracle_connection_string` to the TNS alias defined in your wallet's `tnsnames.ora` file. Example: ``` datasets: - from: oracle:"SALES" name: sales params: oracle_username: admin oracle_password: ${secrets:ORACLE_PASSWORD} oracle_connection_string: 'fgp1tqs1e_low' # TNS alias from tnsnames.ora oracle_wallet: '/path/to/wallet_folder' ``` ### Wallet Auto-Login (SSO) Certificate Provided via Application Secret[​](#wallet-auto-login-sso-certificate-provided-via-application-secret "Direct link to Wallet Auto-Login (SSO) Certificate Provided via Application Secret") If your Oracle Cloud Autonomous Database wallet folder is not available locally, provide the base64-encoded wallet auto-login (SSO) certificate (`cwallet.sso`) using the `oracle_wallet_sso_cert` parameter. Set the `oracle_connection_string` to the Easy Connect string from the *Database connection* section. ``` datasets: - from: oracle:"SALES" name: sales params: oracle_username: admin oracle_password: ${secrets:ORACLE_PASSWORD} oracle_wallet_sso_cert: ${secrets:oracle_wallet_sso_cert} oracle_connection_string: 'tcps://adb.us-sanjose-1.oraclecloud.com:1522/g81f1d1d5c853_fgc1e_low.adb.oraclecloud.com?ssl_server_dn_match=yes' ``` To generate a base64-encoded wallet certificate for use as a secret: ``` base64 -i cwallet.sso > cwallet.b64.txt ``` ### Connecting with Easy Connect string (TLS-only, no wallet required)[​](#connecting-with-easy-connect-string-tls-only-no-wallet-required "Direct link to Connecting with Easy Connect string (TLS-only, no wallet required)") ``` datasets: - from: oracle:"SALES" name: sales params: oracle_username: admin oracle_password: ${secrets:ORACLE_PASSWORD} oracle_connection_string: 'tcps://adb.us-sanjose-1.oraclecloud.com:1522/g81f1d1d5c853_fgc1e_low.adb.oraclecloud.com?ssl_server_dn_match=yes' ``` ## Installation Requirements[​](#installation-requirements "Direct link to Installation Requirements") The Oracle data connector requires the Oracle Instant Client or Oracle Database Client libraries to be installed on the system where Spice is running. Follow the [Oracle installation guide](https://oracle.github.io/odpi/) for your platform. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to connect to and accelerate data from an Oracle database in Spice. [Oracle Data Connector](https://github.com/spiceai/cookbook/blob/trunk/oracle/README.md) --- # PostgreSQL Data Connector PostgreSQL is an advanced open-source relational database management system known for its reliability, extensibility, and support for SQL compliance. The PostgreSQL Server Data Connector enables federated/accelerated SQL queries on data stored in PostgreSQL databases. ``` datasets: - from: postgres:my_table name: my_dataset params: ... ``` ## Quickstart[​](#quickstart "Direct link to Quickstart") Connect to a local PostgreSQL database and accelerate a table for fast local queries: ``` version: v1 kind: Spicepod name: pg_demo datasets: - from: postgres:public.customers name: customers params: pg_host: localhost pg_port: "5432" pg_db: mydb pg_user: spice_reader pg_pass: ${secrets:PG_PASSWORD} acceleration: enabled: true ``` ``` echo "PG_PASSWORD=your_password" > .env ``` Start Spice and query the data: ``` spice run # In another terminal: spice sql sql> SELECT count(*) FROM customers; ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `postgres:my_table` where `my_table` is the table identifer in the PostgreSQL server to read from. The fully-qualified table name (`database.schema.table`) can also be used in the `from` field. ``` datasets: - from: postgres:my_database.my_schema.my_table name: my_dataset params: ... ``` info Unquoted identifiers are normalized to lowercase. To reference a table or schema with mixed-case characters, wrap each case-sensitive part in double quotes: `postgres:my_schema."MixedCaseTable"`. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: postgres:my_database.my_schema.my_table name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` ### `params`[​](#params "Direct link to params") The connection to PostgreSQL can be configured by providing the following `params`: | Parameter Name | Description | | ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `pg_connection_string` | Optional. The connection string to use to connect to the PostgreSQL server. This can be used instead of providing individual connection parameters. | | `pg_host` | The hostname of the PostgreSQL server. | | `pg_port` | The port of the PostgreSQL server. | | `pg_db` | The name of the database to connect to. | | `pg_user` | The username to connect with. | | `pg_pass` | The password to connect with. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load the password from a secret store, e.g. `${secrets:my_pg_pass}`. | | `pg_sslmode` | Optional. Specifies the SSL/TLS behavior for the connection, supported values:
- `verify-full`: (default) This mode requires an SSL connection, a valid root certificate, and the server host name to match the one specified in the certificate.
- `verify-ca`: This mode requires a TLS connection and a valid root certificate.
- `require`: This mode requires a TLS connection.
- `prefer`: This mode will try to establish a secure TLS connection if possible, but will connect insecurely if the server does not support TLS.
- `disable`: This mode will not attempt to use a TLS connection, even if the server supports it. | | `pg_sslrootcert` | Optional. Path to a custom PEM certificate file, or inline PEM content, that the connector will trust when `pg_sslmode` is `verify-ca` or `verify-full`. Both spellings are accepted on the federated read/query path and on the WAL replication transport (`refresh_mode: changes`); a value containing a `-----BEGIN CERTIFICATE-----` block is read as PEM content, and anything else as a path. | | `pg_connection_pool_min_idle` | Optional. The minimum number of idle connections to keep open in the pool. Default is `1`. | | `connection_pool_size` | Optional. The maximum number of connections created in the connection pool. Default is `5`. | #### Replication parameters[​](#replication-parameters "Direct link to Replication parameters") The following parameters configure PostgreSQL [logical replication](https://www.postgresql.org/docs/current/logical-replication.html) (WAL streaming) when using `refresh_mode: changes`: | Parameter Name | Description | | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `pg_replication_slot` | Optional. Name of the replication slot to create/reuse. Must match `[a-z0-9_]{1,63}` and must not be the reserved name `pg_conflict_detection`. Defaults to `spice___`. Each Spice replica MUST have its own unique slot. | | `pg_publication` | Optional. Name of the publication to create/reuse. Defaults to `spice___pub`, or `_pub` when `pg_replication_slot` is set. Shared across replicas for the same dataset. | | `pg_replication_initial_snapshot` | Optional. When `refresh_mode: changes` first loads the table's existing rows: `auto` (default) snapshots a freshly-created replication slot and resumes an existing one without a snapshot; `disabled` streams WAL changes only; `always` snapshots on every start, including slot resume. The legacy booleans `true`/`false` map to `auto`/`disabled`. Default: `auto`. | | `pg_replication_temporary_slot` | **Deprecated and ignored.** The replication slot is always durable. Setting it logs a deprecation warning and has no other effect. To stop an unused slot retaining WAL on the source, drop it with `SELECT pg_drop_replication_slot('')`. | | `pg_replication_status_interval` | Optional. How often to send StandbyStatusUpdate to Postgres (e.g. `10s`). Default: `10s`. | | `pg_replication_ready_lag` | Optional. For `refresh_mode: changes`, the dataset is marked Ready once its replication lag (now minus the newest applied commit's source time) falls below this. Default: `2s`. | | `pg_replication_bootstrap_batch_size` | Optional. Number of rows per emitted batch during the initial replication snapshot. Default: `8192`. Maximum: `1048576`. | | `pg_replication_member_channel_capacity` | Optional. Shared-slot only: envelopes buffered per member table before the shared replication pump back-pressures. Default: `1024`. Maximum: `1048576`. | `pg_sslmode` differs on the WAL replication transport The `pg_sslmode` default of `verify-full` documented above applies to the federated read/query path. On the WAL replication transport used by `refresh_mode: changes`, an unset `pg_sslmode` defaults to **`prefer`**, which negotiates a **plaintext** connection (no certificate verification). Set `pg_sslmode` to `require`, `verify-ca`, or `verify-full` to force TLS on the WAL stream — see [`pg_sslmode` for WAL streaming](/docs/next/features/cdc/postgres-replication#pg_sslmode-for-wal-streaming). This applies to the discrete-parameter path. When the connection is configured with `pg_connection_string` and the string omits `sslmode`, the WAL transport defaults to `verify-full` — see [Connecting with `pg_connection_string`](/docs/next/features/cdc/postgres-replication#connecting-with-pg_connection_string). `pg_sslrootcert` behaves the same on both transports: a value containing a `-----BEGIN CERTIFICATE-----` block is treated as inline PEM content, and any other value as a path to a PEM file. Content that arrives through a channel that does not preserve real newlines — a single-line environment variable, a JSON string pasted verbatim — carries them as the two characters `\n`, which are restored before the certificate is parsed. ## Types[​](#types "Direct link to Types") The table below shows the PostgreSQL data types supported, along with the type mapping to Apache Arrow types in Spice. | PostgreSQL Type | Arrow Type | | --------------- | ----------------------------- | | `int2` | `Int16` | | `int4` | `Int32` | | `int8` | `Int64` | | `money` | `Int64` | | `float4` | `Float32` | | `float8` | `Float64` | | `numeric` | `Decimal128` | | `text` | `Utf8` | | `varchar` | `Utf8` | | `bpchar` | `Utf8` | | `uuid` | `Utf8` | | `bytea` | `Binary` | | `bool` | `Boolean` | | `json` | `Utf8` | | `timestamp` | `Timestamp(Nanosecond, None)` | | `timestamptz` | `Timestamp(Nanosecond, UTC)` | | `date` | `Date32` | | `time` | `Time64(Nanosecond)` | | `interval` | `Interval(MonthDayNano)` | | `point` | `FixedSizeList(Float64[2])` | | `int2[]` | `List(Int16)` | | `int4[]` | `List(Int32)` | | `int8[]` | `List(Int64)` | | `float4[]` | `List(Float32)` | | `float8[]` | `List(Float64)` | | `text[]` | `List(Utf8)` | | `bool[]` | `List(Boolean)` | | `bytea[]` | `List(Binary)` | | `geometry` | `Binary` | | `geography` | `Binary` | | `enum` | `Dictionary(Int8, Utf8)` | | Composite Types | `Struct` | info The Postgres federated queries may result in unexpected result types due to the difference in DataFusion and Postgres size increase rules. Explicitly specify the expected output type of aggregation functions when writing queries involving Postgres tables in Spice. For example, rewrite `SUM(int_col)` into `CAST (SUM(int_col) as BIGINT)`. ## Write Support[​](#write-support "Direct link to Write Support") The PostgreSQL connector supports writing data to PostgreSQL tables using SQL [`INSERT INTO`](/docs/next/reference/sql/dml#insert), `UPDATE`, and `DELETE FROM` statements. To enable writes, set `access: read_write` on the dataset: ``` datasets: - from: postgres:public.events name: events access: read_write params: pg_host: localhost pg_port: '5432' pg_db: mydb pg_user: spice_writer pg_pass: ${secrets:PG_PASSWORD} ``` ``` -- Insert rows INSERT INTO events (id, name, amount) VALUES (1, 'Alice', 100.0), (2, 'Bob', 200.0); -- Update rows UPDATE events SET amount = 150.0 WHERE id = 1; -- Delete rows DELETE FROM events WHERE id = 2; ``` ### Write modes with acceleration[​](#write-modes-with-acceleration "Direct link to Write modes with acceleration") When PostgreSQL is used as the federated source for an accelerated dataset, `acceleration.write_mode` selects how writes propagate between the local accelerator and PostgreSQL: * `write_through` (default) — writes are sent to PostgreSQL synchronously. The client receives an ACK only after the source commits. The local accelerator is updated via the configured refresh path. Choose this for ACID guarantees. * `write_back` — writes are applied to the local accelerator first (fast ACK), then forwarded asynchronously to PostgreSQL. Choose this for write throughput when eventual consistency at the source is acceptable. `acceleration.refresh_mode: changes` is supported for `access: read_write` datasets: writes go to PostgreSQL and the WAL replication stream applies the resulting changes back to the accelerator. ``` datasets: - from: postgres:public.events name: events access: read_write params: pg_host: localhost pg_port: '5432' pg_db: mydb pg_user: spice_writer pg_pass: ${secrets:PG_PASSWORD} # Replication-mode parameters (see Configuration above) pg_publication: spice_pub pg_replication_slot: spice_slot acceleration: engine: duckdb mode: file refresh_mode: changes write_mode: write_through # default; use write_back for fast async writes ``` For more details, see [Data Ingestion](/docs/next/features/data-ingestion). ## Examples[​](#examples "Direct link to Examples") ### Connecting using Username/Password[​](#connecting-using-usernamepassword "Direct link to Connecting using Username/Password") ``` datasets: - from: postgres:my_database.my_schema.my_table name: my_dataset params: pg_host: localhost pg_port: 5432 pg_db: my_database pg_user: my_user pg_pass: ${secrets:my_pg_pass} ``` ### Connect using SSL[​](#connect-using-ssl "Direct link to Connect using SSL") ``` datasets: - from: postgres:my_database.my_schema.my_table name: my_dataset params: pg_host: localhost pg_port: 5432 pg_db: my_database pg_user: my_user pg_pass: ${secrets:my_pg_pass} pg_sslmode: verify-ca pg_sslrootcert: ./custom_cert.pem ``` ### Separate dataset/accelerator secrets[​](#separate-datasetaccelerator-secrets "Direct link to Separate dataset/accelerator secrets") Specify different secrets for a PostgreSQL source and acceleration: ``` datasets: - from: postgres:my_schema.my_table name: my_dataset params: pg_host: localhost pg_port: 5432 pg_db: my_database pg_user: my_user pg_pass: ${secrets:pg1_pass} acceleration: engine: postgres params: pg_host: localhost pg_port: 5433 pg_db: acceleration pg_user: two_user_two_furious pg_pass: ${secrets:pg2_pass} ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure PostgreSQL as a data connector in Spice. [PostgreSQL Data Accelerator](https://github.com/spiceai/cookbook/tree/trunk/postgres/accelerator#readme) * A cookbook recipe to configure AWS RDS for PostgreSQL as a data connector in Spice. [AWS RDS for PostgreSQL](https://github.com/spiceai/cookbook/tree/trunk/postgres/rds#readme) * A cookbook recipe to configure Supabase a data connector in Spice. [Supabase (PostgreSQL Data Connector)](https://github.com/spiceai/cookbook/tree/trunk/postgres/supabase#readme) --- # PostgreSQL Data Connector Deployment Guide Production operating guide for the PostgreSQL data connector covering authentication, connection pool sizing, TLS, metrics, and observability. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The connector uses the native PostgreSQL wire protocol with username/password authentication. | Parameter | Description | | ---------------------- | ------------------------------------------------------------------------- | | `pg_host` | PostgreSQL server hostname. | | `pg_port` | TCP port (default `5432`). | | `pg_db` | Database name. | | `pg_user` | Database user. | | `pg_pass` | Password. Use `${secrets:...}` to resolve from a configured secret store. | | `pg_connection_string` | Alternative to the individual parameters. | Passwords must be sourced from a secret store in production. See [Secret Stores](/docs/next/components/secret-stores) for configuration options (environment variables, file, Kubernetes, AWS Secrets Manager, HashiCorp Vault). ### TLS[​](#tls "Direct link to TLS") TLS is controlled via `pg_sslmode`: | Value | Behavior | | ------------- | ------------------------------------------------------------------- | | `disable` | No TLS. | | `prefer` | Try TLS, fall back to plaintext. Not recommended for production. | | `require` | Require TLS; no server certificate verification. | | `verify-ca` | Require TLS and verify the CA chain. | | `verify-full` | (default) Require TLS, verify CA chain, and verify server hostname. | For production, use `verify-full` with `pg_sslrootcert` pointing to the CA bundle file path. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Connection Pool Sizing[​](#connection-pool-sizing "Direct link to Connection Pool Sizing") The connector maintains a per-dataset connection pool: | Parameter | Default | Description | | ----------------------------- | ------- | ------------------------------------------ | | `pg_connection_pool_min_idle` | `1` | Minimum idle connections held by the pool. | | `connection_pool_size` | `5` | Maximum connections the pool will open. | When `pg_connection_pool_min_idle` exceeds `connection_pool_size`, the pool silently caps idle connections at the pool size. Size the pool to match concurrent query and refresh load for the dataset. The server's `max_connections` (default 100) is a shared budget across Spice datasets, other clients, and server-side background workers — plan accordingly, or front Postgres with PgBouncer. ### Application Name[​](#application-name "Direct link to Application Name") The connector automatically sets `application_name` to the Spice.ai version string, which surfaces in `pg_stat_activity.application_name`. This value is not configurable. ### Retry Behavior[​](#retry-behavior "Direct link to Retry Behavior") Transient query failures are not automatically retried at the connector layer. Dataset refresh retries are controlled by the acceleration refresh policy (see [Data Refresh](/docs/next/features/data-acceleration/data-refresh)). ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Network**: Postgres traffic is TCP. Sum `connection_pool_size` across all Spice datasets sharing the server when sizing `max_connections`. * **Memory**: Result sets are streamed in record batches; memory footprint for federated reads is bounded by DataFusion's batch size (8192 rows default). * **Connection setup cost**: TLS handshake and authentication add latency to cold connections. `connection_pool_min_idle` keeps a warm pool to absorb burst traffic. ## Metrics[​](#metrics "Direct link to Metrics") The PostgreSQL connector exposes observable metrics for its replication pipeline. Every metric below is auto-registered — no configuration is required to export it — **except** `replication_truncates_total` and `replication_bootstrap_rows_total`, which are opt-in and must be listed in the dataset's `metrics` section. To turn an auto-registered metric off for a dataset, set `enabled: false` there. See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. | Metric Name | Type | Description | | --------------------------------------------------- | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `replication_lag_ms` | ObservableGauge | Milliseconds between now and the Postgres commit timestamp of the most recently replicated transaction. Primary freshness signal. | | `replication_lag_bytes` | ObservableGauge | WAL bytes between the server's latest reported position and the last confirmed flush LSN. | | `replication_confirmed_flush_lsn` | ObservableGauge | Most recent LSN acknowledged to Postgres. Matches `pg_replication_slots.confirmed_flush_lsn`. | | `replication_server_wal_end_lsn` | ObservableGauge | Most recent WAL end LSN reported by the Postgres server. | | `replication_reader_input_wait_micros_total` | ObservableCounter | Microseconds the reader spent blocked awaiting the next source event. High relative to the processing counter ⇒ source-bound. | | `replication_reader_processing_micros_total` | ObservableCounter | Microseconds the reader spent decoding WAL and building change batches. | | `replication_transactions_total` | ObservableCounter | Transactions committed and applied to the accelerator. | | `replication_inserts_total` | ObservableCounter | `INSERT` operations received from WAL. | | `replication_updates_total` | ObservableCounter | `UPDATE` operations received from WAL. | | `replication_deletes_total` | ObservableCounter | `DELETE` operations received from WAL. | | `replication_truncates_total` | ObservableCounter | `TRUNCATE` operations received from WAL and applied. **Opt-in.** | | `replication_bootstrap_rows_total` | ObservableCounter | Rows loaded during the initial-snapshot bootstrap. **Opt-in.** | | `replication_bootstrap_rows_expected` | ObservableGauge | Estimated bootstrap row count from schema inference. Absent when no estimate exists; `0` means a known-empty source table. | | `replication_bootstrap_complete` | ObservableGauge | `1` once the initial snapshot finished (or was skipped on resume); `0` while it is still running. | | `replication_decode_errors_total` | ObservableCounter | pgoutput decoding errors encountered while parsing WAL events. | | `replication_schema_mismatch_errors_total` | ObservableCounter | Errors where the source relation no longer matches the declared accelerator schema. | | `replication_recv_errors_total` | ObservableCounter | Transport-level errors while receiving from the replication connection. | | `replication_reconnects_total` | ObservableCounter | Times the stream reconnected after a transient failure. Non-zero with no user-visible error just means it recovered. | | `replication_disconnected_ms_total` | ObservableCounter | Cumulative milliseconds the stream was disconnected across all reconnects, including backoff. | | `replication_member_send_stalled_seconds_total` | ObservableCounter | Seconds the shared-slot pump spent blocked delivering changes into this dataset's channel. Shared slots only. | | `replication_member_send_wait_micros_total` | ObservableCounter | Microseconds the shared-slot pump spent awaiting this dataset's delivery channel. Dedicated-slot datasets export `0`. | | `replication_member_attached` | ObservableGauge | `1` while this dataset is an attached member of its shared slot, `0` once detached. Shared slots only. | | `replication_member_envelopes_delivered_total` | ObservableCounter | Change envelopes delivered to this dataset as distinct units of work. `replication_transactions_total` divided by this is the coalescing factor the apply loop sees. Shared slots only. | | `replication_member_envelope_eager_merges_total` | ObservableCounter | Committed transactions folded into an envelope the shared-slot pump was still holding back, before delivery. Shared slots only. | | `replication_member_envelope_mailbox_merges_total` | ObservableCounter | Committed transactions folded into an envelope already sitting unclaimed in this dataset's delivery buffer — the back-pressure-driven half of coalescing. Shared slots only. | | `replication_member_mailbox_coalesce_limited_total` | ObservableCounter | Times a committed transaction could not be folded into the unclaimed buffer tail because a configured bound refused it. `0` means the bounds never bind. Shared slots only. | Metric instruments are exposed with the prefix `dataset_postgres_`. Each instrument carries a `name` attribute set to the dataset name; `replication_member_attached` also carries a `slot` attribute for grouping shared-slot members. ## Task History[​](#task-history "Direct link to Task History") PostgreSQL operations participate in Spice [task history](/docs/next/reference/task_history) via the shared SQL data-connector spans. Queries executed against Postgres are captured as child spans of the enclosing `sql_query` or `accelerated_table_refresh` task. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * Only TCP connections are supported. Unix sockets are not exposed through Spice configuration. * `pg_sslmode: prefer` silently downgrades to plaintext and is not recommended for production. * `LISTEN/NOTIFY` is not exposed. CDC is supported natively via logical replication (WAL streaming) — see the [replication parameters](/docs/next/components/data-connectors/postgres#replication-parameters) in the connector docs. * Server-side cursors are used for federated reads; long-running queries hold a backend for their duration. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | --------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------- | | `FATAL: password authentication failed` | Incorrect credentials. | Verify credentials via the secret store; test with `psql` using the same credentials. | | `FATAL: too many clients already` | Pool size + other clients exceeds server `max_connections`. | Reduce `connection_pool_size` or raise `max_connections` / front the server with PgBouncer. | | Idle connections never exceed `connection_pool_size` despite a higher `pg_connection_pool_min_idle` | The pool silently caps `min_idle` at the pool size. | Set `pg_connection_pool_min_idle` to `connection_pool_size` or lower for clarity. | | Sustained `active_wait_requests > 0` | Pool saturation. | Increase `connection_pool_size` or reduce concurrent refreshes. | | `certificate verify failed` | `pg_sslmode: verify-ca` / `verify-full` with wrong CA or hostname. | Verify `pg_sslrootcert` matches the server's issuing CA; with `verify-full` ensure hostname matches SAN. | | Sessions lingering with the default app name | Multiple Spice instances share the same version-based name. | The `application_name` is auto-set to the Spice.ai version and is not currently configurable. | --- # Amazon Redshift Data Connector Amazon Redshift is a columnar OLAP database compatible with PostgreSQL. To connect Redshift to Spice, use the [PostgreSQL data connector](/docs/next/components/data-connectors/postgres) and specify the Redshift cluster connection parameters. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Use the format `postgres:schema.table` to reference a Redshift table. The connector parameters should match your Redshift cluster settings. ### Example Spicepod[​](#example-spicepod "Direct link to Example Spicepod") ``` version: v1 kind: Spicepod name: tpch-read datasets: - from: postgres:public.customer name: customer params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.lineitem name: lineitem params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.nation name: nation params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.orders name: orders params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.part name: part params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.partsupp name: partsupp params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.region name: region params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.supplier name: supplier params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true ``` ### Parameters[​](#parameters "Direct link to Parameters") | Parameter Name | Description | | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `pg_connection_string` | Optional. A [PostgreSQL connection string](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNSTRING). Overrides individual connection parameters when provided. | | `pg_host` | Hostname or IP address of the Redshift cluster | | `pg_port` | The PostgreSQL TCP port. Redshift uses port `5439` by default — set this explicitly. | | `pg_db` | Database name | | `pg_user` | Username for authentication | | `pg_pass` | Password for authentication (use secret reference) | | `pg_sslmode` | SSL mode. Default `verify-full`. Supported values: `disable`, `prefer`, `require`, `verify-ca`, `verify-full`. | | `pg_sslrootcert` | Optional. Path to a custom root certificate for SSL verification | | `pg_connection_pool_min_idle` | Optional. The minimum number of idle connections to keep open in the pool. Default is `1`. | | `connection_pool_size` | Optional. The maximum number of connections in the connection pool. Default is `5`. | ## Supported Types[​](#supported-types "Direct link to Supported Types") Redshift types are mapped to PostgreSQL types. See the [PostgreSQL connector documentation](/docs/next/components/data-connectors/postgres) for details on supported types and configuration. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For details, see the [secret stores documentation](/docs/next/components/secret-stores) and [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Amazon Redshift as a data connector in Spice. [Redshift Data Connector](https://github.com/spiceai/cookbook/tree/trunk/redshift#readme) ## References[​](#references "Direct link to References") * [Amazon Redshift Documentation](https://docs.aws.amazon.com/redshift/latest/mgmt/welcome.html) * [PostgreSQL Connector Documentation](/docs/next/components/data-connectors/postgres) --- # S3 Data Connector The S3 Data Connector enables federated SQL querying on files stored in S3 or S3-compatible systems (e.g., MinIO, Cloudflare R2). If a folder path is specified as the dataset source, all files within the folder will be loaded. File formats are specified using the `file_format` parameter, as described in [File Formats](/docs/next/components/data-connectors/#file-formats). ## Quickstart[​](#quickstart "Direct link to Quickstart") Query a public S3 dataset with no authentication: ``` version: v1 kind: Spicepod name: s3_demo datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet ``` For private buckets, add authentication (see [Authentication](#authentication)): ``` datasets: - from: s3://my-private-bucket/data/ name: my_data params: file_format: parquet s3_auth: key s3_key: ${secrets:AWS_ACCESS_KEY_ID} s3_secret: ${secrets:AWS_SECRET_ACCESS_KEY} s3_region: us-west-2 acceleration: enabled: true ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") S3-compatible URI to a folder or file, in the format `s3:///` Example: `from: s3://my-bucket/path/to/file.parquet` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: s3://s3-bucket-name/taxi_sample.csv name: cool_dataset params: file_format: csv ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `file_format` | Specifies the data format. Required if it cannot be inferred from the object URI. Options: `parquet`, `csv`, `json`. Refer to [File Formats](/docs/next/components/data-connectors/#file-formats) for details. | | `s3_endpoint` | S3 endpoint URL (e.g., for MinIO). Default is the region endpoint. E.g. `s3_endpoint: https://my.minio.server` | | `s3_url_style` | URL style used for S3 requests. Options: `vhost` and `path`. When not set, auto-detected: IP endpoints use `path`; other custom endpoints are detected via DNS lookup on the virtual-hosted hostname; standard AWS endpoints use `vhost`, except buckets whose name contains a dot, which use `path`. | | `s3_region` | S3 bucket region. Default: `us-east-1`. | | `client_timeout` | Optional. Timeout for S3 operations. Defaults to `30s` (the underlying object store HTTP client default) when unset. | | `hive_partitioning_enabled` | Enable partitioning using hive-style partitioning from the folder structure. Defaults to `false` | | `s3_auth` | Authentication type. Options: `public`, `key` and `iam_role`. When unset, uses the default AWS credential chain (IAM-based authentication). Set to `public` for unauthenticated access to public buckets. If set to `key` the `s3_key` and `s3_secret` parameters must also be set. If set to `iam_role` the credentials will be loaded from environment variables or IAM roles (see [Authentication](#authentication) for details). | | `s3_iam_role_source` | Optional. IAM role credential source (used when `s3_auth` is `iam_role` or unset). `auto` (default) uses the default AWS credential chain, `metadata` uses only instance/container metadata (IMDS, ECS, EKS/IRSA), `env` uses only environment variables. | | `s3_key` | Access key (e.g. `AWS_ACCESS_KEY_ID` for AWS). Requires `s3_auth` is set to `key`. | | `s3_secret` | Secret key (e.g. `AWS_SECRET_ACCESS_KEY` for AWS). Requires `s3_auth` is set to `key`. | | `s3_session_token` | Session token (e.g. `AWS_SESSION_TOKEN` for AWS) for temporary credentials. Requires `s3_auth` is set to `key`. | | `s3_versioning` | Enables support for S3 buckets with [S3 Versioning](https://docs.aws.amazon.com/AmazonS3/latest/userguide/Versioning.html). Options: `enabled` and `disabled`. Defaults to `enabled`. | | `allow_http` | Enables insecure HTTP connections to `s3_endpoint`. Defaults to `false`. | | `schema_source_path` | Specifies the URL used to infer the dataset schema. Default to the most recently modified file | For additional CSV, JSON, and Parquet specific parameters, see [File Formats](/docs/next/reference/file_format). ### `s3_url_style`[​](#s3_url_style "Direct link to s3_url_style") The S3 connector supports two request URL styles: * `vhost`: `https://./` (virtual-hosted) * `path`: `https:////` (path-style) When `s3_url_style` is not set, the connector auto-detects the correct style: 1. **IP address endpoints** (e.g. `http://192.168.1.100:9000`) → `path` 2. **Custom hostname endpoints** → DNS lookup on `.`. If the name resolves, `vhost` is used; otherwise `path`. 3. **No custom endpoint** (standard AWS S3) → `vhost`, except a bucket whose name contains a dot (e.g. `my.bucket.name`), which uses `path`. A dotted bucket name cannot use virtual-hosted style over HTTPS on standard AWS: the bucket becomes part of a multi-label hostname (`my.bucket.name.s3..amazonaws.com`) that the AWS wildcard certificate (`*.s3..amazonaws.com`) cannot match, so the TLS handshake fails. The connector defaults these buckets to path style, matching [AWS's own recommendation](https://docs.aws.amazon.com/AmazonS3/latest/userguide/VirtualHosting.html). An explicit `s3_url_style` always takes precedence. Set `s3_url_style` explicitly to skip auto-detection. ## Authentication[​](#authentication "Direct link to Authentication") No authentication is required for public endpoints. For private buckets, set `s3_auth` to `key` or `iam_role`. If `s3_auth` is set to `iam_role`, the connector will automatically load credentials from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials (in .aws/credentials) [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile (in .aws/config) [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts ([IRSA](https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html)) * Relies on the environment variables `AWS_WEB_IDENTITY_TOKEN_FILE` and `AWS_ROLE_ARN` to be present. * These environment variables are automatically injected by EKS when the IAM Role is annotated on the pod Service Account. 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint. * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3 permissions (e.g., `s3:ListBucket`, `s3:GetObject`) to access the files. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. kube2iam [`kube2iam`](https://github.com/jtblin/kube2iam) is a project that provides IAM roles to Kubernetes pods based on annotations. It has been superceded by [IAM Roles for service accounts (IRSA)](https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html), which should be preferred for new deployments. Spice requires `kube2iam >= 0.12` - versions prior to [`0.12`](https://github.com/jtblin/kube2iam/releases/tag/0.12.0) only supported IMDSv1. ### Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") Minimum IAM policy for S3 access: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::company-bucketname-datasets" }, { "Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::company-bucketname-datasets/*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | --------------- | ----------------------------------------------------- | | `s3:ListBucket` | Required. Allows scanning all objects from the bucket | | `s3:GetObject` | Required. Allows fetching objects | ## Types[​](#types "Direct link to Types") Refer to [Object Store Data Types](/docs/next/reference/datatypes/object_store) for data type mapping from object store files to arrow data type. ## Examples[​](#examples "Direct link to Examples") ### Public bucket Example[​](#public-bucket-example "Direct link to Public bucket Example") Create a dataset named `taxi_trips` from a public S3 folder. ``` - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet ``` ### MinIO Example[​](#minio-example "Direct link to MinIO Example") Create a dataset named `cool_dataset` from a Parquet file stored in MinIO. ``` - from: s3://s3-bucket-name/path/to/parquet/cool_dataset.parquet name: cool_dataset params: s3_endpoint: http://my.minio.server s3_region: 'us-east-1' # Best practice for MinIO s3_url_style: path allow_http: true ``` ### Hive Partitioning Example[​](#hive-partitioning-example "Direct link to Hive Partitioning Example") Hive partitioning is a data organization technique that improves query performance by storing data in a hierarchical directory structure based on partition column values. This enables efficient data retrieval by skipping unnecessary data scans. For example, a dataset partitioned by year, month, and day might have a directory structure like: ``` s3://bucket/dataset/year=2024/month=03/day=15/data_file.parquet s3://bucket/dataset/year=2024/month=03/day=16/data_file.parquet ``` Spice can automatically infer these partition columns from the directory structure when `hive_partitioning_enabled` is set to `true`. ``` version: v1 kind: Spicepod name: hive_data datasets: - from: s3://spiceai-public-datasets/hive_partitioned_data/ name: hive_data_infer params: file_format: parquet hive_partitioning_enabled: true ``` ### Schema Source Path example[​](#schema-source-path-example "Direct link to Schema Source Path example") Use `schema_source_path` to speed up dataset registration by specifying a URL to use to infer the schema. ``` - from: s3://spiceai-demo-datasets/taxi_trips/ name: taxi_trips params: file_format: parquet schema_source_path: s3://spiceai-demo-datasets/taxi_trips/2014/1/trips_01.parquet # or s3://spiceai-demo-datasets/taxi_trips/2014/1/ ``` ### Metadata Columns Example[​](#metadata-columns-example "Direct link to Metadata Columns Example") Metadata columns expose per-file S3 object metadata (`_location`, `_last_modified`, `_size`) as virtual columns in query results. See [Metadata Columns](/docs/next/components/data-connectors/#metadata-columns) for full details. ``` - from: s3://spiceai-public-datasets/hive_partitioned_data/ name: hive_data params: file_format: parquet hive_partitioning_enabled: true metadata: _location: enabled _last_modified: enabled _size: enabled ``` Query metadata alongside regular data: ``` SELECT id, value, _location, _size, _last_modified FROM hive_data ORDER BY id LIMIT 5; ``` ``` +----+---------+-------------------------------------------------------------------------------------------+-------+----------------------+ | id | value | _location | _size | _last_modified | +----+---------+-------------------------------------------------------------------------------------------+-------+----------------------+ | 0 | value_0 | s3://spiceai-public-datasets/hive_partitioned_data/year=2022/month=1/day=1/data_0.parquet | 2317 | 2024-10-10T05:36:59Z | | 1 | value_1 | s3://spiceai-public-datasets/hive_partitioned_data/year=2022/month=1/day=1/data_0.parquet | 2317 | 2024-10-10T05:36:59Z | | 2 | value_2 | s3://spiceai-public-datasets/hive_partitioned_data/year=2022/month=1/day=1/data_0.parquet | 2317 | 2024-10-10T05:36:59Z | | 3 | value_3 | s3://spiceai-public-datasets/hive_partitioned_data/year=2022/month=1/day=1/data_0.parquet | 2317 | 2024-10-10T05:36:59Z | | 4 | value_4 | s3://spiceai-public-datasets/hive_partitioned_data/year=2022/month=1/day=1/data_0.parquet | 2317 | 2024-10-10T05:36:59Z | +----+---------+-------------------------------------------------------------------------------------------+-------+----------------------+ ``` Filter by specific file: ``` SELECT id, value FROM hive_data WHERE _location = 's3://spiceai-public-datasets/hive_partitioned_data/year=2023/month=2/day=2/data_1.parquet' ORDER BY id; ``` ``` +----+---------+ | id | value | +----+---------+ | 10 | value_0 | | 11 | value_1 | | 12 | value_2 | | 13 | value_3 | | 14 | value_4 | | 15 | value_5 | | 16 | value_6 | | 17 | value_7 | | 18 | value_8 | | 19 | value_9 | +----+---------+ ``` Aggregate per file: ``` SELECT _location, COUNT(*) AS row_count, _size FROM hive_data GROUP BY _location, _size ORDER BY _location; ``` ``` +-------------------------------------------------------------------------------------------+-----------+-------+ | _location | row_count | _size | +-------------------------------------------------------------------------------------------+-----------+-------+ | s3://spiceai-public-datasets/hive_partitioned_data/year=2022/month=1/day=1/data_0.parquet | 10 | 2317 | | s3://spiceai-public-datasets/hive_partitioned_data/year=2022/month=1/day=2/data_4.parquet | 10 | 2319 | | s3://spiceai-public-datasets/hive_partitioned_data/year=2022/month=3/day=3/data_2.parquet | 10 | 2319 | | s3://spiceai-public-datasets/hive_partitioned_data/year=2023/month=2/day=2/data_1.parquet | 10 | 2319 | | s3://spiceai-public-datasets/hive_partitioned_data/year=2023/month=4/day=1/data_3.parquet | 10 | 2319 | +-------------------------------------------------------------------------------------------+-----------+-------+ ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Limitations[​](#limitations "Direct link to Limitations") Performance Considerations When using the S3 Data connector without acceleration, data is loaded into memory during query execution. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. Memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/next/components/data-accelerators/duckdb) and [`sqlite`](/docs/next/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. Each query retrieves data from the S3 source, which might result in significant network requests and bandwidth consumption. This can affect network performance and incur costs related to data transfer from S3. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure S3 as a data connector in Spice. [S3 Data Connector](https://github.com/spiceai/cookbook/tree/trunk/s3#readme) --- # S3 Data Connector Deployment Guide Production operating guide for the S3 data connector covering IAM authentication, credential chains, file-format tuning, metrics, and observability. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") S3 authentication is selected via `s3_auth`: | Value | Behavior | | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | *(unset)* | Default AWS credential chain (IAM-based). Equivalent to `iam_role` with `iam_role_source: auto`. | | `iam_role` | Load credentials from the AWS credential chain; the source is further narrowed by `iam_role_source`. | | `key` | Use the explicit `s3_key` / `s3_secret` pair. Required for S3-compatible stores that do not speak IAM (MinIO, Cloudflare R2 with keys, Backblaze B2, etc.). | | `public` | Unauthenticated access for public buckets. | ### IAM Role Source[​](#iam-role-source "Direct link to IAM Role Source") When `s3_auth` is unset or `iam_role`, the credential source is controlled by `iam_role_source`: | Value | Behavior | | ---------- | ----------------------------------------------------------------------------------------------------------- | | `auto` | Default AWS credential chain (env vars → shared credentials file → IMDS/ECS/IRSA). | | `metadata` | Restrict to instance/container metadata only: IMDS (EC2), ECS task role, EKS IRSA (pod role). | | `env` | Restrict to environment variables only (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN`). | For production on EKS or ECS, prefer `iam_role_source: metadata` to guarantee the runtime only draws credentials from the workload identity, never from ambient environment variables. ### Key Auth for S3-Compatible Stores[​](#key-auth-for-s3-compatible-stores "Direct link to Key Auth for S3-Compatible Stores") For MinIO, R2, B2, or on-prem S3 gateways: ``` params: s3_auth: key s3_key: ${secrets:s3_key} s3_secret: ${secrets:s3_secret} s3_endpoint: https://minio.internal:9000 s3_region: us-east-1 ``` Keys must be sourced from a secret store in production. See [Secret Stores](/docs/next/components/secret-stores). ### Region Validation[​](#region-validation "Direct link to Region Validation") `s3_region` is validated against AWS's known region set. Uppercase regions are auto-corrected to lowercase with a warning. Unrecognized regions produce a startup warning but do not prevent the connector from starting. Custom S3-compatible endpoints still require a valid-looking AWS region code. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Retry Behavior[​](#retry-behavior "Direct link to Retry Behavior") S3 I/O uses the AWS SDK's default retry strategy: standard adaptive backoff with retries on throttling (`SlowDown`, `503`) and transient network errors. Per-operation retry parameters are not currently exposed at the Spice layer. ### Permanent Failures[​](#permanent-failures "Direct link to Permanent Failures") Authentication failures (`401`, `403`) and missing buckets (`404`) surface immediately as query errors. Unlike the Databricks connector, the S3 connector does not permanently disable itself — subsequent queries re-attempt authentication, so transient IAM or network issues self-heal. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Object store throughput**: S3 scales horizontally per prefix. For large Parquet workloads, partition data by date or tenant to maximize parallel reads. * **Hive partitioning**: Enable `hive_partitioning_enabled: true` when listing partitioned datasets so DataFusion can prune irrelevant partitions at plan time instead of listing and filtering at execution time. * **Schema inference cost**: On first registration, Spice samples files to infer schema. Provide an explicit `schema` in the dataset definition for large datasets to avoid repeated list/head operations. * **DataFusion batch size**: Object-store reads yield 8192-row record batches by default. Increase via runtime tuning for CPU-bound scans over compressed formats. ## Metrics[​](#metrics "Direct link to Metrics") The `runtime-object-store` layer that performs S3 I/O is not instrumented, so Spice does not emit S3 transport metrics (request counts, retries, bytes read). See [Component Metrics](/docs/next/features/observability/component_metrics) for the metrics that are available. The connector does not currently register S3-specific dataset-level instruments. Monitor S3 health via: * Standard AWS CloudWatch metrics on the bucket (`AllRequests`, `4xxErrors`, `5xxErrors`, `TotalRequestLatency`). * Spice's query-execution metrics (`query_duration_ms`, `query_returned_rows`) from `runtime.metrics`. ## Task History[​](#task-history "Direct link to Task History") S3 object reads participate in Spice [task history](/docs/next/reference/task_history) through DataFusion's object-store plan nodes. Individual object GETs are attributed to their enclosing `sql_query` or `accelerated_table_refresh` task via the DataFusion execution plan. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * Writes are not supported; the S3 connector is read-only. * S3 Express One Zone directory buckets are supported transparently via `s3://` URIs when the region and endpoint match. * Server-side encryption with customer-provided keys (SSE-C) is not exposed; SSE-S3 and SSE-KMS work transparently when the role/user has KMS decrypt permission. * Requester-pays buckets are not currently supported. * Cross-region access incurs AWS data-transfer charges; place Spice in the same region as the bucket for best cost and latency. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ------------------------------------------------------------------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | | `The request signature we calculated does not match the signature you provided` | Clock skew or wrong `s3_key`/`s3_secret`. | Verify secret values; check system clock (AWS tolerates only \~15 min drift). | | `Access Denied` | IAM policy lacks `s3:GetObject` or `s3:ListBucket`. | Attach a policy granting read on the bucket and prefix. Cross-account buckets also need bucket policy. | | `NoSuchBucket` | Bucket does not exist in the configured region. | Confirm bucket name and `s3_region`. | | `EnvCredentialsNotSet` on EKS | `iam_role_source: env` while running under IRSA. | Set `iam_role_source: metadata` or `auto`. | | `InvalidSignatureException` against MinIO/R2 | `s3_endpoint` not set or AWS SDK trying to sign for AWS S3. | Set `s3_endpoint` and `s3_region` to match the S3-compatible provider. | | Slow queries on large partitioned datasets | Hive partitioning not enabled; every scan lists all files. | Set `hive_partitioning_enabled: true` and encode partitions as `key=value/` in the path. | --- # ScyllaDB Data Connector The ScyllaDB Data Connector enables federated SQL queries on data stored in [ScyllaDB](https://www.scylladb.com/) clusters using CQL (Cassandra Query Language). ``` datasets: - from: scylladb:users name: users params: scylladb_host: localhost scylladb_port: 9042 scylladb_keyspace: my_app ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `scylladb:{table_name}` where `table_name` is the table identifier in the ScyllaDB keyspace to read from. ``` datasets: - from: scylladb:users name: users params: scylladb_keyspace: my_app ... ``` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: scylladb:users name: app_users params: ... ``` ``` SELECT COUNT(*) FROM app_users; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The ScyllaDB data connector can be configured by providing the following `params`. Use the [secret replacement syntax](/docs/next/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:scylladb_pass}`. | Parameter Name | Description | Required | Default | | --------------------- | ------------------------------------------------------------------------------------------------------------------------------- | -------- | ------- | | `scylladb_host` | Hostname(s) of ScyllaDB nodes. Comma-separated for multiple nodes. Either `scylladb_host` or `scylladb_hosts` must be provided. | Yes\* | - | | `scylladb_hosts` | Alternative to `scylladb_host`. Comma-separated list of hostnames. Either `scylladb_host` or `scylladb_hosts` must be provided. | Yes\* | - | | `scylladb_port` | ScyllaDB CQL native transport port. | No | `9042` | | `scylladb_keyspace` | The keyspace to use for queries. | Yes | - | | `scylladb_user` | Username for authentication. | No | - | | `scylladb_pass` | Password for authentication. | No | - | | `scylladb_datacenter` | Preferred datacenter for connection routing. | No | - | | `scylladb_ssl` | Enable SSL/TLS for connections. Not yet implemented — the parameter is accepted but has no effect. | No | - | | `connection_timeout` | Connection timeout in milliseconds. | No | `10000` | ## Types[​](#types "Direct link to Types") The table below shows the CQL data types supported, along with the type mapping to Apache Arrow types in Spice. | CQL Type | Arrow Type | Notes | | ------------------------------- | ------------------------------ | ------------------------------------- | | `boolean` | `Boolean` | | | `tinyint` | `Int8` | | | `smallint` | `Int16` | | | `int` | `Int32` | | | `bigint` | `Int64` | | | `counter` | `Int64` | Cassandra counter type | | `float` | `Float32` | | | `double` | `Float64` | | | `decimal` | `Decimal128(38, 2)` | Arbitrary precision → fixed precision | | `blob` | `Binary` | | | `date` | `Date32` | Days since epoch | | `timestamp` | `Timestamp(Millisecond, None)` | Milliseconds since epoch | | `time` | `Timestamp(Microsecond, None)` | Time of day | | `text`, `varchar`, `ascii` | `Utf8` | | | `uuid`, `timeuuid` | `Utf8` | String representation | | `inet` | `Utf8` | IP address as string | | `varint` | `Utf8` | Arbitrary precision integer as string | | `duration` | `Utf8` | ISO 8601 duration string | | `list`, `set`, `map` | `Utf8` | JSON string representation | | `tuple<...>` | `Utf8` | String representation | | `frozen` | `Utf8` | Same as underlying collection | | User-Defined Types | `Utf8` | String representation | ### Decimal Handling[​](#decimal-handling "Direct link to Decimal Handling") CQL `decimal` is an arbitrary-precision type, while Arrow `Decimal128` has a maximum precision of 38 digits. The connector uses: * **Precision**: 38 (maximum for Decimal128) * **Scale**: 2 (suitable for monetary/financial data) If a CQL `decimal` value cannot be represented exactly at this precision and scale — because rescaling to scale 2 would drop non-zero fractional digits, or the mantissa does not fit in 128 bits — the connector returns a query error rather than silently truncating, rounding, or returning `NULL`. Values that rescale losslessly are converted normally. ### Date/Time Handling[​](#datetime-handling "Direct link to Date/Time Handling") * **CQL `date`**: Stored as days since epoch with a 2^31 offset. Converted to Arrow `Date32` (signed days since 1970-01-01). * **CQL `timestamp`**: Milliseconds since Unix epoch. Directly mapped to Arrow `Timestamp(Millisecond)`. * **CQL `time`**: Nanoseconds since midnight. Converted to Arrow `Timestamp(Microsecond)` with nanosecond truncation. ## Query Execution[​](#query-execution "Direct link to Query Execution") The connector pushes down partition key and clustering key filters to CQL where possible. Filters that CQL cannot express are evaluated locally by DataFusion after data retrieval. Joins, aggregations, and other SQL operations are always performed locally. ### Filter Pushdown[​](#filter-pushdown "Direct link to Filter Pushdown") The following filters are pushed down to ScyllaDB: | Filter type | Operators | Pushdown behavior | | ---------------------------------- | ------------------------- | --------------------------------------------------------------------- | | Partition key equality | `=` | Always pushed down (Exact) | | Clustering key comparison | `=`, `<`, `<=`, `>`, `>=` | Pushed down when a partition key equality filter is present (Inexact) | | Regular column filters | Any | Not pushed down — evaluated locally by DataFusion | | OR conditions, complex expressions | Any | Not pushed down — evaluated locally by DataFusion | Clustering key filters are marked as **Inexact**, meaning DataFusion re-checks them after retrieval to ensure correctness. ### CQL vs SQL[​](#cql-vs-sql "Direct link to CQL vs SQL") CQL lacks many SQL constructs, which is why most filter types cannot be pushed down: | Feature | SQL | CQL | | ---------------- | --- | --- | | CASE WHEN | ✅ | ❌ | | Subqueries | ✅ | ❌ | | Complex JOINs | ✅ | ❌ | | CAST expressions | ✅ | ❌ | | INTERVAL | ✅ | ❌ | | Window functions | ✅ | ❌ | | NULLS FIRST/LAST | ✅ | ❌ | | COUNT(DISTINCT) | ✅ | ❌ | | Arbitrary WHERE | ✅ | ❌ | ### Projection Pushdown[​](#projection-pushdown "Direct link to Projection Pushdown") Projection pushdown is supported — only the columns referenced in the query are fetched from ScyllaDB. ### Streaming Execution[​](#streaming-execution "Direct link to Streaming Execution") Query results are streamed using the scylla driver's paging mechanism in batches of 8192 rows, minimizing memory usage for large result sets. ## Performance Considerations[​](#performance-considerations "Direct link to Performance Considerations") See the dedicated **[Performance](/docs/next/components/data-connectors/scylladb/performance)** page — partition/clustering-key filters, enabling acceleration, datacenter locality, and connection-timeout tuning. ## Examples[​](#examples "Direct link to Examples") ### Basic Federated Query[​](#basic-federated-query "Direct link to Basic Federated Query") ``` datasets: - from: scylladb:users name: users params: scylladb_host: localhost scylladb_port: 9042 scylladb_keyspace: my_app ``` ### With Authentication[​](#with-authentication "Direct link to With Authentication") ``` datasets: - from: scylladb:orders name: orders params: scylladb_host: scylla-cluster.example.com scylladb_keyspace: ecommerce scylladb_user: app_user scylladb_pass: ${secrets:SCYLLA_PASSWORD} ``` ### Multi-Node Cluster[​](#multi-node-cluster "Direct link to Multi-Node Cluster") ``` datasets: - from: scylladb:events name: events params: scylladb_hosts: node1.scylla.local,node2.scylla.local,node3.scylla.local scylladb_keyspace: analytics scylladb_datacenter: us-west-2 ``` ### With Acceleration[​](#with-acceleration "Direct link to With Acceleration") ``` datasets: - from: scylladb:products name: products params: scylladb_host: ${env:SCYLLADB_HOST} scylladb_keyspace: catalog acceleration: enabled: true engine: duckdb refresh_check_interval: 1h ``` ### Using Environment Variables[​](#using-environment-variables "Direct link to Using Environment Variables") ``` datasets: - from: scylladb:customer name: customer params: scylladb_host: ${env:SCYLLADB_HOST} scylladb_port: ${env:SCYLLADB_PORT} scylladb_keyspace: ${env:SCYLLADB_KEYSPACE} ``` ## Limitations[​](#limitations "Direct link to Limitations") ### CQL Limitations[​](#cql-limitations "Direct link to CQL Limitations") The following SQL operations cannot be pushed down to ScyllaDB and are performed locally: * **JOINs**: All joins are performed locally by DataFusion * **Aggregations**: COUNT, SUM, AVG, etc. are computed locally * **Subqueries**: Nested queries are not supported in CQL * **Window functions**: RANK, ROW\_NUMBER, etc. not supported * **Complex WHERE clauses**: Only partition key equality and clustering key comparisons are pushed down; other filters are evaluated locally * **ORDER BY**: Sorting is done locally ### Connector Limitations[​](#connector-limitations "Direct link to Connector Limitations") * **Read-only**: The connector does not support INSERT, UPDATE, or DELETE operations * **Decimal precision**: Fixed at precision=38, scale=2; values that cannot be rescaled to this precision and scale without loss error rather than being silently truncated * **Collection types**: Lists, sets, and maps are converted to JSON string representation * **Large tables**: Without acceleration, large tables cause significant data transfer ### Data Type Limitations[​](#data-type-limitations "Direct link to Data Type Limitations") * **varint**: Arbitrary-precision integers are converted to strings * **duration**: CQL durations are converted to string representation * **UDTs**: User-defined types are converted to string representation * **Nested collections**: Deeply nested collections become complex JSON strings ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure ScyllaDB as a data connector in Spice. [ScyllaDB Data Connector](https://github.com/spiceai/cookbook/tree/trunk/scylladb#readme) ## See Also[​](#see-also "Direct link to See Also") * [ScyllaDB Documentation](https://docs.scylladb.com/) * [ScyllaDB CQL Reference](https://opensource.docs.scylladb.com/stable/cql/) * [Data Acceleration](/docs/next/features/data-acceleration) --- # ScyllaDB Performance Performance considerations for the [ScyllaDB connector](/docs/next/components/data-connectors/scylladb). Partition key and clustering key filters reduce the amount of data transferred from ScyllaDB, but queries without these filters fetch all table data. Consider the following optimizations: ## Enable Acceleration[​](#enable-acceleration "Direct link to Enable Acceleration") For frequently queried data, enable Spice acceleration to cache data locally: ``` datasets: - from: scylladb:products name: products params: scylladb_host: ${env:SCYLLADB_HOST} scylladb_keyspace: catalog acceleration: enabled: true engine: duckdb refresh_check_interval: 1h ``` ## Configure Datacenter Locality[​](#configure-datacenter-locality "Direct link to Configure Datacenter Locality") Set the datacenter preference to route queries to the nearest nodes: ``` params: scylladb_datacenter: us-east-1 ``` ## Adjust Connection Timeout[​](#adjust-connection-timeout "Direct link to Adjust Connection Timeout") Set connection timeouts appropriately for your network: ``` params: connection_timeout: 30000 # 30 seconds ``` --- # SharePoint Data Connector The SharePoint Data Connector enables federated SQL queries on documents and tabular data stored in SharePoint or OneDrive. ``` datasets: - from: sharepoint:drive:Documents/path:/top_secrets/ name: important_documents params: sharepoint_client_id: ${secrets:SPICE_SHAREPOINT_CLIENT_ID} sharepoint_tenant_id: ${secrets:SPICE_SHAREPOINT_TENANT_ID} sharepoint_client_secret: ${secrets:SPICE_SHAREPOINT_CLIENT_SECRET} ``` #### Example[​](#example "Direct link to Example") ``` SELECT * FROM important_documents limit 1; ``` Returns ```` [ { "created_by_id": "cbccd193-f9f1-4603-b01d-ff6f3e6f2108", "created_by_name": "Jack Eadie", "created_at": "2024-09-09T04:57:00", "c_tag": "\"c:{BD4D130F-2C95-4E59-9F93-85BD0A9E1B19},1\"", "e_tag": "\"{BD4D130F-2C95-4E59-9F93-85BD0A9E1B19},1\"", "id": "01YRH3MPAPCNG33FJMLFHJ7E4FXUFJ4GYZ", "last_modified_by_id": "cbccd193-f9f1-4603-b01d-ff6f3e6f2108", "last_modified_by_name": "Jack Eadie", "last_modified_at": "2024-09-09T04:57:00", "name": "ngx_google_perftools_module.md", "size": 959, "web_url": "https://spiceai.sharepoint.com/Shared%20Documents/md/ngx_google_perftools_module.md", "content": "# Module ngx_google_perftools_module\n\nThe `ngx_google_perftools_module` module (0.6.29) enables profiling of nginx worker processes using [Google Performance Tools](https://github.com/gperftools/gperftools). The module is intended for nginx developers.\n\nThis module is not built by default, it should be enabled with the `--with-google_perftools_module` configuration parameter.\n\n> **Note:** This module requires the [gperftools](https://github.com/gperftools/gperftools) library.\n\n## Example Configuration\n\n```nginx\ngoogle_perftools_profiles /path/to/profile;\n```\n\nProfiles will be stored as `/path/to/profile.`.\n\n## Directives\n\n### google_perftools_profiles\n\n- **Syntax:** `google_perftools_profiles file;`\n- **Default:** —\n- **Context:** `main`\n\nSets a file name that keeps profiling information of nginx worker process. The ID of the worker process is always a part of the file name and is appended to the end of the file name, after a dot.\n" } ] ```` The SharePoint connector supports two `from:` URL styles: * **Metadata listing** (`sharepoint:…` — single colon): one row per drive item with optional file content. Best for browsing folders of PDFs, PPTX, DOCX, etc. as document tables. * **Object-store** (`sharepoint://…` — double slash): tabular access via DataFusion's `ListingTable`. Enables `SELECT`, `INSERT INTO`, `COPY TO`, `COPY FROM`, and `CREATE EXTERNAL TABLE` against CSV, JSON, NDJSON, Parquet, and similar formats stored on SharePoint. ## Configuration[​](#configuration "Direct link to Configuration") ### Parameters[​](#parameters "Direct link to Parameters") | Name | Required? | Description | | ------------------------------ | ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `sharepoint_client_id` | Conditional | The client ID of the Azure AD (Entra) application. Required for every flow except `sharepoint_bearer_token`. | | `sharepoint_tenant_id` | Conditional | The tenant ID of the Azure AD (Entra) application. Required for every flow except `sharepoint_bearer_token`. | | `sharepoint_client_secret` | Conditional | The client secret of the Azure AD (Entra) application. Required for client-credentials, authorization-code, and refresh-token flows. | | `sharepoint_bearer_token` | Conditional | A pre-acquired bearer access token. Generally obtained via `spice login sharepoint` (see [docs](/docs/next/cli/reference/login)). | | `sharepoint_auth_code` | Conditional | OAuth2 authorization code (`auth_code` flow). Requires `sharepoint_client_secret` and `sharepoint_redirect_uri`. | | `sharepoint_refresh_token` | Conditional | OAuth2 refresh token. Requires `sharepoint_client_secret`. | | `sharepoint_device_code` | Conditional | A pre-acquired OAuth2 device code (`device_code` flow). | | `sharepoint_saml_assertion` | Conditional | SAML 2.0 bearer assertion ([RFC 7522](https://datatracker.ietf.org/doc/html/rfc7522)) — exchanges a federated IdP assertion for an Azure AD token. | | `sharepoint_redirect_uri` | Conditional | OAuth2 redirect URI. Required when using `sharepoint_auth_code`. | | `sharepoint_scope` | Optional | OAuth2 scope. Defaults to `https://graph.microsoft.com/.default`. | | `sharepoint_conflict_behavior` | Optional | How writes to an existing path are handled. One of `replace` (default; SharePoint stores a new version), `fail` (reject), or `rename` (write under a unique name). Only `replace` is compatible with `INSERT INTO` / `COPY TO`. Applies only to `sharepoint://`. | | `sharepoint_max_put_bytes` | Optional | Hard cap, in bytes, on a single `put`/multipart upload. Writes above this size are rejected rather than silently buffered. Default: `1073741824` (1 GiB). Applies only to `sharepoint://`. | note Exactly one of `sharepoint_client_secret` (alone, for client-credentials), `sharepoint_bearer_token`, `sharepoint_auth_code` (with `sharepoint_client_secret` + `sharepoint_redirect_uri`), `sharepoint_refresh_token` (with `sharepoint_client_secret`), `sharepoint_device_code`, or `sharepoint_saml_assertion` must be supplied. Combining unrelated auth credentials is rejected at startup. When using the `sharepoint://` URL scheme, the standard listing-table parameters (`file_format`, `csv_has_header`, `csv_delimiter`, `json_pointer`, `hive_partitioning_enabled`, etc.) all apply — see [File Formats](/docs/next/components/data-connectors/#file-formats) and the [Object Store File Formats](/docs/next/components/data-connectors/#object-store-file-formats) reference for the full list. ### `from` formats[​](#from-formats "Direct link to from-formats") The SharePoint connector accepts two `from:` URL styles. #### Metadata listing — `sharepoint:` (single colon)[​](#metadata-listing--sharepoint-single-colon "Direct link to metadata-listing--sharepoint-single-colon") Returns one row per drive item, optionally with the parsed `content` column. Use for document workflows over folders of PDF, PPTX, DOCX, XLSX, etc. ``` from: 'sharepoint::/:' ``` `drive_type` supports the following types: | Drive Type | Description | Example | | ---------- | --------------------------- | ---------------------------------------------------------------- | | `drive` | The SharePoint drive's name | `from: sharepoint:drive:Documents/...` | | `driveId` | The SharePoint drive's ID | `from: sharepoint:driveId:b!Mh8opUGD80ec7zGXgX9r/...` | | `site` | A SharePoint site's name | `from: sharepoint:site:MySite/...` | | `siteId` | A SharePoint site's ID | `from: sharepoint:siteId:b!Mh8opUGD80ec7zGXgX9r/...` | | `group` | A SharePoint group's name | `from: sharepoint:group:MyGroup/...` | | `groupId` | A SharePoint group's ID | `from: sharepoint:groupId:b!Mh8opUGD80ec7zGXgX9r/...` | | `user` | A user's drive by user ID | `from: sharepoint:user:48d31887-5fad-4d73-a9f5-3c356e68a038/...` | | `me` | A user's OneDrive | `from: sharepoint:me/...` | note For the `me` drive type the user is identified based on `sharepoint_bearer_token` and cannot be used with `sharepoint_client_secret`. For a name-based `drive_id`, the connector will attempt to resolve the name to an ID at startup. Within a drive, the SharePoint connector can load documents from: | Description | Example | | -------------------------------- | ---------------------------------------------------------------------- | | The root of the drive | `from: sharepoint:me/root` | | A specific path within the drive | `from: sharepoint:drive:Documents/path:/top_secrets` | | A specific folder ID | `from: sharepoint:group:MyGroup/id:01QM2NJSNHBISUGQ52P5AJQ3CBNOXDMVNT` | #### Object-store — `sharepoint://` (double slash)[​](#object-store--sharepoint-double-slash "Direct link to object-store--sharepoint-double-slash") Routes through an `ObjectStore` plus DataFusion's `ListingTable`. Enables `SELECT`, `INSERT INTO`, `COPY TO`, `COPY FROM`, and `CREATE EXTERNAL TABLE` for CSV, JSON, NDJSON, Parquet, and other tabular formats — and binary round-trips for blobs (PDF, etc.) via `(FORMAT binary)`. | URL form | Description | | -------------------------------------------- | --------------------------------- | | `sharepoint://me/{item-path}` | The authenticated user's OneDrive | | `sharepoint://drives/{drive-id}/{item-path}` | A specific drive by ID | | `sharepoint://sites/{site-id}/{item-path}` | A site's default document library | | `sharepoint://users/{user-id}/{item-path}` | A user's default drive | | `sharepoint://groups/{group-id}/{item-path}` | A group's default drive | Path segments are percent-decoded, so site IDs containing `,` (e.g. `contoso.sharepoint.com,abc-def,ghi-jkl`) and file paths containing spaces work without extra escaping beyond standard URL encoding. `file_format` is auto-inferred from the URL extension when omitted, so `from: sharepoint://me/Documents/Q4.xlsx` resolves without specifying `file_format: xlsx`. ## Authentication[​](#authentication "Direct link to Authentication") The SharePoint connector supports six authentication flows. Configure exactly one — the connector picks the flow based on which auth parameter is set. See the [Required Microsoft Graph permissions](#required-microsoft-graph-permissions) section below for the API permissions each flow requires. | Flow | Parameters | Notes | | --------------------------------------------------------------------------- | ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | | Client credentials | `sharepoint_client_secret` | Service principal / daemon workloads. | | Bearer token (passthrough) | `sharepoint_bearer_token` | Short-lived broker-minted token. Typically obtained via `spice login sharepoint`. | | Authorization code | `sharepoint_auth_code` + `sharepoint_client_secret` + `sharepoint_redirect_uri` | Caller has already completed the user-agent redirect and captured the `auth_code`. | | Refresh token | `sharepoint_refresh_token` + `sharepoint_client_secret` | Renewal from a prior grant. | | Device code | `sharepoint_device_code` | Caller has already obtained a device code. | | SAML 2.0 bearer ([RFC 7522](https://datatracker.ietf.org/doc/html/rfc7522)) | `sharepoint_saml_assertion` | Federated IdP (Okta, Ping, ADFS, …) assertion → Azure AD token. | ### Creating an Enterprise Application[​](#creating-an-enterprise-application "Direct link to Creating an Enterprise Application") To use the SharePoint connector with service principal authentication, create an Azure AD application and grant it the necessary permissions. This same app registration also supports the OAuth2 user flows above. 1. Create a new Azure AD application in the [Azure portal](https://portal.azure.com/#view/Microsoft_AAD_IAM/ActiveDirectoryMenuBlade/~/Overview). 2. Under the application's `API permissions`, add the permissions listed in [Required Microsoft Graph permissions](#required-microsoft-graph-permissions). * For service principal authentication, Application permissions are required. * For user authentication, only delegated permissions are required. 3. (For user authentication): Under the application's `Authentication`, add `http://localhost` as a Mobile and desktop applications redirect URI. 4. Add `sharepoint_client_id` (from the `Application (Client) ID` field) and `sharepoint_tenant_id` to the connector configuration. 5. (For service principal authentication): Under the application's `Certificates & secrets`, create a new client secret. Use this for the `sharepoint_client_secret` parameter. ### Required Microsoft Graph permissions[​](#required-microsoft-graph-permissions "Direct link to Required Microsoft Graph permissions") Read-only workflows require: * `Sites.Read.All` * `Files.Read.All` * `User.Read` * `GroupMember.Read.All` Write workflows (`INSERT INTO`, `COPY TO`, `CREATE EXTERNAL TABLE` over `sharepoint://`) additionally require: * `Files.ReadWrite` (for personal drive / specific drive writes), and * `Sites.ReadWrite.All` (for site-scoped writes). ### Default Spice Application[​](#default-spice-application "Direct link to Default Spice Application") For your convenience, Spice AI maintains a default Entra (Azure AD) application that can be used for authentication against your SharePoint instance. This application requires OAuth2 authentication. To use it: ``` datasets: - from: sharepoint:me/root # Set the drive and subpath as needed. name: my_data params: sharepoint_client_id: f2b3116e-b4c4-464f-80ec-73cd9d9886b4 sharepoint_tenant_id: #${env:TENANT_ID} sharepoint_bearer_token: ${secrets:SPICE_SHAREPOINT_BEARER_TOKEN} ``` And set the `SPICE_SHAREPOINT_BEARER_TOKEN` secret via: ``` spice login sharepoint --tenant-id $TENANT_ID --client-id f2b3116e-b4c4-464f-80ec-73cd9d9886b4 ``` ## Read/write examples (`sharepoint://`)[​](#readwrite-examples-sharepoint "Direct link to readwrite-examples-sharepoint") Reading a CSV from a site library: ``` datasets: - from: sharepoint://sites/contoso.sharepoint.com,11111111-2222-3333-4444-555555555555,66666666-7777-8888-9999-aaaaaaaaaaaa/Shared%20Documents/reports/sales.csv name: sales params: sharepoint_client_id: ${secrets:SPICE_SHAREPOINT_CLIENT_ID} sharepoint_tenant_id: ${secrets:SPICE_SHAREPOINT_TENANT_ID} sharepoint_client_secret: ${secrets:SPICE_SHAREPOINT_CLIENT_SECRET} file_format: csv csv_has_header: 'true' ``` Inserting rows: ``` INSERT INTO sales VALUES ('Q2', 123456.78); ``` Copying a query result out as Parquet: ``` COPY (SELECT * FROM orders WHERE year = 2026) TO 'sharepoint://me/Documents/exports/orders-2026.parquet' (FORMAT parquet); ``` Creating an external table over a folder of Parquet files: ``` CREATE EXTERNAL TABLE reports STORED AS PARQUET LOCATION 'sharepoint://sites/{site-id}/Shared%20Documents/reports/'; ``` Round-tripping a binary blob (e.g. a PDF): ``` COPY (SELECT content FROM cache WHERE name = 'Q2-report.pdf') TO 'sharepoint://me/Documents/Q2-report.pdf' (FORMAT binary); ``` Limitations * The `sharepoint:` (metadata-listing) syntax cannot create a dataset from a single file (e.g. an Excel spreadsheet) — datasets must be created from a folder of documents. Use the `sharepoint://` object-store syntax for single-file workflows. * For `INSERT INTO` and `COPY TO`, only `sharepoint_conflict_behavior=replace` is supported. `fail` and `rename` cause writes to be rejected with a clear error. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Sharepoint as a data connector in Spice. [SharePoint Data Connector](https://github.com/spiceai/cookbook/tree/trunk/sharepoint#readme) --- # SMB Data Connector SMB (Server Message Block) is a network file sharing protocol that provides shared access to files, printers, and serial ports. It is commonly used in Windows environments for network shares but is also supported on Linux (via Samba) and macOS. The SMB Data Connector enables federated SQL query across [supported file formats](/docs/next/components/data-connectors#file-formats) stored on SMB/CIFS network shares. It requires the SMB 3.1.1 dialect — the only protocol version the connector negotiates — and is compatible with Windows Server 2016+ and Samba 4.x file shares, NAS devices (Synology, QNAP, etc.) configured for SMB 3.1.1, and Azure Files. ## Quickstart[​](#quickstart "Direct link to Quickstart") Connect to an SMB share and query Parquet files: ``` datasets: - from: smb://fileserver/data/sales/ name: sales params: file_format: parquet smb_user: ${secrets:smb_user} smb_pass: ${secrets:smb_pass} ``` Query the data using SQL: ``` SELECT * FROM sales LIMIT 10; ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Specifies the SMB server, share, and path to connect to. **Format:** `smb:////` * ``: The server hostname or IP address * ``: The share name on the server * ``: Path to a file or directory within the share (optional) When pointing to a directory, Spice loads all files within that directory recursively. **Examples:** ``` # Connect to a specific file from: smb://fileserver/data/reports/quarterly.parquet # Connect to a directory (loads all files) from: smb://fileserver/data/sales/ # Connect to share root from: smb://fileserver/data/ # Using IP address from: smb://192.168.1.100/share/data.parquet ``` ### `name`[​](#name "Direct link to name") The dataset name used as the table name in SQL queries. Cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------ | | `file_format` | Required when connecting to a directory. See [File Formats](/docs/next/components/data-connectors/#file-formats). | | `smb_user` | Username for SMB authentication. Use [secrets](/docs/next/components/secret-stores) syntax: `${secrets:smb_user}`. | | `smb_pass` | Password for SMB authentication. Use [secrets](/docs/next/components/secret-stores) syntax: `${secrets:smb_pass}`. | | `smb_port` | SMB server port. Default: `445`. | | `client_timeout` | Connection timeout duration. E.g. `30s`, `1m`. No timeout when unset. | | `hive_partitioning_enabled` | Enable [Hive-style partitioning](#hive-partitioning) from folder structure. Default: `false`. | ## Examples[​](#examples "Direct link to Examples") ### Basic Connection[​](#basic-connection "Direct link to Basic Connection") Connect to a Windows file share with domain credentials: ``` datasets: - from: smb://fileserver.corp.local/shared/analytics/ name: analytics params: file_format: parquet smb_user: ${secrets:smb_user} smb_pass: ${secrets:smb_pass} ``` ### Domain Authentication[​](#domain-authentication "Direct link to Domain Authentication") For Windows domain environments, include the domain in the username: ``` datasets: - from: smb://fileserver/data/reports/ name: reports params: file_format: csv csv_has_header: true smb_user: DOMAIN\username smb_pass: ${secrets:smb_pass} ``` The domain can be specified as `DOMAIN\user` or `user@domain`. ### Reading a Single File[​](#reading-a-single-file "Direct link to Reading a Single File") When pointing to a specific file, the format is inferred from the file extension: ``` datasets: - from: smb://nas.local/backups/database_export.parquet name: database_export params: smb_user: ${secrets:smb_user} smb_pass: ${secrets:smb_pass} ``` ### Connection with Timeout[​](#connection-with-timeout "Direct link to Connection with Timeout") Configure a timeout for slow or unreliable network connections: ``` datasets: - from: smb://remote-server.example.com/data/ name: remote_data params: file_format: parquet smb_user: ${secrets:smb_user} smb_pass: ${secrets:smb_pass} client_timeout: 60s ``` ### Custom Port Configuration[​](#custom-port-configuration "Direct link to Custom Port Configuration") Connect to SMB servers running on non-standard ports: ``` datasets: - from: smb://custom-server.local/share/ name: custom_data params: file_format: parquet smb_port: 4450 smb_user: ${secrets:smb_user} smb_pass: ${secrets:smb_pass} ``` ### Hive Partitioning[​](#hive-partitioning "Direct link to Hive Partitioning") Enable Hive-style partitioning to automatically extract partition columns from the folder structure: ``` datasets: - from: smb://datalake.corp.local/warehouse/events/ name: events params: file_format: parquet smb_user: ${secrets:smb_user} smb_pass: ${secrets:smb_pass} hive_partitioning_enabled: true ``` Given a folder structure like: ``` /events/ region=us/ year=2024/ data.parquet region=eu/ year=2024/ data.parquet ``` Queries can filter on partition columns: ``` SELECT * FROM events WHERE region = 'us' AND year = '2024'; ``` ### Multiple Shares from One Server[​](#multiple-shares-from-one-server "Direct link to Multiple Shares from One Server") Load different datasets from multiple shares on the same server: ``` datasets: - from: smb://fileserver/sales/ name: sales params: file_format: parquet smb_user: ${secrets:smb_user} smb_pass: ${secrets:smb_pass} - from: smb://fileserver/inventory/ name: inventory params: file_format: csv smb_user: ${secrets:smb_user} smb_pass: ${secrets:smb_pass} ``` ### Accelerated Dataset[​](#accelerated-dataset "Direct link to Accelerated Dataset") Enable local acceleration for faster repeated queries: ``` datasets: - from: smb://archive.corp.local/historical/ name: historical_data params: file_format: parquet smb_user: ${secrets:smb_user} smb_pass: ${secrets:smb_pass} acceleration: enabled: true engine: duckdb refresh_check_interval: 1h ``` Acceleration is recommended for frequently queried data, as SMB operations involve network round-trips for directory listing and file reads. ### TPC-H Benchmark Example[​](#tpc-h-benchmark-example "Direct link to TPC-H Benchmark Example") For benchmark configurations with multiple related tables, use YAML anchors to avoid repeating parameters: ``` datasets: - from: smb://192.168.1.100/data/benchmarks/tpch/customer.parquet name: customer params: &smb_params file_format: parquet smb_user: ${secrets:smb_user} smb_pass: ${secrets:smb_pass} - from: smb://192.168.1.100/data/benchmarks/tpch/lineitem.parquet name: lineitem params: *smb_params - from: smb://192.168.1.100/data/benchmarks/tpch/orders.parquet name: orders params: *smb_params ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores for secure credential management. Store SMB credentials in a secret store and reference them using the `${secrets:key}` syntax. ``` datasets: - from: smb://fileserver/data/ name: secure_data params: file_format: parquet smb_user: ${secrets:smb_username} smb_pass: ${secrets:smb_password} ``` For detailed information, refer to the [secret stores documentation](/docs/next/components/secret-stores). ## Limitations[​](#limitations "Direct link to Limitations") The SMB connector is read-only. Write operations such as `put`, `delete`, and `copy` are not supported. Only username/password authentication is supported. Kerberos and NTLM ticket-based authentication are not available. Direct network access to the SMB server is required; proxy connections are not supported. The firewall must permit SMB traffic (port 445 by default). ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") ### Connection Timeouts[​](#connection-timeouts "Direct link to Connection Timeouts") If connections frequently timeout, increase the `client_timeout` value: ``` params: client_timeout: 120s ``` Verify network connectivity to the server and check that firewall rules permit port 445. ### Authentication Failures[​](#authentication-failures "Direct link to Authentication Failures") Common causes of authentication failures: * **Domain not specified**: For domain-joined servers, include the domain: `DOMAIN\username` or `username@domain` * **Incorrect credentials**: Verify username and password are correctly stored in your secret store * **Permission denied**: Ensure the user has read access to the share and files * **Account locked**: Check if the SMB account is not locked on the server ### Share Access Errors[​](#share-access-errors "Direct link to Share Access Errors") If you receive "share not found" errors: * Verify the share name is correct (share names are case-insensitive on Windows) * Ensure the share exists and is accessible from the network where Spice is running * Check firewall rules: SMB uses TCP port 445 * Confirm the user has permission to access the share ### File Format Errors[​](#file-format-errors "Direct link to File Format Errors") When connecting to a directory, ensure `file_format` is specified and matches the actual file types in the directory. Spice expects all files in a directory to have the same format. ### Debug Logging[​](#debug-logging "Direct link to Debug Logging") Enable debug logging to diagnose SMB connection issues: ``` RUST_LOG=runtime_object_store::store::smb=debug spiced ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure SMB as a data connector in Spice. [SMB Data Connector](https://github.com/spiceai/cookbook/tree/trunk/smb#readme) --- # Snowflake Data Connector The Snowflake Data Connector enables federated SQL queries across datasets in the [Snowflake Cloud Data Warehouse](https://www.snowflake.com/). ``` datasets: - from: snowflake:DATABASE.SCHEMA.TABLE name: table params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin ``` Hint Unquoted identifiers are normalized to lowercase by Spice. Snowflake normalizes unquoted identifiers to uppercase, so unquoted identifiers in the `from` field should be UPPERCASED (e.g. `snowflake:MY_DATABASE.MY_SCHEMA.MY_TABLE`). To reference a table created with mixed-case in Snowflake, wrap it in double quotes: `snowflake:MY_DATABASE.MY_SCHEMA."mixedCaseTable"`. See [Snowflake identifier resolution](https://docs.snowflake.com/en/sql-reference/identifiers-syntax#label-identifier-casing) and [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") A Snowflake fully qualified table name (database.schema.table). For instance `snowflake:SNOWFLAKE_SAMPLE_DATA.TPCH_SF1.LINEITEM` or `snowflake:TAXI_DATA."2024".TAXI_TRIPS` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords) or any of the following keywords that are reserved by Snowflake: * `START` * `CONNECT` * `MATCH_RECOGNIZE` * `SAMPLE` * `TABLESAMPLE` * `FROM` ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `snowflake_warehouse` | Optional, specifies the [Snowflake Warehouse](https://docs.snowflake.com/en/user-guide/warehouses-tasks) to use | | `snowflake_role` | Optional, specifies the role to use for accessing Snowflake data | | `snowflake_account` | Required, specifies the Snowflake account-identifier | | `snowflake_username` | Required, specifies the Snowflake username to use for accessing Snowflake data | | `snowflake_auth_type` | Optional, specifies the authentication type. Accepts `password` (or `snowflake`) and `keypair` (or `snowflake_jwt`); matched case-insensitively. Defaults to password authentication unless only key-pair credentials (`snowflake_private_key` or `snowflake_private_key_path`) are provided, in which case key-pair is selected automatically. | | `snowflake_password` | Required for password authentication. Specifies the Snowflake password for authentication. | | `snowflake_private_key` | Required (one of) for key-pair authentication. Mutually exclusive with `snowflake_private_key_path`. Specifies the private key content as a string. | | `snowflake_private_key_path` | Required (one of) for key-pair authentication. Mutually exclusive with `snowflake_private_key`. Specifies the path to the private key file. | | `snowflake_private_key_passphrase` | Required when the private key is encrypted. Specifies the passphrase to decrypt the private key. | ## Auth[​](#auth "Direct link to Auth") The connector supports password-based and [key-pair](https://docs.snowflake.com/en/user-guide/key-pair-auth) authentication that must be configured using `spice login snowflake` or using [Secrets Stores](/docs/next/components/secret-stores). Login requires the account identifier ('orgname-accountname' format) - use [Finding the organization and account name for an account](https://docs.snowflake.com/en/user-guide/admin-account-identifier#finding-the-organization-and-account-name-for-an-account) instructions. ![](/img/snowflake/ui-snowsight-account-identifier.png) * Env * Kubernetes * Keyring ``` # Password-based SPICE_SNOWFLAKE_ACCOUNT= \ SPICE_SNOWFLAKE_USERNAME= \ SPICE_SNOWFLAKE_PASSWORD= \ spice run # Key-pair using private key file (the `` is optional, used for encrypted keys only) SPICE_SNOWFLAKE_ACCOUNT= \ SPICE_SNOWFLAKE_USERNAME= \ SPICE_SNOWFLAKE_PRIVATE_KEY_PATH= \ SPICE_SNOWFLAKE_PRIVATE_KEY_PASSPHRASE= \ spice run # Key-pair using private key content (the `` is optional, used for encrypted keys only) SPICE_SNOWFLAKE_ACCOUNT= \ SPICE_SNOWFLAKE_USERNAME= \ SPICE_SNOWFLAKE_PRIVATE_KEY= \ SPICE_SNOWFLAKE_PRIVATE_KEY_PASSPHRASE= \ spice run ``` or using the Spice CLI: ``` # Password-based spice login snowflake -a -u -p # Key-pair (the `` is an optional parameter and is used for encrypted private key only) spice login snowflake -a -u -k -s ``` The CLI will create or update an `.env` file that looks like: ``` SPICE_SNOWFLAKE_ACCOUNT="account" SPICE_SNOWFLAKE_PASSWORD="pass" SPICE_SNOWFLAKE_USERNAME="user" ``` Configure the spicepod to load secrets from the `env` secret store: (Note: This is the default setting) `spicepod.yaml` ``` version: v1 kind: Spicepod name: spice-app secrets: - from: env name: env datasets: - from: snowflake:DATABASE.SCHEMA.TABLE name: table params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin snowflake_username: ${env:SPICE_SNOWFLAKE_USERNAME} snowflake_password: ${env:SPICE_SNOWFLAKE_PASSWORD} snowflake_account: ${env:SPICE_SNOWFLAKE_ACCOUNT} ``` Learn more about [Env Secret Store](/docs/next/components/secret-stores/env). ``` # Password-based kubectl create secret generic snowflake \ --from-literal=account='' \ --from-literal=username='' \ --from-literal=password='' # Key-pair using private key file (the `` is optional, used for encrypted keys only) kubectl create secret generic snowflake \ --from-literal=account='' \ --from-literal=username='' \ --from-literal=private_key_path='' \ --from-literal=private_key_passphrase='' # Key-pair using private key content (the `` is optional, used for encrypted keys only) kubectl create secret generic snowflake \ --from-literal=account='' \ --from-literal=username='' \ --from-file=private_key='' \ --from-literal=private_key_passphrase='' ``` `spicepod.yaml` (password-based) ``` version: v1 kind: Spicepod name: spice-app secrets: - from: kubernetes:snowflake name: snowflake datasets: - from: snowflake:DATABASE.SCHEMA.TABLE name: table params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin snowflake_username: ${snowflake:username} snowflake_password: ${snowflake:password} snowflake_account: ${snowflake:account} ``` Learn more about [Kubernetes Secret Store](/docs/next/components/secret-stores/kubernetes). Add new keychain entries (macOS) for user and password: ``` # Password-based security add-generic-password -l "Snowflake Secret" \ -a spiced -s spice_snowflake_password\ -w # Key-pair using private key content (the `` is optional, used for encrypted keys only) security add-generic-password -l "Snowflake Secret" \ -a spiced -s spice_snowflake_private_key\ -w $(cat '') ``` `spicepod.yaml` (key-pair) ``` version: v1 kind: Spicepod name: spice-app secrets: - from: keyring name: keyring datasets: - from: snowflake:DATABASE.SCHEMA.TABLE name: table params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin snowflake_auth_type: keypair snowflake_username: user_name snowflake_private_key: ${keyring:spice_snowflake_private_key} snowflake_account: account_identifier ``` Learn more about [Keyring Secret Store](/docs/next/components/secret-stores/keyring). ## Write Support[​](#write-support "Direct link to Write Support") This connector supports writing data to Snowflake tables using SQL [`INSERT INTO`](/docs/next/reference/sql/dml#insert), `UPDATE`, and `DELETE FROM` statements. To enable writes, set `access: read_write` on the dataset: ``` datasets: - from: snowflake:DATABASE.SCHEMA.TABLE name: table access: read_write params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin ``` ``` -- Insert rows INSERT INTO table (id, name, amount) VALUES (1, 'Alice', 100.0), (2, 'Bob', 200.0); -- Update rows UPDATE table SET amount = 150.0 WHERE id = 1; -- Delete rows DELETE FROM table WHERE id = 2; ``` For more details, see [Data Ingestion](/docs/next/features/data-ingestion). ## Example[​](#example "Direct link to Example") ``` datasets: - from: snowflake:SNOWFLAKE_SAMPLE_DATA.TPCH_SF1.LINEITEM name: lineitem params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin ``` Account Identifier Formats The `snowflake_account` parameter accepts any of the [Snowflake account identifier formats](https://docs.snowflake.com/en/user-guide/admin-account-identifier): * Preferred org/account names (`myorg-myaccount`) * SQL/data-sharing org-qualified names (`myorg.myaccount`) * Full `snowflakecomputing.com` account URLs (`https://myorg-myaccount.snowflakecomputing.com`) * Legacy account locators (`xy12345`, `xy12345.us-east-2.aws`) * SnowGov locators (`xy12345.fhplus.us-gov-west-1.aws`) Invalid identifiers fail at startup with an actionable error. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/next/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/next/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Snowflake as a data connector in Spice. [Snowflake Data Connector](https://github.com/spiceai/cookbook/tree/trunk/snowflake#readme) --- # Apache Spark Connector Apache Spark as a connector for federated SQL query against a Spark Cluster using [Spark Connect](https://spark.apache.org/docs/latest/spark-connect-overview.html) ``` datasets: - from: spark:spiceai.datasets.my_awesome_table name: my_table params: spark_remote: sc://localhost:15002 ``` info Unquoted identifiers are normalized to lowercase. To reference a table with mixed-case characters, wrap each case-sensitive part in double quotes: `spark:my_catalog."MySchema"."MyTable"`. See [Identifier Case Sensitivity](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting). ## Configuration[​](#configuration "Direct link to Configuration") * `spark_remote`: Required. A [spark remote](https://spark.apache.org/docs/latest/spark-connect-overview.html#set-sparkremote-environment-variable) connection URI. Refer to [spark connect client connection string](https://github.com/apache/spark/blob/master/connector/connect/docs/client-connection-string) for parameters in URI. The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ### Auth Examples[​](#auth-examples "Direct link to Auth Examples") Spark clusters configured to accept authenticated requests should not set `spark_remote` as an inline dataset param, as it will contain sensitive data. For this case, use the [secret replacement syntax](/docs/next/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:my_spark_remote}`. Check [Secrets Stores](/docs/next/components/secret-stores) for more details. * Env * Kubernetes * Keyring ``` SPICE_SPARK_REMOTE= \ spice run # Or using the CLI to configure the secrets into an `.env` file spice login spark --spark_remote ``` `.env` ``` SPICE_SPARK_REMOTE= ``` `spicepod.yaml` ``` version: v1 kind: Spicepod name: spice-app secrets: - from: env name: env datasets: - from: spark:spiceai.datasets.my_awesome_table name: my_table params: spark_remote: ${env:SPICE_SPARK_REMOTE} ``` Learn more about [Env Secret Store](/docs/next/components/secret-stores/env). ``` kubectl create secret generic spark \ --from-literal=spark_remote='' ``` `spicepod.yaml` ``` version: v1 kind: Spicepod name: spice-app secrets: - from: kubernetes:spark name: spark datasets: - from: spark:spiceai.datasets.my_awesome_table name: my_table params: spark_remote: ${spark:spark_remote} ``` Learn more about [Kubernetes Secret Store](/docs/next/components/secret-stores/kubernetes). Add new keychain entry (macOS) with the spark remote: ``` security add-generic-password -l "Spark Remote" \ -a spiced -s spice_spark_remote \ -w ``` `spicepod.yaml` ``` version: v1 kind: Spicepod name: spice-app secrets: - from: keyring name: keyring datasets: - from: spark:spiceai.datasets.my_awesome_table name: my_table params: spark_remote: ${keyring:spice_spark_remote} ``` Learn more about [Keyring Secret Store](/docs/next/components/secret-stores/keyring). ## Limitations[​](#limitations "Direct link to Limitations") * Correlated scalar subqueries are only supported in filters, aggregations, projections, and UPDATE/MERGE/DELETE commands. [Spark Docs](https://spark.apache.org/docs/latest/sql-error-conditions-unsupported-subquery-expression-category-error-class.html#unsupported_correlated_scalar_subquery) * The Spark connector does not yet support streaming query results from Spark. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Spark as a data connector in Spice. [Apache Spark Data Connector](https://github.com/spiceai/cookbook/tree/trunk/spark#readme) --- # Spice.ai Data Connector The Spice.ai Data Connector federates SQL queries across **another Spice runtime** over [Arrow Flight](https://arrow.apache.org/docs/format/Flight.html). The same connector targets two architectures: * **Spice → Spice Cloud Platform** — federate datasets hosted on the managed [Spice.ai Cloud Platform](https://docs.spice.ai/building-blocks/datasets). Requires a free [Spice.ai account](https://spice.ai/login). * **Spice → Spice (self-hosted)** — federate to another Spice runtime running anywhere reachable over the network. The most common pattern is a **cluster-sidecar**: a thin Spice runtime co-located with each application instance forwards queries to a heavier-weight upstream Spice (or a cluster of them) that owns the data and acceleration. Both architectures use the same `spice.ai` (or legacy `spiceai`) `from:` URI scheme; the difference is in the URI format and the `endpoint` parameter. ## Quick start[​](#quick-start "Direct link to Quick start") ### Spice → Spice Cloud Platform[​](#spice--spice-cloud-platform "Direct link to Spice → Spice Cloud Platform") ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips params: spiceai_api_key: ${secrets:SPICEAI_API_KEY} spiceai_region: us-east-1 ``` `spice login` writes the API key to `~/.spice/auth` and exposes it as `SPICEAI_API_KEY` via the [env secret store](/docs/next/components/secret-stores/env). The `spiceai_region` parameter selects the Spice Cloud region the API key was created in. ### Spice → Spice (self-hosted, cluster-sidecar)[​](#spice--spice-self-hosted-cluster-sidecar "Direct link to Spice → Spice (self-hosted, cluster-sidecar)") ``` datasets: - from: spice.ai:http://upstream-spice.internal:50051 name: orders # table name on the upstream runtime params: spiceai_api_key: ${secrets:UPSTREAM_API_KEY} # match runtime.auth.api-key on upstream ``` The local sidecar Spice now exposes `orders` as if it were a local dataset, federating every query to `upstream-spice.internal:50051`. Combine with `acceleration.enabled: true` to cache hot data at the sidecar. ## Configuration[​](#configuration "Direct link to Configuration") ### `from:` URI formats[​](#from-uri-formats "Direct link to from-uri-formats") | URI | Mode | Notes | | ------------------------------------------- | ----------- | ------------------------------------------------------------------------------------ | | `spice.ai///datasets/` | Cloud | Path-style. Most common. | | `spice.ai://datasets/` | Cloud | Colon-style. Equivalent to the path-style form. | | `spice.ai:////datasets/` | Cloud | URL-style. Equivalent. | | `spice.ai/
` | Cloud | Short form when the dataset isn't under a 4-segment `//datasets/...` path. | | `spice.ai:http://:` | Self-hosted | Connects without TLS. The local dataset's `name:` field is the upstream table name. | | `spice.ai:https://:` | Self-hosted | TLS-encrypted Flight. | | `spice.ai:grpc+tls://:` | Self-hosted | TLS-encrypted Flight (alias for `https://`). | The legacy `spiceai:` prefix (without the dot) is accepted in every form above for backward compatibility. `grpc://` is not supported Plain `grpc://` (without TLS) is rejected at startup. Use `http://` for unencrypted connections to a local sidecar, or `https://` / `grpc+tls://` for production. ### Parameters[​](#parameters "Direct link to Parameters") | Parameter | Default | Description | | ------------------------------------- | ----------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `spiceai_api_key` | None | API key. **Required** for Cloud. Optional for self-hosted (omit for anonymous; supply to match the upstream's `runtime.auth.api-key.keys`). Use `${secrets:...}`. | | `spiceai_token` | None | Legacy alias for `spiceai_api_key`. Prefer `spiceai_api_key` in new configurations. | | `spiceai_region` | None | Cloud region (e.g. `us-east-1`). **Required** when targeting Cloud without an explicit `spiceai_endpoint`. Used to build `https://-prod-aws-flight.spiceai.io`. | | `spiceai_endpoint` | Built from `spiceai_region` (Cloud) | Override the Flight endpoint URL. Use for VPC / regional Cloud endpoints, or to point at a self-hosted Spice runtime. Schemes: `http://`, `https://`, `grpc+tls://`. | | `spiceai_flight_endpoint` | None | Legacy alias for `spiceai_endpoint`. | | `spiceai_tls_ca_certificate_file` | System cert store | Path to a CA certificate file (PEM) to verify a self-hosted upstream that uses a private CA. Ignored for `http://` endpoints. | | `spiceai_tls_client_certificate_file` | None | Path to a PEM client certificate chain for mutual TLS (mTLS). Must be set together with `spiceai_tls_client_key_file`. Mutually exclusive with `spiceai_tls_client_certificate`. | | `spiceai_tls_client_key_file` | None | Path to the PEM private key matching `spiceai_tls_client_certificate_file`. Must be set together with `spiceai_tls_client_certificate_file`. Mutually exclusive with `spiceai_tls_client_key`. | | `spiceai_tls_client_certificate` | None | Inline PEM client certificate chain for mutual TLS (mTLS). Use the [secret replacement syntax](/docs/next/components/secret-stores) to load from a secret store, e.g. `${secrets:my_cert}`. Must be set together with `spiceai_tls_client_key`. Mutually exclusive with `spiceai_tls_client_certificate_file`. | | `spiceai_tls_client_key` | None | Inline PEM private key for mutual TLS (mTLS). Use the [secret replacement syntax](/docs/next/components/secret-stores) to load from a secret store, e.g. `${secrets:my_key}`. Must be set together with `spiceai_tls_client_certificate`. Mutually exclusive with `spiceai_tls_client_key_file`. | #### Endpoint resolution order[​](#endpoint-resolution-order "Direct link to Endpoint resolution order") The connector picks a Flight endpoint in this order: 1. `spiceai_endpoint` (or its legacy alias `spiceai_flight_endpoint`). 2. The URI in `from:` if it begins with `http://`, `https://`, or `grpc+tls://` (self-hosted forms). 3. `https://-prod-aws-flight.spiceai.io` (Cloud, regional). The legacy `https://flight.spiceai.io` host is rewritten to the regional URL automatically. If both `spiceai_endpoint` and `spiceai_region` are set and they refer to different Cloud regions, the runtime fails fast with a region-mismatch error. For self-hosted endpoints, `spiceai_region` is not validated. ### `name`[​](#name "Direct link to name") The dataset name as exposed on the local runtime. The dataset name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). For self-hosted (`spice.ai:http://...`) `from:` URIs, the connector uses **the `name:` value as the upstream table reference** — so `name: orders` queries the `orders` table on the upstream Spice runtime. For Cloud `from:` URIs, the table reference is parsed from the path. ## Spice → Spice Cloud Platform[​](#spice--spice-cloud-platform-1 "Direct link to Spice → Spice Cloud Platform") ### Authentication[​](#authentication "Direct link to Authentication") API keys are issued in the [Spice.ai Console](https://spice.ai). The `spice login` CLI command writes the active key to a local `.env` file via the [env secret store](/docs/next/components/secret-stores/env), making it available as `SPICEAI_API_KEY` in spicepod parameters: ``` params: spiceai_api_key: ${secrets:SPICEAI_API_KEY} ``` For production, prefer a managed [secret store](/docs/next/components/secret-stores) (Kubernetes Secrets, AWS Secrets Manager, HashiCorp Vault) over the env file. API keys do not expire — rotate manually via the Console and update the secret store. ### Region & endpoint[​](#region--endpoint "Direct link to Region & endpoint") `spiceai_region` is required when targeting Cloud without an explicit endpoint. Spice builds the regional Flight URL automatically: ``` params: spiceai_api_key: ${secrets:SPICEAI_API_KEY} spiceai_region: us-east-1 ``` For VPC-peered or otherwise-routed Cloud endpoints, set `spiceai_endpoint` explicitly. The endpoint must match `spiceai_region` if it's a recognized Cloud regional URL. ### Cloud example[​](#cloud-example "Direct link to Cloud example") ``` datasets: - from: spice.ai/spiceai/tpch/datasets/customer name: tpch.customer params: spiceai_api_key: ${secrets:SPICEAI_API_KEY} spiceai_region: us-east-1 acceleration: enabled: true ``` ## Spice → Spice (self-hosted federation)[​](#spice--spice-self-hosted-federation "Direct link to Spice → Spice (self-hosted federation)") The same connector federates queries to **any** Spice runtime that exposes Arrow Flight on a network endpoint. The most common pattern is a **cluster sidecar**: each application pod runs a small Spice runtime that owns no data but forwards queries to a heavier upstream Spice that owns the acceleration. The sidecar adds local caching, in-process auth, and protocol translation (HTTP / OpenAPI / MCP / gRPC) without each app needing direct access to the upstream cluster. ### When to use it[​](#when-to-use-it "Direct link to When to use it") * **Cluster sidecar**: per-pod Spice next to each application instance, federating to a central cluster Spice. The sidecar handles in-process queries; the cluster Spice handles refresh and storage. * **Edge → core federation**: edge Spice runtimes accelerate local datasets and federate the long-tail to a core Spice in the data center. * **Multi-region read replicas**: regional Spice instances replicate hot datasets and federate cold ones to a central runtime. ### Authentication[​](#authentication-1 "Direct link to Authentication") If the upstream runtime has API key authentication enabled (`runtime.auth.api-key.keys`), set `spiceai_api_key` to a key listed there. The connector sends it via Flight handshake using HTTP Basic — the upstream's `BasicAuthLayer` validates the key and issues a Bearer token for subsequent calls. ``` # Upstream runtime spicepod.yaml runtime: auth: api-key: enabled: true keys: - ${secrets:SIDECAR_API_KEY} ``` ``` # Sidecar runtime spicepod.yaml datasets: - from: spice.ai:https://upstream.cluster.svc:50051 name: events params: spiceai_api_key: ${secrets:SIDECAR_API_KEY} ``` If the upstream has no auth, omit `spiceai_api_key` — the connector falls back to anonymous credentials. ### TLS and private CAs[​](#tls-and-private-cas "Direct link to TLS and private CAs") Production-grade self-hosted federation should use TLS (`https://` or `grpc+tls://`). By default the connector trusts the system certificate store. To pin a private CA — common in cluster-internal deployments with a self-signed or internal-PKI CA: ``` params: spiceai_tls_ca_certificate_file: /etc/spice/upstream-ca.pem ``` For local development and trusted internal networks, `http://` (no TLS) is also supported and avoids cert configuration entirely. ### Mutual TLS (mTLS)[​](#mutual-tls-mtls "Direct link to Mutual TLS (mTLS)") When the upstream Spice runtime enforces mutual TLS (e.g. `runtime.tls.client_auth_mode: required`), the connector can present a client certificate during the handshake. Configure it in one of two mutually exclusive forms: **File-based** — point at PEM files on disk: ``` params: spiceai_endpoint: grpc+tls://upstream.cluster.svc:50051 spiceai_tls_ca_certificate_file: /etc/spice/clients/server-ca.pem spiceai_tls_client_certificate_file: /etc/spice/clients/client.pem spiceai_tls_client_key_file: /etc/spice/clients/client.key ``` **Inline** — supply PEM material directly, typically from a [secret store](/docs/next/components/secret-stores): ``` params: spiceai_endpoint: grpc+tls://upstream.cluster.svc:50051 spiceai_tls_client_certificate: ${secrets:CLIENT_CERT_PEM} spiceai_tls_client_key: ${secrets:CLIENT_KEY_PEM} ``` Within each form, cert and key must be set together — setting only one is rejected at dataset-load time with a clear error naming both fields. The file-based and inline forms are mutually exclusive; mixing them is also rejected. The cached Flight `Channel` is built once at dataset-load time, so certificate rotation requires a runtime restart. ### Append streams (real-time CDC)[​](#append-streams-real-time-cdc "Direct link to Append streams (real-time CDC)") The connector advertises [`supports_append_stream`](/docs/next/features/cdc) — when the upstream Spice exposes a dataset with append-stream support, the sidecar can subscribe over Flight `DoExchange` and receive each new batch as soon as the upstream emits it. Enable on the sidecar via `refresh_mode: append` on an accelerated dataset: ``` datasets: - from: spice.ai:https://upstream.cluster.svc:50051 name: events params: spiceai_api_key: ${secrets:SIDECAR_API_KEY} acceleration: enabled: true refresh_mode: append ``` Append streams are **append-only** — deletes and updates from the upstream are not propagated. Stream reconnection is automatic; persistent loss of connection causes the dataset to enter `Error` state if the lag exceeds the configured acceptable window. See [Data Refresh](/docs/next/features/data-acceleration/data-refresh) for the full append-mode reference. ### Sidecar example[​](#sidecar-example "Direct link to Sidecar example") A sidecar Spice running alongside an application, federating to a cluster Spice over TLS with API key auth and local in-memory acceleration: ``` version: v1 kind: Spicepod name: app-sidecar datasets: - from: spice.ai:https://upstream.cluster.svc:50051 name: orders params: spiceai_api_key: ${secrets:SIDECAR_API_KEY} acceleration: enabled: true refresh_mode: append refresh_check_interval: 30s - from: spice.ai:https://upstream.cluster.svc:50051 name: customers params: spiceai_api_key: ${secrets:SIDECAR_API_KEY} acceleration: enabled: true refresh_mode: full refresh_check_interval: 5m ``` ### Local development sidecar[​](#local-development-sidecar "Direct link to Local development sidecar") To run a sidecar against a local Spice without TLS: ``` datasets: - from: spice.ai:http://localhost:50051 name: events ``` `localhost` and loopback addresses are accepted with `http://` for development. ## Cookbook[​](#cookbook "Direct link to Cookbook") * [Spice.ai Cloud Platform Data Connector](https://github.com/spiceai/cookbook/tree/trunk/spiceai#readme) — end-to-end Cloud connection. ## Limitations[​](#limitations "Direct link to Limitations") * **Read-only.** The connector does not support `INSERT`/`UPDATE`/`DELETE` against Cloud or self-hosted upstreams. Writes happen on the upstream runtime (or via the Spice CLI / Console for Cloud). * **Single endpoint per dataset.** Multi-endpoint failover must be handled at the load-balancer / DNS layer. * **API key auth only.** OIDC / SSO is not supported at the data-plane connector. Use API keys for both Cloud and self-hosted federation. * **Append-only changes stream.** Updates and deletes from the upstream are not propagated; rely on `refresh_mode: full` for datasets that mutate. * **Cloud connections cap at 1000 requests per connection.** When the cap is hit the connection is reset; the Flight client retries automatically. If you see `Connection is reset by the server. Please retry the request.` or the `spiceai-retryable` metadata, the query has been retried already. * **`grpc://` (without TLS) is rejected.** Use `http://` for clear-text or `https://` / `grpc+tls://` for TLS. Memory considerations Without `acceleration.enabled: true`, federated queries that join across multiple Spice instances perform the join in memory on the local runtime. Ensure the local runtime has enough memory for query workspace plus runtime overhead, especially for concurrent queries. For large workloads, accelerate hot datasets locally with a file-mode engine ([`duckdb`](/docs/next/components/data-accelerators/duckdb) or [`sqlite`](/docs/next/components/data-accelerators/sqlite) with `mode: file`) to spill to disk instead of RAM. --- # Spice.ai Data Connector Deployment Guide Production operating guide for the [Spice.ai Data Connector](/docs/next/components/data-connectors/spiceai), covering both the **Spice → Spice Cloud Platform** and **Spice → Spice (self-hosted / cluster-sidecar)** topologies. The connector uses [Arrow Flight](https://arrow.apache.org/docs/format/Flight.html) over gRPC for both. ## Topology decision[​](#topology-decision "Direct link to Topology decision") | Use case | Topology | Endpoint | | --------------------------------------------------------------- | ----------------------- | ---------------------------------------------------- | | Federate datasets hosted on the managed Spice.ai Cloud Platform | Spice → Spice Cloud | `https://-prod-aws-flight.spiceai.io` (auto) | | Per-pod sidecar federating to a heavier upstream Spice runtime | Spice → Spice (sidecar) | `https://upstream.cluster.svc:50051` | | Edge runtime federating cold queries to a core Spice | Spice → Spice | Cluster-internal `https://...` | | Local development against a Spice on `localhost` | Spice → Spice | `http://localhost:50051` | Both topologies use the same `spiceai`-prefixed parameters and the same `spice.ai:` `from:` URI scheme. See the [connector reference](/docs/next/components/data-connectors/spiceai#configuration) for the full parameter list and URI formats. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The connector authenticates to the upstream Spice runtime using `spiceai_api_key`. The same parameter covers both Cloud and self-hosted upstreams: | Topology | Source of the key | Required? | | ----------- | ---------------------------------------------------- | ------------------------------------------------------- | | Cloud | Spice.ai Console; written to `.env` by `spice login` | Yes | | Self-hosted | Listed in the upstream's `runtime.auth.api-key.keys` | Only if upstream has auth enabled (otherwise anonymous) | | Parameter | Description | | ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `spiceai_api_key` | API key. Resolved from any [secret store](/docs/next/components/secret-stores) via `${secrets:...}`. | | `spiceai_token` | Legacy alias for `spiceai_api_key`. | | `spiceai_region` | Cloud region (e.g. `us-east-1`). Required for Cloud unless `spiceai_endpoint` is set. | | `spiceai_endpoint` | Override the Flight endpoint URL. Schemes: `http://`, `https://`, `grpc+tls://`. | | `spiceai_flight_endpoint` | Legacy alias for `spiceai_endpoint`. | | `spiceai_tls_ca_certificate_file` | Path to a CA PEM file for verifying a self-hosted upstream that uses a private CA. Ignored for `http://` endpoints. | | `spiceai_tls_client_certificate_file` | Path to a PEM client certificate chain for mutual TLS (mTLS). Must be set together with `spiceai_tls_client_key_file`. Mutually exclusive with `spiceai_tls_client_certificate`. | | `spiceai_tls_client_key_file` | Path to the PEM private key matching `spiceai_tls_client_certificate_file`. Must be set together with `spiceai_tls_client_certificate_file`. Mutually exclusive with `spiceai_tls_client_key`. | | `spiceai_tls_client_certificate` | Inline PEM client certificate chain for mutual TLS (mTLS). Resolve from a [secret store](/docs/next/components/secret-stores) via `${secrets:...}`. Must be set together with `spiceai_tls_client_key`. Mutually exclusive with `spiceai_tls_client_certificate_file`. | | `spiceai_tls_client_key` | Inline PEM private key for mutual TLS (mTLS). Resolve from a [secret store](/docs/next/components/secret-stores) via `${secrets:...}`. Must be set together with `spiceai_tls_client_certificate`. Mutually exclusive with `spiceai_tls_client_key_file`. | Always source production keys from a managed secret store rather than from a checked-in `.env` file. API keys do not expire — rotate manually in the issuing system and update the secret store. Secret stores that support live reload (Kubernetes, Vault) pick up rotations without restarting the runtime. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Endpoint Verification[​](#endpoint-verification "Direct link to Endpoint Verification") On startup the connector performs a DNS + TCP reachability check against the resolved endpoint before attempting a Flight handshake. Misconfigured endpoints surface as actionable startup errors rather than slow-failure query errors. ### Flight Transport[​](#flight-transport "Direct link to Flight Transport") Data transfer uses Arrow Flight over gRPC. Transient gRPC errors (`UNAVAILABLE`, `DEADLINE_EXCEEDED`) surface to the caller; retries are handled by the Flight client's default policy. For self-hosted upstreams, prefer `https://` or `grpc+tls://` in production. `http://` is supported for local development and trusted networks but transmits Flight payloads unencrypted. Plain `grpc://` is rejected at startup. ### TLS and private CAs[​](#tls-and-private-cas "Direct link to TLS and private CAs") By default the connector trusts the system certificate store. For cluster-internal upstreams that present a private-CA-signed certificate, pin the CA explicitly: ``` params: spiceai_endpoint: https://upstream.cluster.svc:50051 spiceai_tls_ca_certificate_file: /etc/spice/upstream-ca.pem ``` The CA file is loaded once at startup; updates require a runtime restart. Mount it via Kubernetes ConfigMap or Secret in containerized deployments. ### Mutual TLS (mTLS)[​](#mutual-tls-mtls "Direct link to Mutual TLS (mTLS)") When the upstream Spice runtime enforces mutual TLS (e.g. `runtime.tls.client_auth_mode: required`), the connector presents a client certificate during the handshake. Configure it as PEM files mounted into the pod, or as inline PEM material from a managed secret store: ``` # File-based params: spiceai_endpoint: grpc+tls://upstream.cluster.svc:50051 spiceai_tls_ca_certificate_file: /etc/spice/clients/server-ca.pem spiceai_tls_client_certificate_file: /etc/spice/clients/client.pem spiceai_tls_client_key_file: /etc/spice/clients/client.key ``` ``` # Inline / secrets params: spiceai_endpoint: grpc+tls://upstream.cluster.svc:50051 spiceai_tls_client_certificate: ${secrets:CLIENT_CERT_PEM} spiceai_tls_client_key: ${secrets:CLIENT_KEY_PEM} ``` Certificate and key must be set as a pair within each form, and the file-based and inline forms are mutually exclusive. The cached Flight `Channel` is built once at dataset-load time, so cert rotation requires a runtime restart. ### Append Streams[​](#append-streams "Direct link to Append Streams") The connector supports long-lived append streams for real-time CDC. The upstream — whether Cloud or self-hosted — must expose a dataset with append-stream support. The sidecar subscribes over Flight `DoExchange` and receives each new batch as soon as it's emitted. Stream reconnection is automatic; persistent loss of connection causes the dataset to enter `Error` state if the lag exceeds the acceptable window. See [Data Refresh](/docs/next/features/data-acceleration/data-refresh). Append streams are append-only — deletes and updates from the upstream are **not** propagated. Use `refresh_mode: full` for datasets that mutate. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") ### Message Sizing[​](#message-sizing "Direct link to Message Sizing") Arrow Flight record batches can be large for wide or dense schemas. Spice raises the gRPC message limit to `100MiB` by default (well above gRPC's built-in 4 MiB); raise it further with the runtime-level `runtime.flight.max_message_size` setting: | Setting | Default | Description | | --------------------------------- | -------- | --------------------------------------------------------------------------------------------------------------------------------- | | `runtime.flight.max_message_size` | `100MiB` | Maximum inbound and outbound gRPC message size, applied to all Flight clients. Raise for wide result sets or many string columns. | This is a runtime-level setting under `runtime.flight` (not a dataset connector parameter). Accepted units: `B`, `KB`, `MB`, `GB`. The same limit applies to the upstream's Flight server — raising it on the client without raising it on the server still fails. ### Network[​](#network "Direct link to Network") #### Cloud topology[​](#cloud-topology "Direct link to Cloud topology") * Place the Spice runtime in a region geographically close to the Cloud Platform region (`spiceai_region`) to minimize round-trip latency. * Expect typical round-trip latency in the tens of milliseconds plus result streaming time. For interactive dashboards, accelerate (`acceleration.enabled: true`) into a local engine. #### Self-hosted / sidecar topology[​](#self-hosted--sidecar-topology "Direct link to Self-hosted / sidecar topology") * Run the sidecar **in the same network namespace or cluster** as the upstream when possible — Flight is most efficient over single-digit-millisecond RTT. * Size sidecar memory for the local query workspace plus any in-memory acceleration. The sidecar does not need to be sized for the full dataset — only for hot data accelerated locally and the working set of in-flight queries. * Use a Kubernetes `Service` (with stable DNS) or a load balancer in front of multi-replica upstreams. Connection pooling is per-endpoint URL. ### API key lifetime[​](#api-key-lifetime "Direct link to API key lifetime") API keys do not expire. Rotation is manual; coordinate with the secret store used by the runtime. ## Sidecar deployment patterns[​](#sidecar-deployment-patterns "Direct link to Sidecar deployment patterns") ### Per-pod sidecar (Kubernetes)[​](#per-pod-sidecar-kubernetes "Direct link to Per-pod sidecar (Kubernetes)") Co-locate a Spice sidecar with each application pod. The sidecar terminates HTTP / OpenAPI / MCP / gRPC for the app and federates queries to a central upstream Spice cluster. ``` # Sidecar spicepod.yaml mounted into each app pod version: v1 kind: Spicepod name: app-sidecar runtime: http: bind_address: 127.0.0.1:8090 # localhost-only, sidecar talks to the app datasets: - from: spice.ai:https://upstream-spice.spiceai.svc.cluster.local:50051 name: orders params: spiceai_api_key: ${secrets:SIDECAR_API_KEY} spiceai_tls_ca_certificate_file: /etc/spice/cluster-ca.pem acceleration: enabled: true refresh_mode: append refresh_check_interval: 30s ``` The application talks to `127.0.0.1:8090`; the sidecar handles federation and caching. The upstream runs as a `Deployment` or `StatefulSet` with persistent storage for the acceleration files. ### Edge → core federation[​](#edge--core-federation "Direct link to Edge → core federation") Edge Spice runtimes accelerate local datasets and federate the long-tail to a core Spice in the data center. The same connector is used; the difference is that edge runtimes have their own non-federated datasets too: ``` datasets: - from: postgres:public.local_orders # local, accelerated name: local_orders acceleration: enabled: true - from: spice.ai:https://core.example.com:50051 # federated name: historical_orders ``` ## Metrics[​](#metrics "Direct link to Metrics") The Flight client is not instrumented, so this connector emits no transport-level metrics, and it does not currently register Spice.ai-specific dataset-level instruments. Monitor the connector via: * Query execution metrics (`query_duration_ms`, `query_returned_rows`, `query_failures`) from `runtime.metrics`. * Acceleration refresh metrics when the dataset is accelerated (`dataset_acceleration_refresh_duration_ms`, `dataset_acceleration_refresh_errors`). * For Cloud: upstream Console metrics on the source dataset. * For self-hosted: monitor the upstream's own `runtime.metrics` and acceleration metrics. See [Component Metrics](/docs/next/features/observability/component_metrics) for general configuration. ## Task History[​](#task-history "Direct link to Task History") Queries to the upstream Spice runtime participate in [task history](/docs/next/reference/task_history) via Flight client spans. Each Flight request is recorded as a child of the enclosing `sql_query` or `accelerated_table_refresh` task. The upstream runtime records its own task history independently — correlate by request timestamps or by propagated trace IDs. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Read-only.** The connector does not write to the upstream. Cloud writes go through the Spice CLI / Console; self-hosted writes happen on the upstream runtime directly. * **Single endpoint per dataset.** A dataset binds to a single endpoint URL. Multi-endpoint failover lives at the load-balancer / DNS layer. * **API key auth only.** OIDC / SSO is not supported at the data-plane connector. * **Append-only changes stream.** Updates and deletes are not propagated. * **Cloud connections cap at 1000 requests per connection.** When the cap is hit the connection is reset; the Flight client retries automatically. The `spiceai-retryable` metadata flag indicates the retry path. * **No `grpc://` (clear-text gRPC).** Use `http://` for unencrypted Flight or `https://` / `grpc+tls://` for TLS. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ---------------------------------------------------- | ------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------- | | `Failed to connect to SpiceAI endpoint` | DNS, firewall, or TLS issue against the resolved endpoint. | Verify DNS resolution and outbound 443/50051 connectivity. Test with `grpcurl -insecure : list`. | | `UnsupportedEndpointScheme` | Endpoint uses `grpc://` or another unsupported scheme. | Switch to `http://`, `https://`, or `grpc+tls://`. | | `CloudEndpointRegionMismatch` | `spiceai_endpoint` is a Cloud regional URL but `spiceai_region` doesn't match. | Set both to the same region, or remove one and let Spice pick the other. | | `UNAUTHENTICATED` on Flight handshake | Invalid / expired / wrong-environment API key. | For Cloud: regenerate in the Console; update the secret store. For self-hosted: confirm the key is in the upstream's `runtime.auth.api-key.keys`. | | TLS handshake failure with self-signed upstream cert | System cert store doesn't trust the upstream CA. | Set `spiceai_tls_ca_certificate_file` to the upstream's CA PEM, or have the upstream present a publicly-trusted certificate. | | `message size exceeded` / `ResourceExhausted` | Row batch exceeds gRPC message limit. | Increase `runtime.flight.max_message_size` on both client and server, or narrow the query projection. | | Append stream stalled; acceleration lag climbing | Network partition or upstream dataset paused. | Check upstream status; verify the source dataset is healthy; restart the runtime to re-establish the stream. | | Sudden 5xx / `UNAVAILABLE` errors | Transient service-side issue. | Flight client auto-retries; if persistent, check upstream runtime health (or the [Spice.ai status page](https://status.spice.ai)). | | `MissingRequiredParameter: api_key or token` | Targeting a Cloud endpoint with no API key configured. | Set `spiceai_api_key` (Cloud requires authentication; self-hosted endpoints accept anonymous if upstream auth is off). | --- # Embedding Models Embedding models transform raw text into numerical vectors that machine learning models can use. Spice supports running embedding models locally or via hosted services such as OpenAI, Amazon Bedrock, Databricks MosaicAI, or [la Plateforme](https://console.mistral.ai/). Embeddings enable vector-based and similarity search, such as document retrieval. For chat-based large language models, see [Model Providers](/docs/next/components/models). Spice supports a variety of embedding model sources and formats: | Name | Description | Status | Format(s) | | ------------------------------------------------------------- | --------------------------------------- | ----------------- | ------------------------------- | | [`file`](/docs/next/components/embeddings/local) | Local filesystem | Release Candidate | GGUF, GGML, SafeTensor | | [`huggingface`](/docs/next/components/embeddings/huggingface) | Models hosted on HuggingFace | Release Candidate | GGUF, GGML, SafeTensor | | [`openai`](/docs/next/components/embeddings/openai) | OpenAI (or compatible) LLM endpoint | Release Candidate | OpenAI-compatible HTTP endpoint | | [`azure`](/docs/next/components/embeddings/azure) | Azure OpenAI | Alpha | OpenAI-compatible HTTP endpoint | | [`google`](/docs/next/components/embeddings/google) | Google AI embedding models | Alpha | OpenAI-compatible HTTP endpoint | | [`databricks`](/docs/next/components/embeddings/databricks) | Models deployed to Databricks Mosaic AI | Alpha | OpenAI-compatible HTTP endpoint | | [`bedrock`](/docs/next/components/embeddings/bedrock) | Models deployed on Amazon Bedrock | Alpha | OpenAI-compatible HTTP endpoint | | [`model2vec`](/docs/next/components/embeddings/model2vec) | Model2Vec static word embeddings | Alpha | Model2Vec format | ## Overview[​](#overview "Direct link to Overview") Spice provides three ways to handle embedding columns in datasets: 1. **[Just-in-Time (JIT) Embeddings](#jit-embeddings):** Embeddings are computed on demand during query execution, with no precomputation. 2. **[Accelerated Embeddings](#accelerated-embeddings):** Embeddings are precomputed and stored, enabling faster queries and searches. 3. **[Passthrough Embeddings](#passthrough-embeddings):** Pre-existing embeddings in the source dataset are used directly, with no additional computation. ## Configuring Embedding Models[​](#configuring-embedding-models "Direct link to Configuring Embedding Models") Define embedding models in the `spicepod.yaml` file as top-level components. Example configuration in `spicepod.yaml`: ``` embeddings: - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: all_minilm_l6_v2 - from: openai:text-embedding-3-large name: xl_embed params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } - name: my_model from: file:model.safetensors files: - path: config.json - path: models/embed/tokenizer.json ``` Embedding models can be used via: * An OpenAI-compatible [endpoint](/docs/next/api/HTTP/post-embeddings) * Augmenting a dataset with column-level [embeddings](/docs/next/reference/spicepod/datasets#embeddings) for vector-based [search functionality](/docs/next/features/search#vector-search) ### Configuring Embedding Columns on Datasets[​](#configuring-embedding-columns-on-datasets "Direct link to Configuring Embedding Columns on Datasets") To create vector embeddings for specific dataset columns, define them under `columns` in the `spicepod.yaml` file, within the `datasets` section. Example configuration in `spicepod.yaml`: ``` embeddings: - from: openai:text-embedding-3-large name: xl_embed params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } datasets: - from: file:sales_data.parquet name: sales columns: - name: address_line1 description: The first line of the address. embeddings: - from: xl_embed row_id: order_number chunking: enabled: true target_chunk_size: 256 overlap_size: 32 ``` See the [embeddings](/docs/next/reference/spicepod/embeddings) and [datasets](/docs/next/reference/spicepod/datasets#embeddings) reference for more details. ## Embedding Methods[​](#embedding-methods "Direct link to Embedding Methods") ### Just-in-Time (JIT) Embeddings[​](#jit-embeddings "Direct link to Just-in-Time (JIT) Embeddings") JIT embeddings are computed at query time. This is useful when precomputing is impractical (e.g., large or rarely queried datasets, or heavy prefiltering). To add a JIT embedding column, specify it in the dataset's column config. ``` datasets: - name: invoices from: sftp://remote-sftp-server.com/invoices/2024/ columns: - name: line_item_details embeddings: - from: my_embedding_model params: file_format: parquet embeddings: # Or any model you like! - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: my_embedding_model ``` ### Accelerated Embeddings[​](#accelerated-embeddings "Direct link to Accelerated Embeddings") To speed up queries, embeddings can be precomputed and stored in a [data accelerator](/docs/next/components/data-accelerators). Enable this by adding: ``` acceleration: enabled: true ``` to the dataset configuration. All other data accelerator configurations are optional, but can be applied as per their respective [documentation](/docs/next/components/data-accelerators). **Full example:** ``` datasets: - name: invoices from: sftp://remote-sftp-server.com/invoices/2024/ acceleration: enabled: true columns: - name: line_item_details embeddings: - from: my_embedding_model params: file_format: parquet ``` ### Passthrough Embeddings[​](#passthrough-embeddings "Direct link to Passthrough Embeddings") If the dataset already contains embedding columns, Spice can use them for vector search and other embedding features. The schema must match that of Spice-generated embeddings (or be adapted with a [view](/docs/next/reference/spicepod#views)). **Example:** A `sales` table with an `address` column and its embedding: ``` sql> describe sales; +-------------------+-----------------------------------------+-------------+ | column_name | data_type | is_nullable | +-------------------+-----------------------------------------+-------------+ | order_number | Int64 | YES | | quantity_ordered | Int64 | YES | | price_each | Float64 | YES | | order_line_number | Int64 | YES | | address | Utf8 | YES | | address_embedding | FixedSizeList( | NO | | | Field { | | | | name: "item", | | | | data_type: Float32, | | | | nullable: false, | | | | dict_id: 0, | | | | dict_is_ordered: false, | | | | metadata: {} | | | | }, | | | | 384 | | +-------------------+-----------------------------------------+-------------+ ``` The same table if it was chunked: ``` sql> describe sales; +-------------------+-----------------------------------------+-------------+ | column_name | data_type | is_nullable | +-------------------+-----------------------------------------+-------------+ | order_number | Int64 | YES | | quantity_ordered | Int64 | YES | | price_each | Float64 | YES | | order_line_number | Int64 | YES | | address | Utf8 | YES | | address_embedding | List(Field { | NO | | | name: "item", | | | | data_type: FixedSizeList( | | | | Field { | | | | name: "item", | | | | data_type: Float32, | | | | }, | | | | 384 | | | | ), | | | | }) | | +-------------------+-----------------------------------------+-------------+ | address_offset | List(Field { | NO | | | name: "item", | | | | data_type: FixedSizeList( | | | | Field { | | | | name: "item", | | | | data_type: Int32, | | | | nullable: false, | | | | dict_id: 0, | | | | dict_is_ordered: false, | | | | metadata: {} | | | | }, | | | | 2 | | | | ), | | | | }) | | +-------------------+-----------------------------------------+-------------+ ``` Passthrough embedding columns must still be defined in the `spicepod.yaml` file. The Spice instance must also have access to the same embedding model used to generate the embeddings. ``` datasets: - from: sftp://remote-sftp-server.com/sales/2024.csv name: sales columns: - name: address embeddings: - from: local_embedding_model embeddings: - name: local_embedding_model # The model originally used for this column ... ``` #### Requirements[​](#requirements "Direct link to Requirements") To ensure compatibility, embedding columns must meet these requirements: 1. **Underlying Column:** * The original column must exist and be of `string` [Arrow data type](/docs/next/reference/datatypes/accelerators). 2. **Naming Convention:** * The embedding column must be named `_embedding` (e.g., `review_embedding` for a `review` column). 3. **Data Type:** * The embedding column must be: * `FixedSizeList[Float32 or Float64, N]` for unchunked data, where `N` is the embedding vector size. * `List[FixedSizeList[Float32 or Float64, N]]` for chunked data. 4. **Offset Column (for chunked data):** * If chunked, an offset column `_offsets` must exist with type `List[FixedSizeList[Int32, 2]]`, where each pair `[start, end]` maps a chunk to its text segment. * Example: `[[0, 100], [101, 200]]` means two chunks covering indices 0–100 and 101–200. Following these guidelines ensures that the dataset's pre-existing embeddings are fully compatible with Spice. ## Advanced Configuration[​](#advanced-configuration "Direct link to Advanced Configuration") ### Chunking[​](#chunking "Direct link to Chunking") Spice supports chunking large text columns before embedding, which is useful for [Document Tables](/docs/next/components/data-connectors#document-formats). Chunking helps return only the most relevant text during search. Configure chunking in the embedding config: ``` datasets: - from: github:github.com/spiceai/spiceai/issues name: spiceai.issues acceleration: enabled: true columns: - name: body embeddings: - from: local_embedding_model chunking: enabled: true target_chunk_size: 512 ``` The `body` column will be split into chunks of about 512 tokens, preserving sentence and semantic boundaries. See the [API reference](/docs/next/reference/spicepod/datasets#columns-embeddings-chunking) for details. #### Row Identifiers[​](#row-identifiers "Direct link to Row Identifiers") The `row_id` field specifies which column(s) uniquely identify each row, similar to a primary key. This is important for chunked embeddings, so that operations (e.g., [`v1/search`](/docs/next/api/HTTP/post-search)) can map multiple chunked vectors to a single row. Set `row_id` in `columns[*].embeddings[*].row_id`. ``` datasets: - from: github:github.com/spiceai/spiceai/issues name: spiceai.issues acceleration: enabled: true columns: - name: body embeddings: - from: local_embedding_model chunking: enabled: true target_chunk_size: 512 row_id: id ``` ### Multi-Vector Embeddings[​](#multi-vector-embeddings "Direct link to Multi-Vector Embeddings") When the source column is `List` (or `LargeList`), Spice embeds each list element independently and produces a `List>` column. This is the multi-vector (column-of-vectors) mode, useful for rows that carry several independent pieces of text such as tags, section headings, or historical queries. ``` datasets: - from: file:products.parquet name: products acceleration: enabled: true columns: - name: tags # List embeddings: - from: local_embedding_model aggregation: max max_elements_per_row: 64 ``` The `aggregation` field controls how per-element similarities are combined into a per-row score during vector search. `max` (default) is ColBERT-style `MaxSim`; `mean` and `sum` are also supported. The `max_elements_per_row` field caps how many list elements are embedded per row (default `32`, hard limit `1024`). Multi-vector columns also support [ColBERT-style late-interaction search](/docs/next/features/search/multi-vector#late-interaction-multi-query-search) via an array of query strings. See [Multi-Vector Search](/docs/next/features/search/multi-vector) for query usage and [`columns[*].embeddings[*]`](/docs/next/reference/spicepod/datasets#columnsembeddings) for the full field reference. ## [🗃OpenAI](/docs/next/components/embeddings/openai) [1 item](/docs/next/components/embeddings/openai) ## [📄️Azure OpenAI](/docs/next/components/embeddings/azure) [To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page:](/docs/next/components/embeddings/azure) ## [🗃HuggingFace](/docs/next/components/embeddings/huggingface) [1 item](/docs/next/components/embeddings/huggingface) ## [🗃Local](/docs/next/components/embeddings/local) [1 item](/docs/next/components/embeddings/local) ## [📄️Google AI](/docs/next/components/embeddings/google) [To use a hosted Google AI embedding model, specify the google path in the from field of your configuration.](/docs/next/components/embeddings/google) ## [📄️Model2Vec](/docs/next/components/embeddings/model2vec) [Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning.](/docs/next/components/embeddings/model2vec) ## [📄️Amazon Bedrock](/docs/next/components/embeddings/bedrock) [Instructions for using Amazon Bedrock embedding models](/docs/next/components/embeddings/bedrock) ## [📄️Databricks](/docs/next/components/embeddings/databricks) [Instructions for using Databricks Mosaic AI Models](/docs/next/components/embeddings/databricks) --- # Azure OpenAI Embedding Models To use an embedding model hosted on Azure OpenAI, specify the `azure` path in the `from` field and the following parameters from the [Azure OpenAI Model Deployment](https://ai.azure.com/resource/deployments) page: | Param | Description | Default | | ----------------------- | ----------------------------------------------------------------------------------- | ---------- | | `azure_api_key` | The Azure OpenAI API key from the models deployment page. | - | | `azure_api_version` | The API version used for the Azure OpenAI service. | - | | `azure_deployment_name` | The name of the model deployment. | Model name | | `endpoint` | The Azure OpenAI resource endpoint, e.g., `https://resource-name.openai.azure.com`. | - | | `azure_entra_token` | The Azure Entra token for authentication. | - | Only one of `azure_api_key` or `azure_entra_token` can be provided for model configuration. Example: ``` embeddings: - name: embeddings-model from: azure:text-embedding-3-small params: endpoint: ${ secrets:SPICE_AZURE_AI_ENDPOINT } azure_deployment_name: text-embedding-3-small azure_api_version: 2023-05-15 azure_api_key: ${ secrets:SPICE_AZURE_API_KEY } ``` Refer to the [Azure OpenAI Service models](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models) for more details on available models and configurations. Follow the [Azure OpenAI Models Cookbook](https://github.com/spiceai/cookbook/tree/trunk/azure_openai) to try Azure OpenAI models for vector-based search and chat functionalities with structured (taxi trips) and unstructured (GitHub files) data. --- # Amazon Bedrock Model Provider To use an embedding model deployed to [AWS Bedrock service](https://aws.amazon.com/bedrock/), specify the model endpoint name prefixed with `bedrock:` in the `from` field and include the required parameters in the `params` section. ### Parameters[​](#parameters "Direct link to Parameters") #### AWS Parameters[​](#aws-parameters "Direct link to AWS Parameters") | Parameter | Description | | ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `aws_region` | AWS region. Default: `us-east-1`. | | `aws_profile` | Optional. AWS profile to use when loading credentials. | | `aws_access_key_id` | Optional. AWS access key ID for authentication. If not provided, credentials will be loaded from environment variables or IAM roles | | `aws_secret_access_key` | Optional. AWS secret access key for authentication. If not provided, credentials will be loaded from environment variables or IAM roles | | `aws_session_token` | Optional. AWS session token for authentication | | `aws_iam_role_source` | Optional. IAM role credential source. `auto` (default) uses the default AWS credential chain, `metadata` uses only instance/container metadata (IMDS, ECS, EKS/IRSA), `env` uses only environment variables. | | `max_concurrent_invocations` | Optional. The maximum number of concurrent API invocations. Defaults to `40` | | `requests_per_min_limit` | Optional. The maximum number of requests made per minute. Defaults to `1500` | #### AWS Titan Models[​](#aws-titan-models "Direct link to AWS Titan Models") These parameters are used for [Amazon Titan Text](https://docs.aws.amazon.com/bedrock/latest/userguide/titan-embedding-models.html) embedding model | Parameter | Description | | ------------ | ----------------------------------------------------------------------------------------------------------------------- | | `normalize` | Whether or not to normalize the output embedding. Defaults to `true`. | | `dimensions` | The number of dimensions the output embedding should have. The following values are accepted: 1024 (default), 512, 256. | #### Amazon Nova Models[​](#amazon-nova-models "Direct link to Amazon Nova Models") These parameters are used for [Amazon Nova](https://docs.aws.amazon.com/nova/latest/userguide/embeddings-schema.html) multimodal embedding models | Parameter | Description | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `dimensions` | **Required**. The number of dimensions the output embedding should have. Accepted value: 256, 384, 1024 or 3072. | | `truncate_mode` | Optional. Specifies how the API handles inputs longer than the maximum token length. One of: `START`, `END` or `NONE` (default). Also accepted as `truncate`. | | `embedding_purpose` | Optional. Use the Nova embeddings model optimized for different purposes. Default `GENERIC_INDEX`. See reference [docs](https://docs.aws.amazon.com/nova/latest/userguide/embeddings-schema.html) for all options. | #### Cohere Models[​](#cohere-models "Direct link to Cohere Models") | Parameter | Description | | ------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `truncate` | Specifies how the API handles inputs longer than the maximum token length. One of: `START`, `END` (default) or `NONE`. | | `input_type` | Use the Cohere embeddings model optimized for different types of inputs. One of: `search_document` (default), `search_query`, `classification` or `clustering`. | ### Example `spicepod.yaml` configuration, Cohere model[​](#example-spicepodyaml-configuration-cohere-model "Direct link to example-spicepodyaml-configuration-cohere-model") ``` embeddings: - from: bedrock:cohere.embed-english-v3 name: cohere-embeddings params: aws_region: us-east-1 input_type: classification truncate: END aws_access_key_id: ${ secrets:AWS_ACCESS_KEY_ID } aws_secret_access_key: ${ secrets:AWS_SECRET_ACCESS_KEY } ``` ### Example `spicepod.yaml` configuration, Titan model[​](#example-spicepodyaml-configuration-titan-model "Direct link to example-spicepodyaml-configuration-titan-model") ``` - from: bedrock:amazon.titan-embed-text-v2:0 name: titan-embeddings params: dimensions: '256' ``` ### Example `spicepod.yaml` configuration, Amazon Nova model[​](#example-spicepodyaml-configuration-amazon-nova-model "Direct link to example-spicepodyaml-configuration-amazon-nova-model") ``` embeddings: - from: bedrock:amazon.nova-2-multimodal-embeddings-v1:0 name: nova-embeddings params: dimensions: '3072' truncate_mode: START embedding_purpose: GENERIC_RETRIEVAL aws_region: us-east-1 aws_access_key_id: ${ secrets:AWS_ACCESS_KEY_ID } aws_secret_access_key: ${ secrets:AWS_SECRET_ACCESS_KEY } ``` ### Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate bedrock permissions (e.g., `bedrock:InvokeModel`) to access the model. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access DynamoDB tables: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["bedrock:InvokeModel"], "Resource": ["arn:aws:bedrock:us-east-1::foundation-model/amazon.titan-*"] } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | --------------------- | --------------------------------------------- | | `bedrock:InvokeModel` | Required. Used to invoke the embedding model. | ### Additional Information[​](#additional-information "Direct link to Additional Information") Refer to the [Amazon Bedrock documentation](https://docs.aws.amazon.com/bedrock/) for more details on available models and configurations. --- # Databricks Model Provider To use an embedding model deployed to [Databricks Mosaic AI Model Serving](https://docs.databricks.com/aws/en/machine-learning/model-serving/), specify the model endpoint name prefixed with `databricks:` in the `from` field and include the required parameters in the `params` section. ### Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `databricks_endpoint` | The Databricks workspace endpoint, e.g., `dbc-a12cd3e4-56f7.cloud.databricks.com`. | | `databricks_token` | The Databricks API token to authenticate with the Databricks Models API. Use the [secret replacement syntax](/docs/next/components/secret-stores) to reference a secret, e.g., `${secrets:my_databricks_token}`. | | `databricks_client_id` | The Databricks Service Principal Client ID. Can't be used with `databricks_token`. | | `databricks_client_secret` | The Databricks Service Principal Client Secret. Can't be used with `databricks_token`. | ### Example `spicepod.yaml` configuration, using personal access token[​](#example-spicepodyaml-configuration-using-personal-access-token "Direct link to example-spicepodyaml-configuration-using-personal-access-token") To learn more about how to set up personal access tokens, see [Databricks PAT docs](https://docs.databricks.com/aws/en/dev-tools/auth/pat). ``` embeddings: - from: databricks:databricks-gte-large-en name: gte-large-en params: databricks_endpoint: dbc-46470731-42e5.cloud.databricks.com databricks_token: ${ secrets:SPICE_DATABRICKS_TOKEN } ``` ### Example `spicepod.yaml` configuration, using Databricks service principal[​](#example-spicepodyaml-configuration-using-databricks-service-principal "Direct link to example-spicepodyaml-configuration-using-databricks-service-principal") Spice supports the Machine-to-Machine (M2M) OAuth flow with service principal credentials by utilizing the `databricks_client_id` and `databricks_client_secret` parameters. The runtime will automatically refresh the token. The service principal must be granted the "Can Query" permission for model serving. To learn more about how to set up the service principal, see [Databricks M2M OAuth docs](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m). ``` embeddings: - from: databricks:databricks-gte-large-en name: gte-large-en params: databricks_endpoint: dbc-42424242-4242.cloud.databricks.com databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} databricks_client_secret: ${secrets:DATABRICKS_CLIENT_SECRET} ``` ### Additional Information[​](#additional-information "Direct link to Additional Information") Refer to the [Mosaic AI Model Serving documentation](https://docs.databricks.com/aws/en/machine-learning/model-serving/) for more details on available models and configurations. --- # Google AI Embedding Models To use a hosted Google AI embedding model, specify the `google` path in the `from` field of your configuration. Include the model ID in the `from` field; a model ID is required. For example, `google:gemini-embedding-2` selects Google's latest generally available embedding model. The following parameters are specific to Google AI embedding models: | Parameter | Description | Default | | ---------------- | ------------------------------------------------------------------------------------------------ | ------- | | `google_api_key` | The API key for accessing Google AI. | - | | `dimensions` | The output dimensionality of the embeddings. Some embedding models support dynamic output sizes. | - | Below is an example configuration in `spicepod.yaml`: ``` embeddings: - from: google:gemini-embedding-2 name: gemini_embeddings params: google_api_key: ${ secrets:GEMINI_API_KEY } dimensions: 768 # optional parameter - from: google:gemini-embedding-001 name: legacy_embeddings params: google_api_key: ${ secrets:GEMINI_API_KEY } ``` See [Google AI Embedding Models](https://ai.google.dev/gemini-api/docs/models/gemini#text-embedding) for a list of supported embedding models. For detailed instructions and examples on running vector searches, refer to the [Vector-Based Search documentation](/docs/next/features/search/vector-search). --- # HuggingFace Text Embedding Models To use an embedding model from HuggingFace with Spice, specify the `huggingface` path in the `from` field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. The following parameters are specific to HuggingFace models: | Parameter | Description | Default | | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- | | `hf_token` | The Huggingface access token. | - | | `pooling` | The [pooling method](https://huggingface.co/docs/text-embeddings-inference/en/cli_arguments) for embedding models. Supported values are `cls`, `mean`, `splade`, `last_token` | - | Here is an example configuration in `spicepod.yaml`: ``` embeddings: - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: all_minilm_l6_v2 ``` Supported models include: * All models tagged as [text-embeddings-inference](https://huggingface.co/models?other=text-embeddings-inference) on Huggingface * Any Huggingface repository with the correct files to be loaded as a [local embedding model](/docs/next/components/embeddings/local). ## Model Weights[​](#model-weights "Direct link to Model Weights") Spice downloads the repository's `safetensors` weights. When a repository publishes none, it falls back to `pytorch_model.bin` and logs a warning at startup: ``` safetensors weights not found; falling back to `pytorch_model.bin`. Model loading is significantly slower. ``` A repository that publishes neither format fails to load. Prefer a repository with `safetensors` weights where one is available. With the same semantics as [language models](/docs/next/components/models/huggingface#access-tokens), `spice` can run private HuggingFace embedding models: ``` embeddings: - from: huggingface:huggingface.co/secret-company/awesome-embedding-model name: top_secret params: hf_token: ${ secrets:HF_TOKEN } ``` --- # Hugging Face Embedding Deployment Guide Production operating guide for Hugging Face embeddings — loading an embedding model from the Hub and running local inference via the Text Embeddings Inference (TEI) pipeline. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") | Parameter | Description | | ---------- | --------------------------------------------------------------- | | `hf_token` | Hugging Face access token. Required for private or gated repos. | Tokens must be sourced from a [secret store](/docs/next/components/secret-stores) in production. For public models (e.g., `sentence-transformers/all-MiniLM-L6-v2`), the token is optional; for gated models, the token is required. ### Cache Directory[​](#cache-directory "Direct link to Cache Directory") The cache location honors the `HF_HUB_CACHE` environment variable. In containers, set `HF_HUB_CACHE` to a persistent volume to avoid re-downloading on every restart. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Download & Cache[​](#download--cache "Direct link to Download & Cache") Model weights are downloaded on first use into the HF Hub cache and reused on subsequent starts. Download retries follow the shared HTTP client's retry policy on transient failures. ### Revision Pinning[​](#revision-pinning "Direct link to Revision Pinning") Append `:` to the repository id in `from:` to pin a branch, tag, or commit SHA — the same convention the `huggingface:` model source uses. The revision is split off the repository id and passed to the loader, so the pinned revision is what gets downloaded: ``` embeddings: - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2:refs/pr/21 name: pinned_minilm ``` With no `:` suffix, the repository's default revision (`main`) is used. A trailing `:` with nothing after it is treated as no revision rather than as part of the repository name. ### TEI Queue Configuration[​](#tei-queue-configuration "Direct link to TEI Queue Configuration") The TEI pipeline has fixed queue parameters in the current release: * `max_concurrent_requests`: **512** * `max_batch_tokens`: **16384** These are not currently exposed as user-tunable parameters. ### No Automatic Truncation for `embed_pooled`[​](#no-automatic-truncation-for-embed_pooled "Direct link to no-automatic-truncation-for-embed_pooled") Pooled-embed calls do **not** currently auto-truncate inputs longer than the model's max sequence length. Truncate at the caller, or configure `max_seq_length` on the dataset to enforce truncation. ## Pooling[​](#pooling "Direct link to Pooling") | Value | Description | | ------------ | ----------------------------------------------------------------- | | `cls` | Use the `[CLS]` token's embedding. | | `mean` | Mean-pool across tokens. | | `splade` | SPLADE sparse pooling (for sparse retrieval). | | `last_token` | Use the final token's embedding (useful for decoder-only models). | When `pooling` is unset, the loader defaults to `mean` and logs a warning. Set the pooling strategy explicitly for deterministic behavior across Spice versions. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") ### Required Files[​](#required-files "Direct link to Required Files") Local file mode (HF-downloaded) requires: * Model weights (accepted formats: `.gguf`, `.ggml`, `.safetensors`, `pytorch_model.bin`). * `config.json` * `tokenizer.json` If any are missing, load fails with a descriptive error. ### Device Selection[​](#device-selection "Direct link to Device Selection") Device selection follows the local-model pattern: CUDA → Metal → CPU (see the [Hugging Face Model Deployment Guide](/docs/next/components/models/huggingface/deployment#device-selection) for details). ### Memory Footprint[​](#memory-footprint "Direct link to Memory Footprint") Embedding models are typically smaller than LLMs (tens to hundreds of MB). A 384-dim MiniLM consumes \~100 MB on disk and \~200 MB in RAM. Plan for the base model + \~30% for batch buffers. ### Throughput[​](#throughput "Direct link to Throughput") Batched embedding is the dominant throughput driver. With default TEI settings (`max_batch_tokens=16384`), a MiniLM-class model on CPU can process hundreds of inputs per second; on a modern GPU, thousands per second. ## Metrics[​](#metrics "Direct link to Metrics") Shared embedding metrics (see the [OpenAI Embedding Deployment Guide](/docs/next/components/embeddings/openai/deployment#metrics)): * `embeddings_requests` * `embeddings_failures` * `embeddings_internal_request_duration_ms` * `embeddings_load_errors`, `embeddings_active_count`, `embeddings_load_state` See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting metrics. ## Task History[​](#task-history "Direct link to Task History") Embedding requests emit `text_embed` spans in [task history](/docs/next/reference/task_history) with input (truncated), labels, `outputs_produced`, and errors. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **TEI queue limits hardcoded**: `max_concurrent_requests` (512) and `max_batch_tokens` (16384) are not user-tunable in the current release. * **No auto-truncation for pooled embeds**: Inputs longer than `max_seq_length` fail unless truncated by the caller. * **Single process loading**: Models load into the Spice process; no shared inference server across instances. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | --------------------------------------- | --------------------------------------------------- | ---------------------------------------------------------------------------------------- | | `401 Unauthorized` on download | Missing / invalid `hf_token`; gated model. | Set `hf_token`; accept the model's license on Hugging Face. | | `Missing tokenizer.json` at load | Model repo does not ship a fast tokenizer. | Use a model that ships `tokenizer.json`, or convert via `AutoTokenizer.save_pretrained`. | | Input too long errors on `embed_pooled` | No auto-truncation. | Truncate at the caller; or set `max_seq_length` on the dataset. | | `Pooling defaulted to 'mean'` warning | `pooling` not set. | Set pooling explicitly to silence the warning and lock in behavior. | | First request extremely slow | Model downloading / warmup. | Pre-warm: `huggingface-cli download ` into `HF_HUB_CACHE`. | | OOM during batched embedding | Batch size × sequence length exceeds device memory. | Reduce caller batch size; use a smaller model; upgrade device memory. | --- # Local Filesystem Embedding Models Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") To configure an embedding model using local files, you can specify the details in the `spicepod.yaml` file as shown below: ``` embeddings: - name: all_minilm_l6_v2 from: file:model.safetensors files: - path: /Users/jeadie/Github/spiceai/models/embed/config.json - path: models/embed/tokenizer.json ``` ## Required Files[​](#required-files "Direct link to Required Files") * Model file, one of: `model.safetensors`, `pytorch_model.bin`. * A tokenizer file with the filename `tokenizer.json`. * A config file with the filename `config.json`. --- # Local Embedding Deployment Guide Production operating guide for loading an embedding model from the local filesystem and running inference via the Text Embeddings Inference (TEI) pipeline. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The Local embedding provider has no authentication layer. Access control is enforced by the operating system: * The Spice runtime process must have read permission on the model files. * For containers, mount model files as read-only volumes. * For Kubernetes, mount via `PersistentVolumeClaim` or an init-container that downloads into a shared volume. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") The Local embedding provider reads local files synchronously. There is no network layer or retry logic. Failures surface as filesystem errors (`ENOENT`, `EACCES`, `EIO`) and fail the spicepod load at startup. ### TEI Queue Configuration[​](#tei-queue-configuration "Direct link to TEI Queue Configuration") Fixed queue parameters in the current release: * `max_concurrent_requests`: **512** * `max_batch_tokens`: **16384** These are not currently exposed as user-tunable parameters. ### No Automatic Truncation for `embed_pooled`[​](#no-automatic-truncation-for-embed_pooled "Direct link to no-automatic-truncation-for-embed_pooled") Pooled-embed calls do **not** currently auto-truncate inputs longer than the model's max sequence length. Truncate at the caller, or configure `max_seq_length` on the dataset to enforce truncation. ## Pooling[​](#pooling "Direct link to Pooling") | Value | Description | | ------------ | ----------------------------------------------------------------- | | `cls` | Use the `[CLS]` token's embedding. | | `mean` | Mean-pool across tokens. | | `splade` | SPLADE sparse pooling (for sparse retrieval). | | `last_token` | Use the final token's embedding (useful for decoder-only models). | When `pooling` is unset, the loader defaults to `mean` and logs a warning. Set the pooling strategy explicitly for deterministic behavior across Spice versions. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") ### Required Files[​](#required-files "Direct link to Required Files") Local embedding requires all of the following in the model directory: * Model weights (accepted formats: `.gguf`, `.ggml`, `.safetensors`, `pytorch_model.bin`). * `config.json` * `tokenizer.json` If any are missing, load fails with a descriptive error. ### Device Selection[​](#device-selection "Direct link to Device Selection") 1. **CUDA** (CUDA-enabled Spice build + available device) 2. **Metal** (Metal-enabled Spice build — macOS / Apple Silicon) 3. **CPU** fallback ### Memory Footprint[​](#memory-footprint "Direct link to Memory Footprint") Embedding models are typically smaller than LLMs (tens to hundreds of MB). Plan for the base model size + \~30% for batch buffers. ### Throughput[​](#throughput "Direct link to Throughput") Batched embedding dominates throughput. With default TEI settings (`max_batch_tokens=16384`), a MiniLM-class model on CPU can process hundreds of inputs per second; on a modern GPU, thousands per second. ## Metrics[​](#metrics "Direct link to Metrics") Shared embedding metrics (see the [OpenAI Embedding Deployment Guide](/docs/next/components/embeddings/openai/deployment#metrics)): * `embeddings_requests` * `embeddings_failures` * `embeddings_internal_request_duration_ms` * `embeddings_load_errors`, `embeddings_active_count`, `embeddings_load_state` See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting metrics. ## Task History[​](#task-history "Direct link to Task History") Embedding requests emit `text_embed` spans in [task history](/docs/next/reference/task_history) with input (truncated), labels, `outputs_produced`, and errors. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **TEI queue limits hardcoded**: `max_concurrent_requests` (512) and `max_batch_tokens` (16384) are not user-tunable in the current release. * **No auto-truncation for pooled embeds**: Inputs longer than `max_seq_length` fail unless truncated by the caller. * **Single-process loading**: Models load into the Spice process; no shared inference server across instances. * **No hot reload**: Swapping the underlying model file requires a spicepod reload. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ---------------------------------------- | ---------------------------------------------- | ----------------------------------------------------------------------------------- | | `No such file or directory` | Path typo or missing mount. | Verify the files exist in the Spice process filesystem. | | `Permission denied` | Spice user lacks read on the files. | Adjust ACLs or mount with appropriate UID/GID. | | `Missing tokenizer.json` at load | Model directory missing the fast tokenizer. | Add `tokenizer.json`; convert via `AutoTokenizer.save_pretrained`. | | Input too long errors on `embed_pooled` | No auto-truncation. | Truncate at the caller, or set `max_seq_length` on the dataset. | | `Pooling defaulted to 'mean'` warning | `pooling` not set. | Set pooling explicitly. | | Inference falls back to CPU unexpectedly | CUDA / Metal unavailable. | Use a CUDA-enabled Spice build on GPU hosts; on macOS, use the Apple Silicon build. | | OOM during batched embedding | Batch × sequence length exceeds device memory. | Reduce caller batch size; use a smaller model; upgrade device memory. | --- # Model2Vec Embedding Models [Model2Vec](https://huggingface.co/blog/Pringled/model2vec) is a technique that distills embeddings from sentence transformer models into static word embeddings, providing efficient embedding generation, in parallel, without performing external API calls. This can result in sentence transformer models up to 500x faster and 15x smaller. To use a Model2Vec embedding model with Spice, specify the `model2vec` prefix in the `from` field of your configuration. ## Model Compatibility[​](#model-compatibility "Direct link to Model Compatibility") Find models compatible with `model2vec`: * [Model2Vec base models](https://huggingface.co/collections/minishlab/model2vec-base-models-66fd9dd9b7c3b3c0f25ca90e) (ready to use, pre-distilled) * [Sentence Transformers](https://huggingface.co/collections/sentence-transformers/embedding-model-datasets-6644d7a3673a511914aa7552) (needs [distillation](#distilling-your-own-models)) * [Pre-distilled community models](https://huggingface.co/models?search=model2vec) ## Parameters[​](#parameters "Direct link to Parameters") The following parameters are specific to Model2Vec models: | Parameter | Description | Default | | ------------------------- | ------------------------------------------------------------------------------- | ----------------------- | | `hf_token` | The Hugging Face access token for accessing private models. | - | | `normalize` | Whether to normalize embeddings (defaults to the model's configuration). | Model's default setting | | `subfolder` | Optional subfolder path for models that reside in a subfolder of the repo/path. | - | | `parallelism` | Number of parallel threads to use for embedding computation. | System CPU count | | `embed_max_token_length` | Maximum token length for embeddings. | - | | `embed_custom_batch_size` | Custom batch size override for embedding operations. | - | For more details on Model2Vec parameters and functionality, refer to the [model2vec-rs documentation](https://github.com/MinishLab/model2vec-rs). Example configuration in `spicepod.yaml` for [`minishlab/potion-base-8m`](https://huggingface.co/minishlab/potion-base-8M): ``` embeddings: - from: model2vec:minishlab/potion-base-8M name: potion_base_8m ``` ## Local Models[​](#local-models "Direct link to Local Models") Model2Vec models can also be loaded from the local filesystem by specifying a file path: ``` embeddings: - from: model2vec:/path/to/local/model name: local_model2vec ``` ## Revision Pinning Is Not Supported[​](#revision-pinning-is-not-supported "Direct link to Revision Pinning Is Not Supported") Unlike the [Hugging Face embedding source](/docs/next/components/embeddings/huggingface), `model2vec:` cannot load a pinned revision — the underlying loader takes no revision argument. A `from:` value shaped like `model2vec:org/model:revision` is rejected at load with an explicit error rather than being sent to the Hub, where the `:revision` suffix would become part of the repository name and fail as a `401 Unauthorized`: ``` The model 'org/model:revision' pins revision 'revision', but source 'model2vec' cannot load a pinned revision. Remove ':revision' to load the repository's default revision, and try again. ``` Load the repository's default revision instead, or pre-download the revision you want and reference it as a [local model](#local-models). Local paths are unaffected — a path that exists on disk is loaded as a path, even when it happens to look like a pinned repository id. ## Private Models[​](#private-models "Direct link to Private Models") Model2Vec supports private Hugging Face models with authentication: ``` embeddings: - from: model2vec:your-organization/private-model name: private_embeddings params: hf_token: ${ secrets:HF_TOKEN } ``` ## Advanced Configuration[​](#advanced-configuration "Direct link to Advanced Configuration") For performance optimization, configure parallelism and embedding batch sizes: ``` embeddings: - from: model2vec:minishlab/potion-base-8M name: potion_optimized params: parallelism: 8 embed_custom_batch_size: 32 normalize: true ``` ## Distilling Your Own Models[​](#distilling-your-own-models "Direct link to Distilling Your Own Models") Create custom Model2Vec embeddings by distilling existing sentence transformer models. For more detailed instructions, see the [Model2Vec Quickstart guide](https://github.com/MinishLab/model2vec?tab=readme-ov-file#quickstart). Here's how to distill the popular [`sentence-transformers/all-MiniLM-L6-v2`](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) model: 1. **Install the Model2Vec Python library:** ``` pip install model2vec ``` 2. **Create a distillation script:** ``` from model2vec import StaticModel # Load a sentence transformer model and distill it to a static model model = StaticModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2", device="cpu") # Save the static model model.save_pretrained("./all-MiniLM-L6-v2-model2vec") print("Model distilled and saved to ./all-MiniLM-L6-v2-model2vec") ``` 3. **Use the distilled model with Spice:** ``` embeddings: - from: model2vec:./all-MiniLM-L6-v2-model2vec name: distilled_minilm - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: all_minilm_l6_v2 ``` 4. **Race!** Compare the throughput of the distilled embedding model with the full version by declaring both models in the same Spicepod. This uses example [Wikipedia article data from Kaggle](https://www.kaggle.com/datasets/jjinho/wikipedia-20230701): ``` datasets: - from: file://wiki_a.parquet name: wiki_a_distilled acceleration: enabled: true columns: - name: text_trunc embeddings: - from: minilm_distilled - from: file://wiki_a.parquet name: wiki_a_full acceleration: enabled: true refresh_sql: select * from wiki_a_full limit 100; columns: - name: text_trunc embeddings: - from: all_minilm_l6_v2 embeddings: - from: model2vec:./all-MiniLM-L6-v2-model2vec name: distilled_minilm - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: all_minilm_l6_v2 ``` Start Spice with `spice run`: ``` 2025-08-25T15:59:39.033644Z INFO runtime::init::embedding: Embedding Model minilm_distilled ready 2025-08-25T15:59:39.969381Z INFO runtime::init::embedding: Embedding Model all_minilm_l6_v2 ready 2025-08-25T15:59:39.969713Z INFO runtime::init::dataset: Dataset wiki_a_full initializing... 2025-08-25T15:59:39.969713Z INFO runtime::init::dataset: Dataset wiki_a_distilled initializing... 2025-08-25T15:59:39.973287Z INFO runtime::init::dataset: Dataset wiki_a_full registered (file://wiki_a2.parquet), acceleration (arrow), results cache enabled. 2025-08-25T15:59:39.973344Z INFO runtime::init::dataset: Dataset wiki_a_distilled registered (file://wiki_a2.parquet), acceleration (arrow), results cache enabled. 2025-08-25T15:59:39.973637Z INFO runtime::accelerated_table::refresh_task: Loading data for dataset wiki_a_distilled 2025-08-25T15:59:39.973714Z INFO runtime::accelerated_table::refresh_task: Loading data for dataset wiki_a_full 2025-08-25T15:59:50.982854Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 40,960 records 2025-08-25T16:00:02.953192Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 57,344 records 2025-08-25T16:00:14.429262Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 40,960 records 2025-08-25T16:00:26.177483Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 49,152 records 2025-08-25T16:00:36.963177Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 40,960 records 2025-08-25T16:00:49.157552Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 49,152 records 2025-08-25T16:01:09.789714Z INFO runtime::accelerated_table::refresh_task: Loaded 100 rows (904.41 kiB) for dataset wiki_a_full in 1m 29s 816ms. ``` **Performance Results:** Note: The dramatic results are due to `model2vec` embedding execution being parallelized across all of the host's cores (default configuration). Per core, model2vec achieves a throughput of 300/400 rows/sec with this corpus. This specific test machine has 16 cores. Execution of SBERT models is currently not parallelized. | Model Name | Model Type | Records Processed | Throughput (records/sec) | | ---------------------------------------- | --------------------- | ----------------- | ------------------------ | | `sentence-transformers/all-MiniLM-L6-v2` | Model2Vec (Distilled) | 278,528 | \~4,043 | | `sentence-transformers/all-MiniLM-L6-v2` | SBERT (Full) | 100 | \~1.1 | | **Performance Gain** (model2vec) | - | - | **\~3,675x faster** | --- # OpenAI (or Compatible) Embedding Models To use a hosted OpenAI (or compatible) embedding model, specify the `openai` path in the `from` field of your configuration. For a specific model, include its model ID in the `from` field. If no model ID is specified, it defaults to `"text-embedding-3-small"`. The following parameters are specific to OpenAI models: | Parameter | Description | Default | | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------- | | `openai_api_key` | The API key for accessing OpenAI. | - | | `openai_org_id` | The organization ID for OpenAI. | - | | `openai_project_id` | The project ID for OpenAI. | - | | `openai_usage_tier` | The [OpenAI usage tier](https://platform.openai.com/settings/organization/limits) for the account. This parameter sets the maximum number of concurrent requests based on OpenAI's published limits per tier. Valid values are `free`, `tier1`, `tier2`, `tier3`, `tier4`, or `tier5`. | `tier1` | | `endpoint` | The base endpoint for the OpenAI API. | `https://api.openai.com/v1` | Credential parameters must keep the `openai_` prefix, including when `endpoint` points at an OpenAI-compatible provider such as Mistral — an unprefixed `api_key` is ignored with a warning at load time. `endpoint` is the exception and must not be prefixed. Below is an example configuration in `spicepod.yaml`: ``` embeddings: - from: openai:text-embedding-3-large name: xl_embed params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } - name: mistral from: openai:mistral-embed params: endpoint: https://api.mistral.ai/v1 openai_api_key: ${ secrets:SPICE_MISTRAL_API_KEY } ``` For detailed instructions and examples on running vector searches, refer to the [Vector-Based Search documentation](/docs/next/features/search/vector-search). --- # OpenAI Embedding Deployment Guide Production operating guide for the OpenAI embedding provider (and OpenAI-compatible endpoints) covering authentication, usage-tier rate limiting, batching, retries, and observability. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") | Parameter | Description | | ------------------- | --------------------------------------------------------------------------------------------------------------------- | | `openai_api_key` | OpenAI API key. Use `${secrets:...}` to resolve from a configured secret store. | | `openai_org_id` | OpenAI organization ID (optional). | | `openai_project_id` | OpenAI project ID (optional). | | `openai_usage_tier` | OpenAI account [usage tier](https://platform.openai.com/settings/organization/limits). Defaults to `tier1`. | | `endpoint` | Endpoint override. Defaults to `https://api.openai.com/v1`. Set for OpenAI-compatible providers (Azure OpenAI, etc.). | API keys must be sourced from a [secret store](/docs/next/components/secret-stores) in production. Credential parameters require the openai\_ prefix `openai_api_key`, `openai_org_id`, `openai_project_id`, and `openai_usage_tier` are only read under their prefixed names. An unprefixed `api_key`, `org_id`, `project_id`, or `usage_tier` is discarded with an `Ignoring parameter ...: must be prefixed with 'openai_'` warning at load time — the embedding model then starts with no credential and fails authentication at first use. `endpoint` is the one parameter that must **not** be prefixed. ### OpenAI-Compatible Providers[​](#openai-compatible-providers "Direct link to OpenAI-Compatible Providers") Set `endpoint` to route embeddings through any OpenAI-compatible provider (Azure OpenAI, Together, vLLM, Groq, local Ollama with the OpenAI-compat endpoint). Verify the provider implements `/v1/embeddings`. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Usage Tier Rate Limiting[​](#usage-tier-rate-limiting "Direct link to Usage Tier Rate Limiting") Tier selection governs the internal rate controller: | Tier | Max concurrency | Requests / minute | | ------- | --------------- | ----------------- | | `free` | 1 | 100 | | `tier1` | 35 | 3,000 | | `tier2` | 60 | 5,000 | | `tier3` | 60 | 5,000 | | `tier4` | 125 | 10,000 | | `tier5` | 125 | 10,000 | ### Batching[​](#batching "Direct link to Batching") The embeddings client automatically chunks input into batches bounded by: * **256 inputs per batch** (OpenAI's per-request input cap). * **\~512 KiB of string bytes per request batch** (safeguard against oversized requests). Large embedding jobs are transparently split across multiple API calls. ### Retry Behavior[​](#retry-behavior "Direct link to Retry Behavior") Embeddings retry with fibonacci backoff, up to **10 retries**. Retriable conditions: * An OpenAI API error carrying code `429` (rate limit, throttling), `500`, or `503` — or no code at all * A transport-level response with status `429`, `500`, `502`, `503`, or `504` * Transient `reqwest` errors (connect failures, timeouts, request/body errors) * Malformed response bodies that fail JSON deserialization Throttling (429 with rate-limit body) is detected explicitly and surfaces as a structured rate-limit error after retries are exhausted. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Vector dimensions**: Bounded by the selected embedding model (e.g., `text-embedding-3-small`: 1536, `text-embedding-3-large`: 3072). Choose based on downstream storage and retrieval cost. * **Concurrency budget**: Plan for tier-based concurrency × typical per-request latency (\~100-300 ms) to estimate achievable throughput. Embedding requests are IO-bound and scale well with concurrency up to the budget. * **Token limits**: Each input is bounded by the model's context window (8192 tokens for `text-embedding-3-*`). Inputs longer than the window fail with a `400` — truncate or chunk at the caller. ## Metrics[​](#metrics "Direct link to Metrics") Embedding requests use a dedicated metric namespace separate from chat/LLM metrics: | Metric | Type | Labels | Description | | ----------------------------------------- | --------- | ------------------------------------------------------------------ | ---------------------------------- | | `embeddings_requests` | Counter | `model`, `encoding_format`, optional `user`, optional `dimensions` | Total embedding requests issued. | | `embeddings_failures` | Counter | same as above | Total embedding request failures. | | `embeddings_internal_request_duration_ms` | Histogram | same as above | Request latency (client-side). | | `embeddings_load_errors` | Counter | - | Runtime load-time errors. | | `embeddings_active_count` | Gauge | - | Currently-loaded embedding models. | | `embeddings_load_state` | Gauge | - | Load state (0/1). | See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting metrics. ## Task History[​](#task-history "Direct link to Task History") Embedding request operations emit `text_embed` spans in [task history](/docs/next/reference/task_history), with fields: * `input` (truncated) * Labels (`model`, `encoding_format`, optional `user`, optional `dimensions`) * `outputs_produced` (number of vectors returned) * Errors (when applicable) ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **No automatic truncation**: Inputs longer than the model's context window fail with a 400 error; truncate or chunk at the caller. * **No token-level rate limiting**: The rate controller counts requests; token-level TPM limits imposed by OpenAI may still be hit and surface as 429. * **Provider compatibility varies**: OpenAI-compatible providers may not implement every parameter (dimensions, user, encoding\_format). ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ----------------------------------------- | ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `401 Unauthorized` | Wrong / revoked API key. | Rotate the key; update the secret store. | | Sustained `429 rate_limit_exceeded` | Tier budget too low or burst exceeds concurrency. | Raise `openai_usage_tier` to match the account's actual tier, or upgrade the OpenAI tier. Embeddings have no per-model concurrency override — the tier is the only knob. | | `400` with "maximum context length" | Input exceeds model context window. | Truncate or chunk inputs at the caller. | | Embeddings much slower than expected | Single-threaded caller, no batching. | Batch inputs; the client chunks into 256-input / 512 KiB batches but the caller must parallelize embedding jobs. | | Latency spikes every few hundred requests | Transient 429 with fibonacci backoff recovering. | Expected at tier ceiling; raise tier or reduce load. | --- # Model Providers Spice supports various model providers for large language models (LLMs). > **Note**: Support for traditional machine learning (ONNX) models was removed in vNext. See [Machine Learning Models](/docs/features/machine-learning-models) for details. | Name | Description | Status | LLM Format(s)\* | | --------------------------------------------------------- | ---------------------------------------------------------------------- | ----------------- | ------------------------------- | | [`openai`](/docs/next/components/models/openai) | OpenAI (or compatible) LLM endpoint | Stable | OpenAI-compatible HTTP endpoint | | [`bedrock`](/docs/next/components/models/bedrock) | Amazon Bedrock | Alpha | OpenAI-compatible HTTP endpoint | | [`xai`](/docs/next/components/models/xai) | Models hosted on xAI | Alpha | OpenAI-compatible HTTP endpoint | | [`file`](/docs/next/components/models/filesystem) | Local filesystem | Release Candidate | GGUF, GGML, SafeTensor | | [`huggingface`](/docs/next/components/models/huggingface) | Models hosted on HuggingFace | Release Candidate | GGUF, GGML, SafeTensor | | [`spice.ai`](/docs/next/components/models/spiceai) | Models hosted on the Spice.ai Cloud Platform | Release Candidate | OpenAI-compatible HTTP endpoint | | [`azure`](/docs/next/components/models/azure) | Azure OpenAI | Alpha | OpenAI-compatible HTTP endpoint | | [`anthropic`](/docs/next/components/models/anthropic) | Models hosted on Anthropic | Alpha | OpenAI-compatible HTTP endpoint | | [`google`](/docs/next/components/models/google) | Google AI language models | Alpha | OpenAI-compatible HTTP endpoint | | [`databricks`](/docs/next/components/models/databricks) | Models deployed to Databricks Mosaic AI | Alpha | OpenAI-compatible HTTP endpoint | | ~~`perplexity`~~ | ~~Perplexity~~ ([Deprecated](/docs/next/components/models/perplexity)) | Deprecated | - | Spice also tests and evaluates common models and grades their ability to integrate with Spice. See the [Models Grade Report](/docs/next/reference/models). \*LLM Format(s) may require additional files (e.g., `tokenizer_config.json`). The model type is inferred based on the model source and files. For more detail, refer to the `model` [reference specification](/docs/next/reference/spicepod/models). ## Features[​](#features "Direct link to Features") Spice supports a variety of features for large language models (LLMs): * **Custom Tools**: Provide models with tools to interact with the Spice runtime. See [Tools](/docs/next/features/large-language-models/tools). * **System Prompts**: Declaratively define system prompts and default values for [`v1/chat/completion`](/docs/next/api/HTTP/post-chat-completions) parameters. See [Parameter Overrides](/docs/next/features/large-language-models/parameter_overrides). Use Jinja templating to parameterize system prompts per request. See [Parameterized prompts](/docs/next/features/large-language-models/parameterized_prompts). * **Memory**: Provide LLMs with memory persistence tools to store and retrieve information across conversations. See [Memory](/docs/next/features/large-language-models/memory). * **Vector Search**: Perform advanced vector-based searches using embeddings. See [Vector Search](/docs/next/features/search/vector-search). * **Local Models**: Load and serve models locally from various sources, including local filesystems and Hugging Face. See [Local Models](/docs/next/features/large-language-models/serving). For more details, refer to the [Large Language Models documentation](/docs/next/features/large-language-models). ## Model Provider Prefix[​](#model-provider-prefix "Direct link to Model Provider Prefix") The model provider prefix identifies the source or provider of a model in Spice configuration files. This prefix is specified before the model identifier in the `from` field of a model definition, and is used in specifying model [default parameter overrides](#example-setting-default-parameter-overrides). It helps the runtime determine how to load and interact with the model. The following provider prefixes are supported: | Prefix | Description | | ------------ | ------------------------------------- | | `openai` | OpenAI or OpenAI-compatible endpoints | | `azure` | Azure OpenAI | | `xai` | xAI | | `anthropic` | Anthropic | | `google` | Google AI | | `hf` | Hugging Face | | `file` | Local filesystem | | `spiceai` | Spice.ai Cloud Platform | | `databricks` | Databricks Mosaic AI | | `bedrock` | Amazon Bedrock | **Example usage in `spicepod.yaml`:** ``` models: - from: openai:gpt-4o name: openai-model - from: hf:meta-llama/Llama-3-8B-Instruct name: llama3-hf - from: file://absolute/path/to/model.gguf name: local-model ``` ## Model Examples[​](#model-examples "Direct link to Model Examples") The following examples demonstrate how to configure and use various models or model features with Spice. Each example provides a specific use case to help understand the configuration options available. ### Example: Configuring an OpenAI Model[​](#example-configuring-an-openai-model "Direct link to Example: Configuring an OpenAI Model") To use a language model hosted on OpenAI (or compatible), specify the `openai` path and model ID in `from`. For more details, see [OpenAI Model Provider](/docs/next/components/models/openai). Example `spicepod.yml`: ``` models: - from: openai:gpt-4o-mini name: openai params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } - from: openai:llama3-groq-70b-8192-tool-use-preview name: groq-llama params: endpoint: https://api.groq.com/openai/v1 openai_api_key: ${ secrets:SPICE_GROQ_API_KEY } ``` ### Example: Using an OpenAI Model with Tools[​](#example-using-an-openai-model-with-tools "Direct link to Example: Using an OpenAI Model with Tools") To specify tools for an OpenAI model, include them in the `params.tools` field. For more details, see the [Tools documentation](/docs/next/features/large-language-models/tools). ``` models: - name: sql-model from: openai:gpt-4o params: tools: list_datasets, sql, table_schema ``` ### Example: Adding Memory to a Model[​](#example-adding-memory-to-a-model "Direct link to Example: Adding Memory to a Model") To enable memory tools for a model, define a `store` memory dataset and specify `memory` in the model's `tools` parameter. For more details, see the [Memory documentation](/docs/next/features/large-language-models/memory). ``` datasets: - from: memory:store name: llm_memory access: read_write models: - name: memory-enabled-model from: openai:gpt-4o params: tools: memory, sql ``` ### Example: Setting Default Parameter Overrides[​](#example-setting-default-parameter-overrides "Direct link to Example: Setting Default Parameter Overrides") To set default overrides for parameters, use the [model provider prefix](#model-provider-prefix) followed by the parameter name. For more details, see the [Parameter Overrides documentation](/docs/next/features/large-language-models/parameter_overrides). ``` models: - name: pirate-haikus from: openai:gpt-4o params: openai_temperature: 0.1 openai_response_format: { 'type': 'json_object' } ``` ### Example: Configuring a System Prompt[​](#example-configuring-a-system-prompt "Direct link to Example: Configuring a System Prompt") To configure an additional system prompt, use the `system_prompt` parameter. For more details, see the [Parameter Overrides documentation](/docs/next/features/large-language-models/parameter_overrides). ``` models: - name: pirate-haikus from: openai:gpt-4o params: system_prompt: | Write everything in Haiku like a pirate ``` ### Example: Serving a Local Model[​](#example-serving-a-local-model "Direct link to Example: Serving a Local Model") To serve a model from the local filesystem, specify the `from` path as `file` and provide the local path. For more details, see [Filesystem Model Provider](/docs/next/components/models/filesystem). ``` models: - from: file://absolute/path/to/my/model.gguf name: local_fs_model ``` ### Example: Analyzing GitHub Issues with a Chat Model[​](#example-analyzing-github-issues-with-a-chat-model "Direct link to Example: Analyzing GitHub Issues with a Chat Model") This example demonstrates how to pull GitHub issue data from the last 14 days, accelerate the data, create a chat model with memory and tools to access the accelerated data, and use Spice to ask the chat model about the general themes of new issues. #### Step 1: Pull GitHub Issue Data[​](#step-1-pull-github-issue-data "Direct link to Step 1: Pull GitHub Issue Data") First, configure a dataset to pull GitHub issue data from the last 14 days. ``` datasets: - from: github:github.com///issues name: github_issues params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true refresh_mode: append refresh_check_interval: 24h refresh_data_window: 14d ``` #### Step 2: Create a Chat Model with Memory and Tools[​](#step-2-create-a-chat-model-with-memory-and-tools "Direct link to Step 2: Create a Chat Model with Memory and Tools") Next, create a chat model that includes memory and tools to access the accelerated GitHub issue data. ``` datasets: - from: memory:store name: llm_memory access: read_write models: - name: github-issues-analyzer from: openai:gpt-4o params: tools: memory, sql ``` #### Step 3: Query the Chat Model[​](#step-3-query-the-chat-model "Direct link to Step 3: Query the Chat Model") At this step, the `spicepod.yaml` should look like: ``` datasets: - from: github:github.com///issues name: github_issues params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true refresh_mode: append refresh_check_interval: 24h refresh_data_window: 14d - from: memory:store name: llm_memory access: read_write models: - name: github-issues-analyzer from: openai:gpt-4o params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } tools: memory, sql ``` Finally, use Spice to ask the chat model about the general themes of new issues in the last 14 days. The following `curl` command demonstrates how to make this request using the OpenAI-compatible API. ``` curl -X POST http://localhost:8090/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "github-issues-analyzer", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What are the general themes of new issues in the last 14 days?"} ] }' ``` Refer to the [Create Chat Completion API documentation](/docs/next/api/HTTP/post-chat-completions) for more details on making chat completion requests. ## [🗃OpenAI](/docs/next/components/models/openai) [1 item](/docs/next/components/models/openai) ## [📄️Azure OpenAI](/docs/next/components/models/azure) [Instructions for using Azure OpenAI models](/docs/next/components/models/azure) ## [📄️Anthropic](/docs/next/components/models/anthropic) [Instructions for using language models hosted on Anthropic with Spice.](/docs/next/components/models/anthropic) ## [🗃HuggingFace](/docs/next/components/models/huggingface) [1 item](/docs/next/components/models/huggingface) ## [📄️Perplexity (Deprecated)](/docs/next/components/models/perplexity) [Perplexity model support is no longer supported in Spice.](/docs/next/components/models/perplexity) ## [🗃Filesystem](/docs/next/components/models/filesystem) [1 item](/docs/next/components/models/filesystem) ## [📄️Google AI](/docs/next/components/models/google) [Instructions for using language models hosted on Google AI with Spice.](/docs/next/components/models/google) ## [📄️Spice Cloud Platform](/docs/next/components/models/spiceai) [Instructions for using language models served by the Spice.ai Cloud Platform with Spice.](/docs/next/components/models/spiceai) ## [📄️xAI](/docs/next/components/models/xai) [Instructions for using xAI models](/docs/next/components/models/xai) ## [📄️Databricks](/docs/next/components/models/databricks) [Instructions for using Databricks Mosaic AI Models](/docs/next/components/models/databricks) ## [📄️Amazon Bedrock](/docs/next/components/models/bedrock) [How to use Amazon Bedrock models with Spice.](/docs/next/components/models/bedrock) --- # Anthropic Models To use a language model hosted on Anthropic, specify `anthropic` in the `from` field. To use a specific model, include its model ID in the `from` field (see example below). If not specified, the default model is `claude-3-5-sonnet-latest`. The following parameters are specific to Anthropic models: | Parameter | Description | Default | | ---------------------- | --------------------------------------------------------- | ------------------------------ | | `anthropic_api_key` | The Anthropic API key. | - | | `anthropic_auth_token` | The Anthropic Auth Token. | - | | `anthropic_usage_tier` | Anthropic usage tier (1-4). Used for rate limit defaults. | - | | `endpoint` | The Anthropic API base endpoint. | `https://api.anthropic.com/v1` | Example `spicepod.yml` configuration: ``` models: - from: anthropic:claude-sonnet-4-5 name: claude_4_5_sonnet params: anthropic_api_key: ${ secrets:SPICE_ANTHROPIC_API_KEY } ``` See [Anthropic Model Names](https://platform.claude.com/docs/en/about-claude/models/overview) for a list of supported model names. --- # Azure OpenAI Models To use a language model hosted on Azure OpenAI, specify the `azure` path in the `from` field and the following parameters from the [Azure OpenAI Model Deployment](https://ai.azure.com/resource/deployments) page: | Param | Description | Default | | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | | `azure_api_key` | The Azure OpenAI API key from the models deployment page. | - | | `azure_api_version` | The API version used for the Azure OpenAI service. | - | | `azure_deployment_name` | The name of the model deployment. | Model name | | `endpoint` | The Azure OpenAI resource endpoint, e.g., `https://resource-name.openai.azure.com`. | - | | `azure_entra_token` | The Azure Entra token for authentication. | - | | `responses_api` | `enabled` or `disabled`. Controls the Chat Completions backend: `enabled` proxies `/v1/chat/completions` through the backend Responses API; `disabled` uses the standard Chat Completions backend. The `/v1/responses` endpoint is available for all models regardless of this setting. | `disabled` | | `azure_openai_responses_tools` | Comma-separated list of OpenAI-hosted tools exposed via the Responses API for this model. These hosted tools are **not** available from the `/v1/chat/completions` HTTP endpoint. Supported tools: `code_interpreter`, `web_search`, `web_search_preview` (the legacy tool, for older models that do not support `web_search`). | - | Only one of `azure_api_key` or `azure_entra_token` can be provided for model configuration. Example: ``` models: - from: azure:gpt-4o-mini name: gpt-4o-mini params: endpoint: ${ secrets:SPICE_AZURE_AI_ENDPOINT } azure_api_version: 2024-08-01-preview azure_deployment_name: gpt-4o-mini azure_api_key: ${ secrets:SPICE_AZURE_API_KEY } # Responses API configuration responses_api: enabled azure_openai_responses_tools: web_search ``` Refer to the [Azure OpenAI Service models](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models) for more details on available models and configurations. Follow the [Azure OpenAI Models Cookbook](https://github.com/spiceai/cookbook/tree/trunk/azure_openai) to try Azure OpenAI models for vector-based search and chat functionalities with structured (taxi trips) and unstructured (GitHub files) data. --- # Amazon Bedrock Models Spice supports large language models hosted on [Amazon Bedrock](https://aws.amazon.com/bedrock/). Specify the `bedrock:` prefix in the `from` field along with the model ID. ## Supported Models[​](#supported-models "Direct link to Supported Models") Spice supports both Amazon's Nova models and models from other providers that are available on AWS bedrock. Providers include: | Family | Example model IDs | | ---------------- | ----------------------------------------------------------------------------------------------------- | | Amazon Nova | `amazon.nova-micro-v1:0`, `amazon.nova-lite-v1:0`, `amazon.nova-pro-v1:0`, `amazon.nova-premier-v1:0` | | Anthropic Claude | `anthropic.claude-3-5-haiku-20241022-v1:0`, `anthropic.claude-sonnet-4-20250514-v1:0` | | Meta Llama | `meta.llama3-1-70b-instruct-v1:0`, `meta.llama3-2-90b-instruct-v1:0` | | Mistral | `mistral.mixtral-8x7b-instruct-v0:1`, `mistral.mistral-large-2407-v1:0` | | Cohere Command | `cohere.command-r-v1:0`, `cohere.command-r-plus-v1:0` | | AI21 Jamba | `ai21.jamba-1-5-mini-v1:0`, `ai21.jamba-1-5-large-v1:0` | | DeepSeek | `deepseek.r1-v1:0`, `deepseek.v3.2` | Cross-region inference profiles (for example, `us.amazon.nova-lite-v1:0` or `us.meta.llama3-1-70b-instruct-v1:0`) are supported. See the [Amazon Bedrock model IDs documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/model-ids.html) for the latest IDs and availability by region. To request support for additional models, file a [GitHub Issue](https://github.com/spiceai/spiceai/issues). ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Specify the Bedrock model ID in the `from` field: ``` models: - from: bedrock:us.amazon.nova-lite-v1:0 name: novash params: aws_region: us-east-1 aws_access_key_id: ${ secrets:AWS_ACCESS_KEY_ID } aws_secret_access_key: ${ secrets:AWS_SECRET_ACCESS_KEY } ``` ### Parameters[​](#parameters "Direct link to Parameters") #### AWS Authentication[​](#aws-authentication "Direct link to AWS Authentication") | Parameter | Description | Default | | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------- | | `aws_region` | AWS region for Bedrock API requests. | `us-east-1` | | `aws_profile` | AWS profile to use when loading credentials from shared config files. | - | | `aws_access_key_id` | AWS access key ID. If not provided, credentials load from environment variables or IAM roles. | - | | `aws_secret_access_key` | AWS secret access key. If not provided, credentials load from environment variables or IAM roles. | - | | `aws_session_token` | AWS session token for temporary credentials. | - | | `aws_iam_role_source` | IAM role credential source. `auto` uses the default AWS credential chain, `metadata` uses only instance/container metadata (IMDS, ECS, EKS/IRSA), `env` uses only environment variables. | `auto` | #### Guardrails[​](#guardrails "Direct link to Guardrails") Bedrock Guardrails filter model inputs and outputs. See [GuardrailConfiguration](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_GuardrailConfiguration.html). | Parameter | Description | Default | | ------------------------------ | ---------------------------------------------------------------------------------------- | ---------- | | `bedrock_guardrail_identifier` | Guardrail ID or ARN. Example: `arn:aws:bedrock:us-east-1:123456789012:guardrail/abc123`. | - | | `bedrock_guardrail_version` | Guardrail version number or `DRAFT`. | - | | `bedrock_trace` | Trace output for guardrail evaluation. One of: `disabled`, `enabled`, `enabled_full`. | `disabled` | #### Model Parameters[​](#model-parameters "Direct link to Model Parameters") These parameters control model behavior and are passed in the request payload: | Parameter | Description | | --------------- | --------------------------------------------------------------- | | `maxTokens` | Maximum number of tokens to generate. | | `temperature` | Sampling temperature (0.0 to 1.0). Lower is more deterministic. | | `topP` | Nucleus sampling probability (0.0 to 1.0). | | `topK` | Number of highest probability tokens to consider. | | `stopSequences` | Sequences that stop generation when encountered. | See [Parameter Overrides](/docs/next/features/large-language-models/parameter_overrides) for details on setting default values. ## Examples[​](#examples "Direct link to Examples") ### Basic Configuration[​](#basic-configuration "Direct link to Basic Configuration") ``` models: - from: bedrock:amazon.nova-lite-v1:0 name: nova params: aws_region: us-east-1 aws_access_key_id: ${ secrets:AWS_ACCESS_KEY_ID } aws_secret_access_key: ${ secrets:AWS_SECRET_ACCESS_KEY } ``` ### Cross-Region Inference[​](#cross-region-inference "Direct link to Cross-Region Inference") Use cross-region inference profiles for improved availability: ``` models: - from: bedrock:us.amazon.nova-pro-v1:0 name: nova-pro params: aws_region: us-east-1 ``` ### Inference Profile for Models Without On-Demand Throughput[​](#inference-profile-for-models-without-on-demand-throughput "Direct link to Inference Profile for Models Without On-Demand Throughput") Some models (for example, several Anthropic/Meta variants) require inference profile IDs: ``` models: - from: bedrock:us.meta.llama3-1-70b-instruct-v1:0 name: llama31 params: aws_region: us-east-1 - from: bedrock:us.anthropic.claude-opus-4-6-v1 name: claude-opus-46 params: aws_region: us-east-1 ``` ### With Guardrails[​](#with-guardrails "Direct link to With Guardrails") ``` models: - from: bedrock:amazon.nova-lite-v1:0 name: nova-guarded params: aws_region: us-east-1 bedrock_guardrail_identifier: arn:aws:bedrock:us-east-1:123456789012:guardrail/abc123 bedrock_guardrail_version: '1' bedrock_trace: enabled ``` ## Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate bedrock permissions (e.g., `bedrock:InvokeModel`) to access the model. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs permissions to invoke Bedrock models: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"], "Resource": ["arn:aws:bedrock:us-east-1::foundation-model/amazon.nova-*"] } ] } ``` | Permission | Purpose | | --------------------------------------- | --------------------------------------------- | | `bedrock:InvokeModel` | Required. Invoke model for text generation. | | `bedrock:InvokeModelWithResponseStream` | Required. Invoke model with streaming output. | ## Related Resources[​](#related-resources "Direct link to Related Resources") * [Amazon Bedrock Embeddings](/docs/next/components/embeddings/bedrock) - Use Bedrock for text embeddings * [Parameter Overrides](/docs/next/features/large-language-models/parameter_overrides) - Set default model parameters * [Amazon Bedrock User Guide](https://docs.aws.amazon.com/bedrock/latest/userguide/) - AWS documentation * [Bedrock Model IDs](https://docs.aws.amazon.com/bedrock/latest/userguide/model-ids.html) - Available models and inference profiles --- # Databricks Model Provider To use a language model deployed to [Databricks Mosaic AI Model Serving](https://docs.databricks.com/aws/en/machine-learning/model-serving/), specify the model endpoint name prefixed with `databricks:` in the `from` field and include the required parameters in the `params` section. ### Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `databricks_endpoint` | The Databricks workspace endpoint, e.g., `dbc-a12cd3e4-56f7.cloud.databricks.com`. | | `databricks_token` | The Databricks API token to authenticate with the Databricks Models API. Use the [secret replacement syntax](/docs/next/components/secret-stores) to reference a secret, e.g., `${secrets:my_databricks_token}`. | | `databricks_client_id` | The Databricks Service Principal Client ID. Can't be used with `databricks_token`. | | `databricks_client_secret` | The Databricks Service Principal Client Secret. Can't be used with `databricks_token`. | ### Example `spicepod.yaml` configuration, using personal access token[​](#example-spicepodyaml-configuration-using-personal-access-token "Direct link to example-spicepodyaml-configuration-using-personal-access-token") To learn more about how to set up personal access tokens, see [Databricks PAT docs](https://docs.databricks.com/aws/en/dev-tools/auth/pat). ``` models: - from: databricks:databricks-llama-4-maverick name: llama-4-maverick params: databricks_endpoint: dbc-46470731-42e5.cloud.databricks.com databricks_token: ${ secrets:SPICE_DATABRICKS_TOKEN } ``` ### Example `spicepod.yaml` configuration, using Databricks service principal[​](#example-spicepodyaml-configuration-using-databricks-service-principal "Direct link to example-spicepodyaml-configuration-using-databricks-service-principal") Spice supports the Machine-to-Machine (M2M) OAuth flow with service principal credentials by utilizing the `databricks_client_id` and `databricks_client_secret` parameters. The runtime will automatically refresh the token. The service principal must be granted the "Can Query" permission for model serving. To learn more about how to set up the service principal, see [Databricks M2M OAuth docs](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m). ``` models: - from: databricks:databricks-llama-4-maverick name: llama-4-maverick params: databricks_endpoint: dbc-46470731-42e5.cloud.databricks.com databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} databricks_client_secret: ${secrets:DATABRICKS_CLIENT_SECRET} ``` ### Additional Information[​](#additional-information "Direct link to Additional Information") Refer to the [Mosaic AI Model Serving documentation](https://docs.databricks.com/aws/en/machine-learning/model-serving/) for more details on available models and configurations. --- # Filesystem Hosted Models To use a model hosted on a filesystem, specify the path to the model file or folder in the `from` field: ``` models: - from: file://models/llms/llama3.2-1b-instruct/ name: llama3 ``` Supported formats include GGUF, GGML, and SafeTensor for large language models (LLMs). ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") An absolute or relative path to the model file or folder: ``` from: file://absolute/path/models/llms/llama3.2-1b-instruct/ from: file:models/llms/llama3.2-1b-instruct/ ``` ### `params` (optional)[​](#params-optional "Direct link to params-optional") | Param | Description | | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `tools` | Which [tools](/docs/next/features/large-language-models/tools) should be made available to the model. Set to `auto` to use all available tools. | | `system_prompt` | An additional system prompt used for all chat completions to this model. | | `chat_template` | Customizes the transformation of OpenAI chat messages into a character stream for the model. See [Overriding the Chat Template](#overriding-the-chat-template). | | `trust_pickle` | Allow loading pickle-based weight files (`.bin`, `.pt`, `.pth`, `.ckpt`). These formats execute arbitrary code when loaded and are rejected by default. Accepts `true` or `false`; defaults to `false`. Set to `true` only when the weights come from a fully trusted source. | See [Large Language Models](/docs/next/features/large-language-models) for additional configuration options. * [Tools](/docs/next/features/large-language-models/tools) * [Memory](/docs/next/features/large-language-models/memory) * [Parameter overrides](/docs/next/features/large-language-models/parameter_overrides) ### `files` (optional)[​](#files-optional "Direct link to files-optional") The `files` field specifies additional files required by the model, such as tokenizer, configuration, and other files. ``` - name: local-model from: file://models/llms/llama3.2-1b-instruct/model.safetensors files: - path: //models/llms/llama3.2-1b-instruct/tokenizer.json - path: //models/llms/llama3.2-1b-instruct/tokenizer_config.json - path: //models/llms/llama3.2-1b-instruct/config.json ``` ## Examples[​](#examples "Direct link to Examples") ### Loading a GGML Model[​](#loading-a-ggml-model "Direct link to Loading a GGML Model") ``` models: - from: file://absolute/path/to/my/model.ggml name: local_ggml_model files: - path: models/llms/ggml/tokenizer.json - path: models/llms/ggml/tokenizer_config.json - path: models/llms/ggml/config.json ``` ### Example: Loading a SafeTensor Model[​](#example-loading-a-safetensor-model "Direct link to Example: Loading a SafeTensor Model") ``` models: - name: safety from: file:models/llms/llama3.2-1b-instruct/model.safetensors files: - path: models/llms/llama3.2-1b-instruct/tokenizer.json - path: models/llms/llama3.2-1b-instruct/tokenizer_config.json - path: models/llms/llama3.2-1b-instruct/config.json ``` ### Loading LLM from a directory[​](#loading-llm-from-a-directory "Direct link to Loading LLM from a directory") ``` models: - name: llama3 from: file:models/llms/llama3.2-1b-instruct/ ``` Note: The folder provided should contain all the expected files (see examples above). ### Loading a GGUF Model[​](#loading-a-gguf-model "Direct link to Loading a GGUF Model") ``` models: - from: file://absolute/path/to/my/model.gguf name: local_gguf_model ``` ### Overriding the Chat Template[​](#overriding-the-chat-template "Direct link to Overriding the Chat Template") Chat templates convert the OpenAI compatible chat messages (see [format](https://platform.openai.com/docs/api-reference/chat/create#chat-create-messages)) and other components of a request into a stream of characters for the language model. It follows Jinja3 templating [syntax](https://jinja.palletsprojects.com/en/3.1.x/templates/). Further details on chat templates can be found [here](https://huggingface.co/docs/transformers/main/chat_templating#advanced-how-do-chat-templates-work). ``` models: - name: local_model from: file:path/to/my/model.gguf params: chat_template: | {% set loop_messages = messages %} {% for message in loop_messages %} {% set content = '<|start_header_id|>' + message['role'] + '<|end_header_id|>\n\n'+ message['content'] | trim + '<|eot_id|>' %} {{ content }} {% endfor %} {% if add_generation_prompt %} {{ '<|start_header_id|>assistant<|end_header_id|>\n\n' }} {% endif %} ``` #### Templating Variables[​](#templating-variables "Direct link to Templating Variables") * `messages`: List of chat messages, in the OpenAI [format](https://platform.openai.com/docs/api-reference/chat/create#chat-create-messages). * `add_generation_prompt`: Boolean flag whether to add a [generation prompt](https://huggingface.co/docs/transformers/main/chat_templating#what-are-generation-prompts). * `tools`: List of callable tools, in the OpenAI [format](https://platform.openai.com/docs/api-reference/chat/create#chat-create-tools). Limitations * The throughput, concurrency & latency of a locally hosted model will vary based on the underlying hardware and model size. Spice supports Apple Metal and CUDA for accelerated inference. See [CONTRIBUTING.md](https://github.com/spiceai/spiceai/blob/trunk/CONTRIBUTING.md) for build instructions. --- # Filesystem Model Deployment Guide Production operating guide for loading local language models from the filesystem (GGUF, safetensors). ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") The Filesystem model provider has no authentication layer. Access control is enforced by the operating system: * The Spice runtime process must have read permission on the model files. * For containers, mount model files as read-only volumes. * For Kubernetes, mount via `PersistentVolumeClaim` or a model-serving sidecar. For sensitive models (proprietary weights, fine-tunes with PII training data), restrict filesystem ACLs to the Spice process user and encrypt the volume at rest. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") The Filesystem model provider reads local files synchronously. There is no network layer, retry logic, or remote backoff. Failures surface as filesystem errors (`ENOENT`, `EACCES`, `EIO`). Model loading happens once at startup. A missing or unreadable file fails the spicepod load; fix the underlying cause and restart. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") ### Supported Formats[​](#supported-formats "Direct link to Supported Formats") | Format | Extension | Notes | | ------------- | --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | GGUF | `.gguf` | Quantized / unquantized; loaded via the mistral local loader. | | GGML (legacy) | `.ggml` | Legacy llama.cpp format. | | Safetensors | `.safetensors` | Native tensor format; preferred over `.bin` for safety. | | PyTorch | `.bin` / `.pt` / `.pth` / `.ckpt` | Legacy pickle-based checkpoints. **Rejected by default** — these formats execute arbitrary code when loaded. Set [`trust_pickle: true`](/docs/next/components/models/filesystem#params-optional) to load them from a fully trusted source. | ### Device Selection[​](#device-selection "Direct link to Device Selection") Local inference uses the first available backend in order: 1. **CUDA** (if compiled with CUDA support and a device is present) 2. **Metal** (if compiled with Metal support — macOS / Apple Silicon) 3. **CPU** fallback Install the CUDA-enabled Spice build on GPU hosts for materially better throughput on models over a few billion parameters. ### Memory Footprint[​](#memory-footprint "Direct link to Memory Footprint") Model file size on disk is close to RAM / VRAM footprint at load. Quantized GGUF models (Q4, Q5, Q8) reduce footprint roughly proportional to their bit-width. For a 7B-parameter model: * `f16`: \~14 GB * `Q8`: \~7.5 GB * `Q5`: \~5 GB * `Q4`: \~4 GB Add \~20–30% headroom for KV cache and working memory during inference. ### Concurrency[​](#concurrency "Direct link to Concurrency") The runtime rate limiter defaults to **`max_concurrency=1`** for local models — local inference is compute-bound and benefits little from request-level parallelism on a single accelerator. Override via `max_concurrency` for multi-GPU / large-core CPU hosts. ## Metrics[​](#metrics "Direct link to Metrics") Shared LLM metrics apply (see the [OpenAI Model Deployment Guide](/docs/next/components/models/openai/deployment#metrics) for the full metric list): `llm_requests`, `llm_failures`, `llm_internal_request_duration_ms`, `llm_prompt_tokens_total`, `llm_completion_tokens_total`. See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting metrics. ## Task History[​](#task-history "Direct link to Task History") Local inference operations emit `ai_completion` spans (and `health` spans for probes) in [task history](/docs/next/reference/task_history), mirroring the shared model spans. `captured_output` and token usage fields are logged. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Single-process loading**: A model is loaded into the Spice process — it cannot be shared across process instances without a dedicated inference server. * **Format support depends on compile features**: CUDA and Metal support are conditional on the Spice build flavor (default, CUDA). * **No hot reload**: Swapping the underlying model file requires a spicepod reload. * **No integrity check**: The provider trusts the file on disk. Validate checksums out-of-band for supply-chain assurance. * **Model\_type override**: When the loader cannot auto-detect the architecture, `model_type` can force a known architecture. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | --------------------------------------------------- | ----------------------------------------------- | -------------------------------------------------------------------------------------------------- | | `No such file or directory` | Path typo or missing mount. | Verify the file exists in the Spice process's filesystem (`ls` inside the container). | | `Permission denied` | Spice user lacks read on the file. | Adjust ACLs or mount with appropriate UID/GID. | | Model fails to load with `unsupported architecture` | Loader cannot infer architecture from filename. | Set `model_type` explicitly. | | OOM on load | File size exceeds device memory. | Use a smaller / more quantized variant; move to CPU with more RAM; split across GPUs if supported. | | Inference falls back to CPU unexpectedly | CUDA / Metal not available. | Use a CUDA-enabled Spice build on GPU hosts; for macOS, use the Apple Silicon build. | | Slow first inference after startup | JIT / weight-quantization warmup. | Issue a warmup request at startup; subsequent calls are hot. | --- # Google AI Models To use a language model hosted on Google AI, specify `google` in the `from` field. Include a model ID in the `from` field (see example below); a model ID is required. Spice does not apply a default model — the model fails to load if the ID is omitted (`from: google`). The following parameters are specific to Google AI models: | Parameter | Description | Default | | ---------------- | ---------------------- | ------- | | `google_api_key` | The Google AI API key. | - | Example `spicepod.yml` configuration: ``` models: - from: google:gemini-3.5-flash name: flash params: google_api_key: ${ secrets:GEMINI_API_KEY } ``` See [Google AI Models](https://ai.google.dev/gemini-api/docs/models/gemini) for a list of supported model names. See [Large Language Models](/docs/next/features/large-language-models) for additional configuration options: * [Tools](/docs/next/features/large-language-models/tools) * [Memory](/docs/next/features/large-language-models/memory) * [Parameter overrides](/docs/next/features/large-language-models/parameter_overrides) --- # HuggingFace To use a model hosted on HuggingFace, specify the `huggingface.co` path in the `from` field and, when needed, the files to include. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` key takes the form of `huggingface:model_path`. Below shows 2 common example of `from` key configuration. * `huggingface:username/modelname`: Implies the latest version of `modelname` hosted by `username`. * `huggingface:huggingface.co/username/modelname:revision`: Specifies a particular `revision` of `modelname` by `username`, including the optional domain. The `from` key follows the following regex format. ``` \A(huggingface:)(huggingface\.co\/)?(?[\w\-]+)\/(?[\w\-]+)(:(?[\w\d\-\.]+))?\z ``` The `from` key consists of five components: 1. **Prefix:** The value must start with `huggingface:`. 2. **Domain (Optional):** Optionally includes `huggingface.co/` immediately after the prefix. Currently no other Huggingface compatible services are supported. 3. **Organization/User:** The HuggingFace organization (`org`). 4. **Model Name:** After a `/`, the model name (`model`). 5. **Revision (Optional):** A colon (`:`) followed by the git-like revision identifier (`revision`). ### `name`[​](#name "Direct link to name") The model name. This will be used as the model ID within Spice and Spice's endpoints (i.e. `http://localhost:8090/v1/models`). This can be set to the same value as the model ID in the `from` field. ### `params`[​](#params "Direct link to params") | Param | Description | Default | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------- | | `hf_token` | The Huggingface access token. | - | | `model_type` | The architecture to load the model as. Supported text architectures: `mistral`, `gemma`, `mixtral`, `llama`, `phi2`, `phi3`, `qwen2`, `gemma2`, `starcoder2`, `phi3.5moe`, `deepseekv2`, `deepseekv3`, `qwen3`, `glm4`, `glm4moelite`, `glm4moe`, `qwen3moe`, `smollm3`, `granitemoehybrid`, `gpt_oss`, `qwen3next`. Supported multimodal architectures: `phi3v`, `idefics2`, `llava_next`, `llava`, `vllama`, `qwen2vl`, `idefics3`, `minicpmo`, `phi4mm`, `qwen2_5vl`, `gemma3`, `mistral3`, `llama4`, `gemma3n`, `gemma4`, `qwen3vl`, `qwen3vlmoe`, `qwen3_5`, `qwen3_5moe`, `voxtral`. | - | | `tools` | Which \[tools] should be made available to the model. Set to `auto` to use all available tools. | - | | `system_prompt` | An additional system prompt used for all chat completions to this model. | - | ### `files`[​](#files "Direct link to files") The specific file path for Huggingface model. For example, GGUF model formats require a specific file path, other varieties (e.g. `.safetensors`) are inferred. #### Example[​](#example "Direct link to Example") ``` models: - from: huggingface:huggingface.co/lmstudio-community/Qwen2.5-Coder-3B-Instruct-GGUF name: sloth-gguf files: - path: Qwen2.5-Coder-3B-Instruct-Q3_K_L.gguf ``` ## Access Tokens[​](#access-tokens "Direct link to Access Tokens") Access tokens can be provided for Huggingface models in two ways: 1. In the Huggingface token cache (i.e. `~/.cache/huggingface/token`). Default. 2. Via [model params](#params). ``` models: - name: llama_3.2_1B from: huggingface:huggingface.co/meta-llama/Llama-3.2-1B params: hf_token: ${ secrets:HF_TOKEN } ``` ## Examples[​](#examples "Direct link to Examples") ### Load a LLM model to generate text[​](#load-a-llm-model-to-generate-text "Direct link to Load a LLM model to generate text") ``` models: - from: huggingface:huggingface.co/microsoft/Phi-3.5-mini-instruct name: phi ``` ### Load a private model[​](#load-a-private-model "Direct link to Load a private model") ``` models: - name: llama_3.2_1B from: huggingface:huggingface.co/meta-llama/Llama-3.2-1B params: hf_token: ${ secrets:HF_TOKEN } ``` For more details on authentication, see [access tokens](#access-tokens). Limitations * The throughput, concurrency & latency of a locally hosted model will vary based on the underlying hardware and model size. Spice supports Apple Metal and CUDA for accelerated inference. See [CONTRIBUTING.md](https://github.com/spiceai/spiceai/blob/trunk/CONTRIBUTING.md) for build instructions. ## Cookbook[​](#cookbook "Direct link to Cookbook") * Use the Llama family of models locally from HuggingFace using Spice. [Running Llama3 Locally](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md) --- # Hugging Face Model Deployment Guide Production operating guide for loading models from the Hugging Face Hub and running local inference. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") | Parameter | Description | | ---------- | --------------------------------------------------------------- | | `hf_token` | Hugging Face access token. Required for private or gated repos. | | `token` | Alias accepted by some integrations. | Tokens must be sourced from a [secret store](/docs/next/components/secret-stores) in production. For public, non-gated models the token is optional; for private / gated repos (Llama, most Mistral checkpoints), the token is required. ### Token Discovery Fallback[​](#token-discovery-fallback "Direct link to Token Discovery Fallback") When `hf_token` is unset, the local loader falls back to the Hugging Face token cache (typically `~/.cache/huggingface/token` or `HF_TOKEN_PATH`). This makes local development portable but should be explicitly set in production via the secret store to avoid surprise auth behavior across environments. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Download & Cache[​](#download--cache "Direct link to Download & Cache") Models are downloaded on first use into `~/.spice/models///`. Existing files are skipped on subsequent starts (cache-by-file-existence). Download requests use bearer auth when a token is configured. Path-traversal protections ensure all downloaded files stay within the model directory. ### Revision Pinning[​](#revision-pinning "Direct link to Revision Pinning") Model IDs support explicit revision pinning (e.g. `org/model@revision`). `latest` maps to `main`. Revisions are sanitized for path safety before use. Pin revisions in production to guarantee reproducibility — `main` is a moving target. ### Retry Behavior[​](#retry-behavior "Direct link to Retry Behavior") Download retries follow the shared HTTP-client policy with exponential/fibonacci backoff on transient failures. For very large models over slow networks, pre-download into the cache directory with the Hugging Face CLI to avoid first-request latency. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") ### Device Selection[​](#device-selection "Direct link to Device Selection") Local inference uses the first available backend in order: 1. **CUDA** (if compiled with CUDA support and a device is present) 2. **Metal** (if compiled with Metal support — macOS / Apple Silicon) 3. **CPU** fallback Install the CUDA-enabled Spice build on GPU hosts; the standard build uses CPU-only inference which is significantly slower for models over a few billion parameters. ### Memory Footprint[​](#memory-footprint "Direct link to Memory Footprint") Model size on disk is close to RAM/VRAM footprint at load. Quantized GGUF models (Q4, Q5, Q8) reduce footprint roughly proportional to their bit-width. For a 7B parameter model: * `f16`: \~14 GB * `Q8`: \~7.5 GB * `Q5`: \~5 GB * `Q4`: \~4 GB Add \~20–30% headroom for KV cache and working memory during inference. ### Concurrency[​](#concurrency "Direct link to Concurrency") The runtime rate limiter defaults to **`max_concurrency=1`** for local models (HuggingFace, filesystem) — local inference is compute-bound and benefits little from request-level parallelism on a single accelerator. Override via `max_concurrency` for multi-GPU / large-core CPU hosts. ## Metrics[​](#metrics "Direct link to Metrics") Shared LLM metrics apply (see the [OpenAI Model Deployment Guide](/docs/next/components/models/openai/deployment#metrics) for the full metric list): `llm_requests`, `llm_failures`, `llm_internal_request_duration_ms`, `llm_prompt_tokens_total`, `llm_completion_tokens_total`. See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting metrics. ## Task History[​](#task-history "Direct link to Task History") Local inference operations emit `ai_completion` spans (and `health` spans for probes) in [task history](/docs/next/reference/task_history), mirroring the OpenAI-path spans. `captured_output` and token usage fields are logged. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **Single-process loading**: A model is loaded into the Spice process — it cannot be shared across process instances without a dedicated inference server. * **No hot reload**: Switching model revisions requires a spicepod reload. * **Limited Responses API support**: Responses API routing is currently tied to specific providers (OpenAI, xAI); a local HF-loaded model does not serve the Responses API. * **Quantized formats**: Support depends on the local loader (mistral / candle). Verify the format is supported before production deployment. * **Disk-space requirements**: First-run downloads can be multi-GB; ensure `~/.spice/models/` has adequate space. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ---------------------------------------- | ------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | | `401 Unauthorized` on download | Missing or invalid `hf_token`; gated model. | Set `hf_token`; accept the model's license on Hugging Face; verify token has `read` scope. | | OOM on model load | Model size exceeds device memory. | Choose a smaller quantized variant; switch to CPU + larger system RAM; use multi-GPU if supported. | | Inference falls back to CPU unexpectedly | CUDA / Metal unavailable or not detected. | Use a CUDA-enabled Spice build on GPU hosts; verify `nvidia-smi` shows devices; for macOS, use Apple Silicon build. | | Model output changes between restarts | Revision unpinned (`main`). | Pin the revision: `org/model@revision_hash`. | | First request extremely slow | Model downloading on first run. | Pre-warm with `huggingface-cli download` into the Spice model cache, or start with `initial_load: true` if supported. | | Path traversal error on startup | Malformed revision string. | Use a clean revision: alphanumeric + underscores + dashes only; commit SHAs are safe. | --- # OpenAI (or Compatible) Language Models To use a language model hosted on OpenAI (or compatible), specify the `openai` path in the `from` field. For a specific model, include it as the model ID in the `from` field (see example below). The default model is `gpt-4o-mini`. ``` models: - from: openai:gpt-4o-mini name: openai_model params: openai_api_key: ${ secrets:OPENAI_API_KEY } # Required for official OpenAI models tools: auto # Optional. Connect the model to datasets via SQL query/vector search tools system_prompt: 'You are a helpful assistant.' # Optional. # Optional parameters endpoint: https://api.openai.com/v1 # Override to use a compatible provider (i.e. NVidia NIM) openai_org_id: ${ secrets:OPENAI_ORG_ID } openai_project_id: ${ secrets:OPENAI_PROJECT_ID } # Override default chat completion request parameters openai_temperature: 0.1 openai_response_format: { 'type': 'json_object' } # OpenAI Responses API configuration responses_api: enabled openai_responses_tools: web_search, code_interpreter ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `openai:model_id` where `model_id` is the model ID of the OpenAI model, valid model IDs are found in the `{endpoint}/v1/models` API response. Example: ``` curl -H "Authorization: Bearer $OPENAI_API_KEY" https://api.openai.com/v1/models ``` ``` { "object": "list", "data": [ { "id": "gpt-4o-mini", "object": "model", "created": 1727389042, "owned_by": "system" }, ... } ``` ### `name`[​](#name "Direct link to name") The model name. This will be used as the model ID within Spice and Spice's endpoints (i.e. `http://localhost:8090/v1/models`). This can be set to the same value as the model ID in the `from` field. ### `params`[​](#params "Direct link to params") | Param | Description | Default | | ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------- | | `endpoint` | The OpenAI API base endpoint. Can be overridden to use a compatible provider (i.e. NVIDIA NIM). | `https://api.openai.com/v1` | | `tools` | Which [tools](/docs/next/features/large-language-models/tools) should be made available to the model. Set to `auto` to use all available tools. | - | | `system_prompt` | An additional system prompt used for all chat completions to this model. | - | | `openai_api_key` | The OpenAI API key. | - | | `openai_org_id` | The OpenAI organization ID. | - | | `openai_project_id` | The OpenAI project ID. | - | | `openai_temperature` | Set the default temperature to use on chat completions. | - | | `openai_response_format` | An object specifying the format that the model must output, see [structured outputs](https://platform.openai.com/docs/guides/structured-outputs). | - | | `openai_reasoning_effort` | For reasoning models, like `o1`, this parameter specifies the reasoning effort used for the model. | - | | `openai_usage_tier` | The [OpenAI usage tier](https://platform.openai.com/settings/organization/limits) for the account. This parameter sets the maximum number of concurrent requests based on OpenAI's published limits per tier. Valid values are `free`, `tier1`, `tier2`, `tier3`, `tier4`, or `tier5`. | `tier1` | | `responses_api` | `enabled` or `disabled`. Whether to enable invoking this model from the `/v1/responses` HTTP endpoint using [OpenAI's Responses API](https://platform.openai.com/docs/api-reference/responses). When using OpenAI-compatible providers, ensure the provider supports OpenAI's Responses API. | `disabled` | | `openai_responses_tools` | Comma-separated list of OpenAI-hosted tools exposed via the Responses API for this model. These hosted tools are **not** available from the `/v1/chat/completions` HTTP endpoint. Supported tools: `code_interpreter`, `web_search`, `web_search_preview`. | - | `web_search` selects OpenAI's current web search tool. `web_search_preview` selects the legacy tool of the same name, retained for older models that do not support the current one — prefer `web_search` unless the model rejects it. See [Large Language Models](/docs/next/features/large-language-models) for additional configuration options. * [Tools](/docs/next/features/large-language-models/tools) * [Memory](/docs/next/features/large-language-models/memory) * [Parameter overrides](/docs/next/features/large-language-models/parameter_overrides) ## Supported OpenAI Compatible Providers[​](#supported-openai-compatible-providers "Direct link to Supported OpenAI Compatible Providers") Spice supports several OpenAI compatible providers. Specify the appropriate endpoint in the params section. ### Azure OpenAI[​](#azure-openai "Direct link to Azure OpenAI") Follow [Azure AI Models](/docs/next/components/models/azure) instructions. ### Groq[​](#groq "Direct link to Groq") Groq provides OpenAI compatible endpoints. Use the following configuration: ``` models: - from: openai:llama3-groq-70b-8192-tool-use-preview name: groq-llama params: endpoint: https://api.groq.com/openai/v1 openai_api_key: ${ secrets:SPICE_GROQ_API_KEY } ``` ### NVIDIA NIM[​](#nvidia-nim "Direct link to NVIDIA NIM") NVIDIA NIM models are OpenAI compatible endpoints. Use the following configuration: ``` models: - from: openai:my_nim_model_id name: my_nim_model params: endpoint: https://my_nim_host.com/v1 openai_api_key: ${ secrets:SPICE_NIM_API_KEY } ``` View the Spice cookbook for an example of setting up NVIDIA NIM with Spice [here](https://github.com/spiceai/cookbook/tree/trunk/nvidia-nim/ec2). ### Parasail[​](#parasail "Direct link to Parasail") Parasail also offers OpenAI compatible endpoints. Use the following configuration: ``` models: - from: openai:parasail-model-id name: parasail_model params: endpoint: https://api.parasail.com/v1 openai_api_key: ${ secrets:SPICE_PARASAIL_API_KEY } ``` Refer to the respective provider documentation for more details on available models and configurations. --- # OpenAI Model Deployment Guide Production operating guide for the OpenAI model provider (and OpenAI-compatible endpoints) covering authentication, usage-tier-based rate limiting, the Responses API, and observability. ## Authentication & Secrets[​](#authentication--secrets "Direct link to Authentication & Secrets") | Parameter | Description | | ------------------- | --------------------------------------------------------------------------------------------------------------------------- | | `openai_api_key` | OpenAI API key. Use `${secrets:...}` to resolve from a configured secret store. | | `openai_org_id` | OpenAI organization ID (optional). | | `openai_project_id` | OpenAI project ID (optional). | | `endpoint` | Endpoint override. Defaults to `https://api.openai.com/v1`. Set for OpenAI-compatible providers (Azure OpenAI, Groq, etc.). | API keys must be sourced from a [secret store](/docs/next/components/secret-stores) in production. Rotate keys periodically; OpenAI dashboard tracks per-key usage which helps with rotation planning. ### OpenAI-Compatible Providers[​](#openai-compatible-providers "Direct link to OpenAI-Compatible Providers") Set `endpoint` to target any OpenAI-compatible provider (Azure OpenAI, xAI, Groq, Together, on-prem vLLM, etc.). See the [OpenAI model reference](/docs/next/components/models/openai/) for provider-specific configuration examples. The `/v1/responses` endpoint works with all OpenAI-compatible providers through automatic format adaptation. When setting `responses_api: enabled` (which proxies Chat Completions through the Responses API backend), confirm the target provider implements OpenAI's Responses API natively. ## Resilience Controls[​](#resilience-controls "Direct link to Resilience Controls") ### Usage Tier Rate Limiting[​](#usage-tier-rate-limiting "Direct link to Usage Tier Rate Limiting") | Parameter | Default | Description | | ------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `openai_usage_tier` | `tier1` | OpenAI account [usage tier](https://platform.openai.com/settings/organization/limits). Accepted values: `free`, `tier1`, `tier2`, `tier3`, `tier4`, `tier5`. | Tier selection governs the internal rate controller's concurrency + per-minute budget: | Tier | Max concurrency | Requests / minute | | ------- | --------------- | ----------------- | | `free` | 1 | 100 | | `tier1` | 35 | 3,000 | | `tier2` | 60 | 5,000 | | `tier3` | 60 | 5,000 | | `tier4` | 125 | 10,000 | | `tier5` | 125 | 10,000 | Override tier defaults per model via global model parameters: | Parameter | Description | | --------------------------- | ------------------------------------------ | | `max_concurrency` | Override the per-model concurrency budget. | | `requests_per_minute_limit` | Override the per-model RPM budget. | The built-in rate controller queues and paces outbound requests to stay within these budgets, avoiding the tight loop of hitting the OpenAI rate limiter and retrying. ### Retry Behavior[​](#retry-behavior "Direct link to Retry Behavior") **Chat Completions / Responses** requests rely on the rate controller for pacing; there is no application-level retry loop around the chat/responses path. Transient 429 / 5xx responses surface to the caller. **Embeddings** (see the [OpenAI Embedding Deployment Guide](/docs/next/components/embeddings/openai/deployment)) implement an application-level retry with fibonacci backoff and a 10-retry cap. ### Responses API[​](#responses-api "Direct link to Responses API") All configured models are registered for the `/v1/responses` endpoint. For OpenAI-compatible providers, Spice automatically adapts between Chat Completions and Responses API formats, so `/v1/responses` works even when the backend only supports `/v1/chat/completions`. | Parameter | Default | Description | | ------------------------ | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `responses_api` | `disabled` | Controls the Chat Completions backend. `disabled` proxies `/v1/chat/completions` to the backend's `/v1/chat/completions`. `enabled` proxies `/v1/chat/completions` to the backend's `/v1/responses`, which can improve tool-use and reasoning for providers that natively support the Responses API. | | `openai_responses_tools` | - | Comma-separated list of OpenAI-hosted tools (`code_interpreter`, `web_search`, `web_search_preview`) exposed via Responses. | Note: Responses API hosted tools are **not** available from the `/v1/chat/completions` endpoint. `web_search` selects OpenAI's current web search tool. `web_search_preview` selects the legacy tool, retained for older models that do not support the current one. ## Capacity & Sizing[​](#capacity--sizing "Direct link to Capacity & Sizing") * **Context window**: Driven by the selected model (e.g., `gpt-4o-mini`: 128k). Spice does not enforce context trimming — prompt and history management is the caller's responsibility. * **Concurrency**: Set by `openai_usage_tier` or explicit `max_concurrency`. For high-throughput workloads, upgrade the OpenAI tier rather than over-subscribing the rate limiter. * **Latency**: Bounded by OpenAI's server-side latency plus network RTT. Streaming (`stream=true`) begins emitting tokens quickly and is preferred for chat interfaces. ## Metrics[​](#metrics "Direct link to Metrics") All LLM requests (OpenAI, Hugging Face, filesystem) share a common metric namespace: | Metric | Type | Description | | ---------------------------------- | --------- | --------------------------------- | | `llm_requests` | Counter | Total LLM requests issued. | | `llm_failures` | Counter | Total LLM request failures. | | `llm_internal_request_duration_ms` | Histogram | Request latency (client-side). | | `llm_prompt_tokens_total` | Counter | Total prompt tokens sent. | | `llm_completion_tokens_total` | Counter | Total completion tokens received. | Requests routed through the Responses API carry a `responses_api` label for differentiation. See [Component Metrics](/docs/next/features/observability/component_metrics) for enabling and exporting metrics. ## Task History[​](#task-history "Direct link to Task History") Chat and Responses operations emit these [task history](/docs/next/reference/task_history) spans: | Span | Fields | Description | | --------------- | ------------- | ------------------------------------------------------ | | `ai_completion` | \`stream=true | false\`, usage tokens | | `responses` | \`stream=true | false\`, usage tokens | | `health` | - | Health probe against the endpoint (chat or responses). | `captured_output` and token usage fields are logged in task-history entries. ## Known Limitations[​](#known-limitations "Direct link to Known Limitations") * **No automatic token counting for rate limiting**: The rate controller counts requests, not tokens. Token-level rate limits imposed by OpenAI (TPM) are not pre-checked; excess requests surface as 429 errors. * **No chat/responses application retry**: Retries for chat/responses are not implemented at the Spice layer. Pace via `max_concurrency`/`requests_per_minute_limit` to stay under rate limits. * **OpenAI-compatible providers vary**: Not every OpenAI-compatible provider implements every parameter (e.g., `responses_api`, `openai_reasoning_effort`). Test against your specific provider. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | ------------------------------------------ | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | | `401 Unauthorized` | Wrong / revoked API key. | Rotate the key, update the secret store. | | `429 rate_limit_exceeded` | Tier budget too low or burst exceeds concurrency. | Raise `openai_usage_tier`, reduce `max_concurrency`, or upgrade the OpenAI tier. | | `429 tokens_per_min` errors despite pacing | Spice rate-limits requests, not tokens. | Reduce per-request token budget; throttle via `max_concurrency`. | | Responses API returns `404` | Provider does not implement Responses natively and `responses_api: enabled` is set. | Set `responses_api: disabled` (default) to use the automatic Chat Completions adapter for `/v1/responses`. | | Slow first-token latency | `stream=false` waits for full completion. | Use `stream=true` for interactive chat. | --- # Perplexity Models (Deprecated) Deprecated Perplexity model support is no longer supported and was deprecated in [spiceai/spiceai#9910](https://github.com/spiceai/spiceai/pull/9910). For documentation on Perplexity models in previous versions, see the [v1.11.x Perplexity documentation](https://docs.spiceai.org/docs/1.11.x/components/models/perplexity). --- # Spice Cloud Platform To use a large language model served by the Spice.ai Cloud Platform AI gateway — or by another Spice runtime — specify the `spice.ai` path in the `from` field and the associated `spiceai_api_key` parameter. | Param | Description | Default | | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------- | | `spiceai_api_key` | The API key for the Spice.ai Cloud Platform, or for the Spice runtime serving the model. **Required** when the endpoint is the Spice.ai Cloud Platform. | - | | `spiceai_endpoint` | The endpoint serving the model: the Spice.ai Cloud Platform, or another Spice runtime. | `https://data.spiceai.io` | Example: ``` models: - from: spice.ai:openai/gpt-4o name: cloud_llm params: spiceai_api_key: ${secrets:SPICEAI_API_KEY} ``` ## `from` Format[​](#from-format "Direct link to from-format") The `from` field is the source prefix — `spice.ai` or `spiceai` — followed by a `:` or `/` separator and the model identifier: * `spice.ai:openai/gpt-4o` * `spice.ai/openai/gpt-4o` * `spiceai:openai/gpt-4o` * `spiceai/openai/gpt-4o` All four forms resolve to the same model identifier, `openai/gpt-4o`. Both spellings of the prefix are accepted so that a `from` reads the same whether it names a model or a [dataset](/docs/next/components/data-connectors/spiceai). Model identifiers take the form `/` (for example `openai/gpt-4o` or `anthropic/claude-3-5-sonnet`); the identifiers available depend on what the endpoint serves. A `from` that names only the prefix (e.g. `from: spice.ai` or `from: spice.ai:`) carries no model identifier and fails to load. ## Connecting to Another Spice Runtime[​](#connecting-to-another-spice-runtime "Direct link to Connecting to Another Spice Runtime") Because the model is reached over an OpenAI-compatible HTTP API, `spiceai_endpoint` can point at another Spice runtime instead of the Cloud Platform. The endpoint is the root of the deployment — the OpenAI-compatible API is served under `/v1`, and an endpoint that already names `/v1` is used as-is. ``` models: - from: spice.ai:openai/gpt-4o name: upstream_llm params: spiceai_endpoint: http://localhost:8090 ``` An API key is optional for a self-hosted Spice runtime, which may not have authentication enabled. The Spice.ai Cloud Platform always authenticates, so omitting `spiceai_api_key` when the endpoint resolves to the Cloud Platform is rejected at load time. note The `spice.ai` model source serves large language models only. Support for loading and serving traditional machine learning (ONNX) models was removed in vNext — see [Machine Learning Models](/docs/next/features/machine-learning-models). --- # xAI Models To use a language model hosted on xAI, specify `xai` path in the `from` field and the associated `xai_api_key` parameter. When no model is specified in the `from` field (i.e., `from: xai`), the default model is `grok-4.3`. | Param | Description | Default | | ---------------- | --------------------------------------------------- | ------- | | `xai_api_key` | The xAI API key. | - | | `xai_usage_tier` | xAI usage tier (0-4). Used for rate limit defaults. | - | Example: ``` models: - from: xai:grok-4.3 name: xai params: xai_api_key: ${secrets:SPICE_GROK_API_KEY} ``` Refer to the [xAI models documentation](https://docs.x.ai/docs/models) for more details on available models and configurations. note Although the xAI [documentation](https://docs.x.ai/docs/guides/structured-outputs) shows that xAI models can return structured outputs, this is not true. --- # Secret Stores A Secret Store is a secure location where secrets (such as passwords, tokens, and API keys) are stored. Spice retrieves secrets from configured stores at runtime and injects them into component parameters. Supported secret stores include: [`env`](/docs/next/components/secret-stores/env), [`kubernetes`](/docs/next/components/secret-stores/kubernetes), [`keyring`](/docs/next/components/secret-stores/keyring), [`aws_secrets_manager`](/docs/next/components/secret-stores/aws-secrets-manager), [`azure_keyvault`](/docs/next/components/secret-stores/azure-keyvault), and [`hashicorp_vault`](/docs/next/components/secret-stores/hashicorp-vault) (Spice.ai Enterprise). The `env` secret store is loaded by default. ### Default[​](#default "Direct link to Default") The `env` secret store is loaded by default. It reads secrets from environment variables and any `.env.local` or `.env` files in the project directory. ``` secrets: - from: env name: env ``` ## Configured Secret Stores[​](#configured-secret-stores "Direct link to Configured Secret Stores") Secret Stores can be configured using the `secrets` section of the `spicepod.yml` file. The Secret Store type and name are specified using the `from` and `name` fields. The `name` can be referenced by other components, like datasets or models. Some Secret Stores support adding a selector delimited by a colon (`:`), For example, when using the Kubernetes Secret Store, `from: kubernetes:my_secret` selects and enables the `my_secret` secret only to be referenced. Additional parameters may be specified in the `params` field, which are specific to the secret store type. Unknown parameters are rejected with an error listing the supported parameter names, helping catch typos early. Parameter values support `${ env:KEY }` references to load values from environment variables at startup. Example: ``` secrets: - from: kubernetes:my_secret name: k8s - from: env name: env ``` ## Using referenced secrets in component parameters[​](#using-secrets "Direct link to Using referenced secrets in component parameters") Secrets may be used by components with the syntax `${:}`. For example, to reference a secret stored as an environment variable named `MY_SECRET` in the `env` secret store, use `${env:MY_SECRET}`. Example: ``` datasets: - from: postgres:my_table name: my_table params: pg_host: localhost pg_port: 5432 pg_user: ${env:PG_USER} pg_pass: ${env:FOO_PASSWORD} # The environment variable name may differ from the parameter name. ``` This syntax also works within a larger string, like a connection string: ``` datasets: - from: mysql:my_table name: my_table params: mysql_connection_string: mysql://${env:USER}:${env:PASSWORD}@localhost:3306/mysql_db ``` The `` value in `${:}` is the `name` value defined in the secret store configuration. This can be renamed to any value. Example: ``` secrets: - from: env name: my_env ``` ``` datasets: - from: postgres:my_table name: my_table params: pg_host: localhost pg_port: 5432 pg_user: ${my_env:PG_USER} pg_pass: ${my_env:PG_PASS} ``` ## Load secrets from multiple secret stores[​](#load-secrets-from-multiple-secret-stores "Direct link to Load secrets from multiple secret stores") Spice supports configuring multiple secret stores which are loaded in the order they are defined in the `secrets` section of the `spicepod.yml` configuration file. If a secret is defined in multiple secret stores, the secret store defined last will take precedence. To load a secret from any of the configured secret stores in precedence order, use the `${secrets:}` syntax. Example: ``` secrets: - from: env name: env - from: keyring name: keyring ``` ``` datasets: - from: postgres:my_table name: my_table params: pg_host: localhost pg_port: 5432 pg_user: ${secrets:pg_user} pg_pass: ${secrets:pg_pass} ``` In this example, the runtime would look for `pg_user` and `pg_pass` in the `keyring` secret store first and then in the `env` secret store. The `` value in `${secrets:}` is automatically uppercased for the `env` secret store. To override the `keyring` secret store secrets with environment variables, re-order the secret stores in the configuration file: ``` secrets: - from: keyring name: keyring - from: env name: env ``` ## Secret Stores[​](#secret-stores "Direct link to Secret Stores") ## [📄️Environment Secret Store](/docs/next/components/secret-stores/env) [Environment Variables Secret Store Documentation](/docs/next/components/secret-stores/env) ## [📄️AWS Secrets Manager Secret Store](/docs/next/components/secret-stores/aws-secrets-manager) [AWS Secrets Manager Secret Store Documentation](/docs/next/components/secret-stores/aws-secrets-manager) ## [📄️Kubernetes Secret Store](/docs/next/components/secret-stores/kubernetes) [Kubernetes Secret Store Documentation](/docs/next/components/secret-stores/kubernetes) ## [📄️Keyring Secret Store](/docs/next/components/secret-stores/keyring) [Keyring Secret Store Documentation](/docs/next/components/secret-stores/keyring) ## [📄️Azure Key Vault Secret Store](/docs/next/components/secret-stores/azure-keyvault) [Azure Key Vault Secret Store Documentation](/docs/next/components/secret-stores/azure-keyvault) ## [📄️HashiCorp Vault Secret Store](/docs/next/components/secret-stores/hashicorp-vault) [HashiCorp Vault Secret Store Documentation](/docs/next/components/secret-stores/hashicorp-vault) --- # AWS Secrets Manager Secret Store The `aws_secrets_manager` store enables Spice to read secrets from [AWS Secrets Manager](https://aws.amazon.com/secrets-manager/) by specifying the secret’s name with a selector. ``` secrets: from: aws_secrets_manager:my_secret_name name: aws ``` The store reads keys from the secret named in the selector. In the above example `my_secret_name` must be defined in [AWS Secrets Manager](https://console.aws.amazon.com/secretsmanager/listsecrets), and any keys referenced using `${aws:my_key}` will look for a key `my_key` within `my_secret_name`. ![](/img/secrets-aws-secrets-manager-1.png) ![](/img/secrets-aws-secrets-manager-2.png) ## Parameters[​](#parameters "Direct link to Parameters") | Parameter Name | Description | | --------------- | ---------------------------------------------------------------------------------------------------------------- | | `region` | Optional. The AWS region for the Secrets Manager API. Falls back to the SDK default credential chain if not set. | | `endpoint_url` | Optional. Custom endpoint URL for the Secrets Manager API (e.g., for VPC endpoints, FIPS, or LocalStack). | | `key` | Optional. AWS access key ID. Must be set together with `secret`. Overrides the default credential chain. | | `secret` | Optional. AWS secret access key. Must be set together with `key`. Overrides the default credential chain. | | `session_token` | Optional. AWS session token for temporary (STS-issued) credentials. Only used when `key` and `secret` are set. | Parameter values support `${ env:KEY }` references to load values from environment variables. ``` secrets: - from: aws_secrets_manager:my_secret_name name: aws params: region: ${ env:AWS_REGION } key: ${ env:AWS_ACCESS_KEY_ID } secret: ${ env:AWS_SECRET_ACCESS_KEY } ``` note Unknown parameters are rejected with an error listing the supported parameter names. This helps catch typos — e.g., `regoin` instead of `region` will produce an immediate error instead of being silently ignored. ## Example[​](#example "Direct link to Example") A complete spicepod definition with a dataset that uses a secret from AWS Secrets Manager. ``` version: v1 kind: Spicepod name: taxi_trips secrets: - from: aws_secrets_manager:dremio name: dremio datasets: - from: dremio:datasets.taxi_trips name: taxi_trips description: dremio taxi trips params: dremio_endpoint: grpc://20.163.171.81:32010 dremio_username: ${dremio:username} dremio_password: ${dremio:password} ``` ### Authentication[​](#authentication "Direct link to Authentication") Spice will automatically load credentials to connect to AWS Secrets Manager from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate secretsmanager permissions (e.g., `secretsmanager:GetSecretValue`) to access the secrets. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access secrets in Secrets Manager: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "secretsmanager:GetSecretValue" ], "Resource": [ "arn:aws:secretsmanager:us-east-1:123456789012:secret:TestEnv/*" ] } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------------------- | --------------------------------------- | | `secretsmanager:GetSecretValue` | Required. Allows reading secret values. | --- # Azure Key Vault Secret Store The `azure_keyvault` store enables Spice to read secrets from [Azure Key Vault](https://azure.microsoft.com/products/key-vault/) by specifying the vault name with a selector. ``` secrets: - from: azure_keyvault:my-vault name: azure ``` The selector is the Key Vault name. Any keys referenced using `${azure:my_key}` resolve to a secret named `spice-my-key` in `my-vault`, falling back to `my-key` if the prefixed form does not exist. Logical key names use underscores; the store automatically translates them to the hyphen-delimited names that Key Vault requires (`openai_api_key` → `spice-openai-api-key` → `openai-api-key`). ## Parameters[​](#parameters "Direct link to Parameters") | Parameter Name | Description | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `auth_method` | Optional. One of `service_principal`, `managed_identity`, `workload_identity`, `cli`, or `default`. Defaults to `default`: uses `service_principal` if `client_secret` is set, otherwise the Azure CLI / `azd` developer credential. Managed and workload identity are not selected implicitly — set `auth_method` explicitly when running on Azure infrastructure. | | `tenant_id` | Required for `auth_method: service_principal`. Optional for `workload_identity` — when omitted, read from the `AZURE_TENANT_ID` env var injected by the AKS workload-identity webhook. Ignored for `managed_identity`, `cli`, and `default` (developer-tools fallback). The Microsoft Entra tenant ID. | | `client_id` | Required for `auth_method: service_principal`. Optional for `workload_identity` — when omitted, read from the `AZURE_CLIENT_ID` env var injected by the AKS workload-identity webhook. Optional for `managed_identity` to select a user-assigned identity (system-assigned is used when omitted). Ignored for `cli` and `default` (developer-tools fallback). | | `client_secret` | Required for `auth_method: service_principal`. The application secret. | | `endpoint` | Optional. Sovereign cloud override. Accepts either a full URL (`https://my-vault.vault.azure.cn`) or a bare DNS suffix (`vault.azure.cn`). Defaults to the Azure public cloud (`vault.azure.net`). | Parameter values support `${ env:KEY }` references to load values from environment variables. ``` secrets: - from: azure_keyvault:my-vault name: azure params: auth_method: service_principal tenant_id: ${ env:AZURE_TENANT_ID } client_id: ${ env:AZURE_CLIENT_ID } client_secret: ${ env:AZURE_CLIENT_SECRET } ``` note Unknown parameters are rejected with an error listing the supported parameter names. This helps catch typos — e.g., `tenent_id` instead of `tenant_id` will produce an immediate error instead of being silently ignored. ## Example[​](#example "Direct link to Example") A complete Spicepod definition with a dataset that uses a secret from Azure Key Vault. ``` version: v1 kind: Spicepod name: taxi_trips secrets: - from: azure_keyvault:prod-vault name: azure datasets: - from: postgres:public.taxi_trips name: taxi_trips params: pg_host: postgres.example.com pg_user: ${azure:postgres_user} pg_pass: ${azure:postgres_password} ``` With the example above, Spice resolves `${azure:postgres_user}` by reading `spice-postgres-user` from the `prod-vault` Key Vault, falling back to `postgres-user`. ## Authentication[​](#authentication "Direct link to Authentication") The `auth_method` parameter selects the credential source: * **`default`** (the default) — uses an explicit `service_principal` when `client_secret` is set, otherwise falls back to local developer credentials (Azure CLI / `azd`, via the SDK's `DeveloperToolsCredential`). Set `auth_method` explicitly when running on Azure infrastructure. Suitable for local development and for service-principal credentials wired through params. * **`service_principal`** — uses an explicit `tenant_id`, `client_id`, and `client_secret`. Use when Spice must authenticate as a specific Microsoft Entra application. * **`managed_identity`** — uses the [system-assigned managed identity](https://learn.microsoft.com/entra/identity/managed-identities-azure-resources/overview) of the Azure VM, AKS node, Container App, or ACI hosting Spice. Pass `client_id` to select a user-assigned identity. * **`workload_identity`** — uses [federated tokens](https://learn.microsoft.com/azure/aks/workload-identity-overview) injected by AKS or another Kubernetes cluster configured with workload identity. `client_id` and `tenant_id` are typically supplied by the workload-identity admission webhook through the `AZURE_CLIENT_ID`, `AZURE_TENANT_ID`, and `AZURE_FEDERATED_TOKEN_FILE` environment variables. * **`cli`** — uses cached credentials from a local `az login` session. Common during development. For sovereign clouds (Azure China, Azure Government), set `endpoint` to the regional DNS suffix (for example `vault.azure.cn`). ## Required Role Assignments[​](#required-role-assignments "Direct link to Required Role Assignments") The principal used by Spice must have read access to secrets in the target Key Vault. The recommended built-in role is **Key Vault Secrets User** (`4633458b-17de-408a-b874-0445c86b69e6`): ``` az role assignment create \ --assignee-object-id \ --assignee-principal-type ServicePrincipal \ --role "Key Vault Secrets User" \ --scope /subscriptions//resourceGroups//providers/Microsoft.KeyVault/vaults/ ``` The store performs a `list_secret_properties` call on startup to validate connectivity and fail fast on misconfiguration. A `403 Forbidden` on list is tolerated, so principals scoped only to per-secret `Microsoft.KeyVault/vaults/secrets/getSecret/action` (for example, **Key Vault Reader** combined with a custom role) are supported. | Permission | Purpose | | ------------------------------------------------------- | -------------------------------------------------- | | `Microsoft.KeyVault/vaults/secrets/getSecret/action` | Required. Read individual secret values. | | `Microsoft.KeyVault/vaults/secrets/readMetadata/action` | Optional. Powers the startup validation list call. | For Key Vaults configured with the legacy [access policy permission model](https://learn.microsoft.com/azure/key-vault/general/assign-access-policy), grant the `Get` and (optionally) `List` permissions on **Secrets** instead of an RBAC role assignment. ## Caching[​](#caching "Direct link to Caching") The store maintains a per-key in-process cache with a 60-second positive TTL and a 10-second negative TTL, with single-flight coalescing so concurrent lookups for the same key issue at most one network round-trip. To force a refresh, restart the runtime. --- # Environment Secret Store The `env` store type enables Spice to read secrets from environment variables and any `.env.local` or `.env` files in the project directory. This is the default secret store and is loaded automatically as: ``` secrets: - from: env name: env ``` Reference secrets directly in parameters using the syntax `${env:MY_ENV_VAR}`. This will load the value of the environment variable `MY_ENV_VAR` into the parameter. Example: ``` datasets: - from: postgres:my_table name: my_table params: pg_host: localhost pg_port: 5432 pg_user: ${env:MY_PG_USER} pg_pass: ${env:MY_PG_PASSWORD} ``` The `${}` replacement syntax also works within a larger string, like a connection string: ``` datasets: - from: mysql:my_table name: my_table params: connection_string: mysql://${env:MY_USER}:${env:MY_PASSWORD}@localhost:3306/my_db ``` When used with the `${secrets:}` syntax, the `` variable is UPPERCASED to follow the convention of environment variables. Example: ``` datasets: - from: postgres:my_table name: my_table params: pg_host: localhost pg_port: 5432 pg_user: ${secrets:my_pg_user} # same as ${env:MY_PG_USER} pg_pass: ${secrets:my_pg_password} # same as ${env:MY_PG_PASSWORD} ``` ## .env Files[​](#env-files "Direct link to .env Files") The `env` secret store reads secrets from any `.env.local` or `.env` files in the project directory. The `.env.local` file takes precedence over the `.env` file. This enables defining template secrets in the `.env` file which can be checked into source control and overriding them with local secrets in the `.env.local` file. Example `.env` file: ``` MY_PG_USER=postgres MY_PG_PASSWORD=postgres ``` ### Parameters[​](#parameters "Direct link to Parameters") | Parameter Name | Description | | -------------- | ---------------------------------------------------------------------------------------------------------------------------- | | `file_path` | Optional. Path to a specific `.env` file to load. When set, `.env` and `.env.local` in the project directory are not loaded. | To load environment variables from a specific `.env` file, use the `file_path` parameter. When a `file_path` parameter is specified, environment variables from `.env` or `.env.local` will not be loaded. ``` secrets: - from: env name: env params: file_path: ./custom/path/to/.env ``` To continue loading `.env` or `.env.local`, specify them as additional secret stores: ``` secrets: - from: env name: env - from: env name: env params: file_path: ./custom/path/to/.env ``` --- # HashiCorp Vault Secret Store The `hashicorp_vault` store enables Spice to read secrets from a [HashiCorp Vault](https://developer.hashicorp.com/vault) KV secrets engine (v1 or v2). The selector is the path *under the mount* — Spice automatically inserts the `data/` segment for KV v2. The HashiCorp Vault Secret Store is available in the Spice [Enterprise edition](https://docs.spice.ai/docs/enterprise/getting-started/distributions). ``` secrets: - from: hashicorp_vault:myapp/config name: vault params: hashicorp_vault_address: https://vault.example.com:8200 hashicorp_vault_token: ${ env:VAULT_TOKEN } ``` With the default mount (`secret`) and KV v2, the example above reads `/v1/secret/data/myapp/config`. The Vault response body's `data` field is treated as a `key → string` map; each `${vault:my_key}` reference resolves to the corresponding entry. ## Parameters[​](#parameters "Direct link to Parameters") | Parameter Name | Description | | --------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `hashicorp_vault_address` | **Required.** Vault server URL, e.g. `https://vault.example.com:8200`. Plaintext `http://` is allowed only for `localhost` / `127.0.0.1` so local dev with `vault server -dev` works without TLS. | | `hashicorp_vault_namespace` | Optional. Vault Enterprise namespace, sent as the `X-Vault-Namespace` header. Leave unset for OSS Vault. | | `hashicorp_vault_mount` | Optional. Mount path of the KV secrets engine. Defaults to `secret`, matching `vault server -dev`. | | `hashicorp_vault_kv_version` | Optional. `v1` or `v2`. Defaults to `v2`. v2 supports versioning and rolls the read URL through `/data/`; v1 is the legacy layout. | | `hashicorp_vault_auth_method` | Optional. One of `token`, `approle`, `kubernetes`, or `jwt`. Defaults to `token`. | | `hashicorp_vault_token` | Required for `auth_method: token`. Vault client token. Typically `${ env:VAULT_TOKEN }`. | | `hashicorp_vault_role_id` | Required for `auth_method: approle`. AppRole `role_id`. | | `hashicorp_vault_secret_id` | Required for `auth_method: approle`. AppRole `secret_id`. Typically `${ env:VAULT_SECRET_ID }`. | | `hashicorp_vault_role` | Required for `auth_method: kubernetes` and `auth_method: jwt`. Vault role name. | | `hashicorp_vault_jwt` | JWT/OIDC token presented for `auth_method: jwt`, or the Kubernetes service-account JWT for `auth_method: kubernetes` when not reading it from disk. Typically sourced from env. | | `hashicorp_vault_kubernetes_token_path` | Optional. Filesystem path to the service-account JWT for `auth_method: kubernetes`. Defaults to `/var/run/secrets/kubernetes.io/serviceaccount/token`. Ignored when `hashicorp_vault_jwt` is set. | | `hashicorp_vault_auth_mount` | Optional. Mount path of the auth backend, *without* the leading `auth/` segment. Defaults to the auth method name (`approle`, `kubernetes`, `jwt`). Override when the backend has been mounted at a non-default path (e.g. `k8s-prod`). | | `hashicorp_vault_ca_cert` | Optional. Filesystem path to a PEM-encoded CA certificate to add to the TLS trust store. Use for self-signed Vault deployments. | | `hashicorp_vault_tls_skip_verify` | Optional. `true` / `false` (default `false`). Skip TLS certificate verification. Strongly discouraged outside local development. | | `hashicorp_vault_request_timeout` | Optional. Per-request timeout in seconds for Vault HTTP calls. Defaults to `10`. | Parameter values support `${ env:KEY }` references to load values from environment variables. note Unknown parameters are rejected with an error listing the supported parameter names. This helps catch typos — e.g., `hashicorp_vault_addr` instead of `hashicorp_vault_address` will produce an immediate error instead of being silently ignored. ## Authentication[​](#authentication "Direct link to Authentication") The `hashicorp_vault_auth_method` parameter selects the credential source: * **`token`** (the default) — uses the supplied `hashicorp_vault_token` directly. Suitable for local development (`vault server -dev` prints a root token at startup) and for short-lived tokens minted by an external orchestrator. * **`approle`** — performs a [`POST /v1/auth/approle/login`](https://developer.hashicorp.com/vault/docs/auth/approle) using `hashicorp_vault_role_id` and `hashicorp_vault_secret_id`. Production-shaped: long-lived `role_id`, short-lived `secret_id`. * **`kubernetes`** — performs a [`POST /v1/auth/kubernetes/login`](https://developer.hashicorp.com/vault/docs/auth/kubernetes) using `hashicorp_vault_role` and the in-cluster service-account JWT (read from `hashicorp_vault_kubernetes_token_path`, or supplied directly via `hashicorp_vault_jwt`). * **`jwt`** — performs a [`POST /v1/auth/jwt/login`](https://developer.hashicorp.com/vault/docs/auth/jwt) using `hashicorp_vault_role` and `hashicorp_vault_jwt`. Use for OIDC-issued tokens or any other JWT-shaped credential. For `approle`, `kubernetes`, and `jwt`, the returned client token is cached in-process and reused until its lease expires. On a `403 Forbidden` from the data read, the store re-authenticates once and retries transparently. There is no background renewal task. For self-signed Vault deployments, supply the CA bundle path via `hashicorp_vault_ca_cert` (preferred) or, for ad-hoc testing only, set `hashicorp_vault_tls_skip_verify: true`. ## Example[​](#example "Direct link to Example") A complete Spicepod definition with a dataset that uses a secret from Vault, authenticated via AppRole: ``` version: v1 kind: Spicepod name: taxi_trips secrets: - from: hashicorp_vault:databases/postgres/taxi name: vault params: hashicorp_vault_address: https://vault.example.com:8200 hashicorp_vault_auth_method: approle hashicorp_vault_role_id: ${ env:VAULT_ROLE_ID } hashicorp_vault_secret_id: ${ env:VAULT_SECRET_ID } datasets: - from: postgres:public.taxi_trips name: taxi_trips params: pg_host: postgres.example.com pg_user: ${vault:username} pg_pass: ${vault:password} ``` With the example above, Spice reads `/v1/secret/data/databases/postgres/taxi` from Vault and resolves `${vault:username}` and `${vault:password}` against the returned `data` map. ## Required Vault Policy[​](#required-vault-policy "Direct link to Required Vault Policy") The token / role used by Spice needs `read` capability on the configured KV path. For the example above: ``` path "secret/data/databases/postgres/taxi" { capabilities = ["read"] } ``` For KV v1, drop the `data/` segment. ## Caching[​](#caching "Direct link to Caching") The store maintains a per-path in-process cache. The positive TTL honors the `lease_duration` returned by Vault, falling back to 60 seconds when none is provided (typical for KV, which has no lease itself). Confirmed-missing paths (`404`) are negatively cached for 10 seconds. Concurrent lookups for the same path are coalesced behind a single async lock so only one `GET` is in flight per store. To force a refresh, restart the runtime. --- # Keyring Secret Store The `keyring` store enables Spice to access secrets from the secure/credential store of the host operating system: * Linux: The secret-service and kernel keyutils. * macOS: The keychain. * Windows: The Credential Manager. The Keyring Store will read entries where the entry account or user is set to `spiced`. ## Example[​](#example "Direct link to Example") To set the `spiceai` API Key secret using macOS keychain, create a new keychain entry, and set the value with the API Key. `""` ![](/img/secrets-keychain-example.png) The `keyring` store is configured in the Spicepod manifest: ``` secrets: - from: keyring name: keyring ``` And the secret can be referenced in parameters: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips params: spiceai_api_key: ${keyring:spiceai_api_key} # ${secrets:spiceai_api_key} can also be used ``` --- # Kubernetes Secret Store The `kubernetes` store supports reading specific secrets using a selector with the secret's name. ## Example[​](#example "Direct link to Example") ``` secrets: - from: kubernetes:my_secret name: k8s ``` And the secret can be referenced in parameters: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips params: spiceai_api_key: ${k8s:spiceai_api_key} # ${secrets:spiceai_api_key} can also be used to fallback to other secret stores ``` Load secrets from multiple Kubernetes secrets by defining multiple Kubernetes secret stores with the appropriate selectors for the secrets to read: ``` secrets: - from: kubernetes:my_secret name: k8s - from: kubernetes:my_other_secret name: k8s_other ``` ## Parameters[​](#parameters "Direct link to Parameters") | Parameter Name | Description | | -------------- | --------------------------------------------------------------------------------------------------------------------------- | | `namespace` | Optional. The Kubernetes namespace to read the secret from. Defaults to the namespace of the running pod's service account. | ``` secrets: - from: kubernetes:my_secret name: k8s params: namespace: spice ``` note Unknown parameters are rejected with an error listing the supported parameter names. ## Kubernetes Secret Store Configuration[​](#kubernetes-secret-store-configuration "Direct link to Kubernetes Secret Store Configuration") Note: This method requires the Kubernetes service account, which is running the `spiced` pod, to have extended roles for secrets API access. Configure this service account with the necessary permissions to read secrets from the Kubernetes API. Example of Kubernetes role configuration for a custom service account: ``` kind: Role apiVersion: rbac.authorization.k8s.io/v1 metadata: name: spiced-account-role rules: - apiGroups: [''] resources: ['secrets'] verbs: ['get'] ``` --- # LLM Tools (Function Calling) A tool is a function or operation that can be called directly or by a [language model](/docs/next/features/large-language-models) (LLMs). The Spice runtime has several tools available by default, giving LLMs access to various parts of the runtime. Tools can also be added or configured by the user by declaring them in the `tools` section of `spicepod.yaml`. For details about providing LLMs tool access, see [Language Model Tools](/docs/next/features/large-language-models/tools). For details on tool specifications, see the [Tools Spicepod Reference](/docs/next/reference/spicepod/tools). ### Available Tools[​](#available-tools "Direct link to Available Tools") | Name | Description | Default Group | | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------- | | `list_datasets` | List all available datasets in the runtime. | `auto` | | `sql` | Execute SQL queries on the runtime. Write statements (`INSERT`/`UPDATE`/`DELETE`/DDL) are accepted only when the request is authenticated with a ReadWrite API key and the target dataset is configured `access: read_write`; otherwise the tool runs as read-only. | `auto` | | `table_schema` | Get the schema of a specific SQL table. | `auto` | | `search` | Searches a configured dataset based on an input query. | `auto` | | `get_readiness` | Report the readiness state of every runtime component (datasets, accelerators, models, embeddings, catalogs). | `auto` | | `get_current_datetime` | Return the current UTC date and time as an ISO 8601 timestamp. | `auto` | | `sample_distinct_columns` | Generate a synthetic sample of data with distinct values. | `all` | | `random_sample` | Sample random rows from a table. | `all` | | `top_n_sample` | Sample the top N rows from a table based on a specified ordering. | `all` | | `memory:load` | Retrieve all stored memories from the last time period. | `memory` | | `memory:store` | Store information from LLM interaction(s) for future reference. | `memory` | | ~~[`websearch`](/docs/next/components/tools/websearch)~~ | ~~Search the web for information.~~ (Deprecated) | - | The `auto` group is a subset of `all`: the sampling tools (`sample_distinct_columns`, `random_sample`, `top_n_sample`) are provided only by the `all` (and `nsql`) groups, or by naming them individually. ### Tool Groups[​](#tool-groups "Direct link to Tool Groups") Tool groups are predefined sets of tools that can be provided to LLMs in a single tool name. For example, the `auto` tool group provides all default tools to the LLM (see above table). ``` models: - name: full-runtime from: openai:gpt-4o params: tools: auto # Automatically choose direct or registry-based discovery ``` #### Available tool groups[​](#available-tool-groups "Direct link to Available tool groups") | Name | Description | | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | | `auto` | Automatically choose between direct tools and searchable registry discovery based on the number of tools and embedding model availability. | | `all` | All built-in and Spicepod-configured tools, provided directly to the LLM. | | `search_registry` | Use searchable tool registry discovery via `tool_search` and `tool_invoke` meta-tools. Requires an embedding model (see `tool_embedding_model`). | | `memory` | Memory tools for storing and retrieving information across conversations. | | `nsql` | Built-in tools relevant to text-to-SQL: `table_schema`, `sql`, `list_datasets`, `get_current_datetime`, and the sampling tools. | | `disabled` | Provide no tools to the LLM. | | [`MCP`](/docs/next/components/tools/mcp) | Tools provided from an MCP server. Can be run within Spice, or connected to over HTTP(s) SSE | For details on `auto`, `all`, and `search_registry` tool modes, see [Language Model Tools](/docs/next/features/large-language-models/tools#tool-modes). --- # Model Context Protocol Tools Spice integrates with tools and services using the [Model Context Protocol](https://modelcontextprotocol.io/) (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the [Streamable HTTP](https://modelcontextprotocol.io/specification/2025-03-26/basic/transports#streamable-http) transport. ## Overview[​](#overview "Direct link to Overview") MCP helps extend the capabilities of the Spice runtime by enabling integration with external tools and services. This includes: 1. Running stdio-based MCP servers internally. 2. Connecting to external MCP servers over Streamable HTTP. ## Configuring MCP Tools[​](#configuring-mcp-tools "Direct link to Configuring MCP Tools") To configure MCP tools, define them in the `tools` section of your `spicepod.yaml` file. The `from` field specifies the transport mechanism, such as `mcp:npx` for stdio-based tools or an HTTP URL for Streamable HTTP-based tools. ### Example: Adding an MCP Tool (Stdio)[​](#example-adding-an-mcp-tool-stdio "Direct link to Example: Adding an MCP Tool (Stdio)") The following example demonstrates how to configure an MCP tool using `npx` to run a Google Maps MCP server: ``` tools: - name: google_maps from: mcp:npx params: mcp_args: -y @modelcontextprotocol/server-google-maps ``` ### Example: Connecting to an External MCP Server (Streamable HTTP)[​](#example-connecting-to-an-external-mcp-server-streamable-http "Direct link to Example: Connecting to an External MCP Server (Streamable HTTP)") This example shows how to connect to an external MCP server over Streamable HTTP: ``` tools: - name: external_mcp_server from: mcp:http://example.com/v1/mcp ``` ### Example: Connecting to an Auth-Enabled MCP Server (Streamable HTTP)[​](#example-connecting-to-an-auth-enabled-mcp-server-streamable-http "Direct link to Example: Connecting to an Auth-Enabled MCP Server (Streamable HTTP)") Streamable HTTP MCP tools support sending an `Authorization: Bearer` token via `mcp_auth_token`, or arbitrary HTTP headers via `mcp_headers`. Both parameters resolve [secret references](/docs/next/components/secret-stores) before the MCP client is constructed. Sending a bearer token: ``` tools: - name: remote_spice from: mcp:https://my-spice.example.com/v1/mcp params: # Sends: Authorization: Bearer mcp_auth_token: ${ secrets:MCP_SERVER_API_KEY } ``` Sending a custom header (e.g., API key): ``` tools: - name: remote_spice from: mcp:https://my-spice.example.com/v1/mcp params: # Sends: X-API-Key: mcp_headers: 'X-API-Key: ${ secrets:MCP_SERVER_API_KEY }' ``` If both `mcp_auth_token` and a custom `Authorization` header in `mcp_headers` are set, `mcp_auth_token` wins and a warning is logged. ## Using MCP Tools with Models[​](#using-mcp-tools-with-models "Direct link to Using MCP Tools with Models") Once configured, MCP tools can be assigned to models via the `tools` parameter. For example: ``` models: - name: model_with_mcp from: openai:gpt-4o params: tools: google_maps ``` ## Spice as an MCP Server[​](#spice-as-an-mcp-server "Direct link to Spice as an MCP Server") Spice can also act as an MCP server, exposing its tools over Streamable HTTP. This enables other Spice instances or external systems to connect and use the tools. ### Example: Connecting to Another Spice Instance via MCP[​](#example-connecting-to-another-spice-instance-via-mcp "Direct link to Example: Connecting to Another Spice Instance via MCP") ``` tools: - name: spice_instance from: mcp:http://localhost:8090/v1/mcp ``` ### Allowed Hosts[​](#allowed-hosts "Direct link to Allowed Hosts") By default the `/v1/mcp` endpoint only accepts requests with a `Host` header matching `localhost`, `127.0.0.1`, or `::1` to prevent DNS rebinding attacks. To allow additional hosts, configure [`runtime.mcp.allowed_hosts`](/docs/next/reference/spicepod/runtime#runtimemcp): ``` runtime: mcp: allowed_hosts: - localhost - my-host.internal:8090 ``` Set `allowed_hosts: ["*"]` to disable host checking entirely. ## Configuration Options[​](#configuration-options "Direct link to Configuration Options") ### `from`[​](#from "Direct link to from") The `from` field specifies the transport mechanism for the MCP tool: * **Streamable HTTP**: Use an HTTP URL pointing to the MCP endpoint (e.g., `http://localhost:8090/v1/mcp`). * **Stdio**: Use commands like `mcp:npx` or `mcp:docker`. Additional arguments can be passed via `params.mcp_args`. ### `params`[​](#params "Direct link to params") The `params` field provides additional configuration for MCP tools. For stdio-based tools, use `mcp_args` to specify command-line arguments: ``` tools: - name: custom_tool from: mcp:npx params: mcp_args: -y @custom/tool ``` For Streamable HTTP tools, the following auth parameters are supported: * `mcp_auth_token` — Sends `Authorization: Bearer ` on every request to the MCP server. * `mcp_headers` — Sends additional HTTP headers using the same `Header: Value` comma- or semicolon-delimited format as the [HTTP data connector's `http_headers`](/docs/next/components/data-connectors/https). Header values are marked sensitive. Both parameters support [secret expansion](/docs/next/components/secret-stores). When `mcp_auth_token` is set, an `Authorization` header in `mcp_headers` is ignored and a warning is logged to avoid duplicate auth headers. ### `env`[​](#env "Direct link to env") For stdio-based MCP tools, environment variables can be set using the `env` field. ``` tools: - name: tool_with_env from: mcp:docker env: API_KEY: your_api_key ``` ### `description`[​](#description "Direct link to description") The `description` field provides a textual description of the tool. This description is passed to any language model that uses the tool. ``` tools: - name: google_maps from: mcp:npx description: Provides geocoding and mapping capabilities. ``` For more details, see the [MCP Tools Reference](/docs/next/reference/spicepod/tools). --- # Web Search Tool (Deprecated) Deprecated The `websearch` tool is no longer supported and was deprecated in [spiceai/spiceai#9910](https://github.com/spiceai/spiceai/pull/9910). The tool was backed by Perplexity, which is no longer supported. For web search functionality, use [OpenAI's hosted web search tool](https://platform.openai.com/docs/guides/tools-web-search?api-mode=responses) via the [OpenAI Responses API](/docs/features/web-search#web-search-through-openai-hosted-tools). For documentation on the websearch tool in previous versions, see the [v1.11.x websearch documentation](https://docs.spiceai.org/docs/1.11.x/components/tools/websearch). --- # Vector Engines > 🎓 Learn how it works with the [Amazon S3 Vectors with Spice](https://spice.ai/blog/getting-started-with-amazon-s3-vectors-and-spice) engineering blog post. Data sourced by Data Connectors, or views built atop them with vector embedding columns can be indexed and efficiently searched using a vector engine. A vector engine will store all vector embeddings associated with columns in a dataset/view, provide efficient vector search operations and avoid unnecessary recomputation of embeddings. A vector engine is configured by setting the `vectors` configuration. E.g. ``` datasets: - name: dataset_with_embeddings vectors: enabled: true ``` For the complete reference specification see [datasets](/docs/next/reference/spicepod/datasets). Supported Vector engines: | Name | Description | | -------------------------------------------------------------- | ----------------------------------- | | [`s3_vectors`](/docs/next/components/vectors/s3_vectors) | AWS S3 vectors | | [`elasticsearch`](/docs/next/components/vectors/elasticsearch) | Elasticsearch (Spice.ai Enterprise) | | [`duckdb`](/docs/next/components/vectors/duckdb) | DuckDB VSS (HNSW) extension | Limitations * A dataset or view must be accelerated (i.e. `datasets[].acceleration.enabled: true`, see [docs](/docs/next/reference/spicepod/datasets#accelerationenabled)) for a vector engine to be provided the appropriate data to ingest. ## Vector Engine Docs[​](#vector-engine-docs "Direct link to Vector Engine Docs") ## [📄️Amazon S3 Vectors](/docs/next/components/vectors/s3_vectors) [Amazon S3 Vectors Engine Documentation](/docs/next/components/vectors/s3_vectors) ## [📄️Elasticsearch](/docs/next/components/vectors/elasticsearch) [Use Elasticsearch as a vector engine in Spice for kNN vector search, full-text search, and hybrid search.](/docs/next/components/vectors/elasticsearch) ## [📄️DuckDB](/docs/next/components/vectors/duckdb) [Use DuckDB as a vector engine in Spice for HNSW-based vector search via the DuckDB VSS extension.](/docs/next/components/vectors/duckdb) --- # DuckDB Vector Engine DuckDB can be used as a vector engine in Spice to store embeddings and execute vector similarity search using HNSW indexes via the [DuckDB VSS](https://duckdb.org/docs/extensions/vss) extension. This is useful when a dataset or view is already accelerated with DuckDB and a fully embedded, single-process vector store is preferred over an external service. The DuckDB vector engine requires the dataset or view to be accelerated with the [DuckDB accelerator](/docs/next/components/data-accelerators/duckdb). Spice computes embeddings on the configured columns during refresh and write, stores them in the DuckDB accelerator alongside the source data, and creates an HNSW index that is used to answer `vector_search` and `/v1/search` queries. ``` datasets: - from: file:products.parquet name: products acceleration: enabled: true engine: duckdb vectors: enabled: true engine: duckdb params: duckdb_distance_metric: cosine duckdb_hnsw_m: '16' duckdb_hnsw_ef_construction: '128' duckdb_hnsw_ef_search: '64' columns: - name: description embeddings: - from: local_embedding_model embeddings: - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: local_embedding_model ``` ### View example[​](#view-example "Direct link to View example") Accelerated views also support DuckDB HNSW vector indexes. Configure `columns[].embeddings` and `vectors` on the view: ``` views: - name: review_title_view sql: select review_date, review_id, product_title, review_body from amazon_reviews columns: - name: product_title embeddings: - from: local_embedding_model acceleration: enabled: true engine: duckdb primary_key: review_id mode: memory vectors: enabled: true engine: duckdb params: duckdb_distance_metric: cosine ``` ``` SELECT product_title FROM vector_search(review_title_view, 'wireless headphones') LIMIT 10; ``` ## Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | Default | | ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ | | `duckdb_distance_metric` | Optional. Vector similarity metric. Accepts `cosine`, `l2` (or `l2_norm` / `euclidean` / `l2sq`), or `inner_product` (or `ip` / `dot` / `dot_product`). | `cosine` | | `duckdb_metric` | Optional. Alias for `duckdb_distance_metric`. `duckdb_distance_metric` takes precedence when both are set. | — | | `duckdb_hnsw_m` | Optional. HNSW graph parameter `m` — the number of bidirectional links per node. Higher values improve recall at the cost of index size and build time. | DuckDB VSS default | | `duckdb_hnsw_ef_construction` | Optional. HNSW build-time parameter — the size of the dynamic candidate list during index construction. Higher values improve recall at the cost of build time. | DuckDB VSS default | | `duckdb_hnsw_ef_search` | Optional. HNSW query-time parameter — the size of the dynamic candidate list during search. Higher values improve recall at the cost of query latency. | DuckDB VSS default | ## Configuring HNSW Indexes via the `embeddings` Syntax[​](#configuring-hnsw-indexes-via-the-embeddings-syntax "Direct link to configuring-hnsw-indexes-via-the-embeddings-syntax") When a dataset is accelerated with DuckDB and has embedding columns configured, the DuckDB vector engine can be enabled implicitly by placing HNSW parameters directly on the DuckDB accelerator's `params`. This avoids the separate `vectors:` block when an HNSW index is the only vector-engine configuration needed. ``` datasets: - from: file:products.parquet name: products acceleration: enabled: true engine: duckdb params: duckdb_distance_metric: cosine duckdb_hnsw_m: '16' duckdb_hnsw_ef_construction: '128' duckdb_hnsw_ef_search: '64' columns: - name: description embeddings: - from: local_embedding_model ``` Spice detects the HNSW parameters on the accelerator config and automatically attaches a DuckDB vector engine to the dataset. The recognized keys are `duckdb_distance_metric` (or `duckdb_metric`), `duckdb_hnsw_m`, `duckdb_hnsw_ef_construction`, and `duckdb_hnsw_ef_search`; any non-vector accelerator parameters are passed through to DuckDB unchanged. The two configurations are equivalent: * **`embeddings` syntax** — HNSW params on `acceleration.params`. Inferred when the dataset has DuckDB acceleration and at least one recognized HNSW parameter. * **`vectors` block** — `vectors.engine: duckdb` with HNSW params on `vectors.params`. Required if the engine name needs to be set explicitly or to disable the vector engine without removing the HNSW parameters. If both are set, the explicit `vectors:` block takes precedence. ## Overview[​](#overview "Direct link to Overview") When configured as a vector engine, Spice: 1. Reads data from the underlying connector (for example, Parquet on disk or a federated SQL source). 2. Computes embeddings on the configured column(s) using the attached embedding model. 3. Writes vectors and source rows to the DuckDB accelerator alongside the rest of the dataset. 4. Maintains a DuckDB VSS HNSW index on the embedding column. For full-refresh datasets the index is rebuilt after each refresh; for append/CDC datasets the index is auto-maintained by DuckDB VSS as rows are inserted. 5. At query time, routes `vector_search` and `/v1/search` against the DuckDB accelerator, computing similarity natively in DuckDB. The DuckDB VSS extension is installed and loaded automatically by the runtime; no manual setup is required. Limitations * The dataset or view must be accelerated with the DuckDB accelerator (`acceleration.engine: duckdb`) for the DuckDB vector engine to be used. * The dataset or view must have a resolvable primary key, either via the underlying schema or an explicit [`row_id`](/docs/next/reference/spicepod/datasets#columnsembeddingsrow_id). * [Chunking](/docs/next/reference/spicepod/datasets#columns-embeddings-chunking) is not yet supported for the DuckDB vector engine. * `partition_by` is not yet supported for the DuckDB vector engine. * `spill_writes` is not supported for the DuckDB vector engine. * DuckDB VSS uses approximate nearest neighbor search and returns probabilistically closest results. ## Configuration[​](#configuration "Direct link to Configuration") ### Embedding Models[​](#embedding-models "Direct link to Embedding Models") Any embedding model supported by Spice can be used to produce the vectors stored in DuckDB, including local models via [Hugging Face](/docs/next/components/embeddings/huggingface) and hosted models via [OpenAI](/docs/next/components/embeddings/openai), [Bedrock](/docs/next/components/embeddings/bedrock), and others. The vector dimension is inferred from the embedding model and used to size the DuckDB embedding column. ``` embeddings: - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: local_embedding_model ``` ### Primary Keys[​](#primary-keys "Direct link to Primary Keys") Spice requires a primary key to round-trip matches between the HNSW index and the base dataset. If the source dataset does not carry primary key metadata, specify it on the column embedding: ``` columns: - name: description embeddings: - from: local_embedding_model row_id: product_id ``` ### Distance Metric[​](#distance-metric "Direct link to Distance Metric") The distance metric controls how similarity is computed between query and stored vectors. Pick the metric that matches how your embedding model is trained: * `cosine` (default) — cosine similarity. Appropriate for most text embedding models. * `l2` — Euclidean (L2) distance. Aliases: `l2_norm`, `euclidean`, `l2sq`. * `inner_product` — dot-product similarity. Aliases: `ip`, `dot`, `dot_product`, `max_inner_product`. ``` vectors: enabled: true engine: duckdb params: duckdb_distance_metric: inner_product ``` ### HNSW Tuning[​](#hnsw-tuning "Direct link to HNSW Tuning") The `duckdb_hnsw_m`, `duckdb_hnsw_ef_construction`, and `duckdb_hnsw_ef_search` parameters control the trade-off between recall, index size, build time, and query latency. When unset, Spice defers to the DuckDB VSS defaults. See the [DuckDB VSS documentation](https://duckdb.org/docs/extensions/vss) for guidance on tuning these values. ## Querying[​](#querying "Direct link to Querying") Vector search uses the standard Spice search surfaces. When the dataset is backed by the DuckDB vector engine, both `vector_search` and `/v1/search` execute natively in DuckDB using the HNSW index. ### Vector Search[​](#vector-search "Direct link to Vector Search") ``` SELECT product_id, name, score FROM vector_search(products, 'wireless noise cancelling headphones') ORDER BY score DESC LIMIT 10; ``` The query text is embedded with the configured embedding model and used as the probe vector for the HNSW index. ### Search HTTP API[​](#search-http-api "Direct link to Search HTTP API") ``` curl -X POST http://localhost:8090/v1/search \ -H 'Content-Type: application/json' \ -d '{ "datasets": ["products"], "text": "wireless noise cancelling headphones", "additional_columns": ["name"], "limit": 10 }' ``` For the full reference, see [Vector Search](/docs/next/features/search/vector-search) and [Search API Reference](/docs/next/api/HTTP/post-search). --- # Elasticsearch Vector Engine Elasticsearch can be used as a vector engine in Spice to store embeddings and execute kNN similarity search, full-text search (BM25), and hybrid search (RRF) natively in the Elasticsearch cluster. This is useful when Elasticsearch is already the system of record for a workload, or when the operational characteristics of a managed Elasticsearch cluster (replication, sharding, snapshots) are preferred over a dedicated vector store. Unlike the [Elasticsearch Data Connector](/docs/next/components/data-connectors/elasticsearch), which reads an existing Elasticsearch index as a Spice dataset, the Elasticsearch vector engine accepts data from any Spice data connector, generates embeddings using the configured embedding model, and writes vectors (and source fields) to an Elasticsearch index that Spice manages. ``` datasets: - from: file:products.parquet name: products acceleration: enabled: true vectors: enabled: true engine: elasticsearch params: elasticsearch_endpoint: https://localhost:9200 elasticsearch_user: ${secrets:es_user} elasticsearch_pass: ${secrets:es_pass} elasticsearch_index: products-embeddings columns: - name: description embeddings: - from: bedrock_titan embeddings: - from: bedrock:amazon.titan-embed-text-v2:0 name: bedrock_titan params: aws_region: us-east-2 dimensions: '1024' ``` Enterprise edition The Elasticsearch vector engine is available in the Spice [Enterprise edition](https://docs.spice.ai/docs/enterprise/getting-started/distributions). ## Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | Example Value | | ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------- | | `elasticsearch_endpoint` | Required. Cluster URL. | `https://localhost:9200` | | `elasticsearch_user` | Optional. Username for HTTP basic authentication. | `${secrets:es_user}` | | `elasticsearch_pass` | Optional. Password for HTTP basic authentication. | `${secrets:es_pass}` | | `elasticsearch_index` | Optional. Index used to store vectors. Defaults to a sanitized `{dataset}-{column}-{model}` value. | `products-embeddings` | | `elasticsearch_vector_field` | Optional. Name of the `dense_vector` field in Elasticsearch. Defaults to `{column}_embedding`. | `description_embedding` | | `elasticsearch_distance_metric` | Optional. Vector similarity metric for kNN search. One of: `cosine`, `l2_norm`, `dot_product`, `max_inner_product`. | `cosine` | | `elasticsearch_hnsw_m` | Optional. HNSW graph parameter `m` (links per node). Higher values improve recall at the cost of more memory. Elasticsearch default: `16`. | `16` | | `elasticsearch_hnsw_ef_construction` | Optional. HNSW build parameter `ef_construction` (candidate list size at build time). Elasticsearch default: `100`. | `100` | | `client_timeout` | Optional. Total request timeout for the Elasticsearch HTTP client, in time unit format. Default: `30s`. | `30s` | | `connect_timeout` | Optional. Connect timeout for the Elasticsearch HTTP client, in time unit format. Default: `10s`. | `10s` | | `elasticsearch_max_retries` | Optional. Maximum retry attempts for transient Elasticsearch errors (HTTP 429 / 5xx). Default: `3`. | `3` | | `elasticsearch_retry_initial_backoff` | Optional. Initial backoff duration between retries, in time unit format. Default: `200ms`. | `200ms` | | `elasticsearch_batch_write_rows` | Optional. Maximum rows per Elasticsearch `_bulk` request. Controls memory usage and payload size during writes. Default: `1000`. | `1000` | | `elasticsearch_index_settings` | Optional. JSON object passed as Elasticsearch index settings when creating the index. Existing indexes are not recreated. | `{"index":{"codec":"best_compression"}}` | | `elasticsearch_number_of_shards` | Optional. ES `number_of_shards` index setting, applied at index creation only. | `1` | | `elasticsearch_number_of_replicas` | Optional. ES `number_of_replicas` index setting, applied at index creation only. | `0` | | `elasticsearch_refresh_interval` | Optional. ES `refresh_interval` index setting, applied at index creation only. | `1s` | | `elasticsearch_bulk_load_refresh_interval` | Optional. Temporary `refresh_interval` during bulk writes, restored afterward. Set to `-1` to disable refresh during loading. | `-1` | | `elasticsearch_force_merge_after_write` | Optional. Run `_forcemerge` after full/append writes. Default: `false`. | `true` | | `elasticsearch_force_merge_segments` | Optional. Max segments for `_forcemerge`. Setting this also enables force merge. Default when force merge enabled: `1`. | `1` | Not yet supported The Elasticsearch vector engine does **not** currently support: * `partition_by` — Partitioned vector indexes. Setting this parameter (or the dataset-level `vectors.partition_by`) returns a configuration error at startup. Use the [S3 Vectors](/docs/next/components/vectors/s3_vectors) engine for partitioned workloads. * `spill_writes` — Spilling writes to disk for backpressure. Setting `spill_writes: true` returns a configuration error at startup. Remove these parameters from the `vectors:` block, or choose a different vector engine. ## Overview[​](#overview "Direct link to Overview") When configured as a vector engine, Spice: 1. Reads data from the underlying connector (for example, Parquet on disk or a federated SQL source). 2. Computes embeddings on the configured column using the attached embedding model. 3. Writes vectors and source fields to the configured Elasticsearch index, provisioning the index mapping when needed (`dense_vector` of the correct dimension plus text fields for full-text search). 4. At query time, routes `vector_search`, `text_search`, and `rrf` against the Elasticsearch index using native kNN and BM25 queries. Source fields on the dataset are indexed as `text` in Elasticsearch so they can be used as full-text search targets. Primary key columns are indexed as `keyword` and included in kNN results so that matches can be joined back to the Spice base table when additional columns are requested. Limitations * A dataset or view must be accelerated (`datasets[].acceleration.enabled: true`) for the vector engine to be provided the appropriate data to ingest. See [`acceleration.enabled`](/docs/next/reference/spicepod/datasets#accelerationenabled). * The dataset must have a resolvable primary key, either via the underlying schema or an explicit [`row_id`](/docs/next/reference/spicepod/datasets#columnsembeddingsrow_id). * Elasticsearch kNN uses approximate nearest neighbors and returns probabilistically closest results. ## Configuration[​](#configuration "Direct link to Configuration") ### Embedding Models[​](#embedding-models "Direct link to Embedding Models") Any embedding model supported by Spice can be used to produce the vectors written to Elasticsearch, including local models via [Hugging Face](/docs/next/components/embeddings/huggingface), hosted models via [OpenAI](/docs/next/components/embeddings/openai), [Bedrock](/docs/next/components/embeddings/bedrock), and others. The vector dimension is inferred from the embedding model and used to provision the Elasticsearch `dense_vector` field. ``` embeddings: - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: local_embedding_model ``` ### Primary Keys[​](#primary-keys "Direct link to Primary Keys") Spice requires a primary key to round-trip matches between Elasticsearch and the base dataset. If the source dataset does not carry primary key metadata, specify it on the column embedding: ``` columns: - name: description embeddings: - from: local_embedding_model row_id: product_id ``` ### Custom Index and Vector Field Names[​](#custom-index-and-vector-field-names "Direct link to Custom Index and Vector Field Names") By default the index name is a sanitized `{dataset}-{column}-{model}` and the vector field is `{column}_embedding`. Override either with `elasticsearch_index` and `elasticsearch_vector_field`: ``` vectors: enabled: true engine: elasticsearch params: elasticsearch_endpoint: https://localhost:9200 elasticsearch_index: products-vectors-v2 elasticsearch_vector_field: desc_vec ``` ## Querying[​](#querying "Direct link to Querying") Vector, full-text, and hybrid search use the standard Spice UDTFs. When the dataset is backed by the Elasticsearch vector engine, these UDTFs compile to native Elasticsearch queries rather than local computation. ### Vector Search[​](#vector-search "Direct link to Vector Search") ``` SELECT product_id, name, score FROM vector_search(products, 'wireless noise cancelling headphones') ORDER BY score DESC LIMIT 10; ``` The query text is embedded with the configured embedding model and sent to Elasticsearch as a kNN query. By default the number of candidates considered by Elasticsearch is twice the requested `k`. ### Full-Text Search[​](#full-text-search "Direct link to Full-Text Search") Any `Utf8`/`LargeUtf8` column on the dataset is available as a full-text search target: ``` SELECT product_id, name, score FROM text_search(products, 'bluetooth waterproof', description) ORDER BY score DESC LIMIT 10; ``` ### Hybrid Search (RRF)[​](#hybrid-search-rrf "Direct link to Hybrid Search (RRF)") Combine vector and full-text results with [Reciprocal Rank Fusion](/docs/next/reference/sql/search#reciprocal-rank-fusion-rrf): ``` SELECT product_id, name, fused_score FROM rrf( vector_search(products, 'wireless noise cancelling headphones'), text_search(products, 'bluetooth waterproof', description), join_key => 'product_id' ) ORDER BY fused_score DESC LIMIT 10; ``` Advanced RRF options — per-query `rank_weight`, recency decay, and custom smoothing `k` — work identically regardless of the underlying vector engine. See [RRF](/docs/next/reference/sql/search#reciprocal-rank-fusion-rrf) for the full reference. ## Authentication[​](#authentication "Direct link to Authentication") When `elasticsearch_user` and `elasticsearch_pass` are provided, the vector engine uses HTTP basic authentication. Prefer storing credentials in a [secret store](/docs/next/components/secret-stores) and referencing them with `${secrets:...}`. TLS is enabled automatically for `https://` endpoints. ## Comparison with the Data Connector[​](#comparison-with-the-data-connector "Direct link to Comparison with the Data Connector") | Use case | Use | | ----------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | | Query an existing Elasticsearch index (with or without `dense_vector`). | [Elasticsearch Data Connector](/docs/next/components/data-connectors/elasticsearch). | | Ingest data from another source and have Spice manage vectors in ES. | Elasticsearch Vector Engine (this page). | Both paths surface `vector_search`, `text_search`, and `rrf`; pick the one that matches which system owns the data. --- # Amazon S3 Vectors Engine > 🎓 Learn how it works with the [Amazon S3 Vectors with Spice](https://spice.ai/blog/getting-started-with-amazon-s3-vectors-and-spice) engineering blog post. Amazon S3 Vectors, announced in public preview at AWS Summit New York 2025, is a new S3 bucket type designed for storing and querying vector embeddings at scale. It supports billions of vectors with sub-second similarity queries, reducing costs by up to 90% compared to traditional vector databases by separating storage from compute. Spice AI integrates S3 Vectors as a vector index backend, managing embedding indexing, lifecycle, and queries for hybrid search experiences. To use Amazon S3 Vectors as a Vector Engine, specify `s3_vectors` as the `engine`, and configure the associated location and AWS credentials. ``` datasets: - from: spice.ai:dataset.with.embeddings name: my_dataset vectors: enabled: true engine: s3_vectors params: s3_vectors_bucket: my-s3-vector-bucket columns: - name: 'body' embeddings: - from: bedrock_titan embeddings: - name: bedrock_titan # ... Define an embedding model to use. ``` ## Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | Example Value | | ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- | | `s3_vectors_arn` | The S3 vectors index to use. Incompatible with `s3_vectors_bucket` and `s3_vectors_index`. | `arn:aws:s3vectors:us-east-1:123456789012:bucket/a-bucket/index/index-of-important-embeddings` | | `s3_vectors_aws_access_key_id` | Optional. The access key ID for the S3 vectors index. If not specified, credentials will be loaded from the environment. | - | | `s3_vectors_aws_iam_role_source` | Optional. IAM role credential source: `auto` (default AWS credential chain), `metadata` (instance/container metadata only), or `env` (environment variables only). Ignored when an access key and secret are supplied. Default `auto`. | `metadata` | | `s3_vectors_aws_region` | The AWS region for the S3 vectors index. | `us-east-1` | | `s3_vectors_aws_secret_access_key` | Optional. The secret access key for the S3 vectors index. If not specified, credentials will be loaded from the environment. | - | | `s3_vectors_aws_session_token` | Optional. Session token for the S3 vectors index. | - | | `s3_vectors_batch_write_rows` | Optional. Number of rows each record batch is chunked into when writing vectors, to control memory usage during writes. Default `100000`. | `100000` | | `s3_vectors_bucket` | The S3 vectors bucket to use. If `s3_vectors_index` is not specified, an index will be created based on the underlying embedding column. Incompatible with `s3_vectors_arn`. Provided to AWS verbatim — see [Index Naming](#bucket-and-index-naming). | `a-bucket` | | `s3_vectors_index` | The name of the s3 vectors index to use or create. Incompatible with `s3_vectors_arn`. Provided to AWS verbatim — see [Index Naming](#bucket-and-index-naming). | `index-of-important-embeddings` | | `s3_vectors_distance_metric` | The distance metric to be used for similarity search. One of: `euclidean`, `cosine`. Default `cosine`. | `euclidean` | | `s3_vectors_index_poll_interval` | The interval to poll for index updates to avoid excessive API calls. Minimum 5 seconds. Default is to poll on every scan. | `5m` | | `s3_vectors_spill_writes` | Optional. When `true`, writes that exceed S3 Vectors rate limits spill to a separate physical index, which is also queried at read time. Ignored, with a warning, when `partition_by` is set. Default `false`. | `true` | | `client_timeout` | Timeout for S3 operations. Default: `30s`. | `30s`, `9 century`, `1m` | Limitations * `s3_vectors_index` and `s3_vectors_arn` specify a single index for the dataset and therefore should not be used with a dataset containing more than one embedding column. * S3 Vectors uses approximate nearest neighbor (ANN) algorithms for performance, providing probabilistically closest results. * A single vector search can retrieve up to **10,000** results. Spice paginates the underlying `QueryVectors` calls (100 results per page) to reach this limit. When no `LIMIT` is specified, the search returns one page (100 results); a `LIMIT` larger than 10,000 is clamped to 10,000. ## Overview[​](#overview "Direct link to Overview") Amazon S3 Vectors exposes two core operations: upsert vectors (assign a vector to a key with optional metadata) and vector similarity queries (find closest vectors by metrics like cosine or Euclidean distance). It stores vectors durably in S3 and executes queries on transient compute, avoiding always-on servers. This enables petabyte-scale storage at low cost, ideal for semantic search, recommendations, and RAG in AI applications. ## Configuration[​](#configuration "Direct link to Configuration") Annotate dataset schemas to specify columns for embedding and models. Spice supports local or hosted models like Amazon Titan Embeddings. When data ingests, Spice generates embeddings and stores them in the configured S3 Vectors index. Spice handles index creation, updates, and synchronization with data changes. Example with AWS Titan model: ``` datasets: - from: oracle:"CUSTOMER_REVIEWS" name: reviews vectors: enabled: true engine: s3_vectors params: s3_vectors_bucket: my-s3-vector-bucket columns: - name: body embeddings: from: bedrock_titan embeddings: - from: bedrock:amazon.titan-embed-text-v2:0 name: bedrock_titan params: aws_region: us-east-2 dimensions: '256' ``` ## Querying[​](#querying "Direct link to Querying") Perform vector searches via HTTP API or SQL table-valued function `vector_search(dataset, query)`. HTTP example: ``` curl -X POST http://localhost:8090/v1/search \ -H "Content-Type: application/json" \ -d '{ "datasets": ["reviews"], "text": "issues with same day shipping", "additional_columns": ["rating", "customer_id"], "where": "created_at >= now() - INTERVAL '7 days'", "limit": 2 }' ``` SQL example: ``` SELECT review_id, rating, customer_id, body, score FROM vector_search(reviews, 'issues with same day shipping') WHERE created_at >= to_unixtime(now() - INTERVAL '7 days') ORDER BY score DESC LIMIT 2; ``` Results include matching snippets, additional fields, primary keys, scores, and table names. ## Managing Embeddings[​](#managing-embeddings "Direct link to Managing Embeddings") Spice manages the vector lifecycle: embedding data on ingestion, upserting to S3 Vectors with primary keys as identifiers, and handling updates/deletions. Vectors can be pre-stored in S3 for scalability, avoiding memory limits in accelerators or slow just-in-time computations. ## Query Execution[​](#query-execution "Direct link to Query Execution") Spice pushes similarity computations to S3 Vectors, retrieving top matches as (id, score) pairs. These form a temporary table joinable with the dataset for full records, reducing processing to candidates only. ## Handling Filters[​](#handling-filters "Direct link to Handling Filters") Mark columns as filterable metadata to push filters into S3 Vectors queries: ``` columns: - name: created_at metadata: vectors: filterable ``` This ensures results respect constraints like time ranges during similarity search. ### Which Predicates Push Down[​](#which-predicates-push-down "Direct link to Which Predicates Push Down") S3 Vectors accepts a restricted, MongoDB-style metadata filter, so only some predicates translate. A predicate is pushed into the S3 Vectors query when it references a `filterable` metadata column **on the left-hand side** and takes one of these shapes: | Predicate | Pushed down when | | ------------------------ | ---------------------------------------------------------------------------------- | | `=`, `!=` | The literal is a boolean, a string, or a finite number | | `<`, `<=`, `>`, `>=` | The literal is **numeric** — an integer, or a finite float | | `IS NULL`, `IS NOT NULL` | Always | | `IN (…)` | The list is non-empty and every element is a boolean, a string, or a finite number | | `AND`, `OR` | Every sub-expression is itself pushable | Anything else is evaluated by the Spice query engine after candidates come back, which is correct but reads more rows from the index. The cases that most often surprise: * **Range comparisons against strings are not pushed down.** `WHERE category > 'm'` is a valid SQL predicate, but S3 Vectors supports range operators only on numeric metadata, so it is filtered locally. Equality and `IN` on strings do push down. * **The column must be the left operand.** `WHERE rating > 4` pushes down; `WHERE 4 < rating` does not. * **`NaN` and infinite floats never push down**, nor do `NULL` literals — use `IS NULL` instead of `= NULL`. * **Columns that are not `filterable` metadata cannot be pushed at all**, including derived columns such as the similarity distance the index adds to results. ## Optimizations[​](#optimizations "Direct link to Optimizations") Store non-filterable columns as metadata to avoid joins: ``` columns: - name: rating metadata: vectors: non-filterable ``` Queries can then retrieve all needed data directly from the index, improving latency for read-heavy workloads. ## Advanced Features[​](#advanced-features "Direct link to Advanced Features") Spice supports hybrid search (vector + full-text via BM25, merged with RRF), multi-vector queries (weighting columns), and re-ranking (e.g., keyword first-pass then vector on candidates). Compose via SQL CTEs and joins. Hybrid RRF example: ``` WITH vector_results AS ( SELECT review_id, RANK() OVER (ORDER BY score DESC) AS vector_rank FROM vector_search(reviews, 'issues with same day shipping') ), text_results AS ( SELECT review_id, RANK() OVER (ORDER BY score DESC) AS text_rank FROM text_search(reviews, 'issues with same day shipping') ) SELECT COALESCE(v.review_id, t.review_id) AS review_id, (1.0 / (60 + COALESCE(v.vector_rank, 1000)) + 1.0 / (60 + COALESCE(t.text_rank, 1000))) AS fused_score FROM vector_results v FULL OUTER JOIN text_results t ON v.review_id = t.review_id ORDER BY fused_score DESC LIMIT 50; ``` Multi-column example (weighting title higher): ``` WITH body_results AS ( SELECT review_id, score AS body_score FROM vector_search(reviews, 'issues with same day shipping', col => 'body') ), title_results AS ( SELECT review_id, score AS title_score FROM vector_search(reviews, 'issues with same day shipping', col => 'title') ) SELECT COALESCE(body.review_id, title.review_id) AS review_id, COALESCE(body_score, 0) + 2.0 * COALESCE(title_score, 0) AS combined_score FROM body_results FULL OUTER JOIN title_results ON body_results.review_id = title_results.review_id ORDER BY combined_score DESC LIMIT 5; ``` ## Bucket and Index Naming[​](#bucket-and-index-naming "Direct link to Bucket and Index Naming") For S3 vector bucket and index names, AWS allows only lowercase letters, numbers, and hyphens. A name must also be 3 to 63 characters long. When not provided, Spice normalizes the index names it generates, but it passes `s3_vectors_bucket` and an explicit `s3_vectors_index` to AWS unchanged. As a result, invalid names will be rejected. ## Index Partitioning[​](#index-partitioning "Direct link to Index Partitioning") S3 Vectors indexes can be partitioned using an arbitrary logical expression. This enables Spice to compose many actual vector indexes as one logical vector index, enabling elastic scalability for vector storage. To partition your S3 vector indexes: ``` vectors: enabled: true engine: s3_vectors partition_by: - 'bucket(50, PULocationID)' ``` This example uses a `bucket` user-defined function (UDF) to hash the `PULocationID` column and split the associated vectors into one of 50 partitioned indexes. The runtime will use the `s3_vectors_index` parameter as a prefix and generate partition-specific names. The prefix is normalized (`_` and `.` become `-`) and truncated to 45 characters to leave room for the partition suffix within the S3 Vectors 63-character index-name limit — see [Index Naming](#bucket-and-index-naming). Limitations * `partition_by` must have only 1 expression. * Expression must reference exactly one column from the dataset. * Expression must produce a scalar value * Expression cannot contain a subquery ## Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3 Vectors permissions (e.g., `s3vectors:QueryVectors`, `s3vectors:GetVectors`) to access the vectors. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following minimum permissions to access S3 Vectors: ``` { "Version": "2012-10-17", "Statement": [ { "Sid": "AllowApplicationVectorAccess", "Effect": "Allow", "Action": [ "s3vectors:QueryVectors", "s3vectors:GetIndex", "s3vectors:PutVectors", "s3vectors:ListVectors", "s3vectors:GetVectors", "s3vectors:ListIndexes", "s3vectors:DeleteVectors" ], "Resource": [ "arn:aws:s3vectors:aws-region:123456789012:bucket/amzn-s3-demo-vector-bucket/index/*", ] }, { "Sid": "AllowGetVectorBucket", "Effect": "Allow", "Action": "s3vectors:GetVectorBucket", "Resource": "arn:aws:s3vectors:aws-region:123456789012:bucket/*" }, { "Sid": "AllowCreateVectorBucket", "Effect": "Allow", "Action": "s3vectors:CreateVectorBucket", "Resource": "arn:aws:s3vectors:aws-region:123456789012:bucket/*" }, { "Sid": "AllowCreateVectorBucketIndex", "Effect": "Allow", "Action": "s3vectors:CreateIndex", "Resource": "arn:aws:s3vectors:aws-region:123456789012:bucket/amzn-s3-demo-vector-bucket/index/*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------------------ | --------------------------------------------------------------------------------------- | | `s3vectors:DeleteVectors` | Required. Used to remove vectors when rows are deleted from the dataset. | | `s3vectors:GetIndex` | Required. Used to verify if the index already exists or needs to be created. | | `s3vectors:GetVectorBucket` | Required. Used to verify if the vector bucket already exists or needs to be created. | | `s3vectors:GetVectors` | Required. Used to read vector data when the `*_embeddings` column is projected. | | `s3vectors:ListIndexes` | Required when using index partitioning or spill writes, to enumerate physical indexes. | | `s3vectors:ListVectors` | Required. Used to populate the `*_embeddings` column on vector tables in Spice. | | `s3vectors:PutVectors` | Required. Used to populate the vector index with Spice-computed embeddings. | | `s3vectors:QueryVectors` | Required. Used to query for vectors using the `vector_search` table function. | | `s3vectors:CreateIndex` | Optional. Spice can automatically create indexes if this permission is given. | | `s3vectors:CreateVectorBucket` | Optional. Spice can automatically create the vector bucket if this permission is given. | ### `metrics`[​](#metrics "Direct link to metrics") Spice supports the following [S3 Vector engine metrics](/docs/next/features/observability/component_metrics): | Metric Name | Type | Description | | ------------------------------------------------- | --------- | ---------------------------------------------------------------------------- | | `s3_vectors_create_index_errors` | counter | Number of errors returned from create\_index operation. | | `s3_vectors_create_index_latency` | histogram | Total duration of create\_index operation, in milliseconds. | | `s3_vectors_create_index_requests` | counter | Number of requests to create\_index operation. | | `s3_vectors_create_vector_bucket_errors` | counter | Number of errors returned from create\_vector\_bucket operation. | | `s3_vectors_create_vector_bucket_latency` | histogram | Total duration of create\_vector\_bucket operation, in milliseconds. | | `s3_vectors_create_vector_bucket_requests` | counter | Number of requests to create\_vector\_bucket operation. | | `s3_vectors_delete_index_errors` | counter | Number of errors returned from delete\_index operation. | | `s3_vectors_delete_index_latency` | histogram | Total duration of delete\_index operation, in milliseconds. | | `s3_vectors_delete_index_requests` | counter | Number of requests to delete\_index operation. | | `s3_vectors_delete_vector_bucket_errors` | counter | Number of errors returned from delete\_vector\_bucket operation. | | `s3_vectors_delete_vector_bucket_latency` | histogram | Total duration of delete\_vector\_bucket operation, in milliseconds. | | `s3_vectors_delete_vector_bucket_policy_errors` | counter | Number of errors returned from delete\_vector\_bucket\_policy operation. | | `s3_vectors_delete_vector_bucket_policy_latency` | histogram | Total duration of delete\_vector\_bucket\_policy operation, in milliseconds. | | `s3_vectors_delete_vector_bucket_policy_requests` | counter | Number of requests to delete\_vector\_bucket\_policy operation. | | `s3_vectors_delete_vector_bucket_requests` | counter | Number of requests to delete\_vector\_bucket operation. | | `s3_vectors_delete_vectors_errors` | counter | Number of errors returned from delete\_vectors operation. | | `s3_vectors_delete_vectors_latency` | histogram | Total duration of delete\_vectors operation, in milliseconds. | | `s3_vectors_delete_vectors_requests` | counter | Number of requests to delete\_vectors operation. | | `s3_vectors_get_index_errors` | counter | Number of errors returned from get\_index operation. | | `s3_vectors_get_index_latency` | histogram | Total duration of get\_index operation, in milliseconds. | | `s3_vectors_get_index_requests` | counter | Number of requests to get\_index operation. | | `s3_vectors_get_vector_bucket_errors` | counter | Number of errors returned from get\_vector\_bucket operation. | | `s3_vectors_get_vector_bucket_latency` | histogram | Total duration of get\_vector\_bucket operation, in milliseconds. | | `s3_vectors_get_vector_bucket_policy_errors` | counter | Number of errors returned from get\_vector\_bucket\_policy operation. | | `s3_vectors_get_vector_bucket_policy_latency` | histogram | Total duration of get\_vector\_bucket\_policy operation, in milliseconds. | | `s3_vectors_get_vector_bucket_policy_requests` | counter | Number of requests to get\_vector\_bucket\_policy operation. | | `s3_vectors_get_vector_bucket_requests` | counter | Number of requests to get\_vector\_bucket operation. | | `s3_vectors_get_vectors_errors` | counter | Number of errors returned from get\_vectors operation. | | `s3_vectors_get_vectors_latency` | histogram | Total duration of get\_vectors operation, in milliseconds. | | `s3_vectors_get_vectors_requests` | counter | Number of requests to get\_vectors operation. | | `s3_vectors_list_indexes_errors` | counter | Number of errors returned from list\_indexes operation. | | `s3_vectors_list_indexes_latency` | histogram | Total duration of list\_indexes operation, in milliseconds. | | `s3_vectors_list_indexes_requests` | counter | Number of requests to list\_indexes operation. | | `s3_vectors_list_vector_buckets_errors` | counter | Number of errors returned from list\_vector\_buckets operation. | | `s3_vectors_list_vector_buckets_latency` | histogram | Total duration of list\_vector\_buckets operation, in milliseconds. | | `s3_vectors_list_vector_buckets_requests` | counter | Number of requests to list\_vector\_buckets operation. | | `s3_vectors_list_vectors_errors` | counter | Number of errors returned from list\_vectors operation. | | `s3_vectors_list_vectors_latency` | histogram | Total duration of list\_vectors operation, in milliseconds. | | `s3_vectors_list_vectors_requests` | counter | Number of requests to list\_vectors operation. | | `s3_vectors_put_vector_bucket_policy_errors` | counter | Number of errors returned from put\_vector\_bucket\_policy operation. | | `s3_vectors_put_vector_bucket_policy_latency` | histogram | Total duration of put\_vector\_bucket\_policy operation, in milliseconds. | | `s3_vectors_put_vector_bucket_policy_requests` | counter | Number of requests to put\_vector\_bucket\_policy operation. | | `s3_vectors_put_vectors_errors` | counter | Number of errors returned from put\_vectors operation. | | `s3_vectors_put_vectors_latency` | histogram | Total duration of put\_vectors operation, in milliseconds. | | `s3_vectors_put_vectors_requests` | counter | Number of requests to put\_vectors operation. | | `s3_vectors_query_vectors_errors` | counter | Number of errors returned from query\_vectors operation. | | `s3_vectors_query_vectors_latency` | histogram | Total duration of query\_vectors operation, in milliseconds. | | `s3_vectors_query_vectors_requests` | counter | Number of requests to query\_vectors operation. | ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure a dataset with an S3 vectors engine in Spice. [S3 Vectors engine](https://github.com/spiceai/cookbook/tree/trunk/vectors/s3#readme) ## References[​](#references "Direct link to References") * [Spice.ai announcement](https://spice.ai/blog/getting-started-with-amazon-s3-vectors-and-spice) * [Amazon S3 Vectors official page](https://aws.amazon.com/s3/vectors/) --- # Spice.ai Deployment Guide Spice runs as a single binary, a container, a Kubernetes workload, or a fully managed app on the Spice Cloud Platform. This guide helps choose a target environment and a deployment architecture to match an application's latency, scale, and operational requirements. ## Choose a deployment target[​](#choose-a-deployment-target "Direct link to Choose a deployment target") Most users fall into one of three groups: * **Run Spice next to an application** — start with [Docker](/docs/next/deployment/docker) for a local container, or follow [Getting Started](/docs/next/getting-started) to run the binary directly. * **Operate Spice in production on Kubernetes** — use the [Spice Helm chart](/docs/next/deployment/kubernetes/helm). For automated rollouts, see the [CI/CD guide](/docs/next/deployment/ci-cd) for Helm pipelines and GitOps with [Argo CD](/docs/next/deployment/kubernetes/argocd) or [Flux](/docs/next/deployment/kubernetes/flux). * **Use a managed service** — deploy a Spicepod to the [Spice Cloud Platform](/docs/next/deployment/cloud) and connect a [GitHub repository](https://docs.spice.ai/docs/portal/apps/connect-github) for continuous delivery. Self-hosted enterprise deployments For production self-hosted clusters, the [Spice.ai Enterprise Kubernetes Operator](https://docs.spice.ai/docs/enterprise/kubernetes-operator/kubernetes) provides per-replica StatefulSets, automatic PVC resizing, configurable update strategies, crashloop protection, and distributed query execution through `SpicepodSet` and `SpicepodCluster` custom resources. ## Deployment architectures[​](#deployment-architectures "Direct link to Deployment architectures") Architecture refers to where Spice runs in relation to the application and data sources, and how it scales. Pick an architecture before choosing a guide; the same target environment can host any of these patterns. * [Overview](/docs/next/deployment/architectures) — when to choose each architecture. * [Sidecar](/docs/next/deployment/architectures/sidecar) — Spice runs alongside the application for the lowest latency. * [Microservice](/docs/next/deployment/architectures/microservice) — single or multiple replicas behind a load balancer. * [Tiered](/docs/next/deployment/architectures/tiered) — separate read and write tiers for mixed workloads. * [Cluster-Sidecar](/docs/next/deployment/architectures/cluster-sidecar) — combine local and remote Spice instances. * [Hosted](/docs/next/deployment/architectures/hosted) — managed on the Spice Cloud Platform. * [Sharded](/docs/next/deployment/architectures/sharded) — partition data across multiple Spice instances. * [Cluster](/docs/next/deployment/architectures/cluster) — distributed query execution with Spice.ai Enterprise. ## Deployment guides[​](#deployment-guides "Direct link to Deployment guides") Step-by-step instructions for each target environment. | Guide | When to use | | -------------------------------------------------------------------- | -------------------------------------------------------------------------- | | [Kubernetes](/docs/next/deployment/kubernetes) | Self-hosted production deployments. Covers Helm, Argo CD, and Flux. | | [Docker](/docs/next/deployment/docker) | Local development, single-host deployments, and container-based pipelines. | | [Spice Cloud](/docs/next/deployment/cloud) | Fully managed deployments without operating infrastructure. | | [AWS](/docs/next/deployment/aws) | Deployments on AWS using the published CloudFormation template. | | [Azure](/docs/next/deployment/azure) | Deployments on Azure using ARM/Bicep templates. | | [GCP](/docs/next/deployment/gcp) | Deployments on Google Cloud using GKE, Cloud Run, or Compute Engine. | | [CI/CD](/docs/next/deployment/ci-cd) | Automating any of the above through pipelines or GitOps. | | [Read/Write Separation](/docs/next/deployment/read-write-separation) | Production pattern that splits ingest from reads using shared snapshots. | --- # Deployment Architectures ![Spice.ai OSS as a data and AI compute engine over disaggregated storage](https://github.com/user-attachments/assets/da3c0e90-4c48-48ca-b4bd-72eda816cfec) Spice supports multiple deployment architectures: * [Sidecar Deployment](/docs/next/deployment/architectures/sidecar) - Deploy alongside applications * [Microservice Deployment (Single or Multiple Replicas)](/docs/next/deployment/architectures/microservice) - Standalone service deployment * [Tiered Deployment](/docs/next/deployment/architectures/tiered) - Edge, application, and cloud tiers * [Cluster-Sidecar Deployment](/docs/next/deployment/architectures/cluster-sidecar) - Sidecar caching backed by a centralized cluster * [Cloud-Hosted in the Spice Cloud Platform](/docs/next/deployment/architectures/hosted) - Managed cloud deployment * [Sharded Deployment](/docs/next/deployment/architectures/sharded) - Horizontal data partitioning * [Cluster Deployment (Spice.ai Enterprise)](/docs/next/deployment/architectures/cluster) - Distributed cluster architecture --- # Cluster-Based Deployment (Spice.ai Enterprise) A full cluster-based deployment leveraging **Spice.ai Enterprise**, which includes advanced services and integrations for Kubernetes. This method is ideal for organizations requiring large-scale or complex deployments, including specialized clustering capabilities. ![cluster](https://github.com/user-attachments/assets/643e0a5c-6745-40c0-8695-0955c795179b) **Benefits** * Provides **enterprise-grade features**: advanced security, monitoring, and support. * Simplifies **managing multiple nodes** for high availability and large workloads. * Offers **direct integration** with Spice Cloud or on-prem Kubernetes clusters. **Considerations** * **Requires a commercial license** or subscription to Spice Enterprise. * More **complex initial setup**, typically involving specialized DevOps expertise. **Use This Approach When** * You operate at **significant scale** or have stringent availability requirements. * You need **enterprise-level support** and advanced monitoring, security, or compliance features. * Your team can manage a **robust Kubernetes environment** or you plan to integrate with the Spice Cloud at scale. **Example Use Case** A large financial services firm requiring a highly available, secure environment. They run Spice.ai across multiple clusters using Spice Enterprise for advanced monitoring, role-based access control, and dedicated support. --- # Cluster-Sidecar Deployment Modern applications have two fundamentally different data access patterns, and no single deployment model serves both well. Large analytical queries — scanning terabytes of Iceberg data, joining Delta Lake tables, running cross-dataset aggregations — need distributed execution across many nodes. Hot operational queries — serving the working set a microservice actually uses, answering user-facing requests in under 5 milliseconds, feeding fresh context to an AI agent — need data materialized right next to the application, with no network hop. Those workloads share a third requirement: **isolation**. Giving an application — especially an AI agent that writes its own queries — direct credentials to production Postgres, a data lake, or a warehouse means a bad plan, a runaway loop, or a prompt injection can exhaust connection pools, scan petabytes, or touch rows it shouldn't. The retrieval layer needs to be a sandbox, not a passthrough. ![cluster-sidecar](/img/deployment/cluster-sidecar.png) The [cluster-sidecar architecture](https://spice.ai/blog/cluster-sidecar-architecture) addresses all three. Application sidecars handle the hot path with a scoped, locally accelerated working set, while a centralized Spice [cluster](/docs/next/deployment/architectures/cluster) (or the [Spice Cloud Platform](/docs/next/deployment/architectures/hosted)) provides distributed compute for heavy queries, data ingestion, acceleration, and refresh. When a sidecar needs to reach beyond its materialized working set — a historical query, a cross-dataset join, a broad search — it transparently delegates to the cluster over Arrow Flight, which executes the query and returns results. The sidecar can then cache those results for future use. From the application's perspective, everything is `localhost`. From an infrastructure perspective, the system delivers the throughput of a distributed query engine, the latency of an embedded database, and a hard isolation boundary between the application and origin data systems — without ETL between them, sync jobs, or consistency gaps. Think of it as a [CDN for your data](/docs/next/use-cases/data/database-cdn): the cluster is the origin server, the sidecars are the edge nodes, and Spice handles the caching, invalidation, and routing. Origin databases, data lakes, and CDC streams never see the application fleet. Each sidecar is configured declaratively via a `spicepod.yaml` — the datasets, views, acceleration engines, search indices, and AI models it manages. Sidecars start in seconds, typically run on a few hundred megabytes of memory, and scale horizontally with application pods: scale a deployment from 5 to 50 replicas and 50 sidecars come up automatically, each materializing only the working set its spicepod declares. Sidecars never talk to each other; they only talk to the cluster. For the engineering walkthrough of this pattern, including a request-path example and FAQ, see [Localhost Latency at Scale: The Spice Cluster-Sidecar Architecture](https://spice.ai/blog/cluster-sidecar-architecture). ## How it works[​](#how-it-works "Direct link to How it works") ### The sidecar as a sandbox[​](#the-sidecar-as-a-sandbox "Direct link to The sidecar as a sandbox") The most important property of the sidecar is isolation, not just latency. The sidecar is the only data-plane surface the application touches: * **Scoped working set.** A sidecar's `spicepod.yaml` declares exactly which datasets, views, and search indices the application may query. Anything not declared is physically absent from the catalog — not filtered by a policy, not hidden by a row-level rule. Compare this with row-level security, where the underlying data is still present and a single policy misconfiguration can expose it. Even a perfectly crafted prompt injection cannot query a table that is not in the catalog. * **No origin credentials in the application.** The application connects to its sidecar with a local token. The sidecar connects to the cluster over Arrow Flight. Only the cluster holds credentials for Postgres, Snowflake, Databricks, S3, Kafka, and the rest. Compromising an application pod cannot leak origin credentials, because the pod never had them. * **Narrow network surface.** The application's only outbound data dependency is the loopback interface. Network policy can pin the sidecar's egress to the cluster endpoint only. * **Per-application data views.** Sidecars can be specialized per application class or per tenant. A customer-service agent, a fraud-review agent, and an internal dashboard can each run with different spicepods pointing at different slices of the same cluster. This is physical isolation, not policy-based filtering on a shared database. See [Multi-Tenant AI Agents](/docs/use-cases/ai/multi-tenant-agents). * **Bounded resource use.** A rogue query plan or runaway loop exhausts the sidecar's local memory and CPU budget, not the cluster's and not the origin database's. * **Local inference.** Sidecars can serve [LLM inference and tool calls](/docs/next/features/large-language-models) on loopback so sensitive prompts and retrieved context stay in the pod. Heavier models can still be routed to the cluster. * **One audit point.** Every query, search, and inference call flows through the sidecar, so there is one place to log, rate-limit, and enforce policy. Data only moves along one path: **application → sidecar → (optionally) cluster → origin**. It never skips a tier. ### Three latency tiers[​](#three-latency-tiers "Direct link to Three latency tiers") A request is served from the first tier that can answer it: 1. **Sidecar results cache** — repeat queries return from the in-memory [results cache](/docs/next/features/caching) in microseconds (`Results-Cache-Status: HIT`). 2. **Sidecar working set** — novel queries that fit the locally materialized dataset execute on-node in single-digit milliseconds (Arrow, DuckDB, or SQLite). 3. **Cluster delegation** — queries that exceed the working set are forwarded to the cluster. [Apache Ballista](/docs/next/features/distributed-query) distributes execution; [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) accelerates large scans. The sidecar caches the result so subsequent reads drop back to tier 1. The cluster has its own results cache, so one sidecar's miss can become a cluster cache hit for every subsequent sidecar. Typical application-visible latency: p50 at tier 1, p95 at tier 2, tail at tier 3. Useful cache knobs in this topology: * `cache_key_type: plan` (default) shares a cache entry across semantically equivalent SQL, which matters for ORM-generated queries. `sql` is a faster, string-exact lookup. * `stale_while_revalidate_ttl` lets the sidecar serve a stale cached result immediately while a background refresh runs. * Clients can send `Cache-Control` directives per request (`no-cache`, `only-if-cached`, `stale-if-error`). * `encoding: zstd` typically cuts results-cache memory by 50–90%. ### The cluster ingests once[​](#the-cluster-ingests-once "Direct link to The cluster ingests once") Refreshing from upstream sources, running Cayenne acceleration on large Iceberg tables, and keeping CDC streams connected are resource-intensive. Doing any of that *N* times for *N* sidecars is wasteful and often infeasible — source systems have connection limits, and per-pod CDC multiplies cloud cost with fleet size. The cluster ingests each dataset once and produces one authoritative materialization. Source load is bounded by cluster size, not fleet size. A new pod's sidecar pulls its working set from the cluster (or from an [acceleration snapshot](/docs/next/features/data-acceleration/snapshots)) rather than re-scanning the source, so new nodes are operational in seconds. Sidecars stay lightweight because they do **not** own ingest: * They start in seconds — important when application pods autoscale with traffic. * They typically run on a few hundred megabytes of memory. * Scaling from 5 to 50 replicas does **not** add 50 new connections to the source database. * Each sidecar materializes only the working set its spicepod declares, not a full copy of the warehouse. Full accelerated datasets stay on the cluster. ### Query delegation over Arrow Flight[​](#query-delegation-over-arrow-flight "Direct link to Query delegation over Arrow Flight") Sidecar-to-cluster communication is [Arrow Flight](https://arrow.apache.org/docs/format/Flight.html) over gRPC. Results flow as Arrow record batches directly into the sidecar's query engine with no JSON or row-based serialization detour. The sidecar decides locally whether a query can be served from its working set. If not, it forwards and streams results back. The application sees one endpoint and one query. It never knows whether execution happened locally or whether Ballista fanned the query out across cluster executors. Configure delegation with the [Spice.ai Data Connector](/docs/next/components/data-connectors/spiceai): point a sidecar dataset at a Spice Cloud app (`spice.ai///datasets/`) or at a self-hosted cluster (`spice.ai:https://cluster.example:50051`). Combine with local acceleration to materialize the hot working set; leave acceleration off to always delegate. Kubernetes is the most common place to run this topology, but it is not required. The architecture needs only a network path from each sidecar to the cluster. Sidecars run wherever the application runs — Kubernetes, a VPC, on-prem, or at the edge. ### Acceleration snapshots[​](#acceleration-snapshots "Direct link to Acceleration snapshots") The cluster-sidecar split maps onto [acceleration snapshots](/docs/next/features/data-acceleration/snapshots) as a single-writer / many-reader topology: * The **cluster** is the single writer (`snapshots: create_only`). After each refresh it uploads a snapshot to object storage. It never downloads snapshots on startup; it always refreshes from the source. * Each **sidecar** is a reader (`snapshots: bootstrap_only`). On startup — or when ephemeral NVMe is recycled — the sidecar downloads the most recent snapshot and is immediately ready. It never writes snapshots back. This avoids snapshot conflicts, keeps the cluster as the authoritative refresh point, and gives every sidecar a warm start from the same materialization. For CDC-backed datasets with large initial state, the difference between a snapshot bootstrap and a full re-sync can be seconds versus minutes. Use DuckDB or SQLite in `mode: file` on the sidecar when snapshots are in play — snapshots persist and restore the acceleration file itself. Heavyweight [Cayenne](/docs/next/components/data-accelerators/cayenne) acceleration stays on the cluster. For production Spicepod and Helm reference configurations, see [Read/Write Separation](/docs/deployment/read-write-separation). ### Cache coherency is a refresh policy, not a protocol[​](#cache-coherency-is-a-refresh-policy-not-a-protocol "Direct link to Cache coherency is a refresh policy, not a protocol") Sidecars pull from the cluster on a configurable interval using append or full refresh. They do **not** participate in a distributed invalidation protocol. A pull-based model with explicit refresh intervals makes staleness bounded, predictable, and debuggable. For workloads that need sub-second freshness, the cluster consumes [CDC streams](/docs/next/features/cdc) (Postgres logical replication, DynamoDB Streams, Debezium, Kafka) once, and the sidecars pull the resulting accelerated dataset on a short interval. That gives near-real-time propagation without a fleet-wide invalidation bus. If the cluster is temporarily unavailable — a rolling upgrade, a network blip, a zone event — sidecars keep serving cached data and their accelerated working sets. Refreshes pause and resume when connectivity returns. The blast radius of a cluster incident is slightly staler data, not application downtime. A `stale-if-error` cache directive can extend this further for delegated queries. ## Example Spicepods[​](#example-spicepods "Direct link to Example Spicepods") The cluster materializes and accelerates datasets, consumes CDC, and exposes a results cache. It is the only tier with origin credentials. ``` version: v1 kind: Spicepod name: platform-cluster runtime: caching: sql_results: enabled: true max_size: 4GiB item_ttl: 1m stale_while_revalidate_ttl: 30s encoding: zstd snapshots: enabled: true location: s3://my-bucket/spice-snapshots/ params: s3_auth: iam_role datasets: - from: postgres:public.orders name: orders acceleration: enabled: true engine: cayenne mode: file refresh_mode: changes snapshots: create_only primary_key: id on_conflict: id: upsert - from: s3://lakehouse/events/ name: events params: file_format: parquet acceleration: enabled: true engine: cayenne mode: file refresh_mode: append refresh_check_interval: 15m snapshots: create_only ``` The sidecar is much smaller. It pulls from the cluster, keeps a working-set engine local, and caches results. It holds no origin credentials. ``` version: v1 kind: Spicepod name: app-sidecar runtime: caching: sql_results: enabled: true max_size: 256MiB item_ttl: 30s stale_while_revalidate_ttl: 30s snapshots: enabled: true location: s3://my-bucket/spice-snapshots/ bootstrap_on_failure_behavior: warn params: s3_auth: iam_role datasets: - from: spice.ai///datasets/orders name: orders acceleration: enabled: true engine: duckdb mode: file refresh_mode: append refresh_check_interval: 10s snapshots: bootstrap_only - from: spice.ai///datasets/events name: events # No local acceleration — delegate to the cluster on demand. # Queries that match recent hot events still hit the results cache. ``` For a self-hosted cluster, replace the `from:` URIs with `spice.ai:https://cluster.example:50051` and set `name:` to the upstream table name. See the [Spice.ai Data Connector](/docs/next/components/data-connectors/spiceai) for Cloud and self-hosted URI formats. ## A request path[​](#a-request-path "Direct link to A request path") 1. The application (or agent) queries its sidecar on `localhost:8090`: `SELECT ... FROM orders WHERE tenant_id = $1 ORDER BY created_at DESC LIMIT 20`. 2. The sidecar checks its results cache. Hit → return in microseconds. 3. On a miss, the sidecar plans against its local catalog. `orders` is materialized locally (refreshed from the cluster every 10 seconds). It executes against DuckDB, returns in single-digit milliseconds, and populates the results cache. Origin Postgres sees no traffic. 4. A follow-up hybrid search over a multi-gigabyte `tickets` index is not in the sidecar catalog as a local acceleration. The sidecar opens an Arrow Flight stream to the cluster. Ballista distributes the search; Cayenne segment statistics prune most files. 5. Results stream back as Arrow record batches. The sidecar returns them to the application and caches them. A later LLM call can be served by a small local model on loopback, with the large-model call delegated to the cluster. 6. The cluster independently caches its own results. The next replica that asks the same question gets it from the cluster cache. The application sees one endpoint, one wire format, and one latency distribution. ## Benefits[​](#benefits "Direct link to Benefits") * **Kubernetes-native** — designed to run on Kubernetes, leveraging pod-level sidecars with cluster-level orchestration. The same topology also works outside Kubernetes wherever the application has a network path to the cluster. * Sub-millisecond reads via sidecar caching on loopback, with centralized data management in the cluster. * Structural isolation — sidecars expose a scoped catalog, hold no origin credentials, and bound resource use per pod. * Transparent query delegation — sidecars automatically route queries beyond their cached working set to the cluster. * Sidecars remain lightweight — only a working set and a results cache, no ingest or heavy acceleration overhead. * Cluster (or Spice Cloud) handles complex operations: data ingestion, [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) acceleration, [distributed query](/docs/next/features/distributed-query), hybrid search, and refresh from sources. * Works with both self-managed Spice clusters and the managed [Spice Cloud Platform](/docs/next/deployment/architectures/hosted) as the centralized backend. The Spice Cloud cluster-sidecar model is the most common production topology. * Sidecars can run anywhere — in your VPC, on-premises, at the edge, or in any Kubernetes cluster — while connecting securely to the managed cluster. * Horizontal scalability — add sidecars without increasing load on data sources. * Resilience — sidecars serve cached data even if the cluster is temporarily unavailable. * Secure by default — mTLS encryption across all sidecar-to-cluster communication, with data encrypted at rest and in transit. ## Considerations[​](#considerations "Direct link to Considerations") * More complex deployment structure requiring both sidecar and cluster infrastructure. [Spice Cloud](/docs/next/deployment/architectures/hosted) reduces this burden by managing the cluster. * Cache coherency — sidecars must be configured with appropriate refresh intervals or TTLs to balance freshness with performance. There is no distributed invalidation bus. * Requires a Spice cluster deployment or [Spice Cloud Platform](/docs/next/deployment/architectures/hosted) subscription ([Spice.ai Enterprise](https://spice.ai/enterprise) for self-managed clustering with SSO, RBAC, and audit logs). * Network connectivity between sidecars and the cluster must be reliable for cache refreshes and query delegation. ## Use This Approach When[​](#use-this-approach-when "Direct link to Use This Approach When") * Applications or AI agents require sub-millisecond reads, unified retrieval (SQL, full-text, vector, hybrid), and inference on `localhost`, without giving the application direct database credentials. * You need a clear isolation boundary between application code and origin data systems — scoped catalogs, no origin credentials in the pod, and one audit point per replica. * Multiple application instances need fast access to the same datasets without each independently querying data sources. * Reducing load on upstream data sources is a priority — the cluster ingests once, sidecars cache locally. * The system benefits from separating the caching tier (sidecars) from the data processing tier (cluster). * Workloads span both real-time operational queries and large-scale analytical queries on the same data (for example, an operational data lakehouse on S3/Iceberg). * You already run Kubernetes sidecars for other concerns (service mesh, logging, config). ## Not Ideal When[​](#not-ideal-when "Direct link to Not Ideal When") * The application is simple with a single instance and no isolation requirement — the overhead of both sidecar and cluster infrastructure isn't justified. Consider [Sidecar](/docs/next/deployment/architectures/sidecar) or [Microservice](/docs/next/deployment/architectures/microservice). * All queries are batch or analytical with relaxed latency requirements — a [Microservice](/docs/next/deployment/architectures/microservice) deployment is simpler and sufficient. * Network connectivity between sidecars and the cluster is unreliable — query delegation and cache refreshes will fail, leading to stale data. Consider standalone [Sidecar](/docs/next/deployment/architectures/sidecar) deployments with direct source access. ## Example Use Case[​](#example-use-case "Direct link to Example Use Case") A multi-tenant SaaS platform runs an AI support agent. Each tenant's agent pods include a Spice sidecar that materializes that tenant's working set (recent tickets, active customer records, the last 7 days of events, the tenant's private knowledge-base embeddings) into a local DuckDB + vector index. The spicepod exposes exactly those datasets — and no others — to the agent. The sidecar holds no credentials for Postgres, S3, or Snowflake. Agent turns hit the sidecar on `localhost` and return in single-digit milliseconds; repeat retrievals return from the results cache in microseconds. Behind the sidecars, a Spice cluster (often Spice Cloud) ingests from PostgreSQL via CDC, S3 Iceberg tables, and Databricks. Cayenne acceleration and refresh schedules run on the cluster. When an agent asks a broader question — "summarize this tenant's churn signal across the last 12 months" — the sidecar recognizes the query exceeds its working set and delegates over Arrow Flight. The cluster executes it, the sidecar caches the result, and subsequent turns for that tenant return in microseconds. The origin Postgres sees exactly one consumer (the cluster). If a replica is compromised, the attacker gets a loopback endpoint scoped to that tenant's working set, not database credentials and not a query interface to the whole warehouse. The same pattern — single ingestion path, per-pod sandboxing, tiered latency, one isolation boundary — works for any application fleet, not just agents: microservices serving real-time dashboards, services powering search, or internal tools querying operational data. ## See also[​](#see-also "Direct link to See also") * [Localhost Latency at Scale: The Spice Cluster-Sidecar Architecture](https://spice.ai/blog/cluster-sidecar-architecture) — engineering walkthrough, request-path example, and FAQ. * [Read/Write Separation](/docs/deployment/read-write-separation) — production guide for splitting ingest from reads using shared acceleration snapshots, including Spicepod and Helm reference configurations. * [Spice.ai Data Connector](/docs/next/components/data-connectors/spiceai) — how sidecars federate to a cluster or to Spice Cloud over Arrow Flight. * [Caching](/docs/next/features/caching) — results cache, stale-while-revalidate, and `Cache-Control`. * [Acceleration Snapshots](/docs/next/features/data-acceleration/snapshots) — warm-start file-mode accelerations from object storage. * [Distributed Query](/docs/next/features/distributed-query) — Apache Ballista schedulers and executors on the cluster. * [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) — cluster-side acceleration for datasets beyond 1 TB. * [Multi-Tenant AI Agents](/docs/use-cases/ai/multi-tenant-agents) — per-tenant isolation patterns that pair with this architecture. --- # Cloud Hosted The Spice Runtime is deployed on a fully managed service within the Spice Cloud Platform, minimizing the operational burden of managing clusters, upgrades, and infrastructure. ![hosted](https://github.com/user-attachments/assets/a985527b-3481-40f4-a689-f784c893b314) **Benefits** * Reduced overhead for deployment, scaling, and maintenance. * Access to specialized hosting features and quick setup. * Helps reduce operational complexity and cost. **Considerations** * Reliance on external hosting and associated terms or limits. * Potential compliance or data residency considerations for certain industries. * May introduce latency depending on the cloud provider's infrastructure. **Use This Approach When** * Limited DevOps resources are available, or focus on application logic over infrastructure is preferred. * A fully managed environment with minimal setup time is desired. * A single, managed solution is prioritized over running own clusters. * Minimizing operational complexity and cost is the goal. **Example Use Case** A startup or team with limited DevOps support that needs a reliable, managed environment. Quick deployment and minimal in-house infrastructure responsibilities are priorities. --- # Microservice Deployment (Single or Multiple Replicas) The Spice Runtime operates as an independent microservice. Multiple replicas may be deployed behind a load balancer to achieve high availability and handle spikes in demand. ![microservice](https://github.com/user-attachments/assets/b46f050b-e500-4d53-b354-24f0ab30cad3) **Benefits** * Loose coupling between the application and the Spice Runtime. * Independent scaling and upgrades. * Can serve multiple applications or services within an organization. * Helps achieve high availability and redundancy. **Considerations** * Additional network hop introduces latency compared to sidecar. * More complex infrastructure, requiring service discovery and load balancing. * Potentially higher cost due to additional infrastructure components. **Use This Approach When** * A loosely coupled architecture and the ability to independently scale the AI service are desired. * Multiple services or teams need to share the same AI engine. * Heavy or varying traffic is anticipated, requiring independent scaling of the Spice Runtime. * Resiliency and redundancy are prioritized over simplicity. **Example Use Case** A large organization where multiple services (recommendations, analytics, etc.) need to share AI insights. A centralized Spice Runtime microservice cluster helps separate teams consume AI outputs without duplicating efforts. --- # Sharded The Spice Runtime instances can be sharded based on specific criteria, such as by customer, state, or other logical partitions. Each shard operates independently, with a 1:N Application to Spice instances ratio. ![sharded](https://github.com/user-attachments/assets/5730d108-6d22-4ea4-8c14-8e87ad6d0079) **Benefits** * Helps distribute load across multiple instances, improving performance and scalability. * Isolates failures to specific shards, enhancing resiliency. * Allows tailored configurations and optimizations for different shards. **Considerations** * More complex deployment and management due to multiple instances. * Requires effective sharding strategy to balance load and avoid hotspots. * Potentially higher cost due to multiple instances. **Use This Approach When** * Distributing load across multiple instances for better performance is needed. * Isolating failures to specific shards to improve resiliency is desired. * The application can benefit from tailored configurations for different logical partitions. * The complexity of managing multiple instances can be handled. **Example Use Case** A multi-tenant application where each customer has a dedicated Spice Runtime instance. This helps ensure that heavy usage by one customer does not impact others, and allows for customer-specific optimizations. **Sharding vs. partitioning** Sharding splits load across **multiple Spice instances**, each backing a logical slice of the system (a customer, a region, a workload). Each shard runs an independent runtime with its own datasets, accelerations, and resources. Within a single Spice instance, [acceleration partitioning](/docs/next/features/data-acceleration/partitioning) splits a single dataset into multiple physical units (files, tables, or in-memory tables) so that filtered queries only read the relevant subset. The two are complementary: a shard can also use partitioning internally to keep individual datasets pruneable. | Concern | Sharded deployment | Acceleration partitioning | | ----------------- | -------------------------------------------------------- | -------------------------------------------------------------- | | Splits across… | Multiple runtimes/processes | One acceleration on one runtime | | Routing | Application picks which Spice instance to query | Spice prunes partitions automatically based on filter pushdown | | Failure isolation | Per-shard | None — single runtime | | Use when | Tenants/regions have very different load or data volumes | A single dataset is too large to scan whole on every query | --- # Sidecar Deployment Run the Spice Runtime in a separate container or process on the same machine as the main application. For example, in Kubernetes as a [Sidecar Container](https://kubernetes.io/docs/concepts/workloads/pods/sidecar-containers/). This approach minimizes communication overhead as requests to the Spice Runtime are transported over local loopback. ![sidecar](https://github.com/user-attachments/assets/716f7c23-1939-4947-85f5-b0ee2bbd63fc) **Benefits** * Low-latency communication between the application and the Spice Runtime. * Simplified lifecycle management (same pod). * Isolated environment without needing a separate microservice. * Helps ensure resiliency and redundancy by replicating data across sidecars. **Considerations** * Each application pod includes a copy of the Spice Runtime, increasing resource usage. * Updating the Spice Runtime independently requires updating each pod. * Accelerated data is replicated to each sidecar, adding resiliency and redundancy but increasing resource usage and requests to data sources. * May increase overall cost due to resource duplication. **Use This Approach When** * Fast, low-latency interactions between the application and the Spice Runtime are needed (e.g., real-time decision-making). * Scaling needs are small or moderate, making duplication of the Spice Runtime in each pod acceptable. * Keeping the architecture simple without additional services or load balancers is preferred. * Performance and latency are prioritized over cost and complexity. **Example Use Case** A real-time trading bot or a data-intensive application that relies on immediate feedback, where minimal latency is critical. Both containers in the same pod ensure very fast data exchange. --- # Tiered Deployment A hybrid approach combining sidecar deployments for performance-critical tasks and a shared microservice for batch processing or less time-sensitive workloads. ![tiered](https://github.com/user-attachments/assets/e602bad4-bd0d-4069-bc91-5b5678a10710) **Benefits** * Real-time responsiveness where needed (sidecar). * Centralized microservice handles broader or shared tasks. * Balances resource usage by limiting sidecar instances to high-priority operations. * Helps balance performance and latency with cost and complexity. **Considerations** * More complex deployment structure, mixing two patterns. * Must ensure consistent versioning between sidecar and microservice instances. * Potentially higher operational complexity and cost. **Use This Approach When** * Certain application components require ultra-low-latency responses, while others do not. * Centralized AI or analytics is needed, but localized real-time decision-making is also required. * The system can handle the operational complexity of running multiple deployment patterns. * Balancing performance and latency with cost and complexity is the goal. **Example Use Case** A logistics application that calculates routing decisions in real time (sidecar) while a microservice component processes aggregated data for periodic analysis or re-training models. --- # AWS Deployment Options Spice.ai provides multiple deployment options on Amazon Web Services (AWS), enabling data and AI applications to run on AWS's elastic infrastructure. Whether using virtual machines, container orchestration, or managed services, Spice deploys to meet requirements for performance, scalability, and cost efficiency. For a complete list of AWS-compatible data connectors, AI models, vector stores, and secret management, see [AWS Integrations](/docs/next/deployment/aws/integrations). ## Benefits of Deploying on AWS[​](#benefits-of-deploying-on-aws "Direct link to Benefits of Deploying on AWS") * **Scalability**: Easily scale your Spice.ai applications with AWS's elastic infrastructure. * **Global Reach**: Deploy across AWS's [worldwide regions](https://aws.amazon.com/about-aws/global-infrastructure/) for low-latency access. * **Integration**: Connect with other AWS services like [Amazon S3](https://aws.amazon.com/s3/), [Amazon RDS](https://aws.amazon.com/rds/), and [AWS Secrets Manager](https://aws.amazon.com/secrets-manager/). * **Cost Control**: Optimize expenses with various [instance types](https://aws.amazon.com/ec2/instance-types/) and [pricing models](https://aws.amazon.com/pricing/). * **Security and Compliance**: Deploy Spice.ai within your AWS security perimeter using features like [VPC](https://aws.amazon.com/vpc/) isolation, [security groups](https://docs.aws.amazon.com/vpc/latest/userguide/vpc-security-groups.html), [IAM roles](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles.html) to meet organizational compliance requirements. ## Deployment Options[​](#deployment-options "Direct link to Deployment Options") ### Amazon EKS (Elastic Kubernetes Service)[​](#amazon-eks-elastic-kubernetes-service "Direct link to Amazon EKS (Elastic Kubernetes Service)") Run Spice.ai on [Amazon EKS](https://aws.amazon.com/eks/) when the workload benefits from Kubernetes orchestration, multi-replica scale, declarative configuration, or shared cluster tenancy. EKS pairs well with the [Spice Helm chart](https://spiceai.org/docs/deployment/kubernetes/helm) and the [Argo CD](https://spiceai.org/docs/deployment/kubernetes/argocd) or [Flux](https://spiceai.org/docs/deployment/kubernetes/flux) GitOps workflows. #### 1. Provision the cluster[​](#1-provision-the-cluster "Direct link to 1. Provision the cluster") The fastest path is [`eksctl`](https://eksctl.io/), which provisions the VPC, IAM roles, and node groups in a single command: ``` eksctl create cluster \ --name spiceai-prod \ --region us-east-1 \ --version 1.31 \ --nodegroup-name workers \ --node-type m6i.xlarge \ --nodes 3 --nodes-min 2 --nodes-max 6 \ --managed \ --with-oidc ``` `--with-oidc` enables the [OIDC provider](https://docs.aws.amazon.com/eks/latest/userguide/enable-iam-roles-for-service-accounts.html) required for IAM Roles for Service Accounts (IRSA). For production, prefer Terraform or CloudFormation for repeatable provisioning. The community [`terraform-aws-modules/eks`](https://github.com/terraform-aws-modules/terraform-aws-eks) module is a common starting point. For burst or low-utilization workloads, attach an [EKS Fargate profile](https://docs.aws.amazon.com/eks/latest/userguide/fargate-profile.html) so Spice pods run on serverless capacity instead of managed nodes. #### 2. Configure IRSA for AWS access[​](#2-configure-irsa-for-aws-access "Direct link to 2. Configure IRSA for AWS access") Most Spice connectors (S3, DynamoDB, Bedrock, Glue) accept AWS credentials from the environment. Use IRSA so pods receive scoped, short-lived credentials without static keys: ``` # 1. Create an IAM policy with the permissions the Spicepod needs aws iam create-policy \ --policy-name SpiceAIRuntime \ --policy-document file://spiceai-policy.json # 2. Bind the policy to a Kubernetes ServiceAccount via IRSA eksctl create iamserviceaccount \ --name spiceai \ --namespace spiceai \ --cluster spiceai-prod \ --attach-policy-arn arn:aws:iam::123456789012:policy/SpiceAIRuntime \ --approve ``` Reference the service account from the Helm release so Spice pods inherit the role: ``` # values.yaml serviceAccount: create: false name: spiceai ``` For [EKS Pod Identity](https://docs.aws.amazon.com/eks/latest/userguide/pod-identities.html) (the newer alternative to IRSA), associate the role with `aws eks create-pod-identity-association` and skip the OIDC setup step. #### 3. Install Spice.ai[​](#3-install-spiceai "Direct link to 3. Install Spice.ai") ``` helm repo add spiceai https://helm.spiceai.org helm repo update helm upgrade --install spiceai spiceai/spiceai \ --namespace spiceai --create-namespace \ --version 1.11.5 \ -f values.yaml ``` For declarative GitOps, swap this command for an Argo CD `Application` or a Flux `HelmRelease` pointing at the same chart. See the [Argo CD](https://spiceai.org/docs/deployment/kubernetes/argocd) or [Flux](https://spiceai.org/docs/deployment/kubernetes/flux) guides for full manifests. #### 4. Storage and ingress[​](#4-storage-and-ingress "Direct link to 4. Storage and ingress") For stateful acceleration (DuckDB, SQLite, Cayenne): * **Local NVMe (recommended)** — Spice acceleration is latency- and IOPS-sensitive, so the lowest-latency option is a node-local NVMe SSD on an instance-store-backed family (`i4i`, `i7ie`, `m6id`, `m7gd`, `c7gd`, `r7gd` and other `d`-suffixed instances). Provision the [NVMe local volume](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/instance-store-volumes.html) with the [Local Volume Static Provisioner](https://github.com/kubernetes-sigs/sig-storage-local-static-provisioner) or use [Bottlerocket's `local-volume-provisioner`](https://github.com/bottlerocket-os/bottlerocket) to expose it as a `local-storage` StorageClass. Note that local volumes do not survive node replacement, so pair with a refresh strategy or a re-hydration source. * **Amazon EBS io2 Block Express** — when shared / replica-attachable persistence is required and node-local capacity is insufficient, [`io2`](https://docs.aws.amazon.com/ebs/latest/userguide/ebs-volume-types.html#io2-bx) delivers up to 256K IOPS and sub-millisecond latency. Use the [Amazon EBS CSI driver](https://docs.aws.amazon.com/eks/latest/userguide/ebs-csi.html) and a custom StorageClass with `type: io2` and a provisioned `iops` value. * **Amazon EBS gp3** — use `gp3` (with provisioned IOPS bumped above the 3,000 baseline) only when `io2` is unavailable in a region or when cost outweighs the latency improvement. * **Amazon S3 Express One Zone (Cayenne only)** — for Cayenne acceleration that needs to be shared across replicas or persisted independently of the pod lifecycle, [S3 Express One Zone](https://aws.amazon.com/s3/storage-classes/express-one-zone/) provides single-digit-millisecond latency single-AZ object storage. Configure Cayenne to point at an S3 Express directory bucket — see the [Cayenne acceleration documentation](/docs/next/components/data-accelerators/cayenne). * Set `stateful.enabled: true` and `stateful.storageClass: ` in `values.yaml`. tip Amazon EFS works for sharing data across replicas but is not recommended for accelerations: NFS-style latency negates the benefit of using an accelerator. Reserve EFS for stateless artifacts that need to survive pod replacement. Spice.ai Enterprise For production stateful workloads, the [Spice.ai Enterprise](https://spice.ai) Operator's [`SpicepodSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset) provides per-replica `StatefulSet`s with automatic PVC resizing, IRSA-aware ServiceAccount annotations, and configurable update strategies. For distributed query execution across scheduler/executor tiers backed by S3, see [`SpicepodCluster`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodcluster). To expose Spice externally, install the [AWS Load Balancer Controller](https://docs.aws.amazon.com/eks/latest/userguide/aws-load-balancer-controller.html) and front the Spice Service with a [Network Load Balancer](https://docs.aws.amazon.com/eks/latest/userguide/network-load-balancing.html): ``` # values.yaml service: type: LoadBalancer additionalAnnotations: service.beta.kubernetes.io/aws-load-balancer-type: external service.beta.kubernetes.io/aws-load-balancer-nlb-target-type: ip service.beta.kubernetes.io/aws-load-balancer-scheme: internal ``` #### 5. Observability[​](#5-observability "Direct link to 5. Observability") The Spice Helm chart ships a `PodMonitor` resource for the [Prometheus Operator](https://prometheus-operator.dev/). For EKS, the [`kube-prometheus-stack`](https://github.com/prometheus-community/helm-charts/tree/main/charts/kube-prometheus-stack) chart and [Amazon Managed Service for Prometheus](https://aws.amazon.com/prometheus/) are common targets. Set `monitoring.podMonitor.enabled: true` and import the [Spice Grafana dashboard](/docs/next/monitoring/grafana). For comprehensive guidance, refer to the [Amazon EKS User Guide](https://docs.aws.amazon.com/eks/latest/userguide/what-is-eks.html), the [EKS Best Practices Guide](https://aws.github.io/aws-eks-best-practices/), and the [Spice.ai Kubernetes Deployment Guide](https://spiceai.org/docs/deployment/kubernetes). ### EC2 / AWS CloudFormation[​](#ec2--aws-cloudformation "Direct link to EC2 / AWS CloudFormation") Deploy Spice.ai directly on [Amazon EC2](https://aws.amazon.com/ec2/) instances for maximum control over the environment. 1. **Manual EC2 Deployment**: * Launch an [EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/EC2_GetStarted.html) with your preferred Linux distribution * Install [Docker](https://docs.docker.com/engine/install/) * Run [Spice.ai as a Docker Container](https://spiceai.org/docs/deployment/docker#running-spiceai-as-a-docker-container) on your EC2 instance * (Optional) Use Infrastructure as Code (IaC) tools like [AWS CloudFormation](https://aws.amazon.com/cloudformation/) or [Terraform](https://www.terraform.io/) to automate the provisioning, configuration, and management of EC2 resources for repeatable and consistent deployments 2. **Automated EC2 Deployment with CloudFormation**: * Define your infrastructure in a [CloudFormation template](https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGuide/template-guide.html), including EC2 instances (using a [Linux AMI](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/AMIs.html)), security groups, IAM roles, VPC, and subnets * Use EC2 [`UserData`](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html) to automate Docker installation, pull the [Spice.ai Docker image](https://hub.docker.com/r/spiceai/spiceai), retrieve configuration or secrets from [AWS Parameter Store](https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-parameter-store.html) or [Secrets Manager](https://aws.amazon.com/secrets-manager/), and run the container with required environment variables * (Optional) Add [parameters](https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGuide/parameters-section-structure.html) to your template for VPC ID, Subnet ID, [KeyPair](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-key-pairs.html), instance type, and secret names to enable flexible deployments * (Optional) Store sensitive data such as API keys in [Parameter Store](https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-parameter-store.html) or [Secrets Manager](https://aws.amazon.com/secrets-manager/) and reference them securely in `UserData` * (Optional) Deploy and manage your CloudFormation stack using the [AWS Console](https://console.aws.amazon.com/cloudformation/), [CLI](https://docs.aws.amazon.com/cli/latest/reference/cloudformation/), or [CI/CD pipelines](https://aws.amazon.com/devops/continuous-delivery/) for repeatable, version-controlled infrastructure For detailed guidance and best practices, refer to the [AWS CloudFormation User Guide](https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGuide/Welcome.html), [EC2 User Guide for Linux Instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/), and [AWS Systems Manager Parameter Store Documentation](https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-parameter-store.html). ### Amazon ECS (Elastic Container Service)[​](#amazon-ecs-elastic-container-service "Direct link to Amazon ECS (Elastic Container Service)") Run Spice.ai on [Amazon ECS](https://aws.amazon.com/ecs/) when a single managed container is sufficient and operating Kubernetes is not desired. ECS Fargate provides serverless capacity; ECS on EC2 provides full control over the host. Both consume the same task definition. #### 1. Define the task[​](#1-define-the-task "Direct link to 1. Define the task") Create a task definition for the [`spiceai/spiceai` image](https://hub.docker.com/r/spiceai/spiceai), exposing port `8090` (HTTP) and, optionally, `50051` (Arrow Flight) and `9090` (Prometheus). Inject secrets from [AWS Secrets Manager](https://aws.amazon.com/secrets-manager/) or [SSM Parameter Store](https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-parameter-store.html) instead of baking them into the image. `spiceai-task.json`: ``` { "family": "spiceai", "networkMode": "awsvpc", "requiresCompatibilities": ["FARGATE"], "cpu": "1024", "memory": "2048", "executionRoleArn": "arn:aws:iam::123456789012:role/ecsTaskExecutionRole", "taskRoleArn": "arn:aws:iam::123456789012:role/SpiceAITaskRole", "containerDefinitions": [ { "name": "spiceai", "image": "spiceai/spiceai:1.11.5", "essential": true, "portMappings": [ { "containerPort": 8090, "protocol": "tcp" }, { "containerPort": 50051, "protocol": "tcp" } ], "environment": [ { "name": "SPICED_LOG", "value": "INFO" } ], "secrets": [ { "name": "SPICE_SECRET_SPICEAI_KEY", "valueFrom": "arn:aws:secretsmanager:us-east-1:123456789012:secret:spiceai/api-key" } ], "healthCheck": { "command": ["CMD-SHELL", "wget -q --spider http://localhost:8090/health || exit 1"], "interval": 10, "timeout": 3, "retries": 5, "startPeriod": 30 }, "logConfiguration": { "logDriver": "awslogs", "options": { "awslogs-group": "/ecs/spiceai", "awslogs-region": "us-east-1", "awslogs-stream-prefix": "spiceai" } } } ] } ``` Register the task: ``` aws ecs register-task-definition --cli-input-json file://spiceai-task.json ``` The `executionRoleArn` (typically `ecsTaskExecutionRole`) needs `secretsmanager:GetSecretValue` and `ssm:GetParameters` permissions to inject secrets. The `taskRoleArn` is the role Spice itself assumes at runtime — grant it the AWS permissions the Spicepod needs (for example, `s3:GetObject` on referenced buckets, `bedrock:InvokeModel` for Bedrock models). #### 2. Create the service[​](#2-create-the-service "Direct link to 2. Create the service") ``` aws ecs create-service \ --cluster spiceai-cluster \ --service-name spiceai \ --task-definition spiceai \ --launch-type FARGATE \ --desired-count 2 \ --network-configuration "awsvpcConfiguration={subnets=[subnet-aaa,subnet-bbb],securityGroups=[sg-xxx],assignPublicIp=DISABLED}" \ --load-balancers "targetGroupArn=arn:aws:elasticloadbalancing:us-east-1:123456789012:targetgroup/spiceai/abc,containerName=spiceai,containerPort=8090" \ --health-check-grace-period-seconds 60 ``` Front the service with a [Network Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/introduction.html) (low-latency TCP) or an [Application Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/application/introduction.html) (HTTP routing, TLS termination). For internal-only deployments, place the service in private subnets and set `assignPublicIp=DISABLED`. #### 3. Persistent storage[​](#3-persistent-storage "Direct link to 3. Persistent storage") Spice accelerations are latency- and IOPS-sensitive. Choose the storage type based on launch type and sharing requirements: * **ECS on EC2 with local NVMe (recommended for accelerations)** \u2014 launch the cluster on an instance-store-backed family (`i4i`, `i7ie`, `m6id`, `m7gd`, `c7gd`, `r7gd`, etc.) and bind-mount the NVMe device into the task. This delivers the lowest latency and highest IOPS available on AWS but does not survive instance replacement, so pair with a refresh strategy.\n- **Amazon EBS volume attached to an ECS service (EC2 launch type)** \u2014 use the [EBS volume task configuration](https://docs.aws.amazon.com/AmazonECS/latest/developerguide/ebs-volumes.html) with `volumeType: io2` for high-IOPS, low-latency block storage that survives task restarts. Fall back to `gp3` (with provisioned IOPS) when `io2` is unavailable in the region.\n- **Amazon S3 Express One Zone (Cayenne only)** \u2014 for Cayenne acceleration that needs to be shared across tasks or persisted independently of task lifecycle, [S3 Express One Zone](https://aws.amazon.com/s3/storage-classes/express-one-zone/) provides single-digit-millisecond latency. Configure Cayenne against an S3 Express directory bucket \u2014 see the [Cayenne acceleration documentation](/docs/next/components/data-accelerators/cayenne).\n- **Amazon EFS (Fargate-only fallback)** \u2014 EFS is the only persistent storage option supported by Fargate, but its NFS-style latency is not recommended for accelerations. Use it only for stateless artefacts that must survive task replacement, or switch to the EC2 launch type when low-latency local storage is required. ``` "volumes": [ { "name": "spice-data", "efsVolumeConfiguration": { "fileSystemId": "fs-0123456789abcdef0", "rootDirectory": "/spiceai", "transitEncryption": "ENABLED" } } ], "containerDefinitions": [ { "name": "spiceai", "mountPoints": [ { "sourceVolume": "spice-data", "containerPath": "/data" } ] } ] ``` In the Spicepod, point file accelerators at `/data`, for example `duckdb_file: /data/taxi_trips.db`. #### 4. Auto-scaling[​](#4-auto-scaling "Direct link to 4. Auto-scaling") Configure [service auto-scaling](https://docs.aws.amazon.com/AmazonECS/latest/developerguide/service-auto-scaling.html) on average CPU or on custom CloudWatch metrics derived from the Spice `/v1/metrics` endpoint: ``` aws application-autoscaling register-scalable-target \ --service-namespace ecs \ --resource-id service/spiceai-cluster/spiceai \ --scalable-dimension ecs:service:DesiredCount \ --min-capacity 2 --max-capacity 10 ``` For comprehensive details, see the [Amazon ECS Developer Guide](https://docs.aws.amazon.com/ecs/latest/developerguide/Welcome.html) and the [Spice.ai Docker Deployment Guide](https://spiceai.org/docs/deployment/docker). ## Authentication[​](#authentication "Direct link to Authentication") Most AWS services that Spice connects to have explicit parameters for configuring authentication (usually by setting an `access_key_id` and `secret_access_key`). If explicit credentials are not provided, Spice follows the standard AWS SDK behavior for loading credentials from the environment based on the following sources in order: 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate permissions (e.g., `s3:ListBucket`, `s3:GetObject`) to access the service. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Resources[​](#resources "Direct link to Resources") ### Documentation[​](#documentation "Direct link to Documentation") * [AWS Integrations](/docs/next/deployment/aws/integrations) - Complete list of AWS data connectors, AI models, vector stores, and secrets * [AWS Secrets Manager Secret Store](/docs/next/components/secret-stores/aws-secrets-manager) ### AWS Blog Posts[​](#aws-blog-posts "Direct link to AWS Blog Posts") * [Architecting High-Performance AI-Driven Data Applications with Spice.ai and AWS](https://aws.amazon.com/blogs/storage/architecting-high-performance-ai-driven-data-applications-with-spice-ai-and-aws/) - AWS Storage Blog ### Spice.ai Blog Posts[​](#spiceai-blog-posts "Direct link to Spice.ai Blog Posts") * [Amazon S3 Vectors](https://spice.ai/blog/amazon-s3-vectors) - Overview of S3 Vectors integration * [Getting Started with Amazon S3 Vectors and Spice](https://spice.ai/blog/getting-started-with-amazon-s3-vectors-and-spice) - Step-by-step tutorial ### Videos[​](#videos "Direct link to Videos") * [Getting started with Amazon S3 Vectors and Spice](https://www.youtube.com/watch?v=KuWI0yDOnIU) - YouTube walkthrough ### Marketplace[​](#marketplace "Direct link to Marketplace") * [Spice.ai on AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-jmf6jskjvnq7i) - Deploy Spice.ai from AWS Marketplace --- # AWS Integrations ![Spice.ai and AWS](/assets/images/aws-spice-a361b79fcd229ce942248cd4a48f2cff.png) Spice.ai provides deep integrations with Amazon Web Services (AWS), enabling data federation, AI inference, vector search, and secure secret management across the AWS ecosystem. This page consolidates all AWS-compatible components and provides quick access to configuration guides. ## Data Connectors[​](#data-connectors "Direct link to Data Connectors") Data connectors federate SQL queries across AWS data sources without data movement. | Connector | Description | Documentation | | ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- | | **Amazon S3** | Query Parquet, CSV, and JSON files stored in S3 buckets. Supports private buckets with IAM authentication and S3-compatible storage like MinIO. | [S3 Data Connector](/docs/next/components/data-connectors/s3) | | **Amazon S3 Tables** | Query Iceberg tables in [Amazon S3 Tables](https://aws.amazon.com/s3/features/tables/) using the Glue connector with S3 Tables catalog format. | [Glue Data Connector](/docs/next/components/data-connectors/glue) | | **Amazon DynamoDB** | Federated SQL queries on DynamoDB tables with automatic schema inference. | [DynamoDB Data Connector](/docs/next/components/data-connectors/dynamodb) | | **Amazon DynamoDB Streams** | Real-time CDC streaming of table changes via [DynamoDB Streams](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.html). | [DynamoDB Data Connector](/docs/next/components/data-connectors/dynamodb) | | **Amazon Redshift** | Connect to Redshift clusters using the PostgreSQL-compatible connector. | [Redshift Data Connector](/docs/next/components/data-connectors/redshift) | | **Amazon Aurora PostgreSQL** | Connect to Aurora PostgreSQL clusters using the PostgreSQL connector. | [PostgreSQL Data Connector](/docs/next/components/data-connectors/postgres) | | **Amazon Aurora MySQL** | Connect to Aurora MySQL clusters using the MySQL connector. | [MySQL Data Connector](/docs/next/components/data-connectors/mysql) | | **Amazon RDS PostgreSQL** | Connect to RDS PostgreSQL instances using the PostgreSQL connector. | [PostgreSQL Data Connector](/docs/next/components/data-connectors/postgres) | | **Amazon RDS MySQL** | Connect to RDS MySQL instances using the MySQL connector. | [MySQL Data Connector](/docs/next/components/data-connectors/mysql) | | **Amazon MSK** | Stream data from [Amazon MSK](https://aws.amazon.com/msk/) (Managed Streaming for Apache Kafka) topics using the Kafka connector. | [Kafka Data Connector](/docs/next/components/data-connectors/kafka) | | **Debezium (Amazon MSK)** | Change Data Capture (CDC) from databases via Debezium running on Amazon MSK for real-time dataset updates. | [Debezium Data Connector](/docs/next/components/data-connectors/debezium) | | **AWS Glue Data Catalog** | Query Iceberg tables registered in AWS Glue. | [Glue Data Connector](/docs/next/components/data-connectors/glue) | | **Apache Iceberg (AWS)** | Query Iceberg tables stored in S3 with Glue or REST catalog metadata. | [Iceberg Data Connector](/docs/next/components/data-connectors/iceberg) | | **Delta Lake (S3)** | Query Delta Lake tables stored in Amazon S3. | [Delta Lake Data Connector](/docs/next/components/data-connectors/delta-lake) | | **AWS Athena (ODBC)** | Connect to Athena using the ODBC connector with Athena SQL dialect support. | [ODBC Data Connector](/docs/next/components/data-connectors/odbc) | ### Example: Amazon S3[​](#example-amazon-s3 "Direct link to Example: Amazon S3") ``` datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet s3_region: us-east-1 s3_auth: iam_role # Uses IAM credentials from environment ``` ### Example: DynamoDB[​](#example-dynamodb "Direct link to Example: DynamoDB") ``` datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 ``` ### Example: AWS Glue with Amazon S3 Tables[​](#example-aws-glue-with-amazon-s3-tables "Direct link to Example: AWS Glue with Amazon S3 Tables") ``` datasets: - from: glue:my_namespace.orders name: orders params: glue_catalog_id: 123635965758:s3tablescatalog/my-table-bucket glue_region: us-east-2 ``` ## Catalog Connectors[​](#catalog-connectors "Direct link to Catalog Connectors") Catalog connectors provide schema discovery and unified access to tables in AWS data catalogs. | Connector | Description | Documentation | | -------------------- | --------------------------------------------------------------------------------- | ------------------------------------------------------------- | | **AWS Glue Catalog** | Discover and query tables from AWS Glue Data Catalog with glob pattern filtering. | [Glue Catalog Connector](/docs/next/components/catalogs/glue) | ### Example: Glue Catalog[​](#example-glue-catalog "Direct link to Example: Glue Catalog") ``` catalogs: - from: glue name: my_data_lake include: - '*.*' # Include all tables from all databases params: glue_region: us-east-1 ``` ## AI Models (Amazon Bedrock)[​](#ai-models-amazon-bedrock "Direct link to AI Models (Amazon Bedrock)") Spice integrates with [Amazon Bedrock](https://aws.amazon.com/bedrock/) for large language model inference, supporting Amazon Nova and other foundation models. | Provider | Supported Models | Documentation | | ------------------ | ------------------------------------------------------------------------ | ------------------------------------------------------ | | **Amazon Bedrock** | Amazon Nova (Micro, Lite, Pro, Premier), cross-region inference profiles | [Bedrock Models](/docs/next/components/models/bedrock) | ### Example: Amazon Nova[​](#example-amazon-nova "Direct link to Example: Amazon Nova") ``` models: - from: bedrock:us.amazon.nova-lite-v1:0 name: nova params: aws_region: us-east-1 ``` ### Guardrails Support[​](#guardrails-support "Direct link to Guardrails Support") Bedrock Guardrails can filter model inputs and outputs: ``` models: - from: bedrock:amazon.nova-pro-v1:0 name: nova-guarded params: aws_region: us-east-1 bedrock_guardrail_identifier: arn:aws:bedrock:us-east-1:123456789012:guardrail/abc123 bedrock_guardrail_version: '1' ``` ## Embeddings (Amazon Bedrock)[​](#embeddings-amazon-bedrock "Direct link to Embeddings (Amazon Bedrock)") Generate vector embeddings using Amazon Bedrock embedding models for semantic search and RAG applications. | Provider | Supported Models | Documentation | | ------------------ | ------------------------------------------------------------------------ | -------------------------------------------------------------- | | **Amazon Bedrock** | Amazon Titan Embeddings, Amazon Nova Multimodal Embeddings, Cohere Embed | [Bedrock Embeddings](/docs/next/components/embeddings/bedrock) | ### Example: Amazon Titan Embeddings[​](#example-amazon-titan-embeddings "Direct link to Example: Amazon Titan Embeddings") ``` embeddings: - from: bedrock:amazon.titan-embed-text-v2:0 name: titan params: aws_region: us-east-1 dimensions: '256' ``` ### Example: Amazon Nova Multimodal Embeddings[​](#example-amazon-nova-multimodal-embeddings "Direct link to Example: Amazon Nova Multimodal Embeddings") ``` embeddings: - from: bedrock:amazon.nova-2-multimodal-embeddings-v1:0 name: nova_embed params: dimensions: '1024' truncation_mode: START embedding_purpose: GENERIC_RETRIEVAL aws_region: us-east-1 ``` ## Vector Stores (Amazon S3 Vectors)[​](#vector-stores-amazon-s3-vectors "Direct link to Vector Stores (Amazon S3 Vectors)") [Amazon S3 Vectors](https://aws.amazon.com/s3/features/s3-vectors/) is a new S3 bucket type for storing and querying vector embeddings at scale. Spice integrates S3 Vectors as a vector index backend for hybrid search applications. | Engine | Description | Documentation | | --------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- | | **Amazon S3 Vectors** | Sub-second similarity queries on billions of vectors with up to 90% cost reduction compared to traditional vector databases. | [S3 Vectors Engine](/docs/next/components/vectors/s3_vectors) | ### Example: S3 Vectors with Bedrock Embeddings[​](#example-s3-vectors-with-bedrock-embeddings "Direct link to Example: S3 Vectors with Bedrock Embeddings") ``` datasets: - from: oracle:"CUSTOMER_REVIEWS" name: reviews vectors: enabled: true engine: s3_vectors params: s3_vectors_bucket: my-s3-vector-bucket s3_vectors_aws_region: us-east-1 columns: - name: body embeddings: from: bedrock_titan embeddings: - from: bedrock:amazon.titan-embed-text-v2:0 name: bedrock_titan params: aws_region: us-east-1 dimensions: '256' ``` ## Data Accelerators (S3 Express One Zone)[​](#data-accelerators-s3-express-one-zone "Direct link to Data Accelerators (S3 Express One Zone)") Spice Cayenne data accelerator supports [AWS S3 Express One Zone](https://aws.amazon.com/s3/storage-classes/express-one-zone/) for storing accelerated data with single-digit millisecond latency. This is ideal for latency-sensitive query workloads that require persistent storage while maintaining fast access. Storage Recommendation For best performance, store Cayenne data files on local NVMe storage. Use S3 Express One Zone only when persistence of accelerations is required, such as preserving accelerated data across restarts or sharing data between multiple Spice instances. | Accelerator | Description | Documentation | | ----------------- | --------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | | **Spice Cayenne** | High-performance data accelerator using Vortex file format with S3 Express One Zone for sub-10ms latency query performance. | [Cayenne Accelerator](/docs/next/components/data-accelerators/cayenne) | ### Why S3 Express One Zone?[​](#why-s3-express-one-zone "Direct link to Why S3 Express One Zone?") S3 Express One Zone directory buckets provide: * **Single-digit millisecond latency**: 10x faster than S3 Standard for first-byte latency * **High request throughput**: Up to 10x higher request rates than S3 Standard * **Cost efficiency**: Lower per-request costs for high-frequency access patterns * **Durability**: Same 99.999999999% (11 9s) durability as S3 Standard ### Example: Cayenne with S3 Express One Zone[​](#example-cayenne-with-s3-express-one-zone "Direct link to Example: Cayenne with S3 Express One Zone") ``` datasets: - from: s3://source-bucket/events/ name: analytics_events acceleration: engine: cayenne enabled: true mode: file params: # Store accelerated data in S3 Express One Zone bucket cayenne_file_path: s3://my-bucket--usw2-az1--x-s3/cayenne/ cayenne_s3_region: us-west-2 ``` ### Example: Auto-generated Bucket with IAM Role[​](#example-auto-generated-bucket-with-iam-role "Direct link to Example: Auto-generated Bucket with IAM Role") ``` datasets: - from: postgresql://db/events name: fast_events acceleration: engine: cayenne enabled: true mode: file params: # Auto-generates bucket: spice-{spicepod-name}-fast_events--usw2-az1--x-s3 cayenne_s3_zone_ids: usw2-az1 ``` ### Supported AWS Regions[​](#supported-aws-regions "Direct link to Supported AWS Regions") S3 Express One Zone is available in select regions. Spice automatically derives the region from zone IDs: | Zone ID Prefix | Region | | -------------- | -------------- | | `use1` | us-east-1 | | `use2` | us-east-2 | | `usw1` | us-west-1 | | `usw2` | us-west-2 | | `euw1` | eu-west-1 | | `euc1` | eu-central-1 | | `apne1` | ap-northeast-1 | | `apse1` | ap-southeast-1 | See AWS documentation for the complete list of [S3 Express One Zone availability zones](https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-express-Regions-and-Zones.html). ## Secret Management[​](#secret-management "Direct link to Secret Management") Securely store and retrieve credentials using AWS Secrets Manager. | Store | Description | Documentation | | ----------------------- | ----------------------------------------------------- | ------------------------------------------------------------------------------ | | **AWS Secrets Manager** | Read secrets from AWS Secrets Manager by secret name. | [AWS Secrets Manager](/docs/next/components/secret-stores/aws-secrets-manager) | ### Example: Using Secrets Manager[​](#example-using-secrets-manager "Direct link to Example: Using Secrets Manager") ``` secrets: - from: aws_secrets_manager:my_database_creds name: db datasets: - from: postgres:public.users name: users params: pg_host: ${db:host} pg_user: ${db:username} pg_pass: ${db:password} ``` ## Authentication[​](#authentication "Direct link to Authentication") All AWS integrations support the standard AWS SDK credential chain. When credentials are not explicitly configured, Spice loads them from the following sources in order: 1. **Environment Variables**: `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN` 2. **Shared Credentials Files**: `~/.aws/credentials` and `~/.aws/config` 3. **AWS SSO Sessions**: Configured via `aws configure sso` 4. **Web Identity Token**: For OIDC/OAuth (common with EKS IRSA) 5. **ECS Container Credentials**: Automatic IAM role for ECS tasks 6. **EC2 Instance Metadata (IMDSv2)**: Automatic IAM role for EC2 instances ### IAM Permissions[​](#iam-permissions "Direct link to IAM Permissions") Ensure the IAM role or user has appropriate permissions for all AWS services used: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:GetObject", "s3:ListBucket", "dynamodb:Scan", "dynamodb:DescribeTable", "glue:GetTable", "glue:GetTables", "glue:GetDatabase", "glue:GetDatabases", "bedrock:InvokeModel", "secretsmanager:GetSecretValue" ], "Resource": "*" } ] } ``` ## Deployment Options[​](#deployment-options "Direct link to Deployment Options") Deploy Spice on AWS infrastructure for optimal performance and integration: | Option | Description | Documentation | | -------------- | --------------------------------------------------- | -------------------------------------------- | | **Amazon EKS** | Kubernetes orchestration with Helm chart deployment | [AWS Deployment](/docs/next/deployment/aws/) | | **Amazon ECS** | Container service with Fargate or EC2 launch types | [AWS Deployment](/docs/next/deployment/aws/) | | **Amazon EC2** | Direct deployment with Docker or binary | [AWS Deployment](/docs/next/deployment/aws/) | ## Resources[​](#resources "Direct link to Resources") ### AWS Blog Posts[​](#aws-blog-posts "Direct link to AWS Blog Posts") * [Architecting High-Performance AI-Driven Data Applications with Spice.ai and AWS](https://aws.amazon.com/blogs/storage/architecting-high-performance-ai-driven-data-applications-with-spice-ai-and-aws/) - AWS Storage Blog ### Spice.ai Blog Posts[​](#spiceai-blog-posts "Direct link to Spice.ai Blog Posts") * [Amazon S3 Vectors](https://spice.ai/blog/amazon-s3-vectors) - Overview of S3 Vectors integration * [Getting Started with Amazon S3 Vectors and Spice](https://spice.ai/blog/getting-started-with-amazon-s3-vectors-and-spice) - Step-by-step tutorial ### Videos[​](#videos "Direct link to Videos") * [Getting started with Amazon S3 Vectors and Spice](https://www.youtube.com/watch?v=KuWI0yDOnIU) - YouTube walkthrough * [How Spice AI operationalizes data lakes for AI using Amazon S3](https://www.youtube.com/watch?v=KuWI0yDOnIU\&list=PLesJrUXEx3U-WIqfWYfha4zBkyZo9czEJ\&index=2) - Spice presentation at re:Invent ### Marketplace[​](#marketplace "Direct link to Marketplace") * [Spice.ai on AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-jmf6jskjvnq7i) - Deploy Spice.ai from AWS Marketplace ## Quick Start[​](#quick-start "Direct link to Quick Start") Get started with Spice on AWS in minutes: 1. **Install Spice CLI**: ``` curl https://install.spiceai.org | /bin/bash ``` 2. **Configure AWS credentials**: ``` aws configure ``` 3. **Create a Spicepod with S3 data**: ``` # spicepod.yaml version: v1 kind: Spicepod name: aws_quickstart datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet s3_auth: iam_role ``` 4. **Start the runtime**: ``` spice run ``` 5. **Query your data**: ``` spice sql > SELECT COUNT(*) FROM taxi_trips; ``` --- # Azure Deployment Options Spice.ai provides multiple deployment options on Microsoft Azure, enabling data and AI applications to run on Azure's global infrastructure. Whether using virtual machines, container orchestration, or serverless containers, Spice deploys to meet requirements for performance, scalability, and cost efficiency. For a complete list of Azure-compatible data connectors, AI models, and integrations, see [Azure Integrations](/docs/next/deployment/azure/integrations). ## Benefits of Deploying on Azure[​](#benefits-of-deploying-on-azure "Direct link to Benefits of Deploying on Azure") * **Scalability**: Scale Spice.ai workloads with [virtual machine scale sets](https://learn.microsoft.com/azure/virtual-machine-scale-sets/), [AKS](https://azure.microsoft.com/products/kubernetes-service), and [Container Apps](https://azure.microsoft.com/products/container-apps). * **Global Reach**: Deploy across Azure's [worldwide regions](https://azure.microsoft.com/explore/global-infrastructure/geographies/) for low-latency access. * **Integration**: Connect to other Azure services such as [Azure Blob Storage](https://azure.microsoft.com/products/storage/blobs), [Azure SQL Database](https://azure.microsoft.com/products/azure-sql/database/), [Azure Database for PostgreSQL](https://azure.microsoft.com/products/postgresql/), and [Azure Key Vault](https://azure.microsoft.com/products/key-vault). * **Cost Control**: Choose from [VM sizes](https://learn.microsoft.com/azure/virtual-machines/sizes), [reserved instances](https://azure.microsoft.com/pricing/reserved-vm-instances/), and [spot pricing](https://azure.microsoft.com/products/virtual-machines/spot/) to match workload requirements. * **Security and Compliance**: Run Spice.ai inside an Azure security perimeter using [Virtual Networks](https://azure.microsoft.com/products/virtual-network), [private endpoints](https://learn.microsoft.com/azure/private-link/private-endpoint-overview), [Microsoft Entra ID](https://learn.microsoft.com/entra/fundamentals/), and [managed identities](https://learn.microsoft.com/entra/identity/managed-identities-azure-resources/overview). ## Deployment Options[​](#deployment-options "Direct link to Deployment Options") ### Azure Kubernetes Service (AKS)[​](#azure-kubernetes-service-aks "Direct link to Azure Kubernetes Service (AKS)") Run Spice.ai on [Azure Kubernetes Service](https://azure.microsoft.com/products/kubernetes-service) when the workload benefits from Kubernetes orchestration, multi-replica scale, declarative configuration, or shared cluster tenancy. AKS pairs well with the [Spice Helm chart](https://spiceai.org/docs/deployment/kubernetes/helm) and the [Argo CD](https://spiceai.org/docs/deployment/kubernetes/argocd) or [Flux](https://spiceai.org/docs/deployment/kubernetes/flux) GitOps workflows. #### 1. Provision the cluster[​](#1-provision-the-cluster "Direct link to 1. Provision the cluster") Provision an AKS cluster with [workload identity](https://learn.microsoft.com/azure/aks/workload-identity-overview) and the [OIDC issuer](https://learn.microsoft.com/azure/aks/cluster-configuration#oidc-issuer) enabled — both are required for federated credentials to Azure services. ``` RG=spiceai-rg CLUSTER=spiceai-prod LOCATION=eastus az group create --name $RG --location $LOCATION az aks create \ --resource-group $RG \ --name $CLUSTER \ --location $LOCATION \ --kubernetes-version 1.31 \ --node-count 3 \ --node-vm-size Standard_D4s_v5 \ --enable-cluster-autoscaler --min-count 2 --max-count 6 \ --enable-oidc-issuer \ --enable-workload-identity \ --network-plugin azure \ --generate-ssh-keys az aks get-credentials --resource-group $RG --name $CLUSTER ``` For burst or low-utilization workloads, attach [virtual nodes](https://learn.microsoft.com/azure/aks/virtual-nodes) backed by Azure Container Instances. For production, prefer Bicep or Terraform for repeatable provisioning — the [Azure Verified Modules](https://azure.github.io/Azure-Verified-Modules/) library publishes a maintained AKS module. #### 2. Configure workload identity for Azure access[​](#2-configure-workload-identity-for-azure-access "Direct link to 2. Configure workload identity for Azure access") Most Spice connectors (ABFS, Azure SQL, Key Vault, Azure OpenAI) accept Azure credentials from the environment. Use [workload identity](https://learn.microsoft.com/azure/aks/workload-identity-overview) so pods receive scoped, federated tokens without static secrets: ``` # 1. Create a user-assigned managed identity az identity create --resource-group $RG --name spiceai-identity CLIENT_ID=$(az identity show -g $RG -n spiceai-identity --query clientId -o tsv) PRINCIPAL_ID=$(az identity show -g $RG -n spiceai-identity --query principalId -o tsv) # 2. Grant the identity access to Azure resources the Spicepod needs az role assignment create \ --assignee-object-id $PRINCIPAL_ID --assignee-principal-type ServicePrincipal \ --role "Storage Blob Data Reader" \ --scope /subscriptions//resourceGroups/$RG/providers/Microsoft.Storage/storageAccounts/ # 3. Federate the identity with the Kubernetes ServiceAccount ISSUER=$(az aks show -g $RG -n $CLUSTER --query oidcIssuerProfile.issuerUrl -o tsv) az identity federated-credential create \ --name spiceai-fed \ --identity-name spiceai-identity \ --resource-group $RG \ --issuer "$ISSUER" \ --subject system:serviceaccount:spiceai:spiceai \ --audiences api://AzureADTokenExchange ``` Reference the identity from the Helm release so Spice pods inherit federated tokens via the [DefaultAzureCredential](https://learn.microsoft.com/dotnet/api/azure.identity.defaultazurecredential) chain: ``` # values.yaml serviceAccount: create: true name: spiceai annotations: azure.workload.identity/client-id: "" podLabels: azure.workload.identity/use: "true" ``` #### 3. Install Spice.ai[​](#3-install-spiceai "Direct link to 3. Install Spice.ai") ``` helm repo add spiceai https://helm.spiceai.org helm repo update helm upgrade --install spiceai spiceai/spiceai \ --namespace spiceai --create-namespace \ --version 1.11.5 \ -f values.yaml ``` For declarative GitOps, swap this command for an Argo CD `Application` or a Flux `HelmRelease` pointing at the same chart. See the [Argo CD](https://spiceai.org/docs/deployment/kubernetes/argocd) or [Flux](https://spiceai.org/docs/deployment/kubernetes/flux) guides for full manifests. #### 4. Storage and ingress[​](#4-storage-and-ingress "Direct link to 4. Storage and ingress") For stateful acceleration (DuckDB, SQLite, Cayenne): * **Local NVMe (recommended)** — Spice acceleration is latency- and IOPS-sensitive, so the lowest-latency option is a node-local NVMe SSD on an instance family with attached NVMe ([Lsv3 / Lasv3](https://learn.microsoft.com/azure/virtual-machines/lsv3-series), [Ddsv5 / Ddsv6](https://learn.microsoft.com/azure/virtual-machines/ddv5-ddsv5-series), [Edsv5 / Edsv6](https://learn.microsoft.com/azure/virtual-machines/edv5-edsv5-series)). Expose the local NVMe through the [Local Volume Static Provisioner](https://github.com/kubernetes-sigs/sig-storage-local-static-provisioner) as a `local-storage` StorageClass. Local NVMe does not survive node replacement, so pair with a refresh strategy or a re-hydration source. * **Premium SSD v2** — when shared / replica-attachable persistence is required, [Premium SSD v2](https://learn.microsoft.com/azure/virtual-machines/disks-types#premium-ssd-v2) delivers up to 80,000 IOPS and sub-millisecond latency with independently configurable IOPS and throughput. Use the [Azure Disks CSI driver](https://learn.microsoft.com/azure/aks/azure-disk-csi) with a custom StorageClass (`skuName: PremiumV2_LRS`). * **Premium SSD (`managed-csi-premium`)** — use the built-in `managed-csi-premium` storage class only when Premium SSD v2 is unavailable in a region. * **Azure Files (`azurefile-csi`) — not recommended for acceleration** — use only for stateless shared artefacts that need `ReadWriteMany`. SMB/NFS latency negates the benefit of using a local accelerator. * Set `stateful.enabled: true` and `stateful.storageClass: ` in `values.yaml`. Spice.ai Enterprise For production stateful workloads, the [Spice.ai Enterprise](https://spice.ai) Operator's [`SpicepodSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset) provides per-replica `StatefulSet`s with automatic PVC resizing, workload-identity-aware ServiceAccount annotations, and configurable update strategies. For distributed query execution across scheduler/executor tiers backed by Azure Blob Storage, see [`SpicepodCluster`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodcluster). To expose Spice externally, install the [Application Gateway Ingress Controller (AGIC)](https://learn.microsoft.com/azure/application-gateway/ingress-controller-overview) or use a [Standard public Load Balancer](https://learn.microsoft.com/azure/aks/load-balancer-standard): ``` # values.yaml service: type: LoadBalancer additionalAnnotations: service.beta.kubernetes.io/azure-load-balancer-internal: "true" # internal only ``` For internal-only deployments, set `azure-load-balancer-internal: "true"` to bind to the cluster's VNet rather than a public IP. #### 5. Observability[​](#5-observability "Direct link to 5. Observability") The Spice Helm chart ships a `PodMonitor` resource for the [Prometheus Operator](https://prometheus-operator.dev/). On AKS, the [Azure Monitor managed service for Prometheus](https://learn.microsoft.com/azure/azure-monitor/essentials/prometheus-metrics-overview) and [Container insights](https://learn.microsoft.com/azure/azure-monitor/containers/container-insights-overview) are the common targets. Set `monitoring.podMonitor.enabled: true` and import the [Spice Grafana dashboard](/docs/next/monitoring/grafana) into [Azure Managed Grafana](https://azure.microsoft.com/products/managed-grafana). For comprehensive guidance, refer to the [Azure Kubernetes Service documentation](https://learn.microsoft.com/azure/aks/), the [AKS baseline architecture](https://learn.microsoft.com/azure/architecture/reference-architectures/containers/aks/baseline-aks), and the [Spice.ai Kubernetes Deployment Guide](https://spiceai.org/docs/deployment/kubernetes). ### Azure Container Apps[​](#azure-container-apps "Direct link to Azure Container Apps") [Azure Container Apps](https://azure.microsoft.com/products/container-apps) is a serverless container platform suitable for HTTP-driven Spice.ai workloads that benefit from scale-to-zero and request-based autoscaling. Use it when a single managed container is sufficient and operating Kubernetes is not desired. #### 1. Create the environment[​](#1-create-the-environment "Direct link to 1. Create the environment") The environment is the security and networking boundary that hosts one or more container apps: ``` RG=spiceai-rg ENV=spiceai-env LOCATION=eastus az group create --name $RG --location $LOCATION az containerapp env create \ --name $ENV \ --resource-group $RG \ --location $LOCATION \ --logs-destination log-analytics ``` To reach Azure SQL, Storage, or Key Vault behind private endpoints, attach the environment to a [VNet-injected subnet](https://learn.microsoft.com/azure/container-apps/networking) by adding `--infrastructure-subnet-resource-id` and `--internal-only true`. #### 2. Configure managed identity[​](#2-configure-managed-identity "Direct link to 2. Configure managed identity") Container Apps support both [system-assigned and user-assigned managed identities](https://learn.microsoft.com/azure/container-apps/managed-identity). A user-assigned identity is preferred so role assignments survive app recreation: ``` az identity create --resource-group $RG --name spiceai-identity IDENTITY_ID=$(az identity show -g $RG -n spiceai-identity --query id -o tsv) PRINCIPAL_ID=$(az identity show -g $RG -n spiceai-identity --query principalId -o tsv) # Grant access to the resources the Spicepod connects to az role assignment create \ --assignee-object-id $PRINCIPAL_ID --assignee-principal-type ServicePrincipal \ --role "Storage Blob Data Reader" \ --scope /subscriptions//resourceGroups/$RG/providers/Microsoft.Storage/storageAccounts/ ``` #### 3. Deploy Spice.ai[​](#3-deploy-spiceai "Direct link to 3. Deploy Spice.ai") Mount [Azure Files](https://learn.microsoft.com/azure/container-apps/storage-mounts-azure-files) for stateful acceleration, inject secrets from [Key Vault](https://learn.microsoft.com/azure/container-apps/manage-secrets#reference-secret-from-key-vault), and configure HTTP ingress on port `8090`: ``` az containerapp create \ --name spiceai \ --resource-group $RG \ --environment $ENV \ --image spiceai/spiceai:2.0.0 \ --target-port 8090 \ --ingress external \ --transport http \ --user-assigned $IDENTITY_ID \ --min-replicas 1 --max-replicas 5 \ --cpu 1.0 --memory 2.0Gi \ --env-vars \ SPICED_LOG=INFO \ AZURE_CLIENT_ID=$(az identity show -g $RG -n spiceai-identity --query clientId -o tsv) \ --secrets spiceai-key=keyvaultref:https://my-vault.vault.azure.net/secrets/spiceai-key,identityref:$IDENTITY_ID \ --secret-volume-mount /mnt/secrets ``` To run multiple replicas with shared file-based acceleration, define an [Azure Files storage mount](https://learn.microsoft.com/azure/container-apps/storage-mounts-azure-files) and reference it from the app's revision template, then point file accelerators at the mount path (for example, `duckdb_file: /data/taxi_trips.db`). #### 4. Scaling rules[​](#4-scaling-rules "Direct link to 4. Scaling rules") Spice.ai is HTTP-driven, so the default [HTTP scale rule](https://learn.microsoft.com/azure/container-apps/scale-app) (concurrent requests per replica) is usually sufficient. For background workloads (refresh schedules, ingestion) that should not scale to zero, set `--min-replicas 1` and add a [custom scale rule](https://learn.microsoft.com/azure/container-apps/scale-app#custom) backed by a CPU or queue metric: ``` az containerapp update \ --name spiceai --resource-group $RG \ --scale-rule-name http-rule \ --scale-rule-type http \ --scale-rule-http-concurrency 50 ``` #### 5. Health probes and revisions[​](#5-health-probes-and-revisions "Direct link to 5. Health probes and revisions") Configure the [liveness and readiness probes](https://learn.microsoft.com/azure/container-apps/health-probes) to use `/health` and `/v1/ready`. Container Apps creates a new [revision](https://learn.microsoft.com/azure/container-apps/revisions) on each `update`, supporting traffic splitting between revisions for canary upgrades: ``` az containerapp revision set-mode --name spiceai --resource-group $RG --mode multiple az containerapp ingress traffic set --name spiceai --resource-group $RG \ --revision-weight latest=90 spiceai--prev=10 ``` For more details, see the [Azure Container Apps documentation](https://learn.microsoft.com/azure/container-apps/) and the [Spice.ai Docker Deployment Guide](https://spiceai.org/docs/deployment/docker). ### Azure Container Instances (ACI)[​](#azure-container-instances-aci "Direct link to Azure Container Instances (ACI)") Run Spice.ai as a single container without provisioning a cluster using [Azure Container Instances](https://azure.microsoft.com/products/container-instances). Suitable for development environments, scheduled jobs, and low-traffic deployments. ``` az container create \ --resource-group my-rg \ --name spiceai \ --image spiceai/spiceai:latest \ --cpu 2 --memory 4 \ --ports 8090 50051 \ --ip-address public \ --environment-variables SPICED_LOG=INFO \ --azure-file-volume-share-name spice-data \ --azure-file-volume-account-name mystorageacct \ --azure-file-volume-account-key "" \ --azure-file-volume-mount-path /data ``` Refer to the [Azure Container Instances documentation](https://learn.microsoft.com/azure/container-instances/) for advanced networking, [virtual network integration](https://learn.microsoft.com/azure/container-instances/container-instances-virtual-network-concepts), and [managed identity](https://learn.microsoft.com/azure/container-instances/container-instances-managed-identity) configuration. ### Azure Virtual Machines[​](#azure-virtual-machines "Direct link to Azure Virtual Machines") Deploy Spice.ai directly on [Azure Virtual Machines](https://azure.microsoft.com/products/virtual-machines) for maximum control over the environment, GPU access, or large-memory instance types. 1. **Manual VM deployment**: * Provision a Linux VM (Ubuntu, Debian, or Azure Linux) with an appropriate [VM size](https://learn.microsoft.com/azure/virtual-machines/sizes). * Install [Docker Engine](https://docs.docker.com/engine/install/) and run [Spice.ai as a Docker container](https://spiceai.org/docs/deployment/docker), or install the `spice` binary directly. See the [installation guide](https://spiceai.org/docs/installation). * Attach a [managed identity](https://learn.microsoft.com/entra/identity/managed-identities-azure-resources/overview) so Spice can read from Blob Storage, Azure SQL, and Key Vault without static credentials. 2. **Automated deployment with Bicep or Terraform**: * Define infrastructure in a [Bicep template](https://learn.microsoft.com/azure/azure-resource-manager/bicep/) or [Terraform configuration](https://registry.terraform.io/providers/hashicorp/azurerm/latest), including the VM, NIC, NSG, virtual network, and managed identity. * Use [cloud-init](https://learn.microsoft.com/azure/virtual-machines/linux/using-cloud-init) or a [custom script extension](https://learn.microsoft.com/azure/virtual-machines/extensions/custom-script-linux) to install Docker, pull the [Spice.ai image](https://hub.docker.com/r/spiceai/spiceai), retrieve secrets from [Key Vault](https://azure.microsoft.com/products/key-vault), and start the runtime. * Use [VM scale sets](https://learn.microsoft.com/azure/virtual-machine-scale-sets/) for horizontally scaled deployments fronted by an [Azure Load Balancer](https://azure.microsoft.com/products/load-balancer) or [Application Gateway](https://azure.microsoft.com/products/application-gateway). For detailed guidance, refer to the [Linux on Azure documentation](https://learn.microsoft.com/azure/virtual-machines/linux/), [Bicep documentation](https://learn.microsoft.com/azure/azure-resource-manager/bicep/), and [Azure provider for Terraform](https://registry.terraform.io/providers/hashicorp/azurerm/latest/docs). ## Authentication[​](#authentication "Direct link to Authentication") Most Azure services that Spice connects to accept explicit credentials through component parameters (for example, an `azure_storage_account_key` on the [ABFS connector](/docs/next/components/data-connectors/abfs)). When explicit credentials are not provided, Spice follows the standard [Azure Identity DefaultAzureCredential](https://learn.microsoft.com/dotnet/api/azure.identity.defaultazurecredential) chain, attempting credentials in this order: 1. **Environment variables**: * `AZURE_CLIENT_ID`, `AZURE_TENANT_ID`, `AZURE_CLIENT_SECRET` (service principal with secret) * `AZURE_CLIENT_CERTIFICATE_PATH` (service principal with certificate) * `AZURE_USERNAME`, `AZURE_PASSWORD` (resource owner password — not recommended) 2. **Workload Identity** (AKS): federated tokens injected via `AZURE_FEDERATED_TOKEN_FILE`, `AZURE_AUTHORITY_HOST`, `AZURE_CLIENT_ID`, and `AZURE_TENANT_ID`. See [Workload Identity for AKS](https://learn.microsoft.com/azure/aks/workload-identity-overview). 3. **Managed identity**: System-assigned or user-assigned identities on Azure VMs, AKS, Container Apps, and ACI. See [What are managed identities?](https://learn.microsoft.com/entra/identity/managed-identities-azure-resources/overview). 4. **Azure CLI**: Cached credentials from a local `az login` session. Common during development. 5. **Azure Developer CLI** (`azd`) and **Azure PowerShell**: Used when the corresponding CLI is signed in. For services with explicit parameters (Blob Storage, Azure SQL, Cosmos DB, OpenAI), prefer named credentials or managed identity over environment variables in production. Role assignments Regardless of the credential source, the principal must have the appropriate Azure role assignments (for example, `Storage Blob Data Reader` on a storage account, or `SQL DB Contributor` on Azure SQL). When a Spicepod connects to multiple Azure services, the principal must have permissions across all of them. ## Resources[​](#resources "Direct link to Resources") ### Documentation[​](#documentation "Direct link to Documentation") * [Azure Integrations](/docs/next/deployment/azure/integrations) — complete list of Azure data connectors, AI models, and supported services * [Spice.ai Kubernetes Deployment Guide](/docs/next/deployment/kubernetes) — Helm, Argo CD, and Flux options for AKS ### Azure Marketplace[​](#azure-marketplace "Direct link to Azure Marketplace") Spice.ai is not yet published to the [Microsoft Azure Marketplace](https://azuremarketplace.microsoft.com/) (coming soon). In the meantime, deploy using the [`spiceai/spiceai`](https://hub.docker.com/r/spiceai/spiceai) container image or the [Spice Helm chart](https://helm.spiceai.org). --- # Azure Integrations Spice.ai integrates with Microsoft Azure for data federation, AI inference, embeddings, and authentication. This page consolidates Azure-compatible components and links to the relevant configuration guides. ## Data Connectors[​](#data-connectors "Direct link to Data Connectors") Data connectors federate SQL queries across Azure data sources without data movement. | Connector | Description | Documentation | | ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- | | **Azure Blob Storage / ADLS Gen2** | Query Parquet, CSV, and JSON files in [Azure Blob Storage](https://azure.microsoft.com/products/storage/blobs) or [ADLS Gen2](https://azure.microsoft.com/products/storage/data-lake-storage) using the `abfs://` scheme. | [ABFS Data Connector](/docs/next/components/data-connectors/abfs) | | **Azure SQL Database / SQL Server** | Connect to [Azure SQL Database](https://azure.microsoft.com/products/azure-sql/database/), [Azure SQL Managed Instance](https://azure.microsoft.com/products/azure-sql/managed-instance/), and SQL Server VMs. | [MSSQL Data Connector](/docs/next/components/data-connectors/mssql) | | **Azure Database for PostgreSQL** | Connect to flexible server and single server deployments using the PostgreSQL connector. | [PostgreSQL Data Connector](/docs/next/components/data-connectors/postgres) | | **Azure Database for MySQL** | Connect to flexible server deployments using the MySQL connector. | [MySQL Data Connector](/docs/next/components/data-connectors/mysql) | | **Azure Databricks** | Query Databricks tables on Azure using SQL Warehouse or Spark Connect. | [Databricks Data Connector](/docs/next/components/data-connectors/databricks) | | **Apache Iceberg (ADLS)** | Query Iceberg tables stored in ADLS Gen2 with REST or Unity Catalog metadata. | [Iceberg Data Connector](/docs/next/components/data-connectors/iceberg) | | **Delta Lake (ADLS)** | Query Delta Lake tables stored in ADLS Gen2 or Azure Blob Storage. | [Delta Lake Data Connector](/docs/next/components/data-connectors/delta-lake) | | **Microsoft SharePoint** | Index and query documents from SharePoint sites and OneDrive for Business with Microsoft Entra ID authentication. | [SharePoint Data Connector](/docs/next/components/data-connectors/sharepoint) | | **Azure-hosted databases via ODBC** | Connect through ODBC drivers for additional Azure-compatible data sources. | [ODBC Data Connector](/docs/next/components/data-connectors/odbc) | ### Example: Azure Blob Storage (ABFS)[​](#example-azure-blob-storage-abfs "Direct link to Example: Azure Blob Storage (ABFS)") ``` datasets: - from: abfs://container@account.dfs.core.windows.net/path/to/data/ name: events params: file_format: parquet abfs_account: account abfs_use_emulator: 'false' ``` ### Example: Azure SQL Database[​](#example-azure-sql-database "Direct link to Example: Azure SQL Database") ``` datasets: - from: mssql:dbo.orders name: orders params: mssql_connection_string: | Server=tcp:my-server.database.windows.net,1433; Database=mydb; Authentication=Active Directory Default; Encrypt=True; ``` ### Example: Azure Databricks[​](#example-azure-databricks "Direct link to Example: Azure Databricks") ``` datasets: - from: databricks:catalog.schema.table name: orders params: mode: spark_connect databricks_endpoint: my-workspace.azuredatabricks.net databricks_token: ${ secrets:DATABRICKS_TOKEN } ``` ## Catalog Connectors[​](#catalog-connectors "Direct link to Catalog Connectors") Catalog connectors provide schema discovery and unified access to tables in Azure data catalogs. | Connector | Description | Documentation | | ---------------------------- | --------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------- | | **Databricks Unity Catalog** | Discover and query tables governed by Unity Catalog on Azure Databricks. Supports Azure Blob authentication for table data. | [Unity Catalog](/docs/next/components/catalogs/unity-catalog) | | **Databricks Catalog** | Connect to Azure Databricks as a catalog source for federated queries. | [Databricks Catalog](/docs/next/components/catalogs/databricks) | ### Example: Unity Catalog with Azure Blob[​](#example-unity-catalog-with-azure-blob "Direct link to Example: Unity Catalog with Azure Blob") ``` catalogs: - from: unity_catalog name: my_catalog params: unity_catalog_endpoint: https://my-workspace.azuredatabricks.net unity_catalog_token: ${ secrets:DATABRICKS_TOKEN } unity_catalog_azure_storage_account_name: mystorageacct unity_catalog_azure_storage_client_id: ${ secrets:AZURE_CLIENT_ID } unity_catalog_azure_storage_client_secret: ${ secrets:AZURE_CLIENT_SECRET } ``` ## AI Models (Azure OpenAI)[​](#ai-models-azure-openai "Direct link to AI Models (Azure OpenAI)") Spice integrates with [Azure OpenAI Service](https://azure.microsoft.com/products/ai-services/openai-service) for chat completion and reasoning models, including GPT-4 family, GPT-5, and o-series models. | Provider | Supported Models | Documentation | | ---------------- | ------------------------------------------------------ | --------------------------------------------------------- | | **Azure OpenAI** | GPT-4, GPT-4o, GPT-5, o-series, and other deployments. | [Azure OpenAI Models](/docs/next/components/models/azure) | ### Example: Azure OpenAI Chat Model[​](#example-azure-openai-chat-model "Direct link to Example: Azure OpenAI Chat Model") ``` models: - from: azure:gpt-4o name: gpt params: endpoint: ${ secrets:SPICE_AZURE_AI_ENDPOINT } azure_deployment_name: gpt-4o azure_api_version: 2024-08-01-preview azure_api_key: ${ secrets:SPICE_AZURE_API_KEY } ``` For Microsoft Entra ID authentication instead of an API key, set `azure_entra_token` in place of `azure_api_key`. ## Secret Stores[​](#secret-stores "Direct link to Secret Stores") Spice resolves secrets at runtime from configured [secret stores](/docs/next/components/secret-stores). For Azure deployments, the [`azure_keyvault`](/docs/next/components/secret-stores/azure-keyvault) store reads secrets directly from [Azure Key Vault](https://azure.microsoft.com/products/key-vault/), so Spicepods can reference connector and model credentials without baking them into environment variables or `values.yaml`. | Provider | Supported Auth Methods | Documentation | | ------------------- | ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | | **Azure Key Vault** | `service_principal`, `managed_identity`, `workload_identity`, `cli`, `default`. | [Azure Key Vault Secret Store](/docs/next/components/secret-stores/azure-keyvault) | ### Example: Azure Key Vault Secret Store[​](#example-azure-key-vault-secret-store "Direct link to Example: Azure Key Vault Secret Store") ``` secrets: - from: azure_keyvault:prod-vault name: azure params: auth_method: workload_identity datasets: - from: postgres:public.taxi_trips name: taxi_trips params: pg_host: postgres.example.com pg_user: ${azure:postgres_user} pg_pass: ${azure:postgres_password} ``` Logical key names use underscores; the store automatically translates them to Key Vault names like `spice-postgres-user` (with a fallback to `postgres-user`). Pair `azure_keyvault` with [AKS workload identity](/docs/next/deployment/azure) or a [Container Apps managed identity](/docs/next/deployment/azure) so the runtime authenticates without long-lived credentials. ## Embeddings (Azure OpenAI)[​](#embeddings-azure-openai "Direct link to Embeddings (Azure OpenAI)") Generate vector embeddings using Azure OpenAI deployments for semantic search and retrieval-augmented generation (RAG). | Provider | Supported Models | Documentation | | ---------------- | ----------------------------------------------------------------------------- | ----------------------------------------------------------------- | | **Azure OpenAI** | `text-embedding-3-small`, `text-embedding-3-large`, `text-embedding-ada-002`. | [Azure OpenAI Embeddings](/docs/next/components/embeddings/azure) | ### Example: Azure OpenAI Embeddings[​](#example-azure-openai-embeddings "Direct link to Example: Azure OpenAI Embeddings") ``` embeddings: - from: azure:text-embedding-3-small name: azure_embed params: endpoint: ${ secrets:SPICE_AZURE_AI_ENDPOINT } azure_deployment_name: text-embedding-3-small azure_api_version: 2023-05-15 azure_api_key: ${ secrets:SPICE_AZURE_API_KEY } ``` Refer to the [Azure OpenAI Service models](https://learn.microsoft.com/azure/ai-services/openai/concepts/models) for the full list of supported models and regions. ## Authentication[​](#authentication "Direct link to Authentication") All Azure integrations support the standard [Azure Identity DefaultAzureCredential](https://learn.microsoft.com/dotnet/api/azure.identity.defaultazurecredential) chain. When credentials are not explicitly configured, Spice attempts the following in order: 1. **Environment variables** — service principal (`AZURE_CLIENT_ID`, `AZURE_TENANT_ID`, `AZURE_CLIENT_SECRET`), certificate (`AZURE_CLIENT_CERTIFICATE_PATH`), or username/password. 2. **Workload Identity** — federated tokens on AKS via `AZURE_FEDERATED_TOKEN_FILE`. See [Workload Identity for AKS](https://learn.microsoft.com/azure/aks/workload-identity-overview). 3. **Managed Identity** — system-assigned or user-assigned identities on Azure VMs, AKS, Container Apps, and ACI. See [Managed identities for Azure resources](https://learn.microsoft.com/entra/identity/managed-identities-azure-resources/overview). 4. **Azure CLI** — cached credentials from a local `az login` session. 5. **Azure Developer CLI / Azure PowerShell** — used when the corresponding CLI is signed in. For a deployment-side overview of these mechanisms, see the [Authentication](/docs/next/deployment/azure#authentication) section of the Azure deployment guide. ### Role Assignments[​](#role-assignments "Direct link to Role Assignments") Each principal must have the appropriate Azure RBAC role for the services it accesses: | Service | Common role(s) | | ------------------------------ | -------------------------------------------------------------- | | Azure Blob Storage / ADLS Gen2 | `Storage Blob Data Reader` or `Storage Blob Data Contributor` | | Azure Key Vault | `Key Vault Secrets User` (data plane) or RBAC equivalent | | Azure SQL Database | Database-level role assignments granted to the Entra principal | | Azure OpenAI | `Cognitive Services OpenAI User` | | Azure Container Registry | `AcrPull` for image pulls | When a Spicepod connects to multiple Azure services, ensure roles are granted on every resource the runtime touches. ## Cookbooks[​](#cookbooks "Direct link to Cookbooks") * [Azure OpenAI Models](https://github.com/spiceai/cookbook/tree/trunk/azure_openai) — vector search and chat over structured and unstructured data with Azure OpenAI. --- # CI/CD Deployment Spice deployments can be automated through continuous integration and delivery (CI/CD) pipelines. The recommended approach for self-hosted, open-source deployments is the [Spice Helm chart](https://github.com/spiceai/helm-charts), driven either directly from a pipeline runner or declaratively through a GitOps controller. Container and cloud-VM workflows are also supported, as is a managed deploy action for the Spice Cloud Platform. The sections below cover, in order: * [Helm in CI pipelines](#helm-in-ci-pipelines) — push-based deployment from GitHub Actions, GitLab CI, or any runner. * [Kubernetes GitOps](#kubernetes-gitops) — pull-based reconciliation with Argo CD or Flux. * [Containers and cloud VMs](#containers-and-cloud-vms) — Docker, AWS, and Azure pipelines. * [Spice Cloud Platform](#spice-cloud-platform) — Connect Repository from the portal, or the `spicehq/spice-cloud-deploy-action` GitHub Action. Self-hosted enterprise deployments For production self-hosted deployments, the [Spice.ai Enterprise Kubernetes Operator](https://docs.spice.ai/docs/enterprise/kubernetes-operator/kubernetes) is the recommended approach. The operator provides per-replica StatefulSets, automatic PVC resizing, configurable update strategies, crashloop protection, and distributed query execution through `SpicepodSet` and `SpicepodCluster` custom resources, all reconcilable from Git through the same GitOps tooling described below. ## Helm in CI pipelines[​](#helm-in-ci-pipelines "Direct link to Helm in CI pipelines") The [Spice Helm chart](https://github.com/spiceai/helm-charts) is the primary deployment artifact for self-hosted clusters. Any CI runner with `kubectl` and `helm` installed can roll out a release by checking out the repository, authenticating to the target cluster, and running `helm upgrade --install`. The chart loads the Spicepod from a `spicepod` key in the values file. A typical layout keeps a single `values.yaml` that contains both chart configuration and the Spicepod definition: ``` # values.yaml image: repository: spiceai/spiceai tag: '1.10.0' spicepod: name: cayenne version: v1 kind: Spicepod datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet acceleration: enabled: true engine: duckdb ``` For details on chart values, see the [Helm deployment guide](https://spiceai.org/docs/deployment/kubernetes/helm). ### GitHub Actions example[​](#github-actions-example "Direct link to GitHub Actions example") The workflow below deploys the chart to a Kubernetes cluster on every push to `main`. Cluster credentials are provided through a base64-encoded kubeconfig stored in the `KUBE_CONFIG` repository secret. ``` name: Deploy Spice on: push: branches: [main] jobs: deploy: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: azure/setup-helm@v4 with: version: v3.14.0 - name: Configure kubectl run: | mkdir -p "$HOME/.kube" echo "${{ secrets.KUBE_CONFIG }}" | base64 -d > "$HOME/.kube/config" - name: Deploy Spice run: | helm repo add spiceai https://helm.spiceai.org helm repo update helm upgrade --install spiceai spiceai/spiceai \ --namespace spiceai \ --create-namespace \ --values values.yaml \ --atomic \ --wait \ --timeout 5m ``` `--atomic` rolls back on failure, and `--wait` blocks until the release is healthy, so a failed deploy fails the pipeline. ### GitLab CI example[​](#gitlab-ci-example "Direct link to GitLab CI example") The same pattern works in GitLab CI. The job uses the official `alpine/helm` image and reads cluster credentials from a CI/CD variable. ``` deploy: image: alpine/helm:3.14.0 stage: deploy before_script: - apk add --no-cache curl - curl -LO "https://dl.k8s.io/release/$(curl -L -s https://dl.k8s.io/release/stable.txt)/bin/linux/amd64/kubectl" - install -m 0755 kubectl /usr/local/bin/kubectl - mkdir -p ~/.kube && echo "$KUBE_CONFIG" | base64 -d > ~/.kube/config script: - helm repo add spiceai https://helm.spiceai.org - helm repo update - helm upgrade --install spiceai spiceai/spiceai --namespace spiceai --create-namespace --values values.yaml --atomic --wait --timeout 5m only: - main ``` ### Pinning the chart and runtime versions[​](#pinning-the-chart-and-runtime-versions "Direct link to Pinning the chart and runtime versions") Production pipelines should pin both the chart and the Spice runtime image to specific versions. Pass `--version` to `helm upgrade` to pin the chart, and set `image.tag` in `values.yaml` to pin the runtime image: ``` helm upgrade --install spiceai spiceai/spiceai \ --version 1.10.0 \ --values values.yaml ``` Available chart versions are listed in the [helm-charts repository](https://github.com/spiceai/helm-charts/releases). Runtime image tags are published on [GitHub Container Registry](https://github.com/spiceai/spiceai/pkgs/container/spiceai). ### Promoting across environments[​](#promoting-across-environments "Direct link to Promoting across environments") To promote the same artifact across environments, keep a base `values.yaml` and add per-environment overlays such as `values.staging.yaml` and `values.prod.yaml`. Helm merges multiple `-f` flags in order: ``` helm upgrade --install spiceai spiceai/spiceai \ --values values.yaml \ --values values.prod.yaml ``` Each environment can target a different cluster, namespace, or image tag while sharing the same Spicepod definition. ## Kubernetes GitOps[​](#kubernetes-gitops "Direct link to Kubernetes GitOps") GitOps controllers reconcile cluster state from a Git repository, removing the need for the pipeline to hold cluster credentials. The controller runs inside the cluster and pulls changes as they are committed. * [Argo CD](https://spiceai.org/docs/deployment/kubernetes/argocd) — `Application` manifests reconciled by the Argo CD controller. * [Flux](https://spiceai.org/docs/deployment/kubernetes/flux) — `HelmRelease` resources reconciled by the Flux toolkit. Both guides include end-to-end manifests targeting the official chart, including upgrade and rollback patterns. GitOps is the recommended approach for multi-cluster or multi-environment deployments. ## Containers and cloud VMs[​](#containers-and-cloud-vms "Direct link to Containers and cloud VMs") For deployments that target a container runtime or a cloud VM rather than Kubernetes, invoke the standard provider tooling from any pipeline runner: * [Docker](https://spiceai.org/docs/deployment/docker) — build, push, and run the `spiceai/spiceai` image. Pipelines typically run `docker build` and `docker push` against a registry, then `docker compose up -d` or `docker run` on the target host. * [AWS](https://spiceai.org/docs/deployment/aws) — deploy the published CloudFormation template through the AWS CLI or any CloudFormation-aware action. * [Azure](https://spiceai.org/docs/deployment/azure) — deploy through ARM/Bicep templates or the Azure CLI. Each provider guide includes the deployment artifact (image, template, or script) that the pipeline invokes. ## Spice Cloud Platform[​](#spice-cloud-platform "Direct link to Spice Cloud Platform") Deployments targeting the [Spice Cloud Platform](https://spiceai.org/docs/deployment/cloud) can be automated two ways: * **Connect Repository** — link a GitHub repository to a Spice Cloud app from the portal. The app redeploys automatically on each push to the connected branch, with no pipeline configuration required. See [Connect GitHub](https://docs.spice.ai/docs/portal/apps/connect-github). * **GitHub Actions** — use the [`spicehq/spice-cloud-deploy-action`](https://github.com/spicehq/spice-cloud-deploy-action) to deploy from a custom workflow. Use this when the pipeline needs to run tests, build artifacts, or set secrets and tags before deploying. ### GitHub Actions[​](#github-actions "Direct link to GitHub Actions") The `spicehq/spice-cloud-deploy-action` deploys a Spicepod manifest to a Spice Cloud app on each pipeline run. #### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A [Spice Cloud account](https://spice.ai/login). * An OAuth client created from the Spice Cloud Portal. Two repository secrets — `SPICE_CLIENT_ID` and `SPICE_CLIENT_SECRET` — store its credentials. * A `spicepod.yaml` checked into the repository. #### Minimal workflow[​](#minimal-workflow "Direct link to Minimal workflow") ``` name: Deploy Spicepod on: push: branches: [main] jobs: deploy: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: spicehq/spice-cloud-deploy-action@v1 with: client-id: ${{ secrets.SPICE_CLIENT_ID }} client-secret: ${{ secrets.SPICE_CLIENT_SECRET }} app-name: my-app spicepod: spicepod.yaml ``` #### Common options[​](#common-options "Direct link to Common options") | Input | Purpose | | -------------------------------------- | -------------------------------------------------------------------------------------- | | `app-name` or `app-id` | Target Spice Cloud app. One is required. | | `spicepod` | Path to the Spicepod manifest. Defaults to `spicepod.yaml`. | | `region` | Required when `create-app-if-missing` provisions a new app (for example, `us-east-1`). | | `create-app-if-missing` | Boolean. Creates the app on first deploy. | | `secrets` | YAML or JSON map of app-level secrets to set on the deployment. | | `tags` | YAML or JSON map of metadata labels. | | `test-sql`, `test-chat`, `test-search` | Post-deploy smoke checks against the deployed app. | | `wait-for-completion` | Poll until the deployment finishes. Defaults to `true`. | | `timeout-seconds` | Maximum time to wait when polling. Defaults to `600`. | The action emits `app-id`, `app-url`, `deployment-id`, `deployment-status`, and `test-results` outputs that downstream steps can consume. For the full input and output reference, see the [action's README](https://github.com/spicehq/spice-cloud-deploy-action). ## Related[​](#related "Direct link to Related") * [Helm Deployment Guide](https://spiceai.org/docs/deployment/kubernetes/helm) * [Kubernetes Deployment Guide](https://spiceai.org/docs/deployment/kubernetes) * [Argo CD Guide](https://spiceai.org/docs/deployment/kubernetes/argocd) * [Flux Guide](https://spiceai.org/docs/deployment/kubernetes/flux) * [Spice Cloud Platform Deployment](https://spiceai.org/docs/deployment/cloud) * [Spice Helm Chart](https://github.com/spiceai/helm-charts) * [Spice Cloud Deploy Action](https://github.com/spicehq/spice-cloud-deploy-action) --- # Spice Cloud Platform Deployment The Spice Cloud Platform is a managed, cloud-hosted solution designed for deploying data and AI applications and agents. It provides a secure and efficient compute environment powered by Spice.ai OSS, offering building blocks including high-speed SQL queries, LLM inference, vector search, and retrieval-augmented generation (RAG). ## Benefits of the Spice.ai Cloud Platform[​](#benefits-of-the-spiceai-cloud-platform "Direct link to Benefits of the Spice.ai Cloud Platform") * **Simplified Deployment**: Focus on creating data and AI applications without the complexity of managing infrastructure. * **High Performance**: Optimize data queries and AI workflows with cloud-scale compute resources. * **Collaboration**: Share and manage datasets, models, and tools across your team and the enterprise. * **Production-Ready**: Achieve reliability, scalability, and compliance for AI applications. ## Security and Compliance[​](#security-and-compliance "Direct link to Security and Compliance") The Spice.ai Cloud Platform prioritizes security and compliance to ensure the protection of user data and systems. It adheres to SOC 2 Type II compliance standards, providing enterprise-level security validated by third-party audits. The platform employs comprehensive security measures, including encryption of sensitive data both in-transit and at-rest, multi-factor authentication (MFA), and role-based access control (RBAC). Access and usage are logged for auditability, and the principle of least privilege is enforced to minimize unnecessary access. For more details, visit the [Spice.ai Security Documentation](https://docs.spice.ai/security/security). ## Deployment Overview[​](#deployment-overview "Direct link to Deployment Overview") 1. **Sign Up**: Register for an account on the [Spice.ai Cloud Platform](https://spice.ai/login). 2. **Configure Your Application**: Organize datasets, models, and workflows using the platform's cloud portal. 3. **Deploy**: Launch your AI applications and agents with minimal configuration. 4. **Monitor and Scale**: Use built-in monitoring and observability tools to track performance and scale resources as required. ## Learn More[​](#learn-more "Direct link to Learn More") For comprehensive instructions and advanced configuration options, refer to the [Spice.ai Cloud Platform documentation](https://docs.spice.ai/). --- # Docker This guide describes how to run Spice.ai as a Docker container, either directly with `docker run`, with Docker Compose, or by building a custom image that bundles a Spicepod and data files. For Kubernetes deployments, see the [Kubernetes deployment guide](/docs/next/deployment/kubernetes). ## Quickstart[​](#quickstart "Direct link to Quickstart") Run the latest Spice.ai image with a local Spicepod mounted into the container: ``` docker run --rm -it \ -p 8090:8090 \ -p 9090:9090 \ -p 50051:50051 \ -v "$(pwd)":/app \ spiceai/spiceai:latest ``` Spice listens on three ports: * `8090` — HTTP API and `/health` endpoint * `9090` — Prometheus metrics (optional) * `50051` — Arrow Flight (gRPC) for high-throughput query results As of Spice v2.0, AI features (embeddings, local model inference, search) are included in the default image — no separate tag is required. The data-only image, without AI/ML dependencies, is published only in nightly builds for open source and is available as a production-ready distribution through [Spice Cloud and Spice.ai Enterprise](https://spice.ai/pricing). Browse all published tags at [hub.docker.com/r/spiceai/spiceai/tags](https://hub.docker.com/r/spiceai/spiceai/tags). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Docker Engine](https://docs.docker.com/engine/install/) 20.10 or newer (or Docker Desktop). * A `spicepod.yaml` file. See [Spicepods](https://spiceai.org/docs/getting-started/spicepods). * Optional: [Docker Compose v2](https://docs.docker.com/compose/) for multi-container setups. ## Image Tags[​](#image-tags "Direct link to Image Tags") | Tag | Description | | ----------- | ---------------------------------------------------------------------------------------------------------- | | `latest` | Latest stable release. Includes AI features (embeddings, local model inference, vector search) by default. | | `` | A specific stable release, e.g. `2.0.0`. Recommended for production for reproducible deployments. | note The `-models` image tags were removed in Spice v2.0 — AI/ML support is now part of the default image. The data-only image (without AI/ML dependencies) is published only in nightly builds for open source. Pin to a specific version in production to avoid unexpected upgrades: ``` docker run -p 8090:8090 spiceai/spiceai:2.0.0 ``` ## Run with Docker Compose[​](#run-with-docker-compose "Direct link to Run with Docker Compose") `docker-compose.yaml`: ``` services: spiced: image: spiceai/spiceai:latest container_name: spiced ports: - '8090:8090' - '9090:9090' - '50051:50051' volumes: - ./spicepod.yaml:/app/spicepod.yaml:ro - ./data:/app/data env_file: - .env healthcheck: test: ['CMD', 'wget', '-q', '--spider', 'http://localhost:8090/health'] interval: 10s timeout: 3s retries: 5 restart: unless-stopped ``` Start the container: ``` docker compose up ``` ## Build a Custom Image[​](#build-a-custom-image "Direct link to Build a Custom Image") For deployments that ship a Spicepod and data with the runtime, build a custom image that copies them in: ``` # AI features are included in the default image as of Spice v2.0. # See https://hub.docker.com/r/spiceai/spiceai/tags for all tags. FROM spiceai/spiceai:2.0.0 WORKDIR /app # Spicepod definition COPY spicepod.yaml ./ # Optional: bundled data files COPY data ./data # Optional: environment files (.env, .env.local). Avoid baking secrets # into images — prefer runtime --env-file or a secret manager. COPY .env* ./ # --metrics is optional; omit if Prometheus metrics are not needed. CMD ["--http", "0.0.0.0:8090", "--metrics", "0.0.0.0:9090", "--flight", "0.0.0.0:50051"] EXPOSE 8090 9090 50051 ``` Build and run: ``` docker build -t my-spiceai-app . docker run --rm -p 8090:8090 -p 50051:50051 my-spiceai-app ``` Do not bake secrets into images Image layers are cached and distributed. Use `--env-file`, `docker run -e`, or a secret manager such as [Docker secrets](https://docs.docker.com/engine/swarm/secrets/), [HashiCorp Vault](https://www.vaultproject.io/), or [AWS Secrets Manager](https://aws.amazon.com/secrets-manager/) to inject credentials at runtime. ## Environment Variables and Secrets[​](#environment-variables-and-secrets "Direct link to Environment Variables and Secrets") Spice loads secrets from environment variables prefixed with `SPICE_SECRET_`. See the [Environment Secret Store](/docs/next/components/secret-stores/env) for details. Pass secrets at runtime with `--env-file` (preferred) or `-e`: ``` docker run --rm -p 8090:8090 \ -v "$(pwd)/spicepod.yaml":/app/spicepod.yaml:ro \ --env-file .env \ -e SPICED_LOG=INFO \ spiceai/spiceai:latest ``` Common runtime variables: | Variable | Purpose | | --------------------- | --------------------------------------------------------------------- | | `SPICED_LOG` | Log level: `ERROR`, `WARN`, `INFO`, `DEBUG`, `TRACE`. Default `INFO`. | | `SPICE_SECRET_` | Inject a named secret referenced from a Spicepod. | ## Persistence[​](#persistence "Direct link to Persistence") For workloads that use file-based acceleration (for example, DuckDB or SQLite), mount a host directory or named volume so data survives container restarts: ``` docker run --rm -p 8090:8090 \ -v spice-data:/data \ -v "$(pwd)/spicepod.yaml":/app/spicepod.yaml:ro \ spiceai/spiceai:latest ``` In the Spicepod, configure the accelerator to write under the mount path, for example `duckdb_file: /data/taxi_trips.db`. ## Health Checks[​](#health-checks "Direct link to Health Checks") Spice exposes `/health` (process up) and `/v1/ready` (components ready) on the HTTP port. Use these in container orchestrators or load balancers: ``` curl http://localhost:8090/health curl http://localhost:8090/v1/ready ``` A Docker Compose healthcheck example is included in the [Run with Docker Compose](#run-with-docker-compose) section above. ## Cookbook[​](#cookbook "Direct link to Cookbook") * [Running Spice.ai in Docker](https://github.com/spiceai/cookbook/tree/trunk/docker) --- # Docker Sandbox Guide - v1.3.0 ## v1.3.0 Docker Sandbox[​](#v130-docker-sandbox "Direct link to v1.3.0 Docker Sandbox") In the v1.3.0 release, the Docker image changed the sandbox from a script that ran at startup, to being baked into the Docker image itself. Prior to this, the Docker image would start up as a root user, and then set up a sandbox user with restricted permissions before starting the Spice runtime in that restricted context. Additionally, the Docker image includes no standard Linux tools like `bash`. Starting with v1.3.0, the sandboxing logic is baked directly into the final Docker image, and the Docker image starts up as the sandbox user. For most users, this change will be transparent. However, there are a few cases where an action is required to update. ### Building a custom Docker image based on v1.3.0[​](#building-a-custom-docker-image-based-on-v130 "Direct link to Building a custom Docker image based on v1.3.0") Building a custom Docker image based on v1.3.0 that installs additional dependencies will require using the `debian:bookworm-slim` base image and copying the `spiced` binary from the `spiceai/spiceai` image into the custom image. This approach can also be used to restore the previous behavior of including standard Linux tools like `bash`. ``` FROM debian:bookworm-slim # Copy the spiced binary from the spiceai/spiceai image into the custom image. COPY --from=spiceai/spiceai:v1.3.0 /usr/local/bin/spiced /usr/local/bin/spiced # Install any additional dependencies needed for the image. RUN apt update && apt install -y --no-install-recommends # Any other customizations needed for the image. WORKDIR /app ENTRYPOINT ["/usr/local/bin/spiced"] ``` This will restore the previous behavior of starting as the root user and including standard Linux tools, like `bash`. #### Running as a non-root user[​](#running-as-a-non-root-user "Direct link to Running as a non-root user") Spice recommends that custom Docker images based on Spice run as a non-root user. i.e. ``` RUN addgroup -g 1001 -S sandboxgroup && adduser -u 1001 -S -G sandboxgroup sandbox USER sandbox ``` This may require additional configuration of mounted volumes to ensure that the sandbox user has access to the necessary files. i.e. in Kubernetes, this requires adding a `securityContext` to the pod spec. ``` securityContext: runAsUser: 1001 runAsGroup: 1001 # Tells Kubernetes to set the group of the files in the volume to sandboxgroup, # which allows the sandbox user to access the files. fsGroup: 1001 ``` note The `fsGroup` directive does not work for all Kubernetes storage types. For example, it does not work for `hostPath` volumes. In this case, an init container can be used to set the group of the files in the volume. ### Custom Kubernetes deployments[​](#custom-kubernetes-deployments "Direct link to Custom Kubernetes deployments") Kubernetes deployments that do not use the v1.3.0 Helm chart will need to add the following `securityContext` to their pod spec: ``` securityContext: runAsUser: 65534 runAsGroup: 65534 fsGroup: 65534 ``` ### Debugging sandbox container[​](#debugging-sandbox-container "Direct link to Debugging sandbox container") To debug issues with the sandbox container, see the [Debugging Sandbox Container](/docs/next/troubleshooting#debugging-sandbox-container) section of the troubleshooting guide. --- # Google Cloud Deployment Options Spice.ai runs on Google Cloud Platform (GCP) on Kubernetes, serverless containers, or virtual machines. The container image and Helm chart are the same artefacts used in every other environment, so the choice of GCP service is a matter of operational fit rather than packaging. For a complete list of GCP-compatible data connectors, AI models, and supported services, see [GCP Integrations](/docs/next/deployment/gcp/integrations). ## Benefits of deploying on GCP[​](#benefits-of-deploying-on-gcp "Direct link to Benefits of deploying on GCP") * **Scalability**: Scale Spice with [GKE node auto-provisioning](https://cloud.google.com/kubernetes-engine/docs/concepts/node-auto-provisioning), [GKE Autopilot](https://cloud.google.com/kubernetes-engine/docs/concepts/autopilot-overview), and [Cloud Run](https://cloud.google.com/run). * **Global reach**: Deploy across [GCP regions](https://cloud.google.com/about/locations) for low-latency access close to data sources. * **Integration**: Connect to [BigQuery](https://cloud.google.com/bigquery), [Cloud Storage](https://cloud.google.com/storage), [Cloud SQL](https://cloud.google.com/sql), [AlloyDB](https://cloud.google.com/alloydb), and [Secret Manager](https://cloud.google.com/secret-manager). * **Cost control**: Choose from [machine types](https://cloud.google.com/compute/docs/machine-resource), [committed use discounts](https://cloud.google.com/compute/docs/instances/signing-up-committed-use-discounts), and [Spot VMs](https://cloud.google.com/compute/docs/instances/spot). * **Security**: Run inside a VPC with [Private Google Access](https://cloud.google.com/vpc/docs/private-google-access), [VPC Service Controls](https://cloud.google.com/vpc-service-controls), and short-lived credentials via [Workload Identity Federation](https://cloud.google.com/iam/docs/workload-identity-federation). ## Deployment options[​](#deployment-options "Direct link to Deployment options") ### Google Kubernetes Engine (GKE)[​](#google-kubernetes-engine-gke "Direct link to Google Kubernetes Engine (GKE)") Run Spice on [GKE](https://cloud.google.com/kubernetes-engine) when the workload benefits from Kubernetes orchestration, multi-replica scale, or shared cluster tenancy. GKE pairs with the [Spice Helm chart](https://spiceai.org/docs/deployment/kubernetes/helm) and the [Argo CD](https://spiceai.org/docs/deployment/kubernetes/argocd) or [Flux](https://spiceai.org/docs/deployment/kubernetes/flux) GitOps workflows. #### 1. Provision the cluster[​](#1-provision-the-cluster "Direct link to 1. Provision the cluster") The fastest path is `gcloud`. The example below creates a regional Standard cluster with [Workload Identity](https://cloud.google.com/kubernetes-engine/docs/concepts/workload-identity) enabled — required for federated credentials to GCP services. ``` PROJECT=my-project REGION=us-central1 CLUSTER=spiceai-prod gcloud container clusters create $CLUSTER \ --project $PROJECT \ --region $REGION \ --release-channel regular \ --machine-type e2-standard-4 \ --num-nodes 1 \ --enable-autoscaling --min-nodes 2 --max-nodes 6 \ --workload-pool ${PROJECT}.svc.id.goog \ --enable-ip-alias gcloud container clusters get-credentials $CLUSTER --region $REGION --project $PROJECT ``` For burst or low-utilization workloads, use [GKE Autopilot](https://cloud.google.com/kubernetes-engine/docs/concepts/autopilot-overview) — Google manages the nodes, billing is per-pod, and Workload Identity is enabled by default. For production, prefer Terraform for repeatable provisioning. The [`terraform-google-modules/kubernetes-engine`](https://github.com/terraform-google-modules/terraform-google-kubernetes-engine) module is a common starting point. #### 2. Configure Workload Identity for GCP access[​](#2-configure-workload-identity-for-gcp-access "Direct link to 2. Configure Workload Identity for GCP access") Most Spice connectors (Cloud Storage via the [S3 connector](/docs/next/components/data-connectors/s3) with HMAC, BigQuery via [ADBC](/docs/next/components/data-connectors/adbc), Cloud SQL via [PostgreSQL](/docs/next/components/data-connectors/postgres) or [MySQL](/docs/next/components/data-connectors/mysql)) accept GCP credentials from [Application Default Credentials](https://cloud.google.com/docs/authentication/application-default-credentials). Use Workload Identity so pods receive scoped, short-lived tokens without static keys: ``` # 1. Create a Google service account and grant it the roles the Spicepod needs gcloud iam service-accounts create spiceai-runtime --project $PROJECT gcloud projects add-iam-policy-binding $PROJECT \ --member "serviceAccount:spiceai-runtime@${PROJECT}.iam.gserviceaccount.com" \ --role roles/storage.objectViewer gcloud projects add-iam-policy-binding $PROJECT \ --member "serviceAccount:spiceai-runtime@${PROJECT}.iam.gserviceaccount.com" \ --role roles/bigquery.dataViewer # 2. Bind the Google service account to a Kubernetes ServiceAccount gcloud iam service-accounts add-iam-policy-binding \ spiceai-runtime@${PROJECT}.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:${PROJECT}.svc.id.goog[spiceai/spiceai]" ``` Reference the service account from the Helm release so pods inherit federated tokens via the standard ADC chain: ``` # values.yaml serviceAccount: create: true name: spiceai annotations: iam.gke.io/gcp-service-account: spiceai-runtime@my-project.iam.gserviceaccount.com ``` #### 3. Install Spice.ai[​](#3-install-spiceai "Direct link to 3. Install Spice.ai") ``` helm repo add spiceai https://helm.spiceai.org helm repo update helm upgrade --install spiceai spiceai/spiceai \ --namespace spiceai --create-namespace \ --version 1.11.5 \ -f values.yaml ``` For declarative GitOps, swap this command for an Argo CD `Application` or a Flux `HelmRelease` pointing at the same chart. See the [Argo CD](https://spiceai.org/docs/deployment/kubernetes/argocd) or [Flux](https://spiceai.org/docs/deployment/kubernetes/flux) guides for full manifests. #### 4. Storage and ingress[​](#4-storage-and-ingress "Direct link to 4. Storage and ingress") For stateful acceleration (DuckDB, SQLite, Cayenne): * **Local SSD (recommended)** — Spice acceleration is latency- and IOPS-sensitive, so the lowest-latency option is a node-local NVMe SSD on a machine type with attached [Local SSD](https://cloud.google.com/compute/docs/disks/local-ssd) (`n2-standard-*-lssd`, `c3-standard-*-lssd`, `z3` series). Expose Local SSDs through GKE's [Local SSD raw block / ephemeral storage](https://cloud.google.com/kubernetes-engine/docs/how-to/persistent-volumes/local-ssd) provisioner. Local SSDs do not survive node replacement, so pair with a refresh strategy or a re-hydration source. * **Hyperdisk Extreme / Balanced** — when shared, replica-attachable persistence is required, [Hyperdisk](https://cloud.google.com/compute/docs/disks/hyperdisk) provides high IOPS and configurable throughput. Use the [Compute Engine persistent disk CSI driver](https://cloud.google.com/kubernetes-engine/docs/how-to/persistent-volumes/gce-pd-csi-driver) with a custom StorageClass (`type: hyperdisk-balanced` or `hyperdisk-extreme`). * **Persistent Disk SSD (`pd-ssd`, `premium-rwo`)** — use the built-in `premium-rwo` storage class only when Hyperdisk is unavailable in a region. * **Filestore (`filestore-csi`) — not recommended for acceleration** — use only for stateless shared artefacts that need `ReadWriteMany`. NFS latency negates the benefit of using a local accelerator. * Set `stateful.enabled: true` and `stateful.storageClass: ` in `values.yaml`. Spice.ai Enterprise For production stateful workloads, the [Spice.ai Enterprise](https://spice.ai) Operator's [`SpicepodSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset) provides per-replica `StatefulSet`s with automatic PVC resizing, Workload-Identity-aware ServiceAccount annotations, and configurable update strategies. For distributed query execution across scheduler/executor tiers backed by Cloud Storage, see [`SpicepodCluster`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodcluster). To expose Spice externally, install the [GKE Gateway controller](https://cloud.google.com/kubernetes-engine/docs/concepts/gateway-api) or use a [Cloud Load Balancer Service](https://cloud.google.com/kubernetes-engine/docs/how-to/exposing-apps): ``` # values.yaml service: type: LoadBalancer additionalAnnotations: networking.gke.io/load-balancer-type: 'Internal' # internal only ``` For internal-only deployments, set `Internal` to bind to the cluster's VPC rather than a public IP. #### 5. Observability[​](#5-observability "Direct link to 5. Observability") The Spice Helm chart ships a `PodMonitor` resource for the [Prometheus Operator](https://prometheus-operator.dev/). On GKE, [Google Cloud Managed Service for Prometheus](https://cloud.google.com/stackdriver/docs/managed-prometheus) is the common target — it ingests `PodMonitor` resources directly when [managed collection](https://cloud.google.com/stackdriver/docs/managed-prometheus/setup-managed) is enabled. Set `monitoring.podMonitor.enabled: true` and import the [Spice Grafana dashboard](/docs/next/monitoring/grafana) into [Cloud Monitoring](https://cloud.google.com/monitoring) or self-managed Grafana. For comprehensive guidance, refer to the [GKE documentation](https://cloud.google.com/kubernetes-engine/docs), [GKE security best practices](https://cloud.google.com/kubernetes-engine/docs/concepts/security-overview), and the [Spice.ai Kubernetes Deployment Guide](https://spiceai.org/docs/deployment/kubernetes). ### Cloud Run[​](#cloud-run "Direct link to Cloud Run") [Cloud Run](https://cloud.google.com/run) is a serverless container platform suitable for HTTP-driven Spice.ai workloads that benefit from scale-to-zero and request-based autoscaling. Use it when a single managed container is sufficient and operating Kubernetes is not desired. #### 1. Configure a service account[​](#1-configure-a-service-account "Direct link to 1. Configure a service account") Create a service account with the IAM roles the Spicepod requires. Cloud Run attaches it to the service so the runtime authenticates via [Application Default Credentials](https://cloud.google.com/docs/authentication/application-default-credentials) without static keys: ``` gcloud iam service-accounts create spiceai-runtime --project $PROJECT gcloud projects add-iam-policy-binding $PROJECT \ --member "serviceAccount:spiceai-runtime@${PROJECT}.iam.gserviceaccount.com" \ --role roles/storage.objectViewer gcloud projects add-iam-policy-binding $PROJECT \ --member "serviceAccount:spiceai-runtime@${PROJECT}.iam.gserviceaccount.com" \ --role roles/secretmanager.secretAccessor ``` #### 2. Deploy Spice.ai[​](#2-deploy-spiceai "Direct link to 2. Deploy Spice.ai") Cloud Run pulls the [Spice.ai container image](https://hub.docker.com/r/spiceai/spiceai) directly. Mount secrets from [Secret Manager](https://cloud.google.com/run/docs/configuring/secrets) and configure HTTP ingress on port `8090`: ``` gcloud run deploy spiceai \ --project $PROJECT \ --region $REGION \ --image spiceai/spiceai:2.0.0 \ --port 8090 \ --service-account spiceai-runtime@${PROJECT}.iam.gserviceaccount.com \ --min-instances 1 --max-instances 5 \ --cpu 1 --memory 2Gi \ --set-env-vars SPICED_LOG=INFO \ --set-secrets SPICEAI_API_KEY=spiceai-api-key:latest ``` To run multiple replicas with shared file-based acceleration, mount [Cloud Storage with FUSE](https://cloud.google.com/run/docs/configuring/services/cloud-storage-volume-mounts) and point file accelerators at the mount path (for example, `duckdb_file: /data/taxi_trips.db`). Cloud Storage volume latency is significantly higher than local SSD, so prefer GKE for latency-sensitive accelerated workloads. #### 3. Scaling rules[​](#3-scaling-rules "Direct link to 3. Scaling rules") Cloud Run scales by concurrent requests per instance ([default 80](https://cloud.google.com/run/docs/about-concurrency)). For background workloads (refresh schedules, ingestion) that should not scale to zero, set `--min-instances 1`. For workloads with long-running connections (Arrow Flight, streaming refresh), set `--no-cpu-throttling` and tune `--concurrency` to match the runtime's request profile. #### 4. Health probes and revisions[​](#4-health-probes-and-revisions "Direct link to 4. Health probes and revisions") Cloud Run uses [startup and liveness probes](https://cloud.google.com/run/docs/configuring/healthchecks) — point them at `/health` and `/v1/ready`. Each `gcloud run deploy` creates a new [revision](https://cloud.google.com/run/docs/managing/revisions); use [traffic splitting](https://cloud.google.com/run/docs/rollouts-rollbacks-traffic-migration) for canary upgrades: ``` gcloud run services update-traffic spiceai \ --region $REGION \ --to-revisions spiceai-00010-abc=90,spiceai-00009-xyz=10 ``` For more details, see the [Cloud Run documentation](https://cloud.google.com/run/docs) and the [Spice.ai Docker Deployment Guide](https://spiceai.org/docs/deployment/docker). ### Compute Engine[​](#compute-engine "Direct link to Compute Engine") Deploy Spice directly on [Compute Engine](https://cloud.google.com/compute) for maximum control over the environment, GPU access, or large-memory machine types. 1. **Manual VM deployment**: * Provision a Linux VM (Ubuntu, Debian, or Container-Optimized OS) with an appropriate [machine type](https://cloud.google.com/compute/docs/machine-resource). * Install [Docker Engine](https://docs.docker.com/engine/install/) and run [Spice.ai as a Docker container](https://spiceai.org/docs/deployment/docker), or install the `spice` binary directly. See the [installation guide](https://spiceai.org/docs/installation). * Attach a [service account](https://cloud.google.com/compute/docs/access/service-accounts) so Spice can read from Cloud Storage, BigQuery, and Secret Manager without static credentials. 2. **Automated deployment with Terraform or Deployment Manager**: * Define infrastructure in a [Terraform configuration](https://registry.terraform.io/providers/hashicorp/google/latest), including the VM, network, firewall rules, and service account. * Use [startup scripts](https://cloud.google.com/compute/docs/instances/startup-scripts/linux) or [Container-Optimized OS with `cloud-init`](https://cloud.google.com/container-optimized-os/docs/how-to/run-container-instance) to install Docker, pull the [Spice.ai image](https://hub.docker.com/r/spiceai/spiceai), retrieve secrets from [Secret Manager](https://cloud.google.com/secret-manager), and start the runtime. * Use [managed instance groups](https://cloud.google.com/compute/docs/instance-groups) for horizontally scaled deployments fronted by an [external HTTP(S) load balancer](https://cloud.google.com/load-balancing/docs/https) or [internal load balancer](https://cloud.google.com/load-balancing/docs/internal). For detailed guidance, refer to the [Compute Engine documentation](https://cloud.google.com/compute/docs), the [Container-Optimized OS guide](https://cloud.google.com/container-optimized-os/docs), and the [Google provider for Terraform](https://registry.terraform.io/providers/hashicorp/google/latest/docs). ## Authentication[​](#authentication "Direct link to Authentication") Most GCP services that Spice connects to accept explicit credentials through component parameters (for example, `iceberg_gcs_credentials` on the [Iceberg connector](/docs/next/components/data-connectors/iceberg)). When explicit credentials are not provided, Spice follows the standard [Application Default Credentials](https://cloud.google.com/docs/authentication/application-default-credentials) chain: 1. **`GOOGLE_APPLICATION_CREDENTIALS`** — path to a service account JSON key file. Common in local development; not recommended for production. 2. **Attached service account** — the credential of the runtime environment: * Compute Engine, Cloud Run, GKE node default service account. * GKE pods configured with [Workload Identity](https://cloud.google.com/kubernetes-engine/docs/concepts/workload-identity) — federated tokens scoped to a namespaced Kubernetes ServiceAccount, with no static keys on the node. 3. **`gcloud` CLI credentials** — cached credentials from `gcloud auth application-default login`. Common during development. 4. **Workload Identity Federation** — federated identity for workloads running outside GCP (other clouds, on-premises, GitHub Actions). See [Workload Identity Federation](https://cloud.google.com/iam/docs/workload-identity-federation). For services with explicit parameters (Cloud Storage HMAC, BigQuery service account JSON), prefer named credentials or Workload Identity over `GOOGLE_APPLICATION_CREDENTIALS` files in production. IAM role bindings Regardless of the credential source, the principal must have the appropriate IAM role bindings (for example, `roles/storage.objectViewer` on a bucket, or `roles/bigquery.dataViewer` on a BigQuery dataset). When a Spicepod connects to multiple GCP services, the principal must have permissions across all of them. ## Resources[​](#resources "Direct link to Resources") ### Documentation[​](#documentation "Direct link to Documentation") * [GCP Integrations](/docs/next/deployment/gcp/integrations) — complete list of GCP data connectors, AI models, and supported services. * [Spice.ai Kubernetes Deployment Guide](/docs/next/deployment/kubernetes) — Helm, Argo CD, and Flux options for GKE. ### Google Cloud Marketplace[​](#google-cloud-marketplace "Direct link to Google Cloud Marketplace") Spice.ai is not yet published to [Google Cloud Marketplace](https://console.cloud.google.com/marketplace) (coming soon). In the meantime, deploy using the [`spiceai/spiceai`](https://hub.docker.com/r/spiceai/spiceai) container image or the [Spice Helm chart](https://helm.spiceai.org). --- # GCP Integrations Spice.ai integrates with Google Cloud Platform (GCP) for data federation, AI inference, embeddings, and authentication. This page consolidates GCP-compatible components and links to the relevant configuration guides. ## Data Connectors[​](#data-connectors "Direct link to Data Connectors") Data connectors federate SQL queries across GCP data sources without data movement. | Connector | Description | Documentation | | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- | | **BigQuery** (via ADBC) | Query [BigQuery](https://cloud.google.com/bigquery) tables using the [BigQuery ADBC driver](https://docs.adbc-drivers.org/drivers/bigquery/index.html). Includes built-in SQL dialect support for federated queries. | [ADBC Data Connector](/docs/next/components/data-connectors/adbc) | | **Cloud Storage** (S3-compat) | Query Parquet, CSV, and JSON objects in [Cloud Storage](https://cloud.google.com/storage) using the S3 connector with HMAC keys against the GCS [interoperability endpoint](https://cloud.google.com/storage/docs/aws-simple-migration). | [S3 Data Connector](/docs/next/components/data-connectors/s3) | | **Cloud SQL for PostgreSQL** | Connect to [Cloud SQL for PostgreSQL](https://cloud.google.com/sql/docs/postgres) directly or through the [Cloud SQL Auth Proxy](https://cloud.google.com/sql/docs/postgres/sql-proxy). | [PostgreSQL Data Connector](/docs/next/components/data-connectors/postgres) | | **Cloud SQL for MySQL** | Connect to [Cloud SQL for MySQL](https://cloud.google.com/sql/docs/mysql) directly or through the Cloud SQL Auth Proxy. | [MySQL Data Connector](/docs/next/components/data-connectors/mysql) | | **Cloud SQL for SQL Server** | Connect to [Cloud SQL for SQL Server](https://cloud.google.com/sql/docs/sqlserver). | [MSSQL Data Connector](/docs/next/components/data-connectors/mssql) | | **AlloyDB for PostgreSQL** | Connect to [AlloyDB](https://cloud.google.com/alloydb) using the PostgreSQL wire protocol. | [PostgreSQL Data Connector](/docs/next/components/data-connectors/postgres) | | **Apache Iceberg (GCS)** | Query Iceberg tables stored in Cloud Storage with REST or Hive metadata. Native GCS authentication via service account credentials or OAuth tokens. | [Iceberg Data Connector](/docs/next/components/data-connectors/iceberg) | | **Delta Lake (GCS)** | Query Delta Lake tables stored in Cloud Storage. | [Delta Lake Data Connector](/docs/next/components/data-connectors/delta-lake) | | **GCP databases via ODBC** | Connect through ODBC drivers for additional GCP-compatible data sources. | [ODBC Data Connector](/docs/next/components/data-connectors/odbc) | ### Example: BigQuery via ADBC[​](#example-bigquery-via-adbc "Direct link to Example: BigQuery via ADBC") ``` datasets: - from: adbc:my_dataset.orders name: orders params: adbc_driver: bigquery adbc_uri: 'bigquery:///my-gcp-project' adbc_driver_options: | adbc.bigquery.sql.dataset_id=my_dataset adbc.bigquery.sql.auth_type=adbc.bigquery.sql.auth_type.json_credential_file adbc.bigquery.sql.auth_credentials=/var/run/secrets/gcp/key.json ``` When the runtime uses Workload Identity, omit `auth_type` and `auth_credentials` — the BigQuery driver picks up Application Default Credentials automatically. ### Example: Cloud Storage via S3 connector[​](#example-cloud-storage-via-s3-connector "Direct link to Example: Cloud Storage via S3 connector") Cloud Storage exposes an [S3-compatible interoperability endpoint](https://cloud.google.com/storage/docs/aws-simple-migration). Generate an [HMAC key](https://cloud.google.com/storage/docs/authentication/hmackeys) tied to a service account, then point the S3 connector at `storage.googleapis.com`: ``` datasets: - from: s3://my-bucket/path/to/data/ name: events params: file_format: parquet s3_endpoint: https://storage.googleapis.com s3_auth: key s3_key: ${ secrets:GCS_HMAC_ACCESS_ID } s3_secret: ${ secrets:GCS_HMAC_SECRET } ``` For Iceberg or Delta tables stored in GCS, use the connector's native GCS parameters instead, which support service-account credentials and ADC directly. ### Example: Cloud SQL for PostgreSQL[​](#example-cloud-sql-for-postgresql "Direct link to Example: Cloud SQL for PostgreSQL") Run the [Cloud SQL Auth Proxy](https://cloud.google.com/sql/docs/postgres/sql-proxy) as a sidecar (or `127.0.0.1` listener on Compute Engine) and connect over the loopback interface: ``` datasets: - from: postgres:public.orders name: orders params: pg_host: 127.0.0.1 pg_port: '5432' pg_db: app pg_user: ${ secrets:CLOUDSQL_USER } pg_pass: ${ secrets:CLOUDSQL_PASSWORD } pg_sslmode: disable ``` For [Postgres replication-based CDC](/docs/next/features/cdc/postgres-replication), Cloud SQL requires the `cloudsql.logical_decoding = on` flag. ## AI Models (Google AI)[​](#ai-models-google-ai "Direct link to AI Models (Google AI)") Spice integrates with [Google AI Studio](https://aistudio.google.com) for chat completion and reasoning models, including the Gemini family. | Provider | Supported Models | Documentation | | ------------- | ----------------------------------------------------------------------------------------------------------- | ------------------------------------------------------- | | **Google AI** | Gemini 2.0/2.5/Pro, Gemini Flash, and other models from the [Gemini API](https://ai.google.dev/gemini-api). | [Google AI Models](/docs/next/components/models/google) | ### Example: Gemini Chat Model[​](#example-gemini-chat-model "Direct link to Example: Gemini Chat Model") ``` models: - from: google:gemini-2.0-flash-exp name: gemini params: google_api_key: ${ secrets:GEMINI_API_KEY } ``` See [Google AI Models](https://ai.google.dev/gemini-api/docs/models/gemini) for the full list of supported model names. ## Embeddings (Google AI)[​](#embeddings-google-ai "Direct link to Embeddings (Google AI)") Generate vector embeddings using Gemini embedding models for semantic search and retrieval-augmented generation (RAG). | Provider | Supported Models | Documentation | | ------------- | --------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------- | | **Google AI** | `gemini-embedding-2` and other models from [Gemini API embeddings](https://ai.google.dev/gemini-api/docs/models/gemini#text-embedding). | [Google AI Embeddings](/docs/next/components/embeddings/google) | ### Example: Google AI Embeddings[​](#example-google-ai-embeddings "Direct link to Example: Google AI Embeddings") ``` embeddings: - from: google:gemini-embedding-2 name: gemini_embeddings params: google_api_key: ${ secrets:GEMINI_API_KEY } ``` ## Snapshots and shared state[​](#snapshots-and-shared-state "Direct link to Snapshots and shared state") [Snapshots](/docs/next/features/data-acceleration/snapshots) and the [distributed query](/docs/next/features/distributed-query) state location can use Cloud Storage as the shared object store. Configure with the `gs://` scheme: ``` snapshots: location: gs://my-bucket/spiceai/snapshots ``` When no explicit credentials are supplied, Spice reads `GOOGLE_APPLICATION_CREDENTIALS` and the Workload Identity-federated token, in that order. ## Authentication[​](#authentication "Direct link to Authentication") All GCP integrations support the standard [Application Default Credentials](https://cloud.google.com/docs/authentication/application-default-credentials) chain. When credentials are not explicitly configured, Spice attempts the following in order: 1. **`GOOGLE_APPLICATION_CREDENTIALS`** — path to a service account JSON key file. 2. **Attached service account** — Compute Engine, Cloud Run, or GKE node default service account. 3. **GKE Workload Identity** — federated tokens for pods bound to a Google service account via the Kubernetes ServiceAccount. See [Workload Identity for GKE](https://cloud.google.com/kubernetes-engine/docs/concepts/workload-identity). 4. **`gcloud` CLI** — cached credentials from `gcloud auth application-default login`. 5. **Workload Identity Federation** — federated identity for workloads running outside GCP (other clouds, on-premises, GitHub Actions). See [Workload Identity Federation](https://cloud.google.com/iam/docs/workload-identity-federation). For a deployment-side overview of these mechanisms, see the [Authentication](/docs/next/deployment/gcp#authentication) section of the GCP deployment guide. ### IAM role bindings[​](#iam-role-bindings "Direct link to IAM role bindings") Each principal must have the appropriate IAM role for the services it accesses: | Service | Common role(s) | | ------------------------ | ----------------------------------------------------------- | | Cloud Storage | `roles/storage.objectViewer` or `roles/storage.objectAdmin` | | BigQuery | `roles/bigquery.dataViewer` and `roles/bigquery.jobUser` | | Cloud SQL | `roles/cloudsql.client` (proxy) plus database-level grants | | Secret Manager | `roles/secretmanager.secretAccessor` | | Artifact Registry | `roles/artifactregistry.reader` for image pulls | | Cloud Logging/Monitoring | `roles/logging.logWriter`, `roles/monitoring.metricWriter` | When a Spicepod connects to multiple GCP services, ensure roles are granted on every resource the runtime touches. --- # Kubernetes Deployment Spice.ai runs on any Kubernetes cluster — managed (EKS, GKE, AKS) or self-hosted (kubeadm, k3s, Kind, RKE2). The official [Spice Helm chart](https://github.com/spiceai/helm-charts) is the foundation for all three deployment paths covered here. Choose the workflow that matches the operating model: | Path | When to use | | -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | [Helm](/docs/next/deployment/kubernetes/helm) | Direct, imperative deploys with `helm install` / `helm upgrade`. Simplest path for getting started or CI-driven pipelines that already invoke Helm. | | [Argo CD](/docs/next/deployment/kubernetes/argocd) | GitOps with continuous reconciliation. Define the desired state in Git; Argo CD applies and self-heals. | | [Flux](/docs/next/deployment/kubernetes/flux) | GitOps with native Kubernetes-style controllers (`HelmRelease`, `HelmRepository`, `Kustomization`). Lightweight alternative to Argo CD. | All three options consume the same chart and `values.yaml`, so configuration learned in one path transfers directly to the others. Spice.ai Enterprise Operator For production lifecycle management beyond what the Helm chart provides, the [Spice.ai Enterprise](https://spice.ai) Kubernetes Operator introduces two custom resources: * [`SpicepodSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset) — declarative replica management with automatic PVC resizing, rolling/parallel update strategies, crashloop protection, and per-replica `StatefulSet`s when persistent volumes are configured. Use it instead of the chart's `stateful` mode for stateful workloads. * [`SpicepodCluster`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodcluster) — distributed query clusters with dedicated scheduler and executor nodes, automatic mTLS, and shared object-store-backed state. Use it for horizontally scaled query execution and high availability. The operator works alongside Helm, Argo CD, and Flux — install the operator chart and manage `SpicepodSet` / `SpicepodCluster` resources from the same GitOps pipeline. ## CPU sizing[​](#cpu-sizing "Direct link to CPU sizing") Spice sizes its thread pools, query fan-out, and accelerator concurrency from a single CPU entitlement. In Kubernetes that entitlement comes from the pod's own resources, and the default needs no configuration. | Pod resources | Entitlement | | ------------------------------- | ------------------------------------------------------------------------------------------------------------ | | `requests.cpu`, no `limits.cpu` | **Twice the request**, floored at 2 cores and capped by the node — a `requests.cpu: 4` pod sizes for 8 cores | | `limits.cpu` set | The limit | | Neither | Every CPU the pod can see | Sizing above the request is deliberate: a CPU request is a scheduling floor rather than a ceiling, so a burstable pod keeps headroom to burst above it — while not building 64 worker threads for a node it only has a slice of. This requires the pod spec to pass its CPU request in, since the runtime cannot read `resources.requests.cpu` itself. **The Spice Helm chart and the Spice Kubernetes Operator both do this by default** whenever a CPU request is set, so deployments using either — including all three paths above — get it automatically. A hand-written pod spec needs the [downward-API block](/docs/next/reference/spicepod/runtime#sizing-from-a-cpu-request) itself; without it the pod sizes for the whole node, and the runtime warns at startup. ### Bursting across the whole machine[​](#bursting-across-the-whole-machine "Direct link to Bursting across the whole machine") To pack many mostly-idle instances onto a node while letting each burst wide, state that intent once: ``` runtime: cpu: cores: all # every available core, regardless of the CPU request # (a CPU limit, if one is set, is still respected) ``` The [Spice Cloud Platform](https://spice.ai) sets `SPICE_CPU_CORES=all` on hosted instances for exactly this reason — maximum burst capacity, so an instance is never sized down to a fraction of the machine it is scheduled on. Prefer `runtime.cpu.cores` over `resources.limits.cpu` for bounding Spice. A CPU limit is a CFS quota and [throttles](https://home.robusta.dev/blog/stop-using-cpu-limits) even when the node has idle CPU; `runtime.cpu.cores` caps how much machine the runtime organizes itself around without capping how much CPU it may use. See [`runtime.cpu`](/docs/next/reference/spicepod/runtime#runtimecpu) and [Resource Allocation](/docs/next/reference/performance-tuning#resource-allocation). ## Memory sizing[​](#memory-sizing "Direct link to Memory sizing") Unlike CPU, memory is enforced by the kernel: a pod that exceeds `resources.limits.memory` is OOM-killed rather than throttled. Size the limit for the whole process, not just for query execution. A pod's memory divides into a **baseline** that scales with the number of configured datasets, and a **working set** that scales with data volume, per-query cost, and concurrency. The runtime derives `runtime.query.memory_limit` from the pod's cgroup limit with the per-dataset baseline already subtracted, so the common case needs no configuration — leave it unset and set `resources.limits.memory` instead. Note that much of the baseline is bounded buffers that grow toward their limits as traffic drives them there, so a pod's memory shortly after startup understates what it will settle at. Size the limit from a sustained run, not from a fresh pod. | Situation | What to set | | ------------------------------------------------ | ------------------------------------------------------------------------------------------------- | | Standard deployment | `resources.limits.memory` only; let the runtime derive the query pool | | Co-located accelerators managing their own pools | `resources.limits.memory` plus an explicit `runtime.query.memory_limit` that leaves room for them | | Sizing a smaller staging or test environment | Reduce data volume and concurrency; keep the dataset count and accelerator configuration | Two Kubernetes-specific consequences: * **Do not scale the production memory limit down proportionally for a lower environment.** The baseline does not shrink with the data, so a proportionally smaller pod is not a smaller model of production — it is a different regime that predicts nothing about production behavior, and the limits tuned to fit it do not transfer back. * **A pod that plateaus just under its memory limit is sized too tightly**, even while healthy. It has no margin for a refresh spike or a burst of concurrency. Because instances are evicted, preempted, and replaced during rollouts, clients should retry a failed query against the Service before treating it as an outage. See [Managing Memory Usage](/docs/next/reference/memory) for the full sizing model, the load-testing properties that make a memory validation trustworthy, and client resiliency guidance. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to a Kubernetes cluster (v1.25+ recommended). * `kubectl` configured for the target cluster. See the [Kubernetes documentation](https://kubernetes.io/docs/tasks/tools/#kubectl). * For local testing, [Kind](https://kind.sigs.k8s.io/docs/user/quick-start/) or [k3d](https://k3d.io/) provide a quick single-node cluster. ## Cookbook[​](#cookbook "Direct link to Cookbook") * [Running Spice.ai in Kubernetes](https://github.com/spiceai/cookbook/tree/trunk/kubernetes) --- # Kubernetes - Argo CD This guide describes how to deploy Spice.ai on a self-hosted Kubernetes cluster using [Argo CD](https://argo-cd.readthedocs.io/) and the official [Spice Helm chart](https://github.com/spiceai/spiceai/tree/trunk/deploy/chart). Argo CD continuously reconciles the cluster state with manifests stored in Git, providing a GitOps workflow for managing Spice.ai deployments. ## Quickstart[​](#quickstart "Direct link to Quickstart") ``` kubectl apply -n argocd -f - < ``` When the `Application` is managed by Git, prefer reverting the commit that introduced the change so the desired state in Git matches the live state. With `selfHeal: true` enabled, Argo CD will otherwise re-apply the Git state. ## Uninstall[​](#uninstall "Direct link to Uninstall") Delete the `Application` to remove the Spice.ai deployment. With `prune: true`, all resources created by the chart are removed: ``` kubectl delete application spiceai -n argocd ``` PersistentVolumeClaims created by a StatefulSet are retained by default and must be deleted manually if no longer needed: ``` kubectl delete pvc -n spiceai -l app.kubernetes.io/name=spiceai ``` ## Secrets Management[​](#secrets-management "Direct link to Secrets Management") Do not commit plain-text secrets to Git. Reference Kubernetes `Secret` resources from `additionalEnv` and manage the secrets themselves with one of: * [Sealed Secrets](https://github.com/bitnami-labs/sealed-secrets) — encrypted secrets safe to store in Git. * [External Secrets Operator](https://external-secrets.io/) — sync secrets from AWS Secrets Manager, HashiCorp Vault, GCP Secret Manager, and similar. * [SOPS](https://github.com/getsops/sops) with the [Argo CD Vault Plugin](https://argocd-vault-plugin.readthedocs.io/). Example referencing a pre-existing `Secret` named `spice-secrets`: ``` spec: source: helm: valuesObject: additionalEnv: - name: SPICE_SECRET_SPICEAI_KEY valueFrom: secretKeyRef: name: spice-secrets key: spiceai-key ``` See [Environment Variables and Secrets](/docs/next/deployment/kubernetes/helm#environment-variables-and-secrets) in the Helm guide for additional configuration patterns. ## Monitoring[​](#monitoring "Direct link to Monitoring") The Spice Helm chart integrates with the [Prometheus Operator](https://prometheus-operator.dev/) through a `PodMonitor` resource. Enable it with: ``` spec: source: helm: valuesObject: monitoring: podMonitor: enabled: true additionalLabels: release: prometheus-stack ``` Once metrics are collected, import the [Spice Grafana dashboard](/docs/next/monitoring/grafana) for visualization. ## Manage Multiple Spice.ai Releases[​](#manage-multiple-spiceai-releases "Direct link to Manage Multiple Spice.ai Releases") To deploy multiple Spice.ai instances (for example, one per environment or tenant), create one `Application` per release with distinct names, namespaces, and values. The [App of Apps](https://argo-cd.readthedocs.io/en/stable/operator-manual/cluster-bootstrapping/) pattern or [ApplicationSets](https://argo-cd.readthedocs.io/en/stable/operator-manual/applicationset/) can manage these declaratively at scale. Example `ApplicationSet` that deploys Spice.ai to multiple namespaces: ``` apiVersion: argoproj.io/v1alpha1 kind: ApplicationSet metadata: name: spiceai namespace: argocd spec: generators: - list: elements: - name: spiceai-dev namespace: spiceai-dev chartVersion: 1.11.5 - name: spiceai-prod namespace: spiceai-prod chartVersion: 1.11.5 template: metadata: name: '{{name}}' spec: project: default source: repoURL: https://helm.spiceai.org chart: spiceai targetRevision: '{{chartVersion}}' helm: releaseName: '{{name}}' destination: server: https://kubernetes.default.svc namespace: '{{namespace}}' syncPolicy: automated: prune: true selfHeal: true syncOptions: - CreateNamespace=true ``` ## Further Reading[​](#further-reading "Direct link to Further Reading") * [Kubernetes - Helm deployment guide](/docs/next/deployment/kubernetes/helm) — full reference for the Spice Helm chart, including all configurable parameters. * [Argo CD Helm chart documentation](https://argo-cd.readthedocs.io/en/stable/user-guide/helm/) * [Argo CD ApplicationSet controller](https://argo-cd.readthedocs.io/en/stable/operator-manual/applicationset/) * [Spice Helm chart source](https://github.com/spiceai/spiceai/tree/trunk/deploy/chart) ## Cookbook[​](#cookbook "Direct link to Cookbook") * [Running Spice.ai in Kubernetes](https://github.com/spiceai/cookbook/tree/trunk/kubernetes) --- # Kubernetes - Flux This guide describes how to deploy Spice.ai on Kubernetes using [Flux CD](https://fluxcd.io/) and the official [Spice Helm chart](https://github.com/spiceai/helm-charts). Flux is a GitOps toolkit that continuously reconciles cluster state with manifests in a Git repository or a Helm chart repository, using native Kubernetes custom resources. ## Quickstart[​](#quickstart "Direct link to Quickstart") Register the Spice Helm repository and create a `HelmRelease` in a single manifest: ``` kubectl apply -n flux-system -f - < -n spiceai ``` For automatic rollback on failed upgrades, configure `spec.upgrade.remediation` on the `HelmRelease`: ``` spec: upgrade: remediation: retries: 3 strategy: rollback ``` ## Suspend and Resume[​](#suspend-and-resume "Direct link to Suspend and Resume") Suspend reconciliation to make manual changes without Flux reverting them: ``` flux suspend helmrelease spiceai -n flux-system # ... manual changes ... flux resume helmrelease spiceai -n flux-system ``` ## Uninstall[​](#uninstall "Direct link to Uninstall") Delete the `HelmRelease` to remove the Spice.ai deployment. Flux runs `helm uninstall`, which removes all chart-managed resources: ``` kubectl delete helmrelease spiceai -n flux-system ``` PersistentVolumeClaims created by a StatefulSet are retained by default and must be deleted manually if no longer needed: ``` kubectl delete pvc -n spiceai -l app.kubernetes.io/name=spiceai ``` ## Secrets Management[​](#secrets-management "Direct link to Secrets Management") Do not commit plain-text secrets to Git. Reference Kubernetes `Secret` resources from `additionalEnv` and manage the secrets themselves with one of: * [SOPS](https://github.com/getsops/sops) — Flux has built-in support via `kustomize-controller`. See [Manage Kubernetes secrets with SOPS](https://fluxcd.io/flux/guides/mozilla-sops/). * [Sealed Secrets](https://github.com/bitnami-labs/sealed-secrets) — encrypted secrets safe to store in Git. * [External Secrets Operator](https://external-secrets.io/) — sync secrets from AWS Secrets Manager, HashiCorp Vault, GCP Secret Manager, and similar. Example referencing a pre-existing `Secret` named `spice-secrets`: ``` spec: values: additionalEnv: - name: SPICE_SECRET_SPICEAI_KEY valueFrom: secretKeyRef: name: spice-secrets key: spiceai-key ``` See [Environment Variables and Secrets](/docs/next/deployment/kubernetes/helm#environment-variables-and-secrets) in the Helm guide for additional configuration patterns. ## Monitoring[​](#monitoring "Direct link to Monitoring") The Spice Helm chart integrates with the [Prometheus Operator](https://prometheus-operator.dev/) through a `PodMonitor` resource. Enable it in the `HelmRelease` values: ``` spec: values: monitoring: podMonitor: enabled: true additionalLabels: release: prometheus-stack ``` Once metrics are collected, import the [Spice Grafana dashboard](/docs/next/monitoring/grafana) for visualization. ## Manage Multiple Spice.ai Releases[​](#manage-multiple-spiceai-releases "Direct link to Manage Multiple Spice.ai Releases") To deploy multiple Spice.ai instances (for example, one per environment or tenant), create one `HelmRelease` per release with distinct names, target namespaces, and values. Group related releases under a single `Kustomization` so Flux applies them together. Example directory layout in a Flux-managed repository: ``` clusters/ └── production/ ├── flux-system/ │ └── ... └── spiceai/ ├── kustomization.yaml ├── helmrepository.yaml ├── spiceai-app.yaml └── spiceai-analytics.yaml ``` `clusters/production/spiceai/kustomization.yaml`: ``` apiVersion: kustomize.config.k8s.io/v1beta1 kind: Kustomization resources: - helmrepository.yaml - spiceai-app.yaml - spiceai-analytics.yaml ``` Each `HelmRelease` in this directory references the shared `HelmRepository` source. ## Further Reading[​](#further-reading "Direct link to Further Reading") * [Kubernetes - Helm deployment guide](/docs/next/deployment/kubernetes/helm) — full reference for the Spice Helm chart, including all configurable parameters. * [Flux Helm Controller documentation](https://fluxcd.io/flux/components/helm/) * [Flux `HelmRelease` API reference](https://fluxcd.io/flux/components/helm/helmreleases/) * [Flux `HelmRepository` API reference](https://fluxcd.io/flux/components/source/helmrepositories/) * [Flux GitOps guides](https://fluxcd.io/flux/guides/) * [Spice Helm chart source](https://github.com/spiceai/helm-charts) ## Cookbook[​](#cookbook "Direct link to Cookbook") * [Running Spice.ai in Kubernetes](https://github.com/spiceai/cookbook/tree/trunk/kubernetes) --- # Kubernetes - Helm ## Quickstart[​](#quickstart "Direct link to Quickstart") ``` helm repo add spiceai https://helm.spiceai.org helm repo update helm upgrade --install spiceai spiceai/spiceai ``` Deployment Architecture By default, the Spice.ai Helm chart deploys the application as a stateless Kubernetes Deployment. To persist data between restarts (e.g., for file-based acceleration), enable and configure the `stateful` section in the values file. Refer to the [Stateful Configuration](#stateful-configuration) section for details. ## What are Kubernetes and Helm?[​](#what-are-kubernetes-and-helm "Direct link to What are Kubernetes and Helm?") **Kubernetes** is an open-source platform for automating deployment, scaling, and management of containerized applications.
**Helm** is a package manager for Kubernetes that simplifies the installation and configuration of applications using reusable templates called charts. Spice publishes a Helm chart that simplifies the deployment of Spice.ai OSS on Kubernetes. ## Deploy Spice using Helm in Kubernetes[​](#deploy-spice-using-helm-in-kubernetes "Direct link to Deploy Spice using Helm in Kubernetes") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to a Kubernetes cluster. * For local testing, try running a local Kubernetes cluster using [Kind](https://kind.sigs.k8s.io/docs/user/quick-start/). * `kubectl` CLI installed and configured to interact with the target Kubernetes cluster. Visit the [Kubernetes docs](https://kubernetes.io/docs/tasks/tools/#kubectl) for installation instructions. * Helm CLI installed. Visit the [Helm docs](https://helm.sh/docs/intro/install/) for installation instructions. ### Add the Spice Helm repository[​](#add-the-spice-helm-repository "Direct link to Add the Spice Helm repository") A Helm repository (Helm repo) is a storage location where Helm charts are hosted and can be accessed for deployment in Kubernetes clusters. Add the Spice Helm repository to your local Helm client and update the index to get the latest charts: ``` helm repo add {repository-name} https://helm.spiceai.org helm repo update ``` The repository name is customizable and can be set to any preferred value. For example: ``` helm repo add spiceai https://helm.spiceai.org # or helm repo add my-spiceai-repo https://helm.spiceai.org ``` ### Install the Spice Helm chart as a new release[​](#install-the-spice-helm-chart-as-a-new-release "Direct link to Install the Spice Helm chart as a new release") Once the repository with the chart is added, install the chart into the Kubernetes cluster with the following command: ``` helm install {release-name} {repository-name}/spiceai --namespace {namespace} ``` For example: ``` helm install spiceai spiceai/spiceai --namespace default ``` Spice can be installed multiple times in the same cluster by specifying a different release name for each installation. #### Command Breakdown[​](#command-breakdown "Direct link to Command Breakdown") * `helm install`: Installs a new Helm chart. To upgrade an existing release, use `helm upgrade`. Combine both upgrade and install by specifying `helm upgrade --install`. * `spiceai`: The name of the release. This name is customizable and can be set to any preferred value, e.g., `spiceai-my-app-v1` and `spiceai-my-app-v2` are valid release names. Each Helm release is a distinct installation of the same chart. * `spiceai/spiceai`: The chart to install. The first `spiceai` is the repository name added earlier, and the second `spiceai` is the name of the chart to install. While the repository name is customizable, the chart name is not. * `--namespace default`: The Kubernetes namespace to install the chart into. This is optional and defaults to `default`. Another valid command to install the chart is (assuming the repository name is `my-spiceai-repo`): ``` helm upgrade --install spiceai-my-app-1 my-spiceai-repo/spiceai ``` ### Upgrade the Spice Helm chart[​](#upgrade-the-spice-helm-chart "Direct link to Upgrade the Spice Helm chart") To upgrade an existing release, use the `helm upgrade` command: ``` helm upgrade {release-name} {repository-name}/{chart-name} ``` For example: ``` helm upgrade spiceai-my-app-1 my-spiceai-repo/spiceai ``` ### Rollback a Helm release[​](#rollback-a-helm-release "Direct link to Rollback a Helm release") On occasion, you may need to roll back a Spice Helm release to a previous version. To do so, use the `helm rollback` command. This will notify Kubernetes to redeploy Spice back to a previous version of the Helm release: ``` helm rollback {release-name} --namespace {namespace} ``` For example: ``` helm rollback spiceai-my-app-1 --namespace default ``` ### Uninstall a Helm release[​](#uninstall-a-helm-release "Direct link to Uninstall a Helm release") To uninstall a Helm release, use the `helm uninstall` command. This will cause Kubernetes to remove the Spice deployment entirely. Note that any data stored in volumes created by configuring the `stateful` parameter will be preserved and must be manually deleted if desired: ``` helm uninstall {release-name} --namespace {namespace} ``` For example: ``` helm uninstall spiceai-my-app-1 --namespace default ``` ## Customize the Helm release[​](#customize-the-helm-release "Direct link to Customize the Helm release") By default, the Helm release installs a minimal Spice.ai setup with an empty Spicepod. To add a Spicepod and adjust other settings, customize the release as needed by creating a `values.yaml` file or by using the `--set` flag. note `values.yaml` is the configuration file used in Helm to define the user-configurable parameters of a Helm chart. The `--set` flag is used to specify individual values on the command line. Visit the [Helm docs](https://helm.sh/docs/chart_template_guide/values_files/#helm) for more information. Create a `values.yaml` file and override the default values as needed. The full list of configurable parameters and their defaults are specified in the [values.yaml file](https://github.com/spiceai/spiceai/blob/trunk/deploy/chart/values.yaml) in the `spiceai/spiceai` repository. The [Common Parameters](#common-parameters) section below lists the most commonly used configurable parameters and their descriptions. ### Spicepod[​](#spicepod "Direct link to Spicepod") To customize the Spicepod that the Spice.ai runtime will load, define a [Spicepod](https://spiceai.org/docs/getting-started/spicepods) in a new `values.yaml` file. ``` spicepod: name: app version: v1 kind: Spicepod datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet ``` Upgrade or install a new release with the custom Spicepod: ``` helm upgrade --install spiceai spiceai/spiceai -f values.yaml ``` note The Helm convention is to use a file called `values.yaml`, but any file name can be used and passed to the `-f` flag. ## Common Parameters[​](#common-parameters "Direct link to Common Parameters") | **Name** | **Description** | **Value** | | ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------- | | `additionalEnv` | Additional environment variables to set in the Spice.ai container. | `[]` | | `additionalLabels` | Additional labels to add to all resources. | `{}` | | `image.pullSecrets` | Specify Docker registry secret names as an array. | `[]` | | `image.repository` | The repository of the Docker image. | `spiceai/spiceai` | | `image.tag` | Replace with a specific version of Spice.ai to run. | `latest` | | `monitoring.podMonitor.enabled` | Enable Prometheus metrics collection for the Spice pods. Requires the [Prometheus Operator](https://prometheus-operator.dev/docs/operator/api/#monitoring.coreos.com/v1.PodMonitor) CRDs. | `false` | | `replicaCount` | Number of Spice.ai replicas to run. | `1` | | `resources` | Resource requests and limits for the Spice.ai container. See [Container resource examples](https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/#example-1). | `{}` | | `service.type` | Kubernetes service type. Can be null, ClusterIP, NodePort, or LoadBalancer. | `null` | | `serviceAccount.create` | Specifies whether a ServiceAccount should be created. | `false` | | `livenessProbe` | Kubernetes liveness probe configuration. Uses standard [probe shape](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes). | see `values.yaml` | | `readinessProbe` | Kubernetes readiness probe configuration. Uses standard [probe shape](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes). | see `values.yaml` | | `startupProbe` | Kubernetes startup probe configuration. Uses standard [probe shape](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes). | see `values.yaml` | | `spicepod` | Define the [Spicepod](https://spiceai.org/docs/getting-started/spicepods) to be loaded by the Spice.ai runtime. | `{}` | | `stateful.enabled` | Use a StatefulSet with a PVC (Persistent Volume Claim) for the data volume. | `false` | | `stateful.mountPath` | Mount path in container for the persistent volume. | `/data` | | `stateful.size` | Size of each PV in the StatefulSet. | `1Gi` | | `stateful.storageClass` | Storage class for the volume claim template in the StatefulSet. | `standard` | | `tolerations` | List of node taints to tolerate. | `[]` | ## Environment Variables and Secrets[​](#environment-variables-and-secrets "Direct link to Environment Variables and Secrets") Add extra environment variables using the `additionalEnv` property. This can be useful when combining with the [Environment Secret Store](/docs/next/components/secret-stores/env). ``` additionalEnv: - name: SPICED_LOG value: 'DEBUG' - name: SPICE_SECRET_SPICEAI_KEY valueFrom: secretKeyRef: name: spice-secrets key: spiceai-key ``` To create a test secret: ``` kubectl create secret generic spice-secrets --from-literal=spiceai-key="secret-value" ``` Further reading: * [Kubernetes Secrets](https://kubernetes.io/docs/concepts/configuration/secret/) * [Good practices for Kubernetes Secrets](https://kubernetes.io/docs/concepts/security/secrets-good-practices/) ## Monitoring[​](#monitoring "Direct link to Monitoring") The Spice Helm chart includes compatibility with the [Prometheus Operator](https://prometheus-operator.dev/) for collecting Prometheus metrics that can be visualized in the [Spice Grafana dashboard](/docs/next/monitoring/grafana). To enable this feature, set the `monitoring.podMonitor.enabled` value to `true`. This will create a `PodMonitor` resource for the Spice.ai pods that will configure Prometheus to scrape metrics from the Spice.ai pods. Install the Prometheus Operator The easiest way to install the Prometheus Operator along with Grafana is to use the [kube-prometheus-stack](https://github.com/prometheus-community/helm-charts/blob/main/charts/kube-prometheus-stack/README) Helm chart. ``` helm repo add prometheus-community https://prometheus-community.github.io/helm-charts helm repo update helm install prometheus-stack prometheus-community/kube-prometheus-stack \ --set prometheus.prometheusSpec.podMonitorSelectorNilUsesHelmValues=false \ --set prometheus.prometheusSpec.serviceMonitorSelectorNilUsesHelmValues=false ``` Deploy the Spice.ai Helm chart with monitoring enabled: ``` helm upgrade --install spiceai spiceai/spiceai --set monitoring.podMonitor.enabled=true ``` Once the monitoring is enabled, import the [Spice Grafana dashboard](/docs/next/monitoring/grafana) to visualize the Spice.ai metrics. ### Health and Readiness[​](#health-and-readiness "Direct link to Health and Readiness") Spice provides two HTTP endpoints for monitoring the runtime state: `/health` and `/v1/ready`. These endpoints are used for Kubernetes health and readiness probes in the Spice deployment. The Spice Helm chart automatically configures these probes with sensible defaults, and all three probes (`livenessProbe`, `readinessProbe`, `startupProbe`) can be customized via Helm values. #### Health Probe (Liveness + Startup)[​](#health-probe-liveness--startup "Direct link to Health Probe (Liveness + Startup)") The `/health` endpoint indicates whether the Spice process is up and running. The default liveness and startup probes use this endpoint: ``` livenessProbe: httpGet: path: /health port: 8090 timeoutSeconds: 1 periodSeconds: 10 failureThreshold: 3 ``` #### Readiness Probe[​](#readiness-probe "Direct link to Readiness Probe") The `/v1/ready` endpoint indicates **whether the Spice components (datasets, models, etc.) are ready**. While `/health` shows that Spice is running, `/v1/ready` must return `200` to ensure queries will return results: ``` readinessProbe: httpGet: path: /v1/ready port: 8090 timeoutSeconds: 1 periodSeconds: 10 failureThreshold: 3 ``` #### Customizing Probes[​](#customizing-probes "Direct link to Customizing Probes") Probe values can be partially overridden — omitted fields keep the defaults from `values.yaml`. If `exec`, `tcpSocket`, or `grpc` is configured, the chart omits the default `httpGet` handler: ``` livenessProbe: httpGet: path: /health port: 8090 timeoutSeconds: 5 failureThreshold: 5 startupProbe: httpGet: path: /health port: 8090 failureThreshold: 30 periodSeconds: 5 ``` note For more information on how Kubernetes uses probes to determine the health of a pod, see [here](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes). ## Service Configuration[​](#service-configuration "Direct link to Service Configuration") Configure the Kubernetes service for Spice.ai: ``` service: # type can be null, ClusterIP, NodePort, or LoadBalancer type: LoadBalancer additionalAnnotations: {} # If service type is LoadBalancer, specify the source IP ranges to whitelist loadBalancerSourceRanges: - 10.0.0.0/8 - 172.16.0.0/12 # The selector to use for the service selector: custom-selector: value ``` ## Service Account[​](#service-account "Direct link to Service Account") Configure a Kubernetes ServiceAccount for Spice.ai: ``` serviceAccount: # Specifies whether a ServiceAccount should be created create: true # The name of the ServiceAccount to use. # If not set and create is true, a name is generated using the fullname template name: spice-service-account ``` ## Volumes and Volume Mounts[​](#volumes-and-volume-mounts "Direct link to Volumes and Volume Mounts") Define custom volumes and volume mounts for the Spice.ai container: ``` # Define volumes to be mounted to the container volumes: - name: custom-data persistentVolumeClaim: claimName: my-custom-pvc - name: config-volume configMap: name: custom-config # Define volume mounts for the container volumeMounts: - name: custom-data mountPath: /data - name: config-volume mountPath: /config ``` ## Stateful Configuration[​](#stateful-configuration "Direct link to Stateful Configuration") The Spice.ai Helm chart provides two deployment architectures to accommodate different persistence requirements. The default architecture deploys Spice.ai as a stateless Kubernetes Deployment, suitable for workloads that do not require data persistence between pod restarts. For workloads requiring data persistence, the chart supports deploying Spice.ai as a StatefulSet with persistent storage. This architecture becomes essential when implementing file-based acceleration for datasets or when maintaining state between pod restarts is critical. Spice.ai Enterprise For production stateful workloads, the [Spice.ai Enterprise](https://spice.ai) Operator's [`SpicepodSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset) resource offers per-replica StatefulSets with automatic PVC resizing, configurable update strategies, and crashloop protection. For distributed query execution with separate scheduler and executor tiers, see [`SpicepodCluster`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodcluster). Enabling the StatefulSet architecture requires configuration of the `stateful` section: ``` # Use a StatefulSet with a PVC for the data volume stateful: enabled: true # Storage class for the volume claim template storageClass: 'standard' # Size of each PV in the StatefulSet size: 1Gi # Mount path in container mountPath: /data ``` When `stateful.enabled` is set to `true`, the Helm chart creates a StatefulSet instead of a Deployment and provisions a PersistentVolumeClaim for each replica. The persistent volume is mounted at the specified path, so data persists across pod restarts and rescheduling events. ### Storage Class Recommendations[​](#storage-class-recommendations "Direct link to Storage Class Recommendations") Spice accelerations are latency- and IOPS-sensitive. Choose the StorageClass in the following order of preference: 1. **Local NVMe (recommended)** \u2014 the lowest-latency option is a node-local NVMe SSD, exposed via the [Local Volume Static Provisioner](https://github.com/kubernetes-sigs/sig-storage-local-static-provisioner) as a `local-storage` class. On managed clouds use an instance type with attached NVMe \u2014 AWS `i4i` / `m6id` / `c7gd` / `r7gd` (the `d`-suffixed families), Azure [Lsv3](https://learn.microsoft.com/azure/virtual-machines/lsv3-series) / [Ddsv5](https://learn.microsoft.com/azure/virtual-machines/ddv5-ddsv5-series), GCP [`*-lssd`](https://cloud.google.com/compute/docs/disks/local-ssd) machines. Self-hosted clusters can expose local NVMe directly. Local volumes do not survive node replacement, so pair with a refresh strategy or re-hydration source.\n2. **High-IOPS network block storage (shared / replica-attachable)** \u2014 when persistence must survive node replacement, prefer:\n - **AWS**: [Amazon EBS `io2` Block Express](https://docs.aws.amazon.com/ebs/latest/userguide/ebs-volume-types.html#io2-bx) (sub-millisecond latency, up to 256K IOPS), falling back to `gp3` with provisioned IOPS.\n - **Azure**: [Premium SSD v2](https://learn.microsoft.com/azure/virtual-machines/disks-types#premium-ssd-v2) (sub-millisecond latency, up to 80K IOPS), falling back to Premium SSD (`managed-csi-premium`).\n - **GCP**: [Hyperdisk Extreme](https://cloud.google.com/compute/docs/disks/hyperdisks), falling back to Persistent Disk SSD (`pd-ssd`).\n3. **Cayenne shared object storage (Cayenne only, optional)** \u2014 when Cayenne acceleration must be shared across replicas or persisted independently of pod lifecycle, [Amazon S3 Express One Zone](https://aws.amazon.com/s3/storage-classes/express-one-zone/) provides single-digit-millisecond object storage. Configure Cayenne to use an S3 Express directory bucket \u2014 see the [Cayenne acceleration documentation](/docs/next/components/data-accelerators/cayenne).\n\n:::warning\nNetwork file systems (NFS, EFS, Azure Files / SMB, Cloud Filestore) are not recommended for acceleration storage classes. Their latency negates the benefit of using a local accelerator. Reserve them for stateless artefacts that must survive pod replacement.\n::: ## Example values.yaml[​](#example-valuesyaml "Direct link to Example values.yaml") ``` additionalLabels: environment: production app: spice image: repository: spiceai/spiceai tag: latest replicaCount: 1 service: type: ClusterIP additionalAnnotations: service.beta.kubernetes.io/aws-load-balancer-internal: 'true' resources: limits: # cpu: 1000m # Do not set CPU limits - can cause throttling memory: 2Gi requests: cpu: 500m memory: 1Gi additionalEnv: - name: SPICED_LOG value: 'INFO' - name: SPICE_SECRET_SPICEAI_KEY valueFrom: secretKeyRef: name: spice-secrets key: spiceai-key stateful: enabled: true storageClass: 'standard' size: 5Gi mountPath: /data monitoring: podMonitor: enabled: true additionalLabels: release: prometheus spicepod: name: app version: v1 kind: Spicepod datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips description: Demo taxi trips in s3 params: file_format: parquet acceleration: enabled: true engine: duckdb mode: file params: duckdb_file: /data/taxi_trips.db # Uncomment to refresh the acceleration on a schedule # refresh_check_interval: 1h # refresh_mode: full ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * [Running Spice.ai in Kubernetes](https://github.com/spiceai/cookbook/tree/trunk/kubernetes) --- # Read/Write Separation Production data and AI applications typically have two very different workloads on the same data: * **Writes / ingest** — pulling from source systems, normalizing, accelerating, indexing, and refreshing. CPU-, network-, and memory-heavy. Bursty. Doesn't need to be co-located with the application. * **Reads** — answering user-facing requests, feeding context to AI agents, serving dashboards. Latency-sensitive. Often horizontally scaled with the application itself. Running both on the same Spice instance forces a single hardware shape, refresh schedule, and failure domain on workloads that have nothing in common. Read/write separation splits them into two tiers: a centralized **write/ingest cluster** that owns refresh and acceleration, and one or more lightweight **read instances** (typically [sidecars](/docs/next/deployment/architectures/sidecar) next to the application) that serve queries from a local materialized copy. The two tiers communicate through two channels: 1. **Snapshots** in object storage — the cluster periodically writes a compact acceleration file (DuckDB or SQLite) to S3, GCS, or ADLS. Read instances bootstrap from the latest snapshot on startup and (optionally) refresh from snapshots on a schedule. No live network dependency on the cluster. 2. **Live query delegation** — when a read instance needs data outside its materialized working set (a historical query, a cross-dataset join, a broad search), it transparently delegates to the cluster over Arrow Flight. See [Cluster-Sidecar Architecture](/docs/next/deployment/architectures/cluster-sidecar). Most production deployments use both: snapshots for the steady-state working set, and live delegation for the long tail. ## When to use read/write separation[​](#when-to-use-readwrite-separation "Direct link to When to use read/write separation") Use this pattern when: * Application instances need **sub-millisecond reads** but data refresh, ingestion, or acceleration would saturate them. * The same datasets are read by **many replicas**, each currently re-ingesting from source systems. * Read instances need to **start fast** — autoscaling, scale-to-many agent containers, or ephemeral Cloud Run / Knative workloads where cold-starting from source is too slow. * Upstream data sources have **rate or cost limits** that prevent every replica from connecting directly. * Read instances run **outside the cluster's network** — at the edge, in another VPC, on a developer laptop — and cannot maintain a permanent dependency on the source system. It is overkill when one Spice instance is sufficient (start with [Sidecar](/docs/next/deployment/architectures/sidecar)) or when the workload is purely batch/analytical with relaxed latency (use [Microservice](/docs/next/deployment/architectures/microservice)). ## How it works[​](#how-it-works "Direct link to How it works") ### The cluster (write tier)[​](#the-cluster-write-tier "Direct link to The cluster (write tier)") The cluster owns every refresh, acceleration, and search index for the datasets in scope. It runs as a standalone Spice deployment — typically a Kubernetes [`Deployment`](/docs/next/deployment/kubernetes/helm) or [`StatefulSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset), or a managed [Spice Cloud](/docs/next/deployment/cloud) app — and holds the only credentials to the source systems. Cluster Spicepod responsibilities: * Connect to every source: object stores, OLTP databases, lakehouses, search indices, message queues. * Run all refresh schedules, CDC, and stream ingest. * Accelerate to file-mode engines (DuckDB or SQLite) so the materialization can be exported as a snapshot. * Write [snapshots](/docs/next/features/data-acceleration/snapshots) to a shared object store after each refresh. ``` # cluster spicepod.yaml snapshots: enabled: true location: s3://spiceai-snapshots/prod/ params: s3_auth: iam_role datasets: - from: s3://my-lake/orders/ name: orders params: file_format: parquet acceleration: enabled: true engine: duckdb mode: file refresh_check_interval: 5m snapshots: enabled # write a new snapshot after every refresh snapshots_trigger: refresh_complete snapshots_compaction: enabled params: duckdb_file: /data/orders.db - from: postgres:public.customers name: customers params: pg_host: postgres.internal pg_user: ${ secrets:PG_USER } pg_pass: ${ secrets:PG_PASS } acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes # CDC snapshots: enabled snapshots_trigger: time_interval snapshots_trigger_threshold: 10m params: duckdb_file: /data/customers.db ``` Snapshots are partitioned by date and dataset (`month=YYYY-MM/day=YYYY-MM-DD/dataset=/...`), so retention is a normal object-store lifecycle rule. See [Snapshots](/docs/next/features/data-acceleration/snapshots) for the full configuration reference. ### The read instances (read tier)[​](#the-read-instances-read-tier "Direct link to The read instances (read tier)") Read instances run alongside applications — typically as Kubernetes pod sidecars, but the same configuration works in Cloud Run, on bare metal, or on a developer laptop. They never connect to source systems; their only inbound dependencies are the snapshot bucket and (optionally) the cluster's Arrow Flight endpoint. Read Spicepod responsibilities: * Bootstrap each accelerated dataset from the latest snapshot on startup. No source connection required. * Optionally refresh from newer snapshots on a schedule (`bootstrap_only` mode polls for new snapshots without writing them). * Optionally delegate queries that fall outside the materialized working set to the cluster over Arrow Flight. ``` # read instance spicepod.yaml snapshots: enabled: true location: s3://spiceai-snapshots/prod/ bootstrap_on_failure_behavior: fallback # try older snapshots if the newest fails params: s3_auth: iam_role datasets: - from: s3://my-lake/orders/ # same source URL, but never used at runtime name: orders params: file_format: parquet acceleration: enabled: true engine: duckdb mode: file snapshots: bootstrap_only # download only; never write back params: duckdb_file: /local/orders.db - from: postgres:public.customers name: customers acceleration: enabled: true engine: duckdb mode: file snapshots: bootstrap_only params: duckdb_file: /local/customers.db ``` `snapshots: bootstrap_only` is the key setting — read instances **read** snapshots but never **write** them, so multiple replicas don't race to upload. Combine with a periodic refresh trigger to pick up new snapshots without re-querying the source. ### Live delegation for the long tail[​](#live-delegation-for-the-long-tail "Direct link to Live delegation for the long tail") Snapshots cover the working set. For queries that span beyond it — historical analytics, cross-dataset joins, distributed search — read instances delegate to the cluster using a [`spiceai` connector](/docs/next/components/data-connectors/spiceai) entry pointing at the cluster's Arrow Flight endpoint. ``` # read instance spicepod.yaml (continued) datasets: - from: spiceai:orders_history name: orders_history params: endpoint: grpcs://cluster.spice.svc.cluster.local:50051 api_key: ${ secrets:CLUSTER_API_KEY } ``` The application sees a single SQL surface — accelerated tables and delegated tables compose normally in joins and CTEs. See [Cluster-Sidecar Architecture](/docs/next/deployment/architectures/cluster-sidecar) for the conceptual model. ## Operational model[​](#operational-model "Direct link to Operational model") ### Bootstrap and refresh on read instances[​](#bootstrap-and-refresh-on-read-instances "Direct link to Bootstrap and refresh on read instances") When a read instance starts: 1. For each accelerated dataset, Spice checks for the local file (`duckdb_file` / `sqlite_file`). 2. If absent and snapshots are enabled, Spice lists the snapshot prefix, downloads the newest snapshot for that dataset, and the dataset goes ready immediately. 3. If no snapshot is found, behavior is governed by `bootstrap_on_failure_behavior`: * `warn` (default) — boot empty and refresh from the source. Avoid in read-tier instances that should not have source access. * `fallback` — try older snapshots until one loads. * `retry` — keep retrying the newest snapshot. For zero-source-credentials read instances, set `bootstrap_on_failure_behavior: fallback` or `retry` and ensure the dataset is **never** configured with usable source credentials. Steady-state refresh on read instances is configured per dataset: ``` acceleration: refresh_check_interval: 1m # check the snapshot bucket every minute snapshots: bootstrap_only ``` When a newer snapshot is available, the dataset hot-swaps without restarting the pod. ### Snapshot retention and storage[​](#snapshot-retention-and-storage "Direct link to Snapshot retention and storage") Snapshots are written to Hive-partitioned paths so retention is straightforward: ``` s3://spiceai-snapshots/prod/ month=2026-05/day=2026-05-01/dataset=orders/orders_20260501T120000Z.db month=2026-05/day=2026-05-02/dataset=orders/orders_20260502T120000Z.db ``` Apply an object-store lifecycle rule (S3 lifecycle, GCS Object Lifecycle Management, ADLS Lifecycle) to expire old partitions. Most deployments keep 24–72 hours of refresh-triggered snapshots and a daily archive beyond that. The snapshot bucket is the only shared dependency between the tiers, so keep it in the same region as the read instances and apply VPC endpoints / Private Google Access to keep traffic on the private network. ### Versioning the Spicepod[​](#versioning-the-spicepod "Direct link to Versioning the Spicepod") The cluster and the read instances share dataset *names* but not full Spicepods. Two patterns work well: * **Fork two Spicepods from a common base.** Keep `datasets:` definitions in a shared file and merge the cluster-only and read-only fields at deploy time (Helm value overlays, Kustomize, Jsonnet). * **Single Spicepod, role-based behavior.** Use [Spicepod includes](https://docs.spice.ai/docs/reference/spicepod/dependencies) and environment-specific values to switch `snapshots: enabled` (cluster) vs `snapshots: bootstrap_only` (reads) per role. Whichever approach is chosen, treat schema changes as backward-compatible by default — read instances may be running snapshots from a previous cluster version during a rollout. ## Deploy on Kubernetes[​](#deploy-on-kubernetes "Direct link to Deploy on Kubernetes") The reference topology runs the cluster as a `StatefulSet` (or [`SpicepodSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset) on Spice.ai Enterprise) and the read instances as sidecars in application pods. Both use the same [Spice Helm chart](/docs/next/deployment/kubernetes/helm). ### Cluster release[​](#cluster-release "Direct link to Cluster release") ``` # cluster-values.yaml replicaCount: 3 stateful: enabled: true storageClass: gp3 # or hyperdisk-balanced, managed-csi-premium size: 100Gi serviceAccount: create: true name: spiceai-cluster annotations: eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/SpiceAIClusterRole spicepod: # full Spicepod with sources, refresh schedules, and snapshots: enabled ... ``` ``` helm upgrade --install spiceai-cluster spiceai/spiceai \ -n spiceai-cluster --create-namespace \ -f cluster-values.yaml ``` ### Read instance sidecars[​](#read-instance-sidecars "Direct link to Read instance sidecars") Read instances are deployed as a sidecar container in application pods, configured via a `ConfigMap` that holds the read-tier Spicepod. The application points at `127.0.0.1:8090` (HTTP) or `127.0.0.1:50051` (Arrow Flight) — no service discovery needed. ``` # application Deployment spec: template: spec: serviceAccountName: spiceai-read # IRSA / Workload Identity for snapshot bucket volumes: - name: spicepod configMap: name: spiceai-read-spicepod - name: accel emptyDir: {} # ephemeral; bootstrapped from snapshots containers: - name: app image: my-app:1.2.3 env: - name: SPICEAI_HTTP_URL value: http://127.0.0.1:8090 - name: spiceai image: spiceai/spiceai:1.11.5 args: ['--http', '0.0.0.0:8090', '--flight', '0.0.0.0:50051'] volumeMounts: - name: spicepod mountPath: /spicepod readOnly: true - name: accel mountPath: /local readinessProbe: httpGet: { path: /v1/ready, port: 8090 } livenessProbe: httpGet: { path: /health, port: 8090 } ``` The read sidecar's `ServiceAccount` only needs read access to the snapshot bucket. It should **not** be granted source-system credentials — that's what makes the read tier safe to scale to many replicas. ### Spice.ai Enterprise[​](#spiceai-enterprise "Direct link to Spice.ai Enterprise") For production, the [Spice.ai Enterprise Kubernetes Operator](https://docs.spice.ai/docs/enterprise/kubernetes-operator/kubernetes) manages both tiers as custom resources: * [`SpicepodSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset) — per-replica `StatefulSet`s for the cluster, with automatic PVC resizing, configurable update strategies, and crashloop protection. * [`SpicepodCluster`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodcluster) — distributed scheduler/executor tiers when the cluster itself is large enough to need its own internal split. * Sidecar injection via webhook, so application teams add a single annotation to opt in. ## Capacity sizing[​](#capacity-sizing "Direct link to Capacity sizing") Rough first-pass sizing rules: | Tier | Typical shape | | ---------------- | ---------------------------------------------------------------------------------------------------- | | Cluster (writer) | 3+ replicas. Memory sized for the largest accelerated dataset. Network bandwidth for source ingest. | | Read instance | 1 replica per application pod. 0.5–2 vCPU, 512Mi–4Gi memory, 10–50Gi local SSD. | | Snapshot bucket | Standard tier, same region. Lifecycle rule sized to refresh frequency × number of datasets × 24–72h. | Read-tier memory is dominated by the working set of the file-mode acceleration engine. DuckDB compaction (`snapshots_compaction: enabled`) typically reduces snapshot size by 30–60%. ## Security model[​](#security-model "Direct link to Security model") The split simplifies the credential surface area: * **Cluster** — holds source credentials, snapshot **write** credentials, and cluster-internal mTLS. Runs in a private subnet; no public ingress. * **Read instances** — hold snapshot **read** credentials and a per-instance Arrow Flight token to the cluster (for live delegation). No source credentials. * **Application** — talks to its sidecar over loopback. No outbound credentials at all. Compromising a read instance grants the attacker the read tier's snapshot bucket and the delegated query surface — never the source systems. ## Observability[​](#observability "Direct link to Observability") Both tiers expose the same metrics and tracing endpoints. Practical splits: * **Cluster dashboards** — refresh duration, snapshot upload size and latency, source connector errors, ingest queue depth. * **Read dashboards** — bootstrap duration, snapshot age (write-time vs current-time), query latency p50/p99, delegation rate (queries served locally vs forwarded to the cluster). A high delegation rate is a signal to expand the materialized working set. A growing snapshot age is a signal that the cluster is falling behind on refresh. ## Related[​](#related "Direct link to Related") * [Cluster-Sidecar Architecture](/docs/next/deployment/architectures/cluster-sidecar) — the conceptual model and live-delegation pattern. * [Snapshots](/docs/next/features/data-acceleration/snapshots) — full reference for snapshot configuration, triggers, and modes. * [Sidecar Architecture](/docs/next/deployment/architectures/sidecar) — single-instance precursor to this pattern. * [Cluster Architecture](/docs/next/deployment/architectures/cluster) — internal scheduler/executor split for the cluster tier (Spice.ai Enterprise). * [Kubernetes Deployment Guide](/docs/next/deployment/kubernetes) — Helm, Argo CD, and Flux options for the cluster. * [CI/CD](/docs/next/deployment/ci-cd) — automating cluster and read-instance rollouts. * [Spice.ai Enterprise Kubernetes Operator](https://docs.spice.ai/docs/enterprise/kubernetes-operator/kubernetes) — recommended for production self-hosted deployments. --- # Spice.ai FAQ ## 1. What is Spice?[​](#1-what-is-spice "Direct link to 1. What is Spice?") Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven apps and agents. Spice provides four industry standard APIs in a lightweight, portable runtime (single \~140 MB binary): 1. **SQL Query APIs**: Supports HTTP, Arrow Flight, Arrow Flight SQL, ODBC, JDBC, and ADBC. 2. **OpenAI-Compatible APIs**: Provides HTTP APIs for OpenAI SDK compatibility, local model serving (CUDA/Metal accelerated), and hosted model gateway. 3. **Iceberg Catalog REST APIs**: Offers a unified API for Iceberg Catalog. 4. **MCP HTTP+SSE APIs**: Enables integration with external tools via Model Context Protocol (MCP) using HTTP and Server-Sent Events (SSE). Spice embeds [DataFusion](https://datafusion.apache.org/), the fastest single-node Parquet SQL query engine, and [DuckDB](https://duckdb.org), to serve secure, virtualized data views to data-intensive apps, AI, and agents. For a developer-focused walkthrough of when and how to use Spice, see [A Developer's Guide to Understanding Spice.ai](https://spice.ai/blog/a-developers-guide-to-understanding-spice-ai). ## 2. Why should I use Spice?[​](#2-why-should-i-use-spice "Direct link to 2. Why should I use Spice?") Spice is primarily used for: * **Data Federation**: SQL query across any database, data warehouse, or data lake. [Learn More](/docs/next/features/query-federation). * **Data Materialization and Acceleration**: Materialize, accelerate, and cache database queries. [Read the MaterializedView interview - Building a CDN for Databases](https://materializedview.io/p/building-a-cdn-for-databases-spice-ai) * **AI apps and agents**: An AI-database powering retrieval-augmented generation (RAG) and intelligent agents. [Learn More](https://github.com/spiceai/cookbook/tree/trunk/rag#readme). For example, deploy Spice as a sidecar alongside applications served by centralized platforms like Databricks or Snowflake, materializing hot datasets locally to reduce latency and offload query volume from the source. ## 3. How is Spice different?[​](#3-how-is-spice-different "Direct link to 3. How is Spice different?") * **Application-Centric Design:** Spice is designed for 1:1 or 1 :N mappings between applications and Spice instances, making it flexible for tenant-specific or customer-specific configurations. Unlike traditional databases designed for many applications sharing one data system, Spice often runs one instance per application or tenant. * **Dual-Engine Acceleration:** Spice supports both OLAP (DuckDB/Arrow) and OLTP (SQLite/PostgreSQL) databases at the dataset level, providing flexibility for various query workloads. * **Separation of Materialization and Storage/Compute:** Spice enables data to remain close to its source while materializing working sets for fast access, reducing data movement and query latency. * **Deployment Flexibility:** Deployable across infrastructure tiers, including edge, on-prem, and cloud environments. Spice can run as a standalone instance, sidecar, microservice, or cluster. ## 4. What is Data-grounded AI?[​](#4-what-is-data-grounded-ai "Direct link to 4. What is Data-grounded AI?") Data-grounded AI anchors models in accurate, current, domain-specific data rather than relying solely on pre-trained knowledge. Spice unifies enterprise data across databases, data lakes, and APIs, dynamically incorporating real-world context at inference time. This helps minimize hallucinations, reduce operational risk, and build trust in AI by delivering reliable, relevant outputs. ## 5. How is Spice different from Trino/Presto and Dremio?[​](#5-how-is-spice-different-from-trinopresto-and-dremio "Direct link to 5. How is Spice different from Trino/Presto and Dremio?") Spice is purpose-built for data and AI applications and agents, designed with low-latency access, materialization, and proximity to applications. Trino/Presto and Dremio primarily target big data analytics and rely on centralized clusters. Spice's decentralized approach reduces latency, simplifies deployment, and improves efficiency. ## 6. How does Spice compare to Spark?[​](#6-how-does-spice-compare-to-spark "Direct link to 6. How does Spice compare to Spark?") Spark excels at distributed batch processing and large-scale transformations. Spice focuses on real-time, low-latency data access and AI inference. Spice materializes data locally and supports tiered storage, optimizing performance for applications requiring fast access and high concurrency. Starting in v2, Spice also supports multi-node [Distributed Query](/docs/next/features/distributed-query) execution based on [Apache Ballista](https://datafusion.apache.org/ballista/), splitting a single query across scheduler and executor nodes for partitioned data lake sources. This closes much of the gap with Spark for large analytical scans while keeping Spice's low-latency, materialization-first model for application and agent workloads. See [Apache Ballista at Spice AI](https://spice.ai/blog/apache-ballista-at-spice-ai) for the architecture and [Operationalizing Amazon S3 for AI](https://spice.ai/blog/operationalizing-amazon-s3-for-ai) for a data-lake example. ## 7. How does Spice compare to DuckDB?[​](#7-how-does-spice-compare-to-duckdb "Direct link to 7. How does Spice compare to DuckDB?") DuckDB is an embedded analytics database optimized for OLAP queries. Spice integrates DuckDB for data acceleration, combining DuckDB's analytical capabilities with Spice's broader federation, multi-engine support, and flexible deployment. Spice can be considered an enterprise/production productization of DuckDB for data-intensive applications. ## 8. Can Spice handle federated queries?[​](#8-can-spice-handle-federated-queries "Direct link to 8. Can Spice handle federated queries?") Yes. Spice natively supports federated queries across disparate data sources with advanced query push-down capabilities. Spice executes portions of queries directly on source databases, reducing data transfer and improving performance. [Learn More](/docs/next/features/query-federation). ## 9. Can Spice federate joins across many tables and data sources?[​](#9-can-spice-federate-joins-across-many-tables-and-data-sources "Direct link to 9. Can Spice federate joins across many tables and data sources?") Yes, but performance degrades as the number of sources and tables grows because the slowest source bounds end-to-end latency. For workloads that join across many tables or many heterogeneous sources, the recommended pattern is to **materialize the join into an accelerated dataset or [accelerated view](/docs/next/features/views)**. This converts a multi-source federated join into a single local scan, making latency predictable and decoupling query performance from any one source's availability. Plan acceleration of hot joins proactively rather than waiting to discover them under load. See [Query Federation](/docs/next/features/query-federation), [Data Acceleration](/docs/next/features/data-acceleration), and [Performance Tuning](/docs/next/reference/performance-tuning) for tradeoffs and configuration. The [Multi-Tenancy for AI Agents without the Pipelines](https://spice.ai/blog/multi-tenancy-for-ai-agents-without-pipelines) blog post walks through a concrete federation + acceleration example. ## 10. Can Spice query nested fields (JSON / struct columns)?[​](#10-can-spice-query-nested-fields-json--struct-columns "Direct link to 10. Can Spice query nested fields (JSON / struct columns)?") Yes. Spice supports querying nested data through both native nested types (struct, list, map) and JSON-encoded text columns. For nested types, use standard field access (e.g. `column.field.subfield`). For JSON-encoded text columns, Spice provides a family of [JSON SQL functions](/docs/next/reference/sql/json) including `json_get`, `json_get_str`, `json_get_int`, `json_get_float`, `json_get_bool`, `json_get_json`, and `json_get_array`, plus the `->` operator. Spice pushes JSON predicates down to the underlying engine (e.g. DuckDB) where supported. When pushdown is not viable for a particular expression, Spice transparently evaluates the predicate in [DataFusion](/docs/next/features/query-federation) so queries always execute correctly. If a query against a nested field is unexpectedly slow or fails, prefer the typed `json_get_*` variants over generic extraction so the planner can reason about types. ## 11. What query engines does Spice support?[​](#11-what-query-engines-does-spice-support "Direct link to 11. What query engines does Spice support?") Spice uses [Apache DataFusion](https://datafusion.apache.org/) as its primary query execution engine, providing vectorized, multi-threaded query processing with automatic memory management and spilling. DataFusion powers the Arrow and Spice Cayenne (Vortex) accelerators. Spice also supports DuckDB, SQLite, and PostgreSQL as acceleration engines. Developers can select engines based on workload requirements, balancing performance, concurrency, and latency. See [Performance Tuning](/docs/next/reference/performance-tuning) for sizing guidance and the [Memory](/docs/next/reference/memory) reference for how Spice manages query memory and spilling under load. For deeper context, see [How we use Apache DataFusion at Spice AI](https://spice.ai/blog/how-we-use-apache-datafusion-at-spice-ai) and [Vortex at Spice AI](https://spice.ai/blog/vortex-at-spice-ai-the-columnar-format-for-data-intensive-workloads) on the Cayenne columnar format. ## 12. Is Spice a cache?[​](#12-is-spice-a-cache "Direct link to 12. Is Spice a cache?") Not solely. Spice functions as an active cache or working dataset prefetcher. A *working dataset* is a subset of data actively used by an application or model, such as recent records or frequently accessed tables. Unlike traditional caches that fetch data reactively, Spice proactively prefetches and materializes data based on filters, intervals, triggers, or Change Data Capture (CDC), ensuring data readiness for queries. Spice also supports [results caching](/docs/next/features/caching). ## 13. Can Spice be used as a CDN for databases?[​](#13-can-spice-be-used-as-a-cdn-for-databases "Direct link to 13. Can Spice be used as a CDN for databases?") Yes. Spice can be deployed as a CDN for databases by loading and materializing datasets close to applications, reducing latency and improving query efficiency. [Read more](https://materializedview.io/p/building-a-cdn-for-databases-spice-ai). ## 14. Can Spice be used for real-time analytics?[​](#14-can-spice-be-used-for-real-time-analytics "Direct link to 14. Can Spice be used for real-time analytics?") Yes. Spice accelerates data locally using Apache Arrow, Spice Cayenne (Vortex), DuckDB, SQLite, or PostgreSQL, enabling real-time analytics and sub-second query performance for data-intensive applications and dashboards. ## 15. Does Spice support Change Data Capture (CDC)?[​](#15-does-spice-support-change-data-capture-cdc "Direct link to 15. Does Spice support Change Data Capture (CDC)?") Yes. Spice supports streaming ingestion from several sources: * **Native PostgreSQL logical replication** (recommended for Postgres sources). Spice connects directly to the source using Postgres' `wal_level=logical` + pgoutput and streams `INSERT`/`UPDATE`/`DELETE` events into the accelerator. [Learn more](/docs/next/features/cdc/postgres-replication). * **[DynamoDB Streams](/docs/next/components/data-connectors/dynamodb)** for Amazon DynamoDB sources — Spice consumes the table's change stream and applies `INSERT`/`UPDATE`/`DELETE` events to the accelerator with `refresh_mode: changes`. * **[Apache Kafka](/docs/next/components/data-connectors/kafka)** for event-streaming topics — Spice consumes records directly with `refresh_mode: append` for real-time, append-only acceleration. * **[Debezium](/docs/next/components/data-connectors/debezium)** (over Kafka), for sources where Debezium is already deployed, or for databases without a native Spice CDC path (MySQL, SQL Server, etc.). [Learn more](/docs/next/features/cdc). For a real-world architecture using DynamoDB Streams to sync data to thousands of nodes, see [Real-Time Control Plane Acceleration with DynamoDB Streams](https://spice.ai/blog/real-time-acceleration-with-dynamodb-streams). ## 16. How do I keep an accelerated dataset incrementally up-to-date?[​](#16-how-do-i-keep-an-accelerated-dataset-incrementally-up-to-date "Direct link to 16. How do I keep an accelerated dataset incrementally up-to-date?") For sources with a monotonically-increasing version column (e.g. `updated_at`), Spice incrementally ingests new and modified records using [`time_column`](/docs/next/reference/spicepod/datasets#time_column) + [`refresh_mode: append`](/docs/next/features/data-acceleration/data-refresh#append), with [`refresh_append_overlap`](/docs/next/reference/spicepod/datasets#accelerationrefresh_append_overlap) to tolerate clock skew and [`retention_period`](/docs/next/reference/spicepod/datasets#accelerationretention_period) to evict old or soft-deleted records. Pair with `primary_key` + `on_conflict: upsert` to deduplicate re-read rows within the overlap window. See [Data Refresh](/docs/next/features/data-acceleration/data-refresh) for configuration details and examples. ## 17. How can I control the load Spice puts on a source system during refresh?[​](#17-how-can-i-control-the-load-spice-puts-on-a-source-system-during-refresh "Direct link to 17. How can I control the load Spice puts on a source system during refresh?") Spice provides several controls for tuning ingestion pressure: * **Incremental refresh** with [`refresh_mode: append`](/docs/next/features/data-acceleration/data-refresh#append) and a [`time_column`](/docs/next/reference/spicepod/datasets#time_column) to fetch only new or changed rows instead of reloading the full dataset. * **Per-connector concurrency, connection pooling, backoff, and retry**. For example, the [Databricks connector](/docs/next/components/data-connectors/databricks/deployment) exposes `max_concurrent_requests` (shared per SQL Warehouse), and the [GitHub connector](/docs/next/components/data-connectors/github) exposes `github_max_concurrent_connections` (shared per token / app installation). Most connectors also expose pool and timeout parameters. * **Runtime-level parallel dataset loading** via [`runtime.num_of_parallel_loading_at_start_up`](/docs/next/reference/spicepod/runtime) to bound how many datasets initialize concurrently at startup. * **Refresh scheduling controls** — `refresh_check_interval`, `refresh_jitter`, and `refresh_retry_*` parameters in [Data Refresh](/docs/next/features/data-acceleration/data-refresh) — to spread load over time and avoid thundering-herd patterns. Lowering concurrency and increasing intervals reduces source load but increases the time required to reach steady-state freshness; tune iteratively against your source's RPS and concurrency limits. See [Performance Tuning](/docs/next/reference/performance-tuning) and the [Memory](/docs/next/reference/memory) reference for guidance on balancing throughput and resource usage. ## 18. Can Spice load data on demand instead of proactively refreshing every dataset?[​](#18-can-spice-load-data-on-demand-instead-of-proactively-refreshing-every-dataset "Direct link to 18. Can Spice load data on demand instead of proactively refreshing every dataset?") Yes. Several patterns help avoid eagerly refreshing datasets that may never be queried: * **Deferred readiness** with [`ready_state: on_registration`](/docs/next/features/data-acceleration/data-refresh#ready-state) — Spice marks the dataset ready before any data has been loaded, so the dataset is registered immediately and the first refresh is initiated lazily. * **On-demand refresh** via the [`POST /v1/datasets/:name/acceleration/refresh`](/docs/next/api/HTTP/post-dataset-refresh) endpoint to trigger a refresh from an external scheduler, webhook, or first-query signal. * **Stale-while-revalidate caching** for HTTP-backed and connector-backed datasets via [`refresh_mode: caching`](/docs/next/features/data-acceleration/refresh-modes/caching), which fetches on first request and revalidates in the background. * **Federated fallback** with [`on_zero_results: use_source`](/docs/next/features/data-acceleration/data-refresh#behavior-on-zero-results) so queries that miss the acceleration transparently fall through to the source. These patterns combine well in multi-tenant deployments where only a fraction of tenants are active at any given time. See [Multi-Tenancy for AI Agents without the Pipelines](https://spice.ai/blog/multi-tenancy-for-ai-agents-without-pipelines) for a worked example. ## 19. Does Spice support schema evolution?[​](#19-does-spice-support-schema-evolution "Direct link to 19. Does Spice support schema evolution?") Spice infers the schema for datasets and views at startup and does not apply runtime schema changes by default. If the source schema changes while the runtime is running (for example, columns are added, removed, or their types change), data refreshes will fail with a schema mismatch error rather than silently applying the new schema. This behavior is intentional — it protects against unintentional or breaking schema changes propagating into accelerated tables. To pick up a new source schema, restart the Spice runtime. On startup, Spice re-infers the schema from the source, and the accelerated table is re-initialized with the updated schema. Runtime schema evolution controls are planned for a future release but will remain off by default to guard against unexpected schema drift. ## 20. What AI capabilities does Spice provide?[​](#20-what-ai-capabilities-does-spice-provide "Direct link to 20. What AI capabilities does Spice provide?") Spice provides unified APIs for data and AI workflows, including model inference, embeddings, and an AI gateway supporting OpenAI, Anthropic, xAI, and Nvidia NIMs. Spice includes advanced LLM tools such as vector and hybrid search, text-to-SQL, SQL retrieval, data sampling, and context formatting. ## 21. What AI model providers does Spice support?[​](#21-what-ai-model-providers-does-spice-support "Direct link to 21. What AI model providers does Spice support?") Spice supports local model serving (e.g., Llama3) and gateways to hosted AI platforms including OpenAI, Anthropic, xAI, and Nvidia NIMs. [Learn More](/docs/next/features/large-language-models). ## 22. Where should I run embedding models — co-located with each replica or in a central cluster?[​](#22-where-should-i-run-embedding-models--co-located-with-each-replica-or-in-a-central-cluster "Direct link to 22. Where should I run embedding models — co-located with each replica or in a central cluster?") For most deployments, the recommended pattern is to **embed in the central ingestion tier** with GPU-equipped nodes: the embedding model runs once per row during ingestion, vectors are stored in the acceleration, and read-tier replicas access them via [snapshots](/docs/next/features/data-acceleration/snapshots) or direct query. This minimizes data movement and keeps the read tier inexpensive to scale. When embeddings must be computed at query time (e.g. user-supplied queries for [vector search](/docs/next/features/search/vector-search)), dedicated embedding instances co-located with the read tier — or a shared, horizontally-scaled embedding service — provide better latency than per-replica GPUs. See the [Local](/docs/next/components/embeddings/local/deployment) and [Hugging Face TEI](/docs/next/components/embeddings/huggingface/deployment) deployment guides for sizing. ## 23. How does Spice handle vector / hybrid search at scale?[​](#23-how-does-spice-handle-vector--hybrid-search-at-scale "Direct link to 23. How does Spice handle vector / hybrid search at scale?") Spice supports [vector search](/docs/next/features/search/vector-search), [full-text search](/docs/next/features/search/full-text), [multi-vector search](/docs/next/features/search/multi-vector), and [reranking](/docs/next/features/search/rerank) over accelerated datasets. Vectors are stored in the acceleration alongside the source data, so search executes locally without round-tripping to a separate vector database. For best performance, store accelerated vector datasets in **file-mode DuckDB or Spice Cayenne (Vortex)** — file-mode is often faster than in-memory due to columnar compression and dictionary encoding — and configure [acceleration indexes](/docs/next/features/data-acceleration/indexes) on filter columns to narrow brute-force scans. See [Performance Tuning](/docs/next/reference/performance-tuning) and the [Memory](/docs/next/reference/memory) reference for sizing accelerated vector workloads, and [Vortex at Spice AI](https://spice.ai/blog/vortex-at-spice-ai-the-columnar-format-for-data-intensive-workloads) for the storage format underpinning Cayenne. Expanded HNSW and other approximate-nearest-neighbor index types are tracked for future releases. ## 24. What search APIs does Spice provide?[​](#24-what-search-apis-does-spice-provide "Direct link to 24. What search APIs does Spice provide?") Spice exposes search through both an HTTP API and SQL table-valued functions: * **[`POST /v1/search`](/docs/next/api/HTTP/post-search)** — a unified HTTP endpoint that performs vector, full-text, multi-vector, and hybrid search across configured datasets. Requests specify the dataset(s), search text, optional filters, and limit; the response returns ranked rows with relevance scores and any requested passthrough columns. Embeddings are computed automatically when an embedding model is configured on the dataset. * **SQL UDTFs** for in-query search — `vector_search(table, 'query')` and related functions can be composed with standard SQL `JOIN`, `WHERE`, and `ORDER BY` for hybrid filtering and ranking. See [Vector Search](/docs/next/features/search/vector-search), [Full-Text Search](/docs/next/features/search/full-text), [Multi-Vector Search](/docs/next/features/search/multi-vector), and [Reranking](/docs/next/features/search/rerank). * **MCP tools** — the search APIs are also exposed to AI agents through Spice's [MCP server](/docs/next/features/large-language-models), so agents can issue search queries without writing SQL. For an end-to-end example, see the [Search](/docs/next/features/search) overview and the [Amazon S3 Vectors with Spice](https://spice.ai/blog/amazon-s3-vectors-with-spice) blog post. ## 25. What deployment options does Spice support?[​](#25-what-deployment-options-does-spice-support "Direct link to 25. What deployment options does Spice support?") Spice supports multiple deployment configurations: * Standalone binary * Sidecar or microservice * Cluster deployments * Edge, on-prem, and cloud environments Spice Cloud Platform (SCP) provides managed, SOC 2 Type II compliant deployments. [Learn More](https://spiceai.org/docs/deployment/architectures). For moving a deployment to production, see the [Production Readiness](https://docs.spice.ai/docs/enterprise/production/production) guide and the [Enterprise Distribution Comparison](https://docs.spice.ai/docs/enterprise/getting-started/distributions) for choosing between OSS and Enterprise distributions. ## 26. Which deployment patterns are open source vs. Enterprise?[​](#26-which-deployment-patterns-are-open-source-vs-enterprise "Direct link to 26. Which deployment patterns are open source vs. Enterprise?") The sidecar model, cluster model, [snapshot bootstrapping](/docs/next/features/data-acceleration/snapshots), and pointing instances at each other for tiered fallback are all open source. The [Distributed Query](/docs/next/features/distributed-query) feature — splitting a single query across multiple scheduler/executor nodes (Apache Ballista) — is available in v2.x. The [Spice.ai Enterprise](https://spice.ai) Kubernetes Operator additionally provides the [`SpicepodSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset) and [`SpicepodCluster`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodcluster) CRDs for declarative replica management, automatic PVC resizing, mTLS, and scheduler/executor topology. Distributed query is only required for large single-query workloads (TB/PB-scale scans, heavy embedding/indexing). See the [Enterprise Distribution Comparison](https://docs.spice.ai/docs/enterprise/getting-started/distributions) for a feature-by-feature breakdown. ## 27. How does fallback work in tiered (sidecar + central cluster) deployments?[​](#27-how-does-fallback-work-in-tiered-sidecar--central-cluster-deployments "Direct link to 27. How does fallback work in tiered (sidecar + central cluster) deployments?") Reads fall through a layered chain. A request first checks the local sidecar's results cache and acceleration. On a miss, the request is delegated to the central cluster, which checks its own results cache and acceleration. If both miss, the cluster federates back to the original data source. The full chain is **results cache → acceleration → federated source**, applied at each tier. For details, see the [Sidecar](/docs/next/deployment/architectures/sidecar), [Cluster](/docs/next/deployment/architectures/cluster), and [Tiered](/docs/next/deployment/architectures/tiered) architecture guides, along with [Caching](/docs/next/features/caching), [Data Acceleration](/docs/next/features/data-acceleration), and [Query Federation](/docs/next/features/query-federation). See the [Cluster-Sidecar Architecture](https://spice.ai/blog/cluster-sidecar-architecture) blog post for a deeper walkthrough of the pattern. ## 28. How do I run multiple Spice replicas for HA without each replica pulling from the source?[​](#28-how-do-i-run-multiple-spice-replicas-for-ha-without-each-replica-pulling-from-the-source "Direct link to 28. How do I run multiple Spice replicas for HA without each replica pulling from the source?") Separate the **ingestion (write) tier** from the **query (read) tier**: * A small ingestion tier (typically 1–3 instances) connects to the source, builds accelerations, and publishes [snapshots](/docs/next/features/data-acceleration/snapshots). * The read tier scales horizontally and bootstraps from snapshots, so replicas never connect directly to the source. If the ingestion tier is temporarily unavailable, read replicas continue serving (potentially stale) data from snapshots, providing a recovery window measured in hours rather than minutes. The Enterprise [`SpicepodSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset) resource manages replica sets declaratively with rolling/parallel update strategies, automatic PVC resizing, and crashloop protection. See [Cluster](/docs/next/deployment/architectures/cluster) and the [Cluster-Sidecar Architecture](https://spice.ai/blog/cluster-sidecar-architecture) blog post for the open-source pattern, and the [Production Readiness](https://docs.spice.ai/docs/enterprise/production/production) guide for HA checklists. ## 29. How should I shard Spice for very large multi-tenant workloads?[​](#29-how-should-i-shard-spice-for-very-large-multi-tenant-workloads "Direct link to 29. How should I shard Spice for very large multi-tenant workloads?") Shards are independent groups of Spice instances, each with its own Spicepod configuration. Strategies that scale well: * **Bucket multiple tenants per shard** using a deterministic bucket UDF (e.g. `bucket(org_id, N)`) instead of one partition per tenant. ACLs are still enforced at query time, and operational complexity is bounded by `N` rather than tenant count. * **Use [acceleration partitioning](/docs/next/features/data-acceleration/partitioning)** to control physical layout within a shard. * **Manage shards declaratively with the Enterprise Operator** using [`SpicepodSet`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodset) (and [`SpicepodCluster`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodcluster) for distributed query) with rollout strategies, mTLS, and OIDC. See the [Sharded](/docs/next/deployment/architectures/sharded) deployment architecture for guidance, and [Multi-Tenancy for AI Agents without the Pipelines](https://spice.ai/blog/multi-tenancy-for-ai-agents-without-pipelines) for a multi-tenant blueprint. ## 30. Which secret stores does Spice support?[​](#30-which-secret-stores-does-spice-support "Direct link to 30. Which secret stores does Spice support?") Spice supports [Kubernetes Secrets](/docs/next/components/secret-stores/kubernetes), [AWS Secrets Manager](/docs/next/components/secret-stores/aws-secrets-manager), [Azure Key Vault](/docs/next/components/secret-stores/azure-keyvault), [environment variables](/docs/next/components/secret-stores/env), and the local OS [keyring](/docs/next/components/secret-stores/keyring). See the [secret stores overview](/docs/next/components/secret-stores) for configuration and selectors. Additional secret backends — including HashiCorp Vault — are tracked as planned additions. Open or upvote an issue at [github.com/spiceai/spiceai](https://github.com/spiceai/spiceai/issues) to influence prioritization. ## 31. How does Spice handle data privacy and compliance?[​](#31-how-does-spice-handle-data-privacy-and-compliance "Direct link to 31. How does Spice handle data privacy and compliance?") Spice provides secure, auditable data access through sandboxed runtimes, secure endpoint checks, and detailed telemetry and tracing. The Spice Cloud Platform (SCP) is SOC 2 Type II compliant, meeting enterprise security and compliance requirements. See the [Production Readiness](https://docs.spice.ai/docs/enterprise/production/production) guide for hardening recommendations. ## 32. Can Spice integrate with existing BI tools?[​](#32-can-spice-integrate-with-existing-bi-tools "Direct link to 32. Can Spice integrate with existing BI tools?") Yes. Spice integrates with BI tools through standard SQL interfaces (ODBC, JDBC, Arrow Flight SQL), enabling accelerated, real-time analytics for dashboards and reporting. An official [Tableau Connector](/docs/next/clients/tableau) is available and a [BI Acceleration](https://www.youtube.com/watch?v=blEtLgRKu0c) demo using Apache Superset. ## 33. Where can developers find examples and recipes?[​](#33-where-can-developers-find-examples-and-recipes "Direct link to 33. Where can developers find examples and recipes?") The [Spice.ai Cookbook](https://github.com/spiceai/cookbook) provides over 65 quickstarts and examples demonstrating Spice capabilities, including federated queries, RAG, text-to-SQL, and more. ## 34. How can developers get started quickly?[​](#34-how-can-developers-get-started-quickly "Direct link to 34. How can developers get started quickly?") Visit the [Spice.ai Getting Started Guide](/docs/next/getting-started) to install Spice, connect data sources, and begin querying. Spice installs the GPU-accelerated runtime by default (if supported). ## 35. How can developers contribute to Spice?[​](#35-how-can-developers-contribute-to-spice "Direct link to 35. How can developers contribute to Spice?") Developers can contribute by submitting code, documentation, or raising issues on [GitHub](https://github.com/spiceai/spiceai). See [CONTRIBUTING.md](https://github.com/spiceai/spiceai/blob/trunk/CONTRIBUTING) for guidelines. --- # Spice.ai Features Spice provides a set of features for building data-driven applications and AI agents. This page gives an overview of each feature area. ### Data Query and Federation[​](#data-query-and-federation "Direct link to Data Query and Federation") [Query Federation](/docs/next/features/query-federation) connects multiple data sources—databases, data warehouses, and data lakes—through a single SQL interface. Write one query that joins data across PostgreSQL, Snowflake, S3, and other sources. Spice pushes query operations to source databases when possible to reduce data transfer. ### Data Acceleration and Caching[​](#data-acceleration-and-caching "Direct link to Data Acceleration and Caching") [Data Acceleration](/docs/next/features/data-acceleration) materializes remote datasets locally in memory or on disk using engines like Arrow, DuckDB, SQLite, or PostgreSQL. Accelerated datasets stay current through scheduled refreshes, append mode, or [Change Data Capture (CDC)](/docs/next/features/cdc). [Caching](/docs/next/features/caching) stores query and search results in memory with configurable TTLs and eviction policies to avoid redundant computation. ### Views[​](#views "Direct link to Views") [Views](/docs/next/features/views) create virtual tables from SQL queries over other datasets, similar to database views — useful for encapsulating query logic and (when accelerated) materializing precomputed joins or aggregates. ### AI and Language Models[​](#ai-and-language-models "Direct link to AI and Language Models") [Large Language Models](/docs/next/features/large-language-models) provides an OpenAI-compatible API gateway for hosted models (OpenAI, Anthropic, xAI) and locally served models (Llama, Phi) with CUDA and Metal acceleration. Models can call tools to query datasets, run SQL, and retrieve schemas. [Embeddings](/docs/next/features/embeddings) generates vector representations of text for semantic search and RAG workflows. [Workers](/docs/next/features/workers) coordinate interactions between models and tools, supporting load-balancing strategies such as round-robin and fallback across multiple LLM providers. ### Search[​](#search "Direct link to Search") [Search](/docs/next/features/search) supports three methods: vector search (semantic similarity using embeddings), full-text search (keyword matching with BM25 scoring), and hybrid search (combining both with Reciprocal Rank Fusion). All search methods are accessible through SQL UDTFs like `vector_search()` and `text_search()`. ### Functions[​](#functions "Direct link to Functions") [Functions](/docs/next/features/functions) extend SQL with custom scalar functions declared in a Spicepod. Inline SQL bodies run in-process and can use any DataFusion built-in; remote `http://` / `https://` endpoints batch row inputs over JSON for delegating logic to ML models, internal services, or custom code. Every function is automatically callable from SQL and (by default) surfaced as an LLM tool. ### Tool Registry[​](#tool-registry "Direct link to Tool Registry") [Tool Registry](/docs/next/features/tool-registry) keeps per-turn token cost bounded as the runtime's tool catalog grows. It replaces individual tool definitions with searchable `tool_search` and `tool_invoke` meta-tools backed by a hybrid full-text, keyword, schema, and vector search. Applies uniformly to built-in tools, MCP tools, and Functions declared with `as_tool: true` — typically a \~10× reduction in tool-definition tokens for tool-heavy Spicepods. ### Monitoring and Observability[​](#monitoring-and-observability "Direct link to Monitoring and Observability") [Observability](/docs/next/features/observability) exposes Prometheus-compatible metrics, OpenTelemetry metric export, and distributed tracing with Zipkin. Integrations are available for Datadog, Grafana, and other monitoring platforms. ## [🗃Query Federation](/docs/next/features/query-federation) [2 items](/docs/next/features/query-federation) ## [🗃Data Acceleration](/docs/next/features/data-acceleration) [7 items](/docs/next/features/data-acceleration) ## [📄️Caching](/docs/next/features/caching) [Learn how to use Spice in-memory caching](/docs/next/features/caching) ## [📄️Distributed Query](/docs/next/features/distributed-query) [Learn how to run Spice in distributed mode for larger scale queries, including the async queries API.](/docs/next/features/distributed-query) ## [🗃Change Data Capture](/docs/next/features/cdc) [6 items](/docs/next/features/cdc) ## [📄️Data Ingestion](/docs/next/features/data-ingestion) [Learn how to ingest data in Spice.](/docs/next/features/data-ingestion) ## [🗃Large Language Models](/docs/next/features/large-language-models) [7 items](/docs/next/features/large-language-models) ## [📄️Machine Learning Models](/docs/next/features/machine-learning-models) [Deprecated in vNext: Support for loading and serving traditional machine learning (ONNX) models for inference was removed in vNext, along with the /v1/predict and /v1/models//predict prediction endpoints. See the v2.1 docs for documentation of this feature.](/docs/next/features/machine-learning-models) ## [📄️Embedding Datasets](/docs/next/features/embeddings) [Learn how to define, or augment existing datasets with embedding column(s).](/docs/next/features/embeddings) ## [🗃Search](/docs/next/features/search) [4 items](/docs/next/features/search) ## [📄️Functions](/docs/next/features/functions) [Define custom scalar and table SQL functions inline (SQL tier) or by calling remote HTTP services (Remote tier), automatically exposed as SQL functions and LLM tools.](/docs/next/features/functions) ## [📄️Semantic Model](/docs/next/features/semantic-model) [Attach descriptions and metadata to datasets, views, and columns in Spice so LLMs, SQL functions, and humans share the same understanding of your data.](/docs/next/features/semantic-model) ## [📄️Tool Registry](/docs/next/features/tool-registry) [Reduce per-turn token cost and improve LLM tool selection accuracy by replacing individual tool definitions with searchable tool\_search and tool\_invoke meta-tools backed by hybrid full-text, keyword, schema, and vector search.](/docs/next/features/tool-registry) ## [🗃Observability](/docs/next/features/observability) [1 item](/docs/next/features/observability) ## [📄️Web Search](/docs/next/features/web-search) [Learn how Spice can perform web search](/docs/next/features/web-search) ## [📄️Views](/docs/next/features/views) [Documentation for defining Views in Spice](/docs/next/features/views) ## [📄️Workers](/docs/next/features/workers) [Configure workers in the Spice runtime to coordinate interactions between LLMs and tools, with load-balancing, round-robin, and fallback strategies.](/docs/next/features/workers) --- # Caching Spice supports in-memory caching for SQL query results and search results, which are both enabled by default when querying or searching via the HTTP (`/v1/sql`, `/v1/search`) and Arrow Flight APIs. Results caching improves performance for repeated requests and non-accelerated results, such as refresh data returned [on zero results](/docs/next/features/data-acceleration/data-refresh#behavior-on-zero-results). The cache uses a [least-recently-used (LRU)](https://en.wikipedia.org/wiki/Cache_replacement_policies#LRU) replacement policy. You can configure the cache to set an item expiration duration, which defaults to 1 second. ``` version: v1 kind: Spicepod name: app runtime: caching: sql_results: enabled: true max_size: 1GiB # Default 128 MiB item_ttl: 1m # Default 1s stale_while_revalidate_ttl: 30s # Default 0s (disabled) search_results: enabled: true max_size: 1GiB # Default 128 MiB item_ttl: 1m # Default 1s ``` ## `caching` Parameters[​](#caching-parameters "Direct link to caching-parameters") | Parameter name | Optional | Description | | ---------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `sql_results` | Yes | Enabled by default. Configures the Runtime cache for results from SQL queries. See the [SQL Results Parameters](#cachingsql_results-parameters) for cache parameter details. | | `search_results` | Yes | Enabled by default. Configures the Runtime cache for results from searches. See the [Common Caching Parameters](#common-caching-parameters) for cache parameter details. | | `embeddings` | Yes | Enabled by default. Configures the Runtime cache for embeddings requests. See the [Common Caching Parameters](#common-caching-parameters) for cache parameter details. | ## Common Caching Parameters[​](#common-caching-parameters "Direct link to Common Caching Parameters") Every cache type (`sql_results`, `search_results`, `embeddings`) supports the following parameters: | Parameter name | Optional | Default | Description | | ------------------- | -------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `enabled` | Yes | `true` | Defaults to `true`. | | `max_size` | Yes | `128MiB` | Maximum cache size. Defaults to `128MiB`. | | `eviction_policy` | Yes | `lru` | Cache replacement policy when the cache reaches `max_size`. Defaults to `lru`. Supports `lru` (Least Recently Used) and `tiny_lfu` (Tiny Least Frequently Used, higher hit rate for skewed access patterns). | | `item_ttl` | Yes | `1s` | Cache entry expiration duration (Time to Live). Defaults to 1 second. | | `hashing_algorithm` | Yes | `xxh3` | Selects which hashing algorithm is used to hash the cache keys when storing the results. Defaults to `xxh3`. Supports `xxh3`, `ahash`, `siphash`, `blake3`, `xxh32`, `xxh64`, or `xxh128`. | ## `caching.sql_results` Parameters[​](#cachingsql_results-parameters "Direct link to cachingsql_results-parameters") In addition to the common caching parameters, `sql_results` also supports additional parameters: | Parameter name | Optional | Default | Description | | ---------------------------- | -------- | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cache_key_type` | Yes | `plan` | Determines how cache keys are generated. Defaults to `plan`. `plan` uses the query's logical plan, while `sql` uses the raw SQL query string. | | `encoding` | Yes | `none` | Compression algorithm for cached results. Defaults to `none`. Supports `none` or `zstd`. | | `stale_while_revalidate_ttl` | Yes | `0s` | Duration to serve stale cache entries while revalidating in the background. When set to a non-zero value, expired cache entries continue to be served while a background refresh occurs. Defaults to `0s` (disabled). | ### Choosing a `cache_key_type`[​](#choosing-a-cache_key_type "Direct link to choosing-a-cache_key_type") * **`plan` (Default):** Uses the query's logical plan as the cache key. This approach matches semantically equivalent queries, even if their SQL syntax differs. However, it requires query parsing, which introduces some overhead. * **`sql`:** Uses the raw SQL string as the cache key. This method provides faster lookups but requires exact string matches. Queries with dynamic functions, such as `NOW()`, may produce unexpected results because the cache key changes with each execution. Use `sql` only when query results are predictable and consistent. Use `sql` for the lowest latency with identical queries that do not include dynamic functions. Use `plan` for greater flexibility and semantic matching of queries. ## Choosing a `hashing_algorithm`[​](#choosing-a-hashing_algorithm "Direct link to choosing-a-hashing_algorithm") The hashing algorithm determines how cache keys are hashed before being stored, impacting both lookup speed and protection against potential DOS attacks. * **`xxh3` (Default):** Uses the [XXH3](https://cyan4973.github.io/xxHash/) algorithm for hashing the cache keys. XXH3 is a fast, non-cryptographic hash algorithm that provides high performance and good distribution. It is suitable for scenarios where speed is critical and cryptographic security is not required. * **`siphash`:** Uses the SipHash1-3 algorithm for hashing the cache keys, the [default hashing algorithm of Rust](https://github.com/rust-lang/rust/commit/db1b1919baba8be48d997d9f70a6a5df7e31612a). This hashing algorithm is a secure algorithm that implements verified protections against ["hash flooding"](https://v8.dev/blog/hash-flooding) denial of service (DoS) attacks. Reasonably performant, and provides a high level of security. * **`ahash`:** Uses the [AHash](https://github.com/tkaitchuck/ahash) algorithm for hashing the cache keys. The AHash algorithm is a [high quality](https://github.com/tkaitchuck/aHash/blob/master/compare/readme#Quality) hashing algorithm, and has claimed resistance against hashing DoS attacks. AHash has higher performance than SipHash1-3, especially when used with `cache_key_type: plan`. * **`blake3`:** Uses the [BLAKE3](https://github.com/BLAKE3-team/BLAKE3) cryptographic hash function. BLAKE3 is a fast, parallelizable hash function that provides cryptographic security while maintaining high performance. It is suitable for scenarios requiring both speed and cryptographic guarantees. * **`xxh32`, `xxh64`, `xxh128`:** Variants of the XXH hashing algorithm with different output sizes. These algorithms offer a balance between speed and collision resistance, with larger hash sizes providing better collision resistance at the cost of performance. Use `xxh3` (the default) for its superior speed in most scenarios. Use `ahash`, `xxh64` or `xxh128` for reduced collision probability when caching a large number of queries. Use `blake3` when cryptographic security is required. Use `siphash` when protection against hash flooding attacks is a priority. ### Choosing an `encoding`[​](#choosing-an-encoding "Direct link to choosing-an-encoding") The encoding algorithm determines how cached results are compressed in memory, trading CPU for memory efficiency. Currently supported for SQL results only. * **`none` (Default):** Stores query results uncompressed. Uses more memory but has zero compression overhead. Best for small result sets or when memory is not a constraint. * **`zstd`:** Uses the [Zstandard compression algorithm](https://facebook.github.io/zstd/) to compress cached query results. Provides high compression ratios (often 50-90% reduction) with fast decompression speeds. Recommended when caching large result sets to maximize cache capacity. Use `zstd` when maximizing cache efficiency is important, especially for large queries that would otherwise quickly fill the cache. Use `none` for the lowest latency when memory is not constrained. ## Per-Principal Cache Isolation[​](#per-principal-cache-isolation "Direct link to Per-Principal Cache Isolation") When [authentication](/docs/next/api/auth) is enabled, all cache layers (SQL results, search results, and caching-mode acceleration storage) are automatically scoped per principal. Each authenticated caller has an isolated cache namespace — one caller's cached output is never served to a different caller. | Scenario | Cache namespace | | -------------------------------------------------- | -------------------------------------------------------------- | | Auth disabled or anonymous request | `public` — all callers share the cache | | Authenticated principal (e.g., API key) | Per-principal — isolated by a hash of the principal's identity | | Background / system tasks (e.g., SWR revalidation) | Inherits the originating principal's namespace | There are no user-facing knobs to disable isolation. Scope follows authentication presence by construction. Breaking change for caching accelerator The caching accelerator storage schema gains an internal `__spice_cache_namespace` column. Existing accelerator storage from earlier versions must be deleted (e.g., remove the `duckdb_file`, drop the backing table) before upgrading. The runtime errors with a clear message at startup if it encounters a non-extended schema. The SQL results cache and search cache are in-memory and require no migration. ## Cached Responses[​](#cached-responses "Direct link to Cached Responses") Responses from HTTP APIs include headers that indicate the cache status and scope: | Cache | Status Header | Scope Header | | ---------------- | ----------------------------- | ---------------------------- | | `sql_results` | `Results-Cache-Status` | `Results-Cache-Scope` | | `search_results` | `Search-Results-Cache-Status` | `Search-Results-Cache-Scope` | The status header indicates the cache status: | Header value | Description | | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | `HIT` | The query result was served from the cache. | | `MISS` | The cache was checked, but the result was not found. | | `BYPASS` | The cache was bypassed for this query (e.g., when `cache-control: no-cache` is specified). | | `STALE` | A stale cache entry was served while the cache is being revalidated in the background (when `stale_while_revalidate_ttl` is configured). | | *header not present* | The cache did not apply to this query (e.g., when caching is disabled or querying a system table). | The scope header indicates the cache namespace: | Scope value | Description | | ----------- | ------------------------------------------------------------------------------------------------------------------------------------- | | `shared` | The cache entry is in the public namespace (auth disabled or anonymous). | | `user` | The cache entry is scoped to the authenticated principal. When the scope is `user`, the response also includes `Vary: Authorization`. | | `system` | The cache entry is in the system/background namespace. | ### Examples[​](#examples "Direct link to Examples") #### Cached Response[​](#cached-response "Direct link to Cached Response") ``` $ curl -XPOST -i http://localhost:8090/v1/sql -d 'select * from taxi_trips limit 1;' HTTP/1.1 200 OK content-type: text/plain; charset=utf-8 results-cache-status: HIT results-cache-scope: shared vary: origin, access-control-request-method, access-control-request-headers content-length: 416 date: Thu, 13 Feb 2025 03:05:39 GMT ``` #### Uncached Response[​](#uncached-response "Direct link to Uncached Response") ``` $ curl -XPOST -i http://localhost:8090/v1/sql -d 'select * from taxi_trips limit 1;' HTTP/1.1 200 OK content-type: text/plain; charset=utf-8 results-cache-status: MISS results-cache-scope: shared vary: origin, access-control-request-method, access-control-request-headers content-length: 416 date: Thu, 13 Feb 2025 03:13:19 GMT ``` #### Bypassed Cache with `cache-control: no-cache`[​](#bypassed-cache-with-cache-control-no-cache "Direct link to bypassed-cache-with-cache-control-no-cache") ``` $ curl -H "cache-control: no-cache" -XPOST -i http://localhost:8090/v1/sql -d 'select * from taxi_trips limit 1;' HTTP/1.1 200 OK content-type: text/plain; charset=utf-8 results-cache-status: BYPASS vary: origin, access-control-request-method, access-control-request-headers content-length: 416 date: Thu, 13 Feb 2025 03:14:00 GMT ``` #### Stale Cache Response (Stale-While-Revalidate)[​](#stale-cache-response-stale-while-revalidate "Direct link to Stale Cache Response (Stale-While-Revalidate)") ``` $ curl -XPOST -i http://localhost:8090/v1/sql -d 'select * from taxi_trips limit 1;' HTTP/1.1 200 OK content-type: text/plain; charset=utf-8 results-cache-status: STALE vary: origin, access-control-request-method, access-control-request-headers content-length: 416 date: Thu, 13 Feb 2025 03:15:30 GMT ``` ## Cache Control[​](#cache-control "Direct link to Cache Control") You can control caching behavior for specific requests using HTTP headers. The `Cache-Control` header helps skip the cache for a request while caching the results for subsequent requests. ### Stale-While-Revalidate[​](#stale-while-revalidate "Direct link to Stale-While-Revalidate") The `stale_while_revalidate_ttl` parameter configures a grace period during which stale cache entries continue to be served while a background refresh occurs. This technique reduces latency for end users by serving cached data immediately, even after `item_ttl` expires, while the system fetches fresh data asynchronously. When `stale_while_revalidate_ttl` is set to a non-zero value: 1. Cache entries are served normally until `item_ttl` expires. 2. After `item_ttl` expires but before `item_ttl + stale_while_revalidate_ttl` expires, the stale entry is served immediately with a `STALE` cache status. 3. Simultaneously, a background task refreshes the cache entry. 4. Once the background refresh completes, subsequent requests receive the fresh data with a `HIT` cache status. 5. After `item_ttl + stale_while_revalidate_ttl` expires, the entry is evicted and the next request results in a `MISS`. #### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` runtime: caching: sql_results: enabled: true item_ttl: 10s stale_while_revalidate_ttl: 10s ``` With this configuration: * Fresh cache entries are served for 10 seconds after creation. * Between 10-20 seconds after creation, stale entries are served while being refreshed in the background. * After 20 seconds, the entry is evicted if not refreshed. This approach is particularly useful for queries that take significant time to execute, providing a better user experience by reducing perceived latency while keeping data reasonably fresh. Conflict with Caching Accelerator SWR When using a dataset with `refresh_mode: caching`, you cannot configure both the results cache's `stale_while_revalidate_ttl` and the caching accelerator's `caching_stale_while_revalidate_ttl` for the same dataset. These parameters control similar behavior at different layers. Choose one approach: * **Results cache SWR**: Configure `runtime.caching.sql_results.stale_while_revalidate_ttl` for SQL query results caching * **Caching accelerator SWR**: Configure `acceleration.params.caching_stale_while_revalidate_ttl` for [HTTP-based dataset caching](/docs/next/features/data-acceleration/refresh-modes/caching) ### HTTP/Flight API[​](#httpflight-api "Direct link to HTTP/Flight API") The following endpoints support the standard HTTP [`Cache-Control` header](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control): * SQL query (HTTP and Arrow Flight) * Search (HTTP) The following `Cache-Control` directives are supported: | Directive | Description | | ---------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [`no-cache`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control#no-cache) | Skips the cache for the current request but caches the results for future requests. | | [`min-fresh`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control#min-fresh) | Specifies the minimum time (in seconds) that a cached response must remain fresh. For example, `min-fresh=60` requires the cached entry to be fresh for at least 60 more seconds. | | [`max-stale`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control#max-stale) | Indicates the client will accept a stale response. An optional value in seconds specifies the maximum staleness allowed. For example, `max-stale=30` accepts responses stale for up to 30 seconds. | | [`only-if-cached`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control#only-if-cached) | Returns only cached responses. If no cached response is available, returns an error instead of fetching fresh data. | | [`stale-if-error`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control#stale-if-error) | Serves stale cached responses if an error occurs while fetching fresh data. An optional value in seconds specifies how stale the response can be. For example, `stale-if-error=600` serves responses stale for up to 10 minutes if fetching fails. | #### HTTP Example[​](#http-example "Direct link to HTTP Example") ``` # Default behavior (uses cache) curl -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' # Skip cache for this query, but cache the results for future queries curl -H "cache-control: no-cache" -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' # Only use cached response if it will be fresh for at least 30 more seconds curl -H "cache-control: min-fresh=30" -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' # Accept cached responses that are stale for up to 60 seconds curl -H "cache-control: max-stale=60" -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' # Only return cached responses, fail if cache miss curl -H "cache-control: only-if-cached" -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' # Serve stale cache (up to 300 seconds old) if fetching fresh data fails curl -H "cache-control: stale-if-error=300" -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' ``` #### Arrow FlightSQL Example[​](#arrow-flightsql-example "Direct link to Arrow FlightSQL Example") The following example skips the cache for a specific query using FlightSQL in Rust: ``` let sql_command = arrow_flight::sql::CommandStatementQuery { query: "SELECT 1".to_string(), transaction_id: None, }; let sql_command_bytes = sql_command.as_any().encode_to_vec(); let mut request = FlightDescriptor::new_cmd(sql_command_bytes).into_request(); request .metadata_mut() .insert("cache-control", "no-cache"); // Send the request ``` The cache can be controlled using JDBC properties. For example, ``` Properties props = new Properties(); props.setProperty("cache-control", "no-cache"); Connection conn = DriverManager.getConnection("jdbc:arrow-flight-sql://localhost:50051", props); ``` ### `spice` CLI[​](#spice-cli "Direct link to spice-cli") The `spice sql` and `spice search` commands accept a `--cache-control` flag that supports all cache-control directives: ``` # Default behavior (use cache if available) spice sql # Same as above spice sql --cache-control cache # Skip cache for this query, but cache the results for future queries spice sql --cache-control no-cache # Only use cached response if fresh for at least 30 more seconds spice sql --cache-control min-fresh=30 # Accept cached responses stale for up to 60 seconds spice sql --cache-control max-stale=60 # Only return cached responses, fail if cache miss spice sql --cache-control only-if-cached # Serve stale cache (up to 300 seconds) if fetching fails spice sql --cache-control stale-if-error=300 # Default behavior (use cache if available) spice search # Same as above spice search --cache-control cache # Skip cache for this search, but cache the results for future searches spice search --cache-control no-cache # Accept stale search results up to 60 seconds old spice search --cache-control max-stale=60 ``` ## Custom Cache Keys[​](#custom-cache-keys "Direct link to Custom Cache Keys") Set the `Spice-Cache-Key` header to supply a custom cache key. When set, a supplied cache key takes precedence over `caching.sql_results.cache_key_type`. Info A valid cache key consists of up to 128 alphanumeric characters (and the characters `-` and `_`). ### HTTP Example[​](#http-example-1 "Direct link to HTTP Example") Consider the case of two semantically equivalent queries: ``` Time: 0.0251325 seconds. 2 rows. sql> select * from users where org_id = 1; +----+--------+-------+----------------+ | id | org_id | name | email | +----+--------+-------+----------------+ | 1 | 1 | Jane | jane@spice.ai | | 2 | 1 | Sarah | sarah@spice.ai | +----+--------+-------+----------------+ Time: 0.008993042 seconds. 2 rows. sql> select * from users where split_part(email, '@', 2) = 'spice.ai'; +----+--------+-------+----------------+ | id | org_id | name | email | +----+--------+-------+----------------+ | 1 | 1 | Jane | jane@spice.ai | | 2 | 1 | Sarah | sarah@spice.ai | +----+--------+-------+----------------+ ``` To share a cache key for these queries, set `Spice-Cache-Key`. The first request is a cache miss: ``` $ curl -i -XPOST http://localhost:8090/v1/sql -H"spice-cache-key: users_spiceai" -d "select * from users where org_id = 1;" HTTP/1.1 200 OK content-type: application/json x-cache: Miss from spiceai results-cache-status: MISS vary: Spice-Cache-Key vary: origin, access-control-request-method, access-control-request-headers content-length: 119 date: Thu, 24 Jul 2025 14:15:53 GMT [{"id":1,"org_id":1,"name":"Jane","email":"jane@spice.ai"},{"id":2,"org_id":1,"name":"Sarah","email":"sarah@spice.ai"}] ``` The subsequent request with the different (but semantically equivalent) query is a cache hit: ``` $ curl -i -XPOST http://localhost:8090/v1/sql -H"spice-cache-key: users_spiceai" -d "select * from users where split_part(email, '@', 2) = 'spice.ai';" HTTP/1.1 200 OK content-type: application/json x-cache: Hit from spiceai results-cache-status: HIT vary: Spice-Cache-Key vary: origin, access-control-request-method, access-control-request-headers content-length: 119 date: Thu, 24 Jul 2025 14:18:00 GMT ``` Note When supplying a custom cache key, **ensure the semantic equivalence of queries**. For example, this is expected behavior: ``` $ curl -i -XPOST http://localhost:8090/v1/sql -H"spice-cache-key: users_spiceai" -d "select 1" HTTP/1.1 200 OK content-type: application/json x-cache: Hit from spiceai results-cache-status: HIT vary: Spice-Cache-Key vary: origin, access-control-request-method, access-control-request-headers content-length: 119 date: Thu, 24 Jul 2025 14:21:32 GMT [{"id":1,"org_id":1,"name":"Jane","email":"jane@spice.ai"},{"id":2,"org_id":1,"name":"Sarah","email":"sarah@spice.ai"}] ``` ## Metrics[​](#metrics "Direct link to Metrics") Cache metrics can be monitored using the [Prometheus-compatible Metrics Endpoint](/docs/next/features/observability). The following metrics are available for each cache type: | Metric | Type | Description | | ------------------------ | ------- | -------------------------------------------------------------------------- | | `*_cache_max_size_bytes` | Gauge | Maximum configured cache size in bytes. | | `*_cache_requests` | Counter | Total number of cache lookup requests. | | `*_cache_hits` | Counter | Total number of cache hits. | | `*_cache_items_count` | Gauge | Current number of items in the cache. | | `*_cache_size_bytes` | Gauge | Current cache size in bytes. | | `*_cache_evictions` | Counter | Total number of entries removed from the cache, split by a `reason` label. | | `*_cache_hit_ratio` | Gauge | Current cache hit ratio (hits / total requests). | `*_cache_evictions` carries a `reason` label with one of three values: | `reason` | Meaning | | ------------- | ----------------------------------------------------------------------------- | | `size` | The cache exceeded `max_size` and reclaimed an entry. | | `expired` | The entry outlived `item_ttl`. | | `invalidated` | A dataset refresh or a DML write dropped the entries that referenced a table. | On an accelerated dataset with a periodic refresh, `invalidated` is usually the dominant — often the only — reason, which is why it is a separate label value rather than folded into an unlabelled total: an alert on cache pressure should watch `size` and `expired`. Every cache counter is published at zero when the runtime starts, so a counter that has not yet fired still appears in a scrape as a zero series rather than being absent. The SQL results cache additionally emits `results_cache_stale_rejections`, a counter of lookups that found an entry but refused to serve it because a table the result read had since been invalidated. These are also counted in `results_cache_misses`, so the two together separate "nothing was cached" from "something was cached but had gone stale". The `*` prefix corresponds to the cache type: * `results_*` - SQL query results cache metrics * `search_results_*` - Search results cache metrics * `embeddings_*` - Embeddings cache metrics Example metrics output: ``` # HELP results_cache_evictions Number of cache evictions, by reason. # TYPE results_cache_evictions counter results_cache_evictions{reason="size"} 2 results_cache_evictions{reason="expired"} 5 results_cache_evictions{reason="invalidated"} 41 # HELP results_cache_hit_ratio Cache hit ratio (hits / total requests). # TYPE results_cache_hit_ratio gauge results_cache_hit_ratio 0.625 # HELP results_cache_hits Cache hit count. # TYPE results_cache_hits counter results_cache_hits 14 # HELP results_cache_items_count Number of items currently in the cache. # TYPE results_cache_items_count gauge results_cache_items_count 1 # HELP results_cache_max_size_bytes Maximum allowed size of the cache in bytes. # TYPE results_cache_max_size_bytes gauge results_cache_max_size_bytes 134217728 # HELP results_cache_misses Cache miss count. # TYPE results_cache_misses counter results_cache_misses 4 # HELP results_cache_requests Number of requests to get a key from the cache. # TYPE results_cache_requests counter results_cache_requests 18 # HELP results_cache_size_bytes Size of the cache in bytes. # TYPE results_cache_size_bytes gauge results_cache_size_bytes 7776 # HELP results_cache_stale_rejections Number of lookups that found an entry but refused to serve it because a table it read had since been invalidated. # TYPE results_cache_stale_rejections counter results_cache_stale_rejections 0 ``` --- # Change Data Capture (CDC) Change Data Capture (CDC) captures insert, update, and delete events from a database's transaction log and delivers them to consumers with low latency. This technique enables Spice to keep [locally accelerated](/docs/next/features/data-acceleration) datasets synchronized with the source data in near real-time. CDC is efficient because it transfers only changed rows instead of re-fetching the entire dataset. ## Benefits[​](#benefits "Direct link to Benefits") Using locally accelerated datasets configured with CDC enables Spice to provide high-performance accelerated queries and efficient real-time updates. ## Example Use Case[​](#example-use-case "Direct link to Example Use Case") Consider a fraud detection application that needs to determine whether a pending transaction is likely fraudulent. The application queries a Spice-accelerated, real-time updated table of recent transactions to check if a pending transaction resembles known fraudulent ones. With CDC, the table is kept up-to-date, so the application can quickly identify potential fraud. ## Considerations[​](#considerations "Direct link to Considerations") When configuring datasets to be accelerated with CDC, ensure that the [data connector](/docs/next/components/data-connectors) supports CDC and can return a stream of row-level changes. See the [Supported Data Connectors](#supported-data-connectors) section for more information. The startup time for CDC-accelerated datasets may be longer than for non-CDC-accelerated datasets due to the initial synchronization. tip It is recommended to use CDC-accelerated datasets with persistent data accelerator configurations (i.e., `file` mode for [`DuckDB`](/docs/next/components/data-accelerators/duckdb)/[`SQLite`](/docs/next/components/data-accelerators/sqlite) or [`PostgreSQL`](/docs/next/components/data-accelerators/postgres)). This ensures that when Spice restarts, it can resume from the last known state of the dataset instead of re-fetching the entire dataset. ## Tuning ingestion[​](#tuning-ingestion "Direct link to Tuning ingestion") Spice applies CDC events through a single apply loop that coalesces a contiguous run of buffered change events ("envelopes") into one accelerator write. The coalescing behavior is controlled by the following `runtime.params`, set once under the top-level `runtime.params` and applied to every CDC-accelerated dataset in the instance. Each parameter also accepts a `SPICE_`-prefixed environment variable; the `runtime.params` value takes precedence, falling back to the environment variable, then the default. Cayenne-accelerated datasets can additionally override any of these values per-dataset — see [Per-dataset overrides](#per-dataset-overrides) below. | Parameter | Description | Default | | ----------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------- | | `cdc_prefetch_buffer` | Number of source change events buffered ahead of the apply loop. Range `1`–`16384`. | `128` | | `cdc_max_coalesced_envelopes` | Maximum number of change events combined into a single accelerator write. Range `1`–`16384`. | `256` | | `cdc_max_coalesced_bytes` | Maximum in-memory Arrow size (bytes) of a single coalesced write. Range `1`–`1073741824` (1 GiB). | `134217728` (128 MiB) | | `cdc_max_coalesce_age_ms` | Apply-loop linger window in milliseconds. When `> 0`, the loop keeps accumulating change events into one write until the envelope cap, the byte budget, or this window elapses — whichever comes first. The window is measured from the start of the previous apply. `0` disables lingering, so each buffered event is applied as soon as it arrives. | `0` (no linger) | | `cdc_commit_timeout_ms` | Maximum time to wait for the previous source-side commit before surfacing ingestion as stalled. Range `1`–`3600000` (1 hour). | `30000` (30s) | ``` runtime: params: cdc_max_coalesce_age_ms: 250 # linger up to 250ms to coalesce slowly-arriving events into fewer writes ``` Out-of-range or unparseable values are rejected with a warning and fall back to the default. ### Per-dataset overrides[​](#per-dataset-overrides "Direct link to Per-dataset overrides") For [Cayenne](/docs/next/components/data-accelerators/cayenne)-accelerated datasets, any of the five parameters above can also be set per-dataset under the dataset's `acceleration.params` to override the instance-wide value for that dataset only. A per-dataset value layers on top of the resolved global configuration (`runtime.params` → environment variable → default): a dataset overrides only the parameters it sets and inherits the global value for the rest. Out-of-range or unparseable per-dataset values are rejected with a warning and keep the global value. ``` runtime: params: cdc_max_coalesce_age_ms: 250 # global default for every CDC-accelerated dataset datasets: - from: postgres:public.orders name: orders acceleration: engine: cayenne refresh_mode: changes params: cdc_max_coalesce_age_ms: 1000 # this dataset lingers longer than the global 250ms cdc_prefetch_buffer: 1024 # ...and buffers more aggressively ``` ## Supported Data Connectors[​](#supported-data-connectors "Direct link to Supported Data Connectors") Enabling CDC by setting `refresh_mode: changes` in the acceleration settings requires support from the data connector to provide a stream of row-level changes. Spice currently supports streaming ingestion via: * **[PostgreSQL Logical Replication](/docs/next/features/cdc/postgres-replication)** — **recommended** for PostgreSQL sources. Spice connects directly to the source using Postgres' native logical replication protocol (`wal_level=logical` + pgoutput) and streams `INSERT`/`UPDATE`/`DELETE` events into the accelerator. No Kafka, no Debezium, no external services. * **[MySQL Binlog Replication](/docs/next/features/cdc/mysql-replication)** — **recommended** for MySQL sources. Spice subscribes to the source's binary log (`binlog_format=ROW`) as a replica and streams `INSERT`/`UPDATE`/`DELETE` events into the accelerator. No Kafka, no Debezium, no external services. * **[DynamoDB Streams](/docs/next/features/cdc/dynamodb-streams)** — for Amazon DynamoDB sources. Spice consumes the table's DynamoDB Streams directly and applies `INSERT`/`UPDATE`/`DELETE` events to the accelerator. * **[MongoDB Change Streams](/docs/next/features/cdc/mongodb-streams)** — for MongoDB replica sets and sharded clusters. Spice opens a native Change Stream on the source collection and applies inserts, updates, replaces, and deletes to the accelerator. * **[Apache Kafka](/docs/next/components/data-connectors/kafka)** — for event-streaming topics. Spice consumes records directly with `refresh_mode: append` for real-time, append-only acceleration (no separate CDC connector required). * **[Debezium](/docs/next/features/cdc/debezium)** (over Kafka) — for sources where Debezium + Kafka is already deployed, or for databases without a native Spice CDC path (SQL Server, etc.). * **[Debezium Push Ingest](/docs/next/features/cdc/debezium-ingest)** (no Kafka) — for any database with a Debezium source plugin (Oracle, SQL Server, Db2, …). The Debezium plugin POSTs change events (JSON or Avro) directly to `spiced` via `from: cdc:…`, with no Kafka bus. ## Example[​](#example "Direct link to Example") See an example of configuring a dataset to use CDC with Debezium by following the recipe at [Streaming changes in real-time with Debezium CDC](https://github.com/spiceai/cookbook/tree/trunk/cdc-debezium#readme). ``` version: v1 kind: Spicepod name: cdc-debezium datasets: - from: debezium:cdc.public.customer_addresses name: cdc params: debezium_transport: kafka debezium_message_format: json kafka_bootstrap_servers: localhost:19092 kafka_security_protocol: PLAINTEXT acceleration: enabled: true engine: sqlite mode: file refresh_mode: changes ``` --- # Debezium (CDC over Kafka) Consume [Debezium](https://debezium.io/) change events from a Kafka topic and apply them to a Spice-accelerated dataset. Use Debezium when: * You already operate Debezium + Kafka for change data capture; or * The source database does not have a native Spice CDC path (e.g. MySQL, SQL Server, Oracle). For sources with a native CDC path, prefer the dedicated connector — [PostgreSQL Logical Replication](/docs/next/features/cdc/postgres-replication), [DynamoDB Streams](/docs/next/features/cdc/dynamodb-streams), or [MongoDB Change Streams](/docs/next/features/cdc/mongodb-streams) — to avoid the extra Kafka + Debezium hop. ## How it works[​](#how-it-works "Direct link to How it works") ``` ┌────────────────┐ Debezium connector ┌───────────┐ Spice consumes ┌───────────────────┐ ChangeBatch ┌───────────────┐ │ Source DB │ ────────────────────▶│ Kafka │ ────────────────────▶│ Spice runtime │──────────────────▶│ Accelerator │ │ (MySQL, │ WAL → JSON events │ topic │ one consumer group │ (debezium │ (INSERT/ │ DuckDB / │ │ SQL Server, │ │ │ per Spice replica │ connector) │ UPDATE / │ SQLite / │ │ Oracle, …) │ │ │ │ │ DELETE) │ Postgres │ └────────────────┘ └───────────┘ └───────────────────┘ └───────────────┘ ``` On startup, Spice subscribes to the configured Debezium-managed Kafka topic using either a uniquely generated consumer group or one specified via `kafka_consumer_group_id`. With a persistent acceleration engine (`mode: file`), data is fetched starting from the **last committed offset**, so restarts resume without reprocessing historical events. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running **Debezium connector** publishing change events to a **Kafka topic** for the source table. * A reachable Kafka cluster (one or more `bootstrap.servers`). * A Spice acceleration engine that supports CDC: `duckdb`, `sqlite`, or `postgres`. ## Minimal configuration[​](#minimal-configuration "Direct link to Minimal configuration") ``` datasets: - from: debezium:my_kafka_topic_with_debezium_changes name: customer_addresses params: debezium_transport: kafka # Optional. Only `kafka` is currently supported. debezium_message_format: json # Optional. Only `json` is currently supported. kafka_bootstrap_servers: localhost:9092 kafka_security_protocol: PLAINTEXT acceleration: enabled: true # Required. engine: duckdb # duckdb / sqlite / postgres mode: file # Persist Kafka offsets so restarts resume. refresh_mode: changes # Required. ``` The `from` field takes the form `debezium:`. The topic must contain Debezium-formatted change events for a single source table. ## SASL/SSL authentication[​](#saslssl-authentication "Direct link to SASL/SSL authentication") For Kafka clusters with SASL/SSL enabled: ``` datasets: - from: debezium:my_kafka_topic_with_debezium_changes name: orders params: kafka_bootstrap_servers: broker1:9092,broker2:9092,broker3:9092 kafka_security_protocol: sasl_ssl # Default kafka_sasl_mechanism: SCRAM-SHA-512 # PLAIN / SCRAM-SHA-256 / SCRAM-SHA-512 kafka_sasl_username: kafka_user kafka_sasl_password: ${secrets:kafka_sasl_password} kafka_ssl_ca_location: ./certs/kafka_ca_cert.pem acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes ``` The full set of `kafka_*` parameters is documented in the [Debezium connector reference](/docs/next/components/data-connectors/debezium#params). ## Consumer-group management[​](#consumer-group-management "Direct link to Consumer-group management") The connector manages Kafka consumer groups so offsets persist across restarts: * **Default** — Spice auto-generates a unique consumer group ID, stores it in the acceleration metadata, and reuses it on subsequent restarts. * **Custom** — Pass `kafka_consumer_group_id` to use your own group ID. The same ID must be used on every restart; if Spice detects a mismatch against the stored ID, it returns an error to prevent data inconsistency. To recover from a deliberate consumer-group change, reset the acceleration data so Spice starts fresh. See the full description in the [Debezium connector reference](/docs/next/components/data-connectors/debezium#consumer-group-management). ## Schema evolution[​](#schema-evolution "Direct link to Schema evolution") Debezium emits change events whose schema may evolve as the upstream table is altered. Set `schema_evolution: true` to have Spice peek at the latest Kafka message on reload and detect schema changes: ``` params: schema_evolution: true # Default: false ``` ## Batching[​](#batching "Direct link to Batching") Two parameters control how many events Spice groups into a single CDC batch before applying it to the accelerator: | Parameter | Default | Description | | -------------------- | ------- | ---------------------------------------------------------------- | | `batch_max_size` | `10000` | Max number of change events to batch together before processing. | | `batch_max_duration` | `1s` | Max time to wait for a batch to fill before processing. | Larger batches improve throughput at the cost of higher per-batch latency. ## Metrics[​](#metrics "Direct link to Metrics") The connector exposes the following [component metrics](/docs/next/features/observability/component_metrics): | Metric Name | Type | Description | | ------------------------ | ------- | ------------------------------------------------------------------------------------ | | `bytes_consumed_total` | Counter | Total number of bytes consumed from the Kafka topic | | `records_consumed_total` | Counter | Total number of records (messages) consumed from Kafka topics | | `records_lag` | Gauge | Total consumer lag across all topic partitions (number of messages not yet consumed) | These metrics are opt-in; see the [Debezium connector reference](/docs/next/components/data-connectors/debezium#metrics) for an example `metrics:` block. ## Limitations[​](#limitations "Direct link to Limitations") * Only `kafka` is supported as the Debezium transport. * Only `json` is supported as the message format. * Acceleration is required — Debezium cannot be used as a federated, non-accelerated dataset. ## See also[​](#see-also "Direct link to See also") * [Debezium Data Connector](/docs/next/components/data-connectors/debezium) — complete parameter reference. * [Streaming changes in real-time with Debezium CDC](https://github.com/spiceai/cookbook/tree/trunk/cdc-debezium#readme) — cookbook recipe. * [Streaming changes with Debezium and SASL/SCRAM authentication](https://github.com/spiceai/cookbook/tree/trunk/cdc-debezium/sasl-scram#readme) — authenticated-Kafka recipe. * [`refresh_mode: changes`](/docs/next/features/data-acceleration/refresh-modes/changes) — refresh-mode reference. --- # Debezium Push Ingest (CDC without Kafka) Stream change events from **any** [Debezium](https://debezium.io/) source plugin directly into a Spice-accelerated dataset over HTTP, with **no Kafka bus**. The Debezium source (Debezium Server or the Embedded Engine) POSTs change events to `spiced`, which decodes them and applies them to the accelerator. Spice supports a **dual-path** CDC model: | Path | When to use | | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | **Native CDC** | Postgres WAL, MySQL binlog, MongoDB change streams, DynamoDB Streams — Spice captures directly from the database. Prefer this when a native path exists. See [Change Data Capture](/docs/next/features/cdc). | | **Debezium push ingest** | Any database with a Debezium source plugin (Oracle, SQL Server, Db2, …) where you want change capture without running Kafka. | This page covers the second path: the `cdc:` connector plus `POST /v1/datasets/{name}/cdc`. If you already operate Debezium **over Kafka**, continue using the [`debezium:` connector](/docs/next/features/cdc/debezium) — the Kafka consumer path is unchanged. ## How it works[​](#how-it-works "Direct link to How it works") ``` OLTP source (Oracle, SQL Server, Db2, Postgres, …) │ ▼ Debezium source plugin (Debezium Server / Embedded Engine) │ HTTP POST (JSON or Avro) ▼ spiced POST /v1/datasets/{name}/cdc │ decode → change batch ▼ accelerator (Cayenne / DuckDB / SQLite / Arrow / …) ``` The Debezium plugin runs as a separate process and pushes change events into Spice. The HTTP request blocks until the batch has been applied to the accelerator (ack), so the capturer can safely commit its source offsets only after a successful response. ## Spicepod configuration[​](#spicepod-configuration "Direct link to Spicepod configuration") Configure a dataset with `from: cdc:` and an accelerator running in `refresh_mode: changes`: ``` datasets: - from: cdc:orders name: orders columns: - name: id type: int64 - name: customer_id type: int64 - name: amount type: float64 acceleration: enabled: true engine: cayenne # or duckdb, sqlite, arrow, … mode: file refresh_mode: changes primary_key: id on_conflict: id: upsert params: # Avro only — Confluent-compatible Schema Registry base URL: # cdc_schema_registry_url: https://schema-registry:8081 # Or, for raw (non-Confluent) Avro bodies, embed the Avro schema: # cdc_avro_schema: | # { "type": "record", "name": "Envelope", ... } ``` ### Requirements[​](#requirements "Direct link to Requirements") * `acceleration.enabled: true` * `acceleration.refresh_mode: changes` * Explicitly declared `columns` with types — there is no upstream bus or catalog to infer the schema from. * `primary_key` and an `on_conflict` upsert mapping, except for the append-only [Arrow](/docs/next/components/data-accelerators/arrow) accelerator. ### Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `cdc_schema_registry_url` | Confluent-compatible Schema Registry base URL for Avro CDC ingest (Confluent wire format). Avro only. | | `cdc_avro_schema` | Avro schema JSON used when the request body is raw Avro (not Confluent wire format). Can also be sent per-request via the `X-Avro-Schema` header. Avro only. | ## HTTP API[​](#http-api "Direct link to HTTP API") ``` POST /v1/datasets/{name}/cdc ``` The endpoint accepts JSON or Avro, selected by the `Content-Type` header: | Content-Type | Body | | --------------------------------------------------- | -------------------- | | `application/json`, `application/vnd.debezium+json` | JSON change event(s) | | `application/avro`, `application/vnd.debezium+avro` | Avro change event | Writing to this endpoint requires write access — see [API authentication](/docs/next/api/auth). ### JSON body shapes[​](#json-body-shapes "Direct link to JSON body shapes") A JSON body may be: * A single Debezium change event (schemaless, or with an embedded `schema`). * A JSON array of change events. * Newline-delimited JSON (NDJSON). Example create event: ``` { "before": null, "after": { "id": 1, "customer_id": 9, "amount": 12.5 }, "source": { "connector": "oracle", "ts_ms": 1710000000000 }, "op": "c", "ts_ms": 1710000000000 } ``` The `op` field maps to a row-level change: `c` create, `u` update, `d` delete, `r` snapshot read, `t` truncate. Internal Debezium `m` (message) events are skipped. ### Avro body[​](#avro-body "Direct link to Avro body") | Mode | How | | --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | | Confluent wire format | Body begins with the magic byte `0` and a 4-byte schema id; set `cdc_schema_registry_url` on the dataset so Spice can resolve the schema. | | Raw Avro | Body is a single Avro datum; provide the schema via the `cdc_avro_schema` dataset param or the `X-Avro-Schema` request header. | ### Response[​](#response "Direct link to Response") On success the endpoint returns the number of applied changes: ``` { "applied": 1, "dataset": "orders" } ``` | Status | Meaning | | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | | `200 OK` | Batch applied | | `400 Bad Request` | Malformed body, decode failure, or apply failure | | `403 Forbidden` | Write access required (see [API authentication](/docs/next/api/auth)) | | `404 Not Found` | Dataset is not registered for CDC ingest (not yet initialized, or not configured with `from: cdc:…` + `refresh_mode: changes`) | | `501 Not Implemented` | This `spiced` build was compiled without the `debezium` feature | | `503 Service Unavailable` | Ingest channel closed (stream stopped) | | `504 Gateway Timeout` | Apply timed out | Because the request blocks until the change is applied, a Debezium Server or Embedded sink should commit source offsets only after receiving `200 OK`. ## Debezium Server sink[​](#debezium-server-sink "Direct link to Debezium Server sink") Point a [Debezium Server](https://debezium.io/documentation/reference/stable/operations/debezium-server.html) HTTP sink at the Spice endpoint. Example `application.properties`: ``` debezium.sink.type=http debezium.sink.http.url=http://spiced:8090/v1/datasets/orders/cdc debezium.sink.http.headers.Content-Type=application/json # Source — any Debezium connector debezium.source.connector.class=io.debezium.connector.oracle.OracleConnector debezium.source.database.hostname=oracle # ... connector-specific settings ... ``` For Avro payloads, configure the sink's serializer for Avro and either set `cdc_schema_registry_url` on the Spice dataset (Confluent wire format) or provide `cdc_avro_schema`. A ready-to-copy example lives in the [`cdc-debezium-ingest`](https://github.com/spiceai/spiceai/tree/trunk/examples/cdc-debezium-ingest) directory of the `spiceai/spiceai` repository. ## Native vs. push[​](#native-vs-push "Direct link to Native vs. push") | | Native (`postgres` / `mysql` / `mongodb` / `dynamodb`) | Push (`cdc`) | | ------------------ | ------------------------------------------------------ | ------------------------------------------ | | Where capture runs | Inside `spiced` | Debezium plugin process | | Kafka required | No | No | | Formats | Internal | JSON + Avro | | Sources | Those Spice implements natively | Any database with a Debezium source plugin | Prefer a native path when one exists. Use push ingest for Oracle, SQL Server, Db2, and other engines where Debezium already ships a source connector but Spice has no native capture. ## Feature availability[​](#feature-availability "Direct link to Feature availability") Push ingest requires the `debezium` feature, which is included in the default `spiced` builds and official container images. A build compiled without it returns `501 Not Implemented` from the ingest endpoint. --- # DynamoDB Streams (Native CDC) Stream every `INSERT`, `UPDATE`, and `DELETE` from an Amazon DynamoDB table directly into a Spice-accelerated dataset using [DynamoDB Streams](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.html). This is the recommended way to keep a Spice accelerator ([DuckDB](/docs/next/components/data-accelerators/duckdb), [SQLite](/docs/next/components/data-accelerators/sqlite), [PostgreSQL](/docs/next/components/data-accelerators/postgres), Cayenne) continuously in sync with a DynamoDB source — no Kafka, no Debezium, no Lambda required. ## How it works[​](#how-it-works "Direct link to How it works") ``` ┌──────────────────┐ DynamoDB Streams ┌───────────────────┐ ChangeBatch ┌───────────────┐ │ DynamoDB │──────────────────────▶│ Spice runtime │──────────────────▶│ Accelerator │ │ table │ shard iterators │ (dynamodb │ (INSERT/ │ DuckDB / │ │ + stream │ │ connector) │ UPDATE / │ SQLite / │ │ │ │ │ DELETE) │ Postgres / │ └──────────────────┘ └───────────────────┘ │ Cayenne │ └───────────────┘ ``` On first start the connector: 1. Bootstraps the accelerator with a full `Scan` of the source table so initial state is captured. 2. Subscribes to each open **stream shard** and begins polling for records. 3. Applies each `INSERT` / `MODIFY` / `REMOVE` event as a row-level change to the accelerator. The dataset reports `Ready` once stream lag drops below `dynamodb_replication_ready_lag` (default `2s`). On subsequent restarts, file-backed accelerators resume from the persisted shard checkpoint instead of re-scanning the source table. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") ### 1. Enable Streams on the source table[​](#1-enable-streams-on-the-source-table "Direct link to 1. Enable Streams on the source table") Streams must be enabled with view type **`NEW_AND_OLD_IMAGES`** so Spice receives both the old and new image of every modified item. ``` aws dynamodb update-table \ --table-name orders \ --stream-specification StreamEnabled=true,StreamViewType=NEW_AND_OLD_IMAGES ``` ### 2. IAM permissions[​](#2-iam-permissions "Direct link to 2. IAM permissions") The credentials Spice uses need both table and stream actions: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "dynamodb:DescribeTable", "dynamodb:Scan", "dynamodb:DescribeStream", "dynamodb:GetShardIterator", "dynamodb:GetRecords" ], "Resource": [ "arn:aws:dynamodb:*:*:table/orders", "arn:aws:dynamodb:*:*:table/orders/stream/*" ] } ] } ``` ### 3. Acceleration must be enabled with `refresh_mode: changes`[​](#3-acceleration-must-be-enabled-with-refresh_mode-changes "Direct link to 3-acceleration-must-be-enabled-with-refresh_mode-changes") Using DynamoDB Streams **requires** acceleration with `refresh_mode: changes`. Supported engines are `duckdb`, `sqlite`, `postgres`, and `cayenne`. ## Minimal configuration[​](#minimal-configuration "Direct link to Minimal configuration") ``` datasets: - from: dynamodb:orders name: orders_stream params: dynamodb_aws_region: us-east-1 dynamodb_aws_access_key_id: ${secrets:aws_access_key_id} dynamodb_aws_secret_access_key: ${secrets:aws_secret_access_key} acceleration: enabled: true engine: duckdb mode: file # Persistence is recommended so restarts skip the initial Scan refresh_mode: changes ``` ## Tuning[​](#tuning "Direct link to Tuning") ``` datasets: - from: dynamodb:orders name: orders_stream params: scan_interval: 100ms # Poll DynamoDB Streams every 100 ms (default 0s) dynamodb_replication_ready_lag: 1s # Report Ready when stream lag drops below 1s (default 2s) dynamodb_replication_invalid_checkpoint_behavior: restart # Re-bootstrap on > 24h lag (default error) acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes snapshots: enabled snapshots_trigger: stream_batches snapshots_trigger_threshold: 5 # Snapshot every 5 batch updates ``` * `scan_interval` — Polling frequency for new records. Lower values give lower latency at the cost of more `GetRecords` API calls. * `dynamodb_replication_ready_lag` — Maximum stream lag before the dataset is reported `Ready` for queries. Defaults to `2s`. (Previously `ready_lag`, still accepted as a deprecated alias.) * `dynamodb_replication_invalid_checkpoint_behavior` — What to do when the persisted checkpoint can no longer be honored because stream lag exceeds the DynamoDB shard retention window (\~24h). One of `error` (default) or `restart`. Replaces the deprecated `lag_exceeds_shard_retention_behavior` (`ready_after_load` → `restart`; `ready_before_load` removed → `restart`). See the [connector parameter reference](/docs/next/components/data-connectors/dynamodb#params) for the full description. * `snapshots_trigger: stream_batches` and `snapshots_trigger_threshold` let you trigger acceleration snapshots based on stream-batch counts rather than wall time. See [Acceleration Snapshots](/docs/next/features/data-acceleration/snapshots). ## Metrics[​](#metrics "Direct link to Metrics") The connector exposes the following [component metrics](/docs/next/features/observability/component_metrics) for monitoring streaming health: | Metric | Type | Description | | -------------------------------------------------------- | ------- | -------------------------------------------------------------------------- | | `shards_active` | Gauge | Current number of active shards in the stream | | `records_consumed_total` | Counter | Total number of records consumed from the stream | | `lag_ms` | Gauge | Current lag in milliseconds between stream watermark and now | | `errors_transient_total` | Counter | Total number of transient errors encountered while polling from the stream | | `reinitializations_on_lag_exceeds_shard_retention_total` | Counter | Total rebootstrap operations triggered due to expired shards | These metrics are opt-in. See the [DynamoDB Streams Metrics](/docs/next/components/data-connectors/dynamodb#metrics) section of the connector docs for an example `metrics:` block and a sample Grafana dashboard. ## Limitations[​](#limitations "Direct link to Limitations") * DynamoDB Streams shards are retained for 24 hours. If Spice falls behind by more than that, the connector follows `dynamodb_replication_invalid_checkpoint_behavior` (default: `error`). * `refresh_sql` is not supported with DynamoDB Streams. ## See also[​](#see-also "Direct link to See also") * [DynamoDB Data Connector](/docs/next/components/data-connectors/dynamodb) — complete parameter reference, deployment guide, and supported AWS regions. * [DynamoDB Streams cookbook recipe](https://github.com/spiceai/cookbook/tree/trunk/dynamodb/streams#readme). * [`refresh_mode: changes`](/docs/next/features/data-acceleration/refresh-modes/changes) — refresh-mode reference. --- # MongoDB Change Streams (Native CDC) Stream every insert, update, replace, and delete from a MongoDB collection directly into a Spice-accelerated dataset using native [MongoDB Change Streams](https://www.mongodb.com/docs/manual/changeStreams/). This is the recommended way to keep a Spice accelerator ([DuckDB](/docs/next/components/data-accelerators/duckdb), [SQLite](/docs/next/components/data-accelerators/sqlite), [PostgreSQL](/docs/next/components/data-accelerators/postgres), [Turso](/docs/next/components/data-accelerators/turso), Cayenne) continuously in sync with a MongoDB source — no Kafka, no Debezium, no external services. ## How it works[​](#how-it-works "Direct link to How it works") ``` ┌──────────────────┐ Change Streams ┌───────────────────┐ ChangeBatch ┌───────────────┐ │ MongoDB │ ───────────────────────▶│ Spice runtime │──────────────────▶│ Accelerator │ │ replica set / │ fullDocument= │ (mongodb │ (INSERT/ │ DuckDB / │ │ sharded │ updateLookup │ connector) │ UPDATE / │ SQLite / │ │ cluster │ + resume tokens │ │ DELETE) │ Postgres / │ └──────────────────┘ └───────────────────┘ │ Turso / │ │ Cayenne │ └───────────────┘ ``` On first start the connector: 1. Opens a Change Stream on the source collection with `fullDocument=updateLookup`. 2. Emits a CDC `TRUNCATE` and applies a full snapshot of the collection as upsert rows. 3. Signals readiness, then processes Change Stream events in batches. Opening the Change Stream **before** the snapshot prevents gaps between the snapshot and the live stream. For file-backed accelerators (acceleration `mode: file` / `file_create` / `file_update`, or `engine: postgres`), Spice persists the most recent Change Stream resume token in a sidecar table named `spice_sys_mongodb` alongside the accelerator data. The token is committed only after the downstream accelerator write succeeds (at-least-once semantics). On restart, Spice resumes from the persisted token and skips the snapshot. In-memory accelerators do not persist a resume token; restarts re-bootstrap from a fresh snapshot. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * **MongoDB 4.0+** with Change Streams enabled. MongoDB requires a **replica set** or **sharded cluster** — single-node `mongod` deployments do not support Change Streams. * The MongoDB user must have the `changeStream` privilege on the source collection. * The accelerator must support upsert behavior — use `duckdb`, `sqlite`, `postgres`, `turso`, or `cayenne`. * `acceleration.primary_key: _id` is required. Delete events only include the document key, so Spice needs `_id` to route deletes. * `acceleration.on_conflict` must specify `upsert` on `_id` so update and replace events overwrite existing rows. ## Minimal configuration[​](#minimal-configuration "Direct link to Minimal configuration") ``` datasets: - from: mongodb:users name: users params: mongodb_host: localhost mongodb_port: '27017' mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} acceleration: enabled: true engine: duckdb mode: file # Persist resume tokens so restarts skip the snapshot refresh_mode: changes primary_key: _id on_conflict: _id: upsert ``` ## Tuning[​](#tuning "Direct link to Tuning") These optional runtime parameters live under dataset `params:`. Defaults are reasonable; tune only when you have a specific batching or oplog-window concern. | Parameter Name | Default | Description | | --------------------------------------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `change_stream_batch_max_size` | `1000` | Max number of Change Stream events to group into one CDC batch before applying it. | | `change_stream_batch_max_duration` | `1s` | Max time to wait for a Change Stream batch to fill before applying it. Accepts [fundu](https://docs.rs/fundu) duration strings. | | `change_stream_max_await_time` | `1s` | Max time MongoDB waits for new events before returning an empty server batch. Accepts fundu duration strings. | | `change_stream_batch_size` | `1000` | Number of Change Stream events MongoDB should request from the server per batch. | | `mongodb_resume_token_invalid_behavior` | `error` | Behavior when a persisted resume token is rejected (e.g. past the oplog window). `error` surfaces the failure; `rebootstrap` drops the token and re-snapshots. | The existing `mongodb_unnest_depth` parameter applies to Change Stream documents too, so nested BSON is flattened the same way as normal MongoDB reads. ## Event mapping[​](#event-mapping "Direct link to Event mapping") | MongoDB event | Applied as | Notes | | ---------------------------------------------- | --------------- | -------------------------------------------------------------------------------------------- | | `insert` | create / upsert | Uses `fullDocument`. | | `update` | update / upsert | Uses `fullDocument` from `fullDocument=updateLookup`. | | `replace` | update / upsert | Uses `fullDocument`. | | `delete` | delete | Uses `documentKey`; non-key columns are `null`. | | `drop`, `rename`, `dropDatabase`, `invalidate` | truncate | Collection continuity is no longer guaranteed; the accelerator is reset and re-bootstrapped. | If MongoDB does not include `fullDocument` for an update or replace event, Spice fails the stream with a clear error instead of applying a partial row. ## Resumability across restarts[​](#resumability-across-restarts "Direct link to Resumability across restarts") For file-accelerated datasets, the persisted resume token lets Spice resume from where it left off without re-snapshotting. When MongoDB rejects the token (typical codes `ChangeStreamHistoryLost` 286 or `ChangeStreamFatalError` 280 — usually when the oplog window has rolled past the token's position), the behavior is governed by `mongodb_resume_token_invalid_behavior`: * `error` (default) — Spice surfaces a clear error and stops; the operator decides what to do. * `rebootstrap` — Spice drops the persisted token and re-snapshots the collection. Re-snapshotting a large collection is opt-in by default to prevent silent expensive rebootstraps. ## Limitations[​](#limitations "Direct link to Limitations") * Change Streams require a replica set or sharded cluster — they do not work against a single-node `mongod`. * `refresh_sql` is not supported with Change Streams. * In-memory accelerators do not persist resume tokens; every restart re-snapshots. ## See also[​](#see-also "Direct link to See also") * [MongoDB Data Connector](/docs/next/components/data-connectors/mongodb) — complete parameter reference and connection options. * [`refresh_mode: changes`](/docs/next/features/data-acceleration/refresh-modes/changes) — refresh-mode reference. --- # MySQL Binlog Replication (Native CDC) Stream every `INSERT`, `UPDATE`, and `DELETE` from a MySQL table directly into a Spice-accelerated dataset by subscribing to the source's binary log (binlog) — the MySQL analog of [PostgreSQL Logical Replication](/docs/next/features/cdc/postgres-replication). This is the recommended way to keep a Spice accelerator ([DuckDB](/docs/next/components/data-accelerators/duckdb), [SQLite](/docs/next/components/data-accelerators/sqlite), [PostgreSQL](/docs/next/components/data-accelerators/postgres), [Cayenne](/docs/next/components/data-accelerators/cayenne), [Turso](/docs/next/components/data-accelerators/turso), Arrow) continuously in sync with a MySQL source. No Kafka, no Debezium, no external CDC infrastructure — one `spicepod.yml` entry gives you a locally materialized, continuously updated copy of a MySQL table. ## How it works[​](#how-it-works "Direct link to How it works") ``` ┌────────────┐ binlog (ROW events) ┌──────────────────────────────┐ │ MySQL │ ──────────────────────► │ spiced │ │ (source) │ │ decode → ChangeBatch → │ │ │ │ accelerator upsert/delete │ └────────────┘ └──────────────────────────────┘ ▲ │ └── resume position persisted in the ────┘ accelerator (spice_sys_mysql_binlog) ``` On first start the connector: 1. **Validates and discovers.** Spice validates the source server settings (see [Prerequisites](#prerequisites)) and reads the table's column layout from `information_schema` (binlog row events are positional — they carry no column names). 2. **Captures the head position.** The current binlog file and offset is captured *before* the snapshot, so changes racing the snapshot are delivered at least once and converge via the primary-key upsert. 3. **Snapshots.** The table's existing rows stream in over a `START TRANSACTION WITH CONSISTENT SNAPSHOT` read, in batches of `mysql_replication_bootstrap_batch_size`. A truncate barrier is applied first so a re-bootstrap over a persistent accelerator starts clean. 4. **Persists the position.** Once the snapshot is durably applied, the captured head position is written to the accelerator's `spice_sys_mysql_binlog` sidecar table. This is Spice's replacement for a Postgres replication slot — MySQL keeps no per-replica cursor server-side. 5. **Streams.** Spice attaches as a replica (`COM_BINLOG_DUMP`) and turns committed transactions into insert/update/delete/truncate changes. The committed position checkpoints to the sidecar every `mysql_replication_checkpoint_interval`. On subsequent restarts, if a persisted position exists, the connector resumes from it directly — no snapshot — and marks the dataset ready immediately. Because the position lives inside the accelerator itself, data and cursor share one lifecycle: a non-persistent accelerator (`arrow`, or `mode: memory`) boots with no position and naturally re-snapshots — no special configuration needed. Delivery is **at-least-once**: a crash between applying a change and checkpointing its position replays up to one checkpoint interval of history, which the primary-key upsert absorbs idempotently. This is the same contract as the [PostgreSQL connector's](/docs/next/features/cdc/postgres-replication) snapshot/WAL boundary. ## Minimal configuration[​](#minimal-configuration "Direct link to Minimal configuration") ``` datasets: - from: mysql:mydb.orders name: orders params: mysql_host: db.internal mysql_tcp_port: '3306' mysql_user: replicator mysql_pass: ${secrets:mysql_pass} mysql_db: mydb acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes primary_key: id on_conflict: id: upsert ``` `primary_key` + `on_conflict: upsert` are **required** (except on the append-only `arrow` engine): UPDATE events apply as upserts keyed on the primary key and DELETE events are routed by it. The connector fails fast at startup with an actionable message if either is missing. All upsert-capable accelerator engines are supported — `duckdb`, `sqlite`, `cayenne`, `postgres`, and `turso` — and each persists the binlog resume position in its `spice_sys_mysql_binlog` sidecar when file-backed. The `arrow` engine works append-only (UPDATEs insert new rows). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") The following source server settings are required. Spice validates them at startup and fails fast with an actionable message if any is wrong. | Setting | Required value | Notes | | ---------------------------- | ---------------------------------- | ---------------------------------------------------------------- | | `log_bin` | `ON` | Default on MySQL 8.0+. | | `binlog_format` | `ROW` | Default on MySQL 8.0+. Validated at startup. | | `binlog_row_image` | `FULL` | Default. Validated at startup; `MINIMAL` images are rejected. | | `binlog_row_value_options` | `''` (empty) | Validated at startup; partial JSON row images cannot be applied. | | `binlog_expire_logs_seconds` | ≥ longest expected spiced downtime | See [When the position is purged](#when-the-position-is-purged). | The connecting user needs: ``` GRANT SELECT ON mydb.orders TO 'replicator'@'%'; -- snapshot + layout discovery GRANT REPLICATION SLAVE, REPLICATION CLIENT ON *.* TO 'replicator'@'%'; ``` `REPLICATION SLAVE` (streams the binlog) and `REPLICATION CLIENT` (reads its position) are **global-only** privileges in MySQL — they cannot be granted per-database, which is why they are granted `ON *.*`. On MariaDB 10.5+ the same two privileges are named `REPLICATION REPLICA` and `BINLOG MONITOR`; either spelling satisfies Spice, as does `ALL PRIVILEGES`. Spice pre-flights these grants via `SHOW GRANTS FOR CURRENT_USER()` before it validates the server settings above — an account without `REPLICATION CLIENT` can still read those settings, so checking them first would report a healthy server and defer the real problem to an opaque `Access denied` at stream start. A definitively missing privilege fails the dataset with the account name, the privileges it lacks, and a ready-to-paste `GRANT`: ``` MySQL account replicator@% is missing privileges required for change data capture: REPLICATION SLAVE, REPLICATION CLIENT. Grant them with: GRANT REPLICATION SLAVE, REPLICATION CLIENT, SELECT ON *.* TO 'replicator'@'%'; ``` The suggested `GRANT` includes `SELECT ON *.*` so it can be pasted as-is; narrow the `SELECT` to the replicated tables (as in the block above) if you prefer least privilege. Two cases deliberately do **not** fail the pre-flight: * **`SELECT` is not audited.** It is grantable at database, table, and column scope, and the check runs before the dataset's table is known — so a valid per-table grant would read as missing. A missing `SELECT` instead surfaces at snapshot time, named against the table that needs it. * **An inconclusive `SHOW GRANTS` defers to the server.** If the statement cannot be read, or lists a role grant (whose constituent privileges `SHOW GRANTS` does not expand), no conclusion is drawn and startup proceeds rather than blocking a dataset that would replicate correctly. ## Parameters[​](#parameters "Direct link to Parameters") Configure replication behavior with the following `params` on the MySQL dataset: | Parameter | Default | Description | | ----------------------------------------------- | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mysql_replication_server_id` | derived | The `server_id` this replica registers with. Must be unique among all replicas attached to the source. The default is derived from the **connection identity** (host, port, user, credentials, TLS) mixed with a per-process nonce, so datasets sharing a connection coalesce onto one binlog connection while two spiced instances don't collide, and derived values stay at or above `100000` to avoid hand-assigned replica ids. Setting distinct explicit values is how to opt a dataset out of [connection sharing](#sharing-one-binlog-connection). | | `mysql_replication_initial_snapshot` | `auto` | When existing rows load: `auto` snapshots when no resumable position exists and resumes without a snapshot when one does; `disabled` streams changes only; `always` re-snapshots on every start, discarding any persisted position. | | `mysql_replication_checkpoint_interval` | `10s` | How often the committed position persists to the sidecar. Bounds crash-replay volume. | | `mysql_replication_bootstrap_batch_size` | `8192` | Rows per emitted snapshot batch. Maximum: `1048576`. | | `mysql_replication_invalid_checkpoint_behavior` | `error` | What to do when the persisted position cannot be resumed losslessly — it was purged from the source, the source's GTID history diverged from the checkpoint, or the source table's column layout drifted incompatibly with the recorded position: `error`, or `restart` (drop the position and re-snapshot). | | `mysql_replication_ready_lag` | `2s` | For `refresh_mode: changes`, the dataset is marked Ready once its replication lag (now minus the newest applied commit's binlog-header timestamp) falls below this. The dataset stays not-ready while snapshotting or draining a backlog on resume, so it never serves stale data. | The runtime-level CDC apply tunables (`cdc_prefetch_buffer`, `cdc_max_coalesced_envelopes`, `cdc_max_coalesced_bytes`, `cdc_max_coalesce_age_ms`, `cdc_commit_timeout_ms`) apply to this connector the same way they do to the PostgreSQL one — see [Tuning ingestion](/docs/next/features/cdc#tuning-ingestion). ## Resume identity: GTID or file + offset[​](#resume-identity-gtid-or-file--offset "Direct link to Resume identity: GTID or file + offset") The connector picks its resume identity from the source. When the source reports `@@GLOBAL.gtid_mode = ON`, Spice resumes with **GTID auto-positioning** and persists the executed GTID set in the `spice_sys_mysql_binlog` sidecar alongside the binlog file and offset. Otherwise it resumes from **file + offset**: | Source `@@GLOBAL.gtid_mode` | Resume identity | | ------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------- | | `ON` | GTID auto-positioning | | `OFF`, `ON_PERMISSIVE`, `OFF_PERMISSIVE` | File + offset — a mixed topology can still emit anonymous transactions, which GTID auto-positioning cannot resume from. | | Variable not supported (MariaDB, pre-GTID MySQL) | File + offset | A GTID cursor survives a source failover: a promoted replica inherits the same global GTIDs, so the persisted set still resolves against the new primary. File + offset positions do not — binlog coordinates are per-server. ## When the source is reset or rebuilt[​](#when-the-source-is-reset-or-rebuilt "Direct link to When the source is reset or rebuilt") On a GTID resume, Spice requires the persisted GTID set to be a **subset of the source's current `@@gtid_executed`**. A `RESET MASTER`, a rebuilt server with a fresh `server_uuid`, a different source, or a diverged history reports an executed set that no longer contains the checkpoint — resuming from it would silently serve pre-reset data. When the check fails, `mysql_replication_invalid_checkpoint_behavior` decides the response, exactly as for a purged position: `error` (the default) stops with an actionable error, and `restart` drops the stale position and re-snapshots. On the file + offset path the equivalent guard is the presence of the persisted binlog file in the source's log index. ## When the position is purged[​](#when-the-position-is-purged "Direct link to When the position is purged") MySQL expires binary logs on its own schedule (`binlog_expire_logs_seconds`, default 30 days; much shorter on some managed services). If spiced is down long enough for the persisted position's file to be purged, the stream cannot resume losslessly. By default this surfaces as an error naming the fix; set: ``` params: mysql_replication_invalid_checkpoint_behavior: restart ``` to instead drop the stale position, truncate the accelerator, and re-snapshot the table automatically. ## Sharing one binlog connection[​](#sharing-one-binlog-connection "Direct link to Sharing one binlog connection") `COM_BINLOG_DUMP` has no server-side table filter — every subscriber receives the whole server binlog — so a connection per dataset would duplicate the entire stream for no benefit. Sharing is therefore **always on and requires no configuration**: every `refresh_mode: changes` dataset that connects the same way joins a single shared source (one dump connection, one `server_id`), and decoded transactions are routed to each member's accelerator by `(database, table)`. Sharing is keyed by connection identity — host, port, user, password, TLS settings, and `mysql_replication_server_id`. The database is **not** part of the key, so datasets on the same server but different databases still share one connection. Datasets that connect differently (a different user, or an explicit distinct `mysql_replication_server_id`) get their own connection. Because MySQL keeps no server-side cursor, each member persists its own committed position in its own `spice_sys_mysql_binlog` sidecar row, and the shared connection resumes from the **minimum** committed position across members; members already ahead of it suppress the replay. A member taking its initial snapshot is held out of routing with its floor pinned at the position it joined at, so a long snapshot never back-pressures the others. Delivery into a member is must-deliver: the pump has to place each envelope in that dataset's bounded channel before it reads more of the stream, so one member's stalled apply loop stops the dump socket being drained for every member on the connection. `COM_BINLOG_DUMP` is one-way once started — the server is never waiting on Spice — so the source's `net_write_timeout` counts purely against data Spice has not read, and at the MySQL default of 60 seconds one long apply cycle is enough for the source to abort the dump and make the whole group reconnect. Spice therefore raises the timeout on the dump session before starting the dump: ``` SET SESSION net_write_timeout = GREATEST(@@SESSION.net_write_timeout, 180) ``` It is a **floor**, not an assignment — raising `net_write_timeout` on the source is the manual workaround for this symptom, and a value already higher than 180 seconds is left alone. If the source rejects the statement (for example the replication user may not set session variables), Spice logs a warning and continues on the server's value; grant that permission or raise `net_write_timeout` on the source itself. Time the pump spends blocked on a member accrues per dataset in `replication_member_send_stalled_seconds_total`. A stalled member can force the group to re-bootstrap Unlike a PostgreSQL replication slot, a held binlog position pins nothing server-side. If a member detaches — a schema change classified as DDL is member-fatal, detaching that one dataset while the group keeps running — its held floor holds back only the shared resume position. Should the source then purge binlogs past that point (`binlog_expire_logs_seconds`), the whole group must re-bootstrap. A detach logs an `ERROR` and flips that dataset's `dataset_mysql_replication_member_attached` gauge to `0`, which identifies exactly which dataset is holding the group back. Recovery is bounded per member by `mysql_replication_invalid_checkpoint_behavior` on the next resume. A slow-but-live member is handled by back-pressure and is never detached. ## When the source layout changes[​](#when-the-source-layout-changes "Direct link to When the source layout changes") Binlog row events are positional — each event carries column values in source-ordinal order, not by name — so the connector tracks the source table's column layout and refuses to apply events it can't line up against that layout. Compatible changes are adopted automatically: an additive `ALTER TABLE` is picked up from `information_schema` at the schema-change boundary and the stream continues without interruption — subject to the cross-check in [Layout adopted under lag](#layout-adopted-under-lag). An **incompatible** change — one where the recorded position would replay row images against a layout that no longer matches (for example resuming after the source table's shape drifted while spiced was down, in a way the stream cannot reconcile with the events it still needs to replay) — cannot be applied without risking silent column misalignment. Rather than corrupt the accelerator, the connector stops and leaves the last known-good position in place. As with a purged position, `mysql_replication_invalid_checkpoint_behavior` controls the response: `error` (the default) surfaces an actionable error, and `restart` drops the position and re-snapshots the table from the current layout. ### Layout adopted under lag[​](#layout-adopted-under-lag "Direct link to Layout adopted under lag") A layout re-read from `information_schema` describes the source table **as it is now**, which under replication lag can already be a *later* DDL than the events still in flight. If the source applies a second `ALTER TABLE` with the same column count — a reorder (`MODIFY ... FIRST` / `AFTER`) or a rename swap — before Spice reaches the first one, the adopted layout maps ordinals the in-flight row images do not use. Column counts still agree, every name still resolves, and values still convert, so nothing would fail on its own. To catch this, every routed change is cross-checked against the column types the binlog's own `TableMap` event carries — the one description of the row image that travels *with* it. A disagreement is member-fatal: that dataset detaches with an error naming the column and its ordinal, and the rest of the shared binlog group keeps running. Recover by letting replication catch up before the next schema change, then re-bootstrapping the detached dataset with `mysql_replication_invalid_checkpoint_behavior: restart`. The comparison is deliberately coarse, because a false positive would break a healthy stream. It compares only type *classes*, and skips types whose wire encoding is ambiguous — `DATETIME`/`TIMESTAMP`/`TIME`, `TEXT`/`BLOB`, and `REAL` are not compared at all, and `CHAR`/`VARCHAR`/`BINARY`/`VARBINARY`/`ENUM`/`SET` are treated as a single class. What it does detect decisively is a reorder that moves columns of genuinely different types across each other (for example `INT` ↔ `VARCHAR`, `INT` ↔ `BIGINT`, or `DECIMAL` ↔ `INT`). Each detection increments `replication_schema_mismatch_errors_total`. ## Semantics and type notes[​](#semantics-and-type-notes "Direct link to Semantics and type notes") * **TRUNCATE TABLE** on the source applies as a truncate on the accelerator. * **Primary-key updates** (`UPDATE ... SET id = ...`) apply as a delete of the old key plus an upsert of the new row, so no orphan rows linger. * **TIMESTAMP columns** replicate as UTC (that is how the binlog stores them), and the snapshot pins its session to UTC to match. Set `mysql_time_zone: '+00:00'` (the default) so federated reads agree. * **Zero dates** (`0000-00-00`) coerce to `NULL`, matching the read connector's default `mysql_zero_date_behavior: null`. * **ENUM / SET** columns resolve to their label strings using the definition in `information_schema`. * **UNSIGNED integers** map to the same signed Arrow types the read connector uses; values above the signed maximum fail loudly rather than wrap. * **Negative `TIME` values** and spatial/geometry types are not supported — exclude such columns from the dataset schema. ## Schema changes[​](#schema-changes "Direct link to Schema changes") On `ALTER TABLE` against the replicated table, Spice re-fetches the table's layout from `information_schema` and keeps streaming as long as every dataset column still exists on the source — the same tolerance the PostgreSQL connector's block mode has: * **Columns added on the source** are not replicated (a warning names them); add them to the dataset schema and restart to capture them. * **Dropping or renaming a dataset column** (or `RENAME TABLE` / `DROP TABLE`) stops the stream with an actionable error. * **Retyping a dataset column** keeps streaming while values remain convertible to the dataset's Arrow type; an unconvertible value stops the stream with a decode error. If the stream stopped across a DDL boundary with an un-checkpointed tail, the restart may be unable to decode pre-DDL events — re-bootstrap with `mysql_replication_invalid_checkpoint_behavior: restart`. Quiescing writes to the table around DDL avoids that case entirely. Issuing a second `ALTER TABLE` before replication has caught up on the first can also detach the dataset — see [Layout adopted under lag](#layout-adopted-under-lag). ## Metrics[​](#metrics "Direct link to Metrics") Exposed under `dataset_mysql_*` alongside the connection-pool [component metrics](/docs/next/features/observability/component_metrics): | Metric | Meaning | | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `replication_gtid_enabled` | `1` while the stream positions by GTID auto-positioning, `0` for binlog file + offset — see [Resume identity](#resume-identity-gtid-or-file--offset). | | `replication_lag_ms` | Now minus the newest applied source commit timestamp (1s granularity). | | `replication_lag_bytes` | Binlog bytes between the source head and the resume position. Reported only while both are in the same binlog file; absent otherwise. | | `replication_source_head_file` / `replication_source_head_pos` | The source server's binlog head, polled every checkpoint interval. | | `replication_committed_binlog_file` / `replication_committed_binlog_pos` | The checkpointed resume position. | | `replication_transactions_total` | Source transactions observed. | | `replication_inserts_total` / `replication_updates_total` / `replication_deletes_total` / `replication_truncates_total` | Row events applied. | | `replication_bootstrap_rows_total` / `replication_bootstrap_rows_expected` / `replication_bootstrap_complete` | Snapshot progress. | | `replication_decode_errors_total` | Binlog decode failures. | | `replication_schema_mismatch_errors_total` | Mid-stream DDL detections. | | `replication_recv_errors_total` / `replication_reconnects_total` | Transport health. | | `replication_checkpoint_persists_total` / `replication_checkpoint_persist_errors_total` | Sidecar checkpoint writes. | | `replication_member_attached` | `1` while the dataset is an attached member of its shared binlog group, `0` once detached — see [Sharing one binlog connection](#sharing-one-binlog-connection). | | `replication_member_send_stalled_seconds_total` | Seconds the shared dump pump spent blocked delivering changes into this dataset's channel. Rising means this member's apply loop is holding up the stream for every member on the connection — see [Sharing one binlog connection](#sharing-one-binlog-connection). | ## Limitations[​](#limitations "Direct link to Limitations") * **Failover only survives on a GTID source.** With `gtid_mode = ON` the persisted GTID set resolves against a promoted replica; on a file + offset source, resume positions do not survive a failover to a different primary (see [Resume identity](#resume-identity-gtid-or-file--offset)). * **One table per dataset.** Every dataset replicates a single source table, though datasets that connect the same way share one binlog connection — see [Sharing one binlog connection](#sharing-one-binlog-connection). * **Schema evolution is block-mode only** — compatible `ALTER TABLE` is tolerated (see [Schema changes](#schema-changes)), but `on_schema_change` policies that *adopt* new columns (`append_new_columns` / `sync_all_columns`) are not yet wired to this connector. * **XA (two-phase) transactions are not supported.** An XA transaction that touches the replicated table stops the stream with an error; XA activity on other tables logs a warning and is ignored. * Not supported source types: geometry/spatial, vectors, negative `TIME`. --- # PostgreSQL Logical Replication (Native CDC) Stream every `INSERT`, `UPDATE`, and `DELETE` from a PostgreSQL table directly into a Spice-accelerated dataset over Postgres' native logical replication protocol. This is the recommended way to keep a Spice accelerator ([DuckDB](/docs/next/components/data-accelerators/duckdb), [SQLite](/docs/next/components/data-accelerators/sqlite), [PostgreSQL](/docs/next/components/data-accelerators/postgres), Cayenne, Arrow) continuously in sync with a PostgreSQL source. ## How it works[​](#how-it-works "Direct link to How it works") ``` ┌────────────────┐ WAL (pgoutput) ┌───────────────────┐ ChangeBatch ┌───────────────┐ │ PostgreSQL │──────────────────▶│ Spice runtime │────────────────▶│ Accelerator │ │ wal_level= │ replication │ (postgres │ (INSERT/ │ DuckDB / │ │ logical │ slot │ connector) │ UPDATE / │ SQLite / │ │ │ │ │ DELETE) │ Postgres / │ └────────────────┘ └───────────────────┘ │ Cayenne │ └───────────────┘ ``` On first start the connector: 1. Creates a **publication** (default name `spice___pub`) containing the source table. 2. Creates a **replication slot** (default `spice___`). The `` gives each Spice replica its own slot. 3. Runs a **REPEATABLE READ snapshot** of the source table so the accelerator starts with all existing rows. 4. Starts streaming WAL changes from the slot. Each committed transaction is delivered as a `ChangeBatch` (grouped `INSERT`/`UPDATE`/`DELETE`) and applied to the accelerator. On subsequent restarts the connector compares the slot against the position it recorded locally for this acceleration, and either resumes streaming or rebuilds the accelerated table — see [Recovering from a lost replication slot](#recovering-from-a-lost-replication-slot). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") ### 1. Enable logical replication on the source Postgres[​](#1-enable-logical-replication-on-the-source-postgres "Direct link to 1. Enable logical replication on the source Postgres") This requires a server restart. ``` # postgresql.conf wal_level = logical max_replication_slots = 10 # at least one per Spice replica per dataset max_wal_senders = 10 ``` Verify: ``` SHOW wal_level; -- must be 'logical' SHOW max_replication_slots; ``` On managed Postgres services: | Service | How to enable | | ----------------- | --------------------------------------------------------------------- | | AWS RDS | Set `rds.logical_replication = 1` in the parameter group and restart. | | Aurora PostgreSQL | Set `rds.logical_replication = 1`; wait for DB reboot. | | GCP Cloud SQL | Flag: `cloudsql.logical_decoding = on`. | | Azure Database | Under **Replication**, set *Replication support* to `LOGICAL`. | | Supabase / Neon | Logical replication is enabled by default. | ### 2. The source table must have a replica identity[​](#2-the-source-table-must-have-a-replica-identity "Direct link to 2. The source table must have a replica identity") Spice needs the primary key columns in every `UPDATE`/`DELETE` event, so one of the following must be true: * The table has a **primary key** (default — nothing to do). * Or the table has `REPLICA IDENTITY FULL`: ``` ALTER TABLE public.users REPLICA IDENTITY FULL; ``` Tables with `REPLICA IDENTITY NOTHING` are rejected at startup. ### 3. The Postgres role needs these privileges[​](#3-the-postgres-role-needs-these-privileges "Direct link to 3. The Postgres role needs these privileges") ``` GRANT CONNECT ON DATABASE mydb TO spice; GRANT USAGE ON SCHEMA public TO spice; GRANT SELECT ON public.users TO spice; ALTER ROLE spice WITH REPLICATION; -- or be a superuser -- If you let Spice create the publication (default): GRANT CREATE ON DATABASE mydb TO spice; ``` ## Minimal configuration[​](#minimal-configuration "Direct link to Minimal configuration") ``` datasets: - from: postgres:public.users name: users params: pg_host: pg.internal pg_port: '5432' pg_user: spice pg_pass: ${secrets:pg_pass} pg_db: myapp pg_sslmode: verify-full # or: disable | prefer | require | verify-ca pg_sslrootcert: /etc/ssl/pg-ca.pem # optional; omit to use system root CAs acceleration: enabled: true engine: duckdb # or: sqlite | postgres | cayenne | arrow refresh_mode: changes # <-- triggers WAL streaming primary_key: id on_conflict: id: upsert # required for UPDATE to become an upsert ``` Start the runtime. Spice will: * Auto-create publication `spice_users__pub`. * Auto-create replication slot `spice_users__`. * Snapshot `public.users` into the DuckDB accelerator. * Stream every subsequent change as it commits on Postgres. Use a persistent accelerator Pair with `mode: file` on DuckDB/SQLite (or the PostgreSQL accelerator) so restarts resume from the last acknowledged LSN instead of re-snapshotting. ## Full configuration reference[​](#full-configuration-reference "Direct link to Full configuration reference") All replication-specific parameters live under `params:` on the dataset and start with `pg_`: | Parameter | Default | Description | | ---------------------------------------- | ------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `pg_replication_slot` | `spice___` | Name of the replication slot to create/reuse — see [Replication slot naming](#replication-slot-naming) for the accepted characters. Datasets on the same connection that name the same slot **share** it — one slot, one publication, one replication connection — see [Sharing one slot across datasets](#sharing-one-slot-across-datasets). Each Spice replica must still use its own unique slot. | | `pg_publication` | `spice___pub` | Publication name. Defaults to `_pub` when `pg_replication_slot` is set explicitly (so datasets sharing a slot agree on it). Shared across replicas. Auto-created if missing. | | `pg_replication_initial_snapshot` | `auto` | When `refresh_mode: changes` first loads existing rows: `auto` (default) snapshots a freshly-created slot and resumes an existing one without a snapshot (a non-persistent accelerator still re-snapshots on every start); `disabled` streams WAL changes only; `always` snapshots on every start, including slot resume. The legacy booleans `true`/`false` are deprecated and map to `auto`/`disabled`. | | `pg_replication_ready_lag` | `2s` | For `refresh_mode: changes`, the dataset is marked Ready once its replication lag (now minus the newest applied commit's source time) falls below this. It stays not-ready while snapshotting or draining a backlog on resume, so it never serves stale data. | | `pg_replication_temporary_slot` | — | **Deprecated and ignored.** The slot is always durable. A temporary slot belongs to the Postgres session that creates it, and Spice creates the slot on a short-lived setup connection, so Postgres dropped it before `START_REPLICATION` could attach and the stream could never start. Setting it logs a deprecation warning and has no other effect — remove it. To stop an unused slot retaining WAL on the source, drop it with `SELECT pg_drop_replication_slot('')`. | | `pg_replication_status_interval` | `10s` | How often `StandbyStatusUpdate` (LSN acknowledgement) is sent back to Postgres. Lower values free WAL faster; higher values reduce network chatter. Accepts any duration string (`500ms`, `30s`, `2m`). | | `pg_replication_bootstrap_batch_size` | `8192` | Rows per emitted batch during the initial-snapshot bootstrap. Larger batches reduce per-batch overhead at the cost of more memory per batch. Maximum: `1048576`. | | `pg_replication_member_channel_capacity` | `1024` | Shared-slot only: envelopes buffered per member table before the shared replication pump back-pressures. Too small a value lets one member's transient stall block the whole slot (head-of-line blocking). Maximum: `1048576`. | All existing `pg_host`, `pg_port`, `pg_user`, `pg_pass`, `pg_db`, `pg_sslmode`, `pg_connection_string` parameters continue to apply — see the [PostgreSQL Data Connector](/docs/next/components/data-connectors/postgres) reference. ### Replication slot naming[​](#replication-slot-naming "Direct link to Replication slot naming") PostgreSQL restricts replication slot names to `[a-z0-9_]` — lowercase letters, digits, and underscores only — with a maximum length of 63 bytes, and reserves `pg_conflict_detection` for its own conflict-detection slot. An explicit `pg_replication_slot` is validated against these rules while the dataset's replication parameters are built, so a name with (for example) a hyphen or an uppercase letter fails immediately with an error naming the parameter, rather than surfacing later as a refresh-task failure from the server. The generated default (`spice___`) already conforms: the dataset name is sanitized and truncated to keep the whole identifier within the 63-byte limit. ### Connecting with `pg_connection_string`[​](#connecting-with-pg_connection_string "Direct link to connecting-with-pg_connection_string") `refresh_mode: changes` accepts `pg_connection_string` in place of the discrete `pg_host` / `pg_port` / `pg_user` / `pg_pass` / `pg_db` parameters, in both libpq `key=value` and `postgresql://` URI form, following the same rules as the federated read path: ``` datasets: - from: postgres:public.users name: users params: pg_connection_string: postgresql://spice:${secrets:pg_pass}@pg.internal:5432/myapp acceleration: enabled: true engine: duckdb refresh_mode: changes primary_key: id on_conflict: id: upsert ``` * The connection string **takes precedence** over discrete host/user/database parameters when both are set. * `pg_sslmode` and `pg_sslrootcert` are the exception: set discretely, they override whatever the connection string carries. * A connection string that omits `sslmode` defaults to **`verify-full`** on the replication transport — unlike the discrete-parameter path, where an unset `pg_sslmode` defaults to `prefer` (see below). * Unix-socket hosts are not supported for replication; the connection string must name a TCP host. ### `pg_sslmode` for WAL streaming[​](#pg_sslmode-for-wal-streaming "Direct link to pg_sslmode-for-wal-streaming") `verify-full` is the recommended production default. | `pg_sslmode` | Replication transport | Cert chain verified | Hostname verified | | ------------------------------------------- | --------------------- | ------------------- | ----------------- | | `disable` | plaintext | — | — | | `prefer` (default with discrete parameters) | plaintext | — | — | | `require` | TLS | ❌ | ❌ | | `verify-ca` | TLS | ✅ | ❌ | | `verify-full` | TLS | ✅ | ✅ | info `prefer` behaves as plaintext on the replication transport because the replication client does not expose a safe "try TLS, fall back to plaintext" path. Set `require`, `verify-ca`, or `verify-full` to force TLS on the WAL stream. When the connection is configured with [`pg_connection_string`](#connecting-with-pg_connection_string) and the string omits `sslmode`, the replication transport defaults to `verify-full` instead of `prefer`. ### Accelerator engines[​](#accelerator-engines "Direct link to Accelerator engines") | Engine | `INSERT` | `UPDATE` | `DELETE` | Notes | | ---------- | -------- | ---------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `duckdb` | ✅ | ✅ (upsert) | ✅ | Recommended for most workloads. | | `sqlite` | ✅ | ✅ (upsert) | ✅ | Great for small/medium datasets. | | `postgres` | ✅ | ✅ (upsert) | ✅ | Use when the accelerator is another Postgres. | | `cayenne` | ✅ | ✅ (upsert) | ✅ | S3-backed Vortex format, good for read-heavy analytics. | | `arrow` | ✅ | ✅ (upsert with primary key) | ✅ | Arrow's in-memory engine uses a hash index for primary-key upserts. Without a primary key, `UPDATE`s are appended as new rows. `DELETE` and `TRUNCATE` are applied via Arrow's `DeletionTableProvider`. | For Arrow workloads that need true upsert semantics (so `UPDATE`s replace existing rows instead of duplicating them), configure a `primary_key`. DuckDB, SQLite, PostgreSQL, and Cayenne also support upsert behavior. ## Sharing one slot across datasets[​](#sharing-one-slot-across-datasets "Direct link to Sharing one slot across datasets") By default each changes-mode dataset gets its own replication slot and publication. On the source database that costs one logical slot **and** one walsender decoder over the full WAL stream per dataset, so mirroring many tables can exhaust `max_replication_slots` and multiply decode work. When several datasets on the **same connection** name the same `pg_replication_slot`, Spice multiplexes them onto **one slot, one publication, and one replication connection**, routing decoded changes to each dataset's accelerator by `(schema, table)`. Sharing is implicit — name the same slot and the datasets share it: ``` datasets: - from: postgres:public.users name: users params: &repl pg_host: db.internal pg_db: app pg_user: spice pg_pass: ${secrets:pg_pass} pg_replication_slot: spice_app_cdc # same slot ⇒ shared acceleration: enabled: true engine: duckdb refresh_mode: changes primary_key: id on_conflict: id: upsert - from: postgres:public.orders name: orders params: *repl # same connection + slot ⇒ shares the slot above acceleration: enabled: true engine: duckdb refresh_mode: changes primary_key: id on_conflict: id: upsert ``` Notes: * A slot named by only one dataset behaves exactly as before (a single member). * Datasets without an explicit `pg_replication_slot` keep their dedicated per-dataset slot and publication. * Members of a shared slot must agree on the publication. The default becomes `_pub`; an explicit `pg_publication` still wins and is validated for consistency across members. * Each source table can back **at most one dataset per shared slot**. Pointing two datasets at the same `(schema, table)` through one slot is rejected — give the second dataset a different `pg_replication_slot` (or remove the param for a dedicated slot). * Sharing is per Spice instance. Across replicas, each replica must still use its own unique slot — see [Multi-replica deployments](#multi-replica-deployments). ### Envelope coalescing[​](#envelope-coalescing "Direct link to Envelope coalescing") Committed changes reach each member of a shared slot as **change envelopes**, not one unit of work per source transaction. A workload that commits constantly in small transactions would otherwise put an envelope per commit into each member's buffer, filling it long before the buffered rows are worth an apply — and while the shared pump is blocked delivering, it is not reading from the replication connection, so the back-pressure reaches the source walsender. The pump folds consecutive commits for the same member table together in two stages, both operating on raw pgoutput message chunks with no tuple decode: 1. **Eager hold** — the throughput stage. The pump holds one unpublished envelope per member and folds later commits for that table into it, until the envelope reaches the row limit or the age limit elapses. The age is measured from the *first* commit the envelope absorbed, so a low-traffic table is never held indefinitely. 2. **Mailbox tail fold** — the back-pressure stage. Publishing folds into the unclaimed tail of the member's buffer, with no age limit, so a member whose accelerator has stopped draining collapses envelopes instead of multiplying them. Folding never crosses a correctness boundary: changes destined for a different acknowledgement position, relation generation, or working schema are always kept in separate envelopes. The defaults need no tuning. They are process-wide operator escape hatches, set by environment variable rather than as dataset `params`: | Environment variable | Default | Maximum | Effect | | ------------------------------------------------------- | ------------------- | --------- | -------------------------------------------------------------------------------------------------------------------- | | `SPICE_POSTGRES_CDC_MAX_ENVELOPE_AGE_MS` | `10` | `60000` | How long the pump may hold a member's envelope open. `0` publishes every commit straight through, disabling stage 1. | | `SPICE_POSTGRES_CDC_MAX_ROWS_PER_ENVELOPE` | `8192` | `1048576` | Rows at which a held envelope is published. | | `SPICE_POSTGRES_CDC_MAX_BACKPRESSURE_ROWS_PER_ENVELOPE` | `2048` | `1048576` | Rows at which mailbox-tail folding seals an envelope. | | `SPICE_POSTGRES_CDC_MAX_MAILBOX_BYTES` | `33554432` (32 MiB) | 8 GiB | Estimated Arrow bytes one member's buffer may hold across every buffered envelope. | A value that is unparseable or above the maximum logs a warning and the default is used instead. The two stage-2 bounds ship deliberately low, because mailbox folding absorbs back-pressure rather than adding throughput. Raise them only on evidence: `dataset_postgres_replication_member_mailbox_coalesce_limited_total` rising alongside `dataset_postgres_replication_member_envelope_mailbox_merges_total` means the bounds are binding while folding is still paying off, whereas a `dataset_postgres_replication_member_mailbox_coalesce_limited_total` that stays at `0` means the bounds never bind and there is nothing to tune. See [Metrics](#metrics). ## Multi-replica deployments[​](#multi-replica-deployments "Direct link to Multi-replica deployments") Every Spice replica must have its own replication slot. Spice hashes the replica's identity into the default slot name: | Source | Used for | | ------------------------------- | ---------------------------------------------------------------------- | | `SPICE_INSTANCE_ID` env | Preferred — set it explicitly per replica. | | `HOSTNAME` / `COMPUTERNAME` env | Fallback — works on Kubernetes where each pod has a distinct hostname. | ### Example: Kubernetes StatefulSet[​](#example-kubernetes-statefulset "Direct link to Example: Kubernetes StatefulSet") ``` apiVersion: apps/v1 kind: StatefulSet metadata: name: spice spec: replicas: 3 serviceName: spice template: spec: containers: - name: spice env: - name: SPICE_INSTANCE_ID valueFrom: fieldRef: fieldPath: metadata.name # spice-0, spice-1, spice-2 ``` ### Example: explicit slot names[​](#example-explicit-slot-names "Direct link to Example: explicit slot names") ``` # Replica A params: pg_replication_slot: spice_users_a # Replica B params: pg_replication_slot: spice_users_b ``` Each Spice replica can use a different `pg_replication_slot` while sharing a publication (`pg_publication`). ## Operations[​](#operations "Direct link to Operations") ### Monitoring replication lag[​](#monitoring-replication-lag "Direct link to Monitoring replication lag") ``` SELECT slot_name, active, confirmed_flush_lsn, pg_wal_lsn_diff(pg_current_wal_lsn(), confirmed_flush_lsn) AS lag_bytes FROM pg_replication_slots WHERE slot_name LIKE 'spice_%'; ``` ### Decommissioning a replica[​](#decommissioning-a-replica "Direct link to Decommissioning a replica") Drop unused slots A permanent replication slot **holds on to WAL** until dropped. If you retire a Spice replica without cleaning up its slot, Postgres will keep accumulating WAL indefinitely and can run out of disk. After removing a Spice replica, drop its slot: ``` SELECT pg_drop_replication_slot('spice_users_'); ``` ### Recovering from a lost replication slot[​](#recovering-from-a-lost-replication-slot "Direct link to Recovering from a lost replication slot") A replication slot that is dropped or invalidated takes the server-side `confirmed_flush_lsn` with it, so the source can no longer say which changes a resumed stream still owes. Spice therefore records the position **locally** as well: each dataset keeps an applied-LSN watermark — the LSN its accelerated table is complete as of — in a `spice_sys_postgres_replication` sidecar table inside its own accelerator, the same way MySQL replication uses `spice_sys_mysql_binlog`. The watermark is written by the same call that acknowledges the slot, so it can never claim rows that are not durable yet. On each start the decision is arithmetic rather than an inference: | Recorded watermark | Slot state | Action | | ------------------------------------------------------ | --------------------------------------------------------------- | --------------------------------------------------------------------------------------------- | | None, on an accelerator that does not survive restarts | any | **First bootstrap** — snapshot, then stream (the accelerator boots empty every start) | | None, on a durable acceleration that can record one | any | **Rebuild** — a table that outlives the process may already hold rows this start did not load | | Present | Slot's `restart_lsn` is at or before the watermark | **Resume** — the WAL in between is still retained and is replayed | | Present | Slot's `restart_lsn` is past the watermark, or the slot is gone | **Rebuild** — the missing changes no longer exist on the source | | Recorded against a different source | any | **Rebuild** — LSNs are only comparable within one source's history | A rebuild replaces the accelerated table's contents through the ordinary full-refresh write path, so it is atomic: on Cayenne, readers keep seeing the pre-rebuild table until the new snapshot swaps in. Why a rebuild rather than another snapshot A snapshot bootstrap emits only insert events, and nothing clears a durable accelerated table first — so re-snapshotting over the existing rows is an upsert merge. A row deleted at the source while the slot was gone has no change event left to replay, and would survive in the acceleration and be returned by every later query. Rebuilding re-reads the table instead. The first start after upgrading a durable `refresh_mode: changes` dataset to a version that records watermarks has no recorded position, so it rebuilds once and records one from then on. If a durable acceleration has nowhere to record a watermark, Spice logs a warning at startup naming the dataset: slot loss cannot be detected for it, and rows deleted at the source while the slot was gone would survive in the acceleration. ### Rebuilding an accelerator from scratch[​](#rebuilding-an-accelerator-from-scratch "Direct link to Rebuilding an accelerator from scratch") Delete the accelerator's local storage (DuckDB file, SQLite file, etc.) and drop the replication slot. On next start, Spice will create a fresh slot, snapshot the table, and resume streaming. ### Resilience[​](#resilience "Direct link to Resilience") * **Network blips / Postgres restarts**: transient — retried with exponential backoff (500 ms → 30 s, ±20 % jitter). The slot's server-side state is the source of truth, so reconnects resume from the last acknowledged LSN — no data loss. * **Dropped or invalidated slot**: recovered automatically — the acceleration is rebuilt from the source rather than resumed on a gap. See [Recovering from a lost replication slot](#recovering-from-a-lost-replication-slot). * **Auth failures, schema mismatch**: fatal — surfaced as a stream-level error so operators can fix the configuration. * **Watch `dataset_postgres_replication_reconnects_total`** to detect flaky networks. ## Metrics[​](#metrics "Direct link to Metrics") Spice emits OpenTelemetry metrics for every replicated Postgres dataset. Metric names follow the pattern `dataset_postgres_replication_` with a `name=` attribute. Core freshness signals (auto-registered): | Metric | Type | Description | | ------------------------------------------------------------------------------------------------------------------------------------------ | ------- | ------------------------------------------------------------------------------------ | | `dataset_postgres_replication_lag_ms` | Gauge | `now() − commit_time(latest ingested txn)`. Primary CDC freshness signal. | | `dataset_postgres_replication_lag_bytes` | Gauge | `server_wal_end_lsn − confirmed_flush_lsn`. Unacknowledged WAL held by Spice's slot. | | `dataset_postgres_replication_transactions_total` | Counter | Committed transactions applied. | | `dataset_postgres_replication_inserts_total` / `dataset_postgres_replication_updates_total` / `dataset_postgres_replication_deletes_total` | Counter | Row-level events from WAL. | | `dataset_postgres_replication_reconnects_total` | Counter | Number of times the stream reconnected after a transient failure. | Shared-slot delivery and coalescing (auto-registered; reported only for datasets on a shared, explicitly-named slot — see [Envelope coalescing](#envelope-coalescing)): | Metric | Type | Description | | -------------------------------------------------------------------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `dataset_postgres_replication_member_envelopes_delivered_total` | Counter | Change envelopes delivered to this dataset as distinct units of work. Divide `dataset_postgres_replication_transactions_total` by this for the coalescing factor the accelerator's apply loop actually sees. | | `dataset_postgres_replication_member_envelope_eager_merges_total` | Counter | Committed transactions folded into an envelope the pump was still holding back, before it crossed into this dataset's buffer (stage 1). | | `dataset_postgres_replication_member_envelope_mailbox_merges_total` | Counter | Committed transactions folded into an envelope already sitting unclaimed in this dataset's buffer (stage 2). Rising alongside a flat `dataset_postgres_replication_member_send_stalled_seconds_total` means back-pressure is being absorbed rather than stalling the slot. | | `dataset_postgres_replication_member_mailbox_coalesce_limited_total` | Counter | Times a committed transaction could not be folded into the unclaimed buffer tail because a configured bound refused it, rather than because the changes were not foldable. `0` means the bounds never bind. | ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Cause and fix | | -------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Error: *`Table public.X has REPLICA IDENTITY NOTHING`* | Run `ALTER TABLE public.X REPLICA IDENTITY FULL;` (or add a primary key). | | Error: *`replication slot "..." already exists`* on startup | Another Spice replica is using the same slot name. Set `pg_replication_slot` uniquely, or ensure `SPICE_INSTANCE_ID` differs. | | Error mentioning *permission denied for database* during setup | The role needs `CREATE` on the database, or pre-create the publication/slot yourself. | | `pg_replication_slots.active` is `true` but the accelerator isn't updating | Check Spice logs for schema-mismatch errors. The replication task holds the slot even after failure — restart after fixing. | | `wal` on the source disk growing forever | An abandoned slot. Drop it with `pg_drop_replication_slot`. | | The whole table is re-read on a restart that used to resume | The slot's WAL no longer reaches the recorded position, or none was recorded (first start after an upgrade). See [Recovering from a lost replication slot](#recovering-from-a-lost-replication-slot). | | `UPDATE`s on Arrow-engine dataset don't replace rows | Configure a `primary_key` so Arrow can use its hash index for upserts, or switch to `duckdb`, `sqlite`, `postgres`, or `cayenne`. | | Huge `TEXT`/`JSONB` columns show as `NULL` after `UPDATE` | Unchanged TOASTed columns are omitted by pgoutput. Run `ALTER TABLE ... REPLICA IDENTITY FULL;` if you need them in every event. | ## Limitations[​](#limitations "Direct link to Limitations") * **One table per dataset.** Each Spice dataset replicates exactly one source table; each dataset gets its own slot and publication. * **No DDL replication.** Schema changes on the source are not propagated automatically. Add new columns as nullable on the source first, update the Spice dataset, then reload the Spicepod. * **Arrow engine** supports `on_conflict` upserts when a `primary_key` is configured. Without a primary key, `UPDATE`s appear as additional inserts rather than replacing existing rows. `DELETE` and `TRUNCATE` are applied either way. ## Comparison with Debezium + Kafka[​](#comparison-with-debezium--kafka "Direct link to Comparison with Debezium + Kafka") | Aspect | Debezium + Kafka | Native WAL streaming (this feature) | | ------------------------ | -------------------------------------------- | ------------------------------------------ | | External services | Kafka + Schema Registry + Debezium + Connect | None — Spice connects to Postgres directly | | Deployment footprint | JVM stack + ZooKeeper/KRaft | Zero extra pods | | Setup complexity | Multiple topics, connector configs, ACLs | One connector config | | Operational model | Consumer groups, topic retention | One replication slot per replica | | Schema registry required | Yes (Avro/Protobuf) | No — schema derived from Postgres catalog | | Latency | Kafka-bound (\~100 ms+) | Commit-driven, typically <100 ms | For greenfield Postgres → Spice CDC, prefer native WAL streaming. If Kafka is already deployed for other reasons, the [Debezium](/docs/next/components/data-connectors/debezium) path continues to work. ## See also[​](#see-also "Direct link to See also") * [Change Data Capture overview](/docs/next/features/cdc) * [PostgreSQL Data Connector](/docs/next/components/data-connectors/postgres) * [PostgreSQL: Logical Replication](https://www.postgresql.org/docs/current/logical-replication.html) --- # Data Acceleration Datasets and views can be locally accelerated by the Spice runtime, pulling data from any [Data Connector](/docs/next/components/data-connectors) and storing it locally in a [Data Accelerator](/docs/next/components/data-accelerators) for faster access. The data can be kept up-to-date in real-time or on a refresh schedule, ensuring deployments maintain the latest data locally for querying. ![Spice.ai Open Source Query Federation with Acceleration](/assets/images/data-acceleration-8a4401ca91dc42fc622704dd11d962c1.png) ## Benefits[​](#benefits "Direct link to Benefits") Local data acceleration stores data alongside the application, providing faster query times by eliminating network latency. This is especially beneficial for large query results, as data transfer over the network is avoided. Depending on the [Acceleration Engine](/docs/next/components/data-accelerators) used, data can also be stored in-memory, further reducing query times. [Indexes](/docs/next/features/data-acceleration/indexes) can be applied to speed up certain queries. Locally accelerated datasets can also have [primary key constraints](/docs/next/features/data-acceleration/constraints) applied. This feature supports specifying actions when a constraint is violated, such as dropping the violating row or upserting it into the accelerated table. [Acceleration snapshots](/docs/next/features/data-acceleration/snapshots) (preview) help file-mode accelerations become ready in seconds by bootstrapping from managed snapshots stored in object storage such as Amazon S3. For larger datasets, [partitioning](/docs/next/features/data-acceleration/partitioning) splits the acceleration into smaller physical units (Hive-style files, per-partition tables, or in-memory tables) keyed by an expression. Queries that filter on the partitioning column read only the relevant partitions, dramatically reducing scan size. ## Example Use Case[​](#example-use-case "Direct link to Example Use Case") Consider a high-volume e-trading frontend application backed by an AWS RDS database containing a table of trades. To retrieve all trades over the last 24 hours, the application would need to query the remote database and transfer the data over the network. By accelerating the trades table locally using the [AWS RDS Data Connector](https://github.com/spiceai/cookbook/tree/trunk/mysql/rds-aurora#readme), the data is brought to the application, saving round trip time and data transfer time. ## Considerations[​](#considerations "Direct link to Considerations") **Storage Capacity**: Accelerated datasets consume local storage. In-memory engines (Arrow) require sufficient RAM; file-based engines (DuckDB, SQLite, Cayenne) require sufficient disk space. As a guideline, allocate at least 1.5x the source dataset size to account for indexing and temporary refresh overhead. Check current usage by querying `runtime.metrics`. **Data Security**: Accelerating a dataset copies data from the source to the local runtime. Assess whether the data sensitivity is appropriate for the deployment environment. Secure network connections between the runtime and data source using TLS (`pg_sslmode: verify-full` for PostgreSQL, `s3_auth: iam_role` for S3). Encrypt data at rest when using file-based accelerators in production. **Refresh Latency**: The `refresh_check_interval` controls how frequently the runtime checks for new data. Shorter intervals increase load on the source database. For real-time requirements, use [Change Data Capture (CDC)](/docs/next/features/cdc) instead of polling. **Schema Changes**: Spice infers the schema for each accelerated dataset at startup and does not apply schema changes at runtime. If the source schema changes (columns added, removed, or types changed) while the runtime is running, data refreshes will fail rather than silently applying the new schema. To pick up source schema changes, restart the runtime, or use [`mode: file_update`](/docs/next/reference/spicepod/datasets#accelerationmode) which automatically detects schema changes on refresh and recreates the acceleration when incompatible changes are found. See [Schema Inference](/docs/next/components/data-connectors#schema-inference) for details. **Engine Selection**: Choose the acceleration engine based on workload characteristics: | Engine | Best For | Mode | | ---------- | -------------------------------------------------------------- | ------------------------------------------------- | | `arrow` | Read-heavy analytics, in-memory speed | `memory` | | `duckdb` | Complex analytical queries under 10 GB, file-based persistence | `memory`, `file`, `file_create`, or `file_update` | | `sqlite` | OLTP-style point lookups, concurrent reads/writes | `memory`, `file`, `file_create`, or `file_update` | | `postgres` | When a full SQL database is needed as accelerator | External | | `cayenne` | Datasets 10 GB and above, high-performance columnar | `file`, `file_create`, or `file_update` | | `turso` | Embedded libSQL, lightweight file-based caching | `memory`, `file`, `file_create`, or `file_update` | ## Example[​](#example "Direct link to Example") ### Locally Accelerating taxi\_trips[​](#locally-accelerating-taxi_trips "Direct link to Locally Accelerating taxi_trips") * Start Spice with the following dataset: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true refresh_mode: full refresh_check_interval: 10s ``` * The dataset `taxi_trips` is accelerated locally by the Spice runtime. The data refreshes every 10 seconds. * Query times can be compared against the Spice platform: ``` curl \ --url 'https://data.spiceai.io/v1/sql?api_key=[API_KEY]' \ --data 'select * from taxi_trips' ``` The locally accelerated dataset can then be queried locally: ``` spice sql select * from taxi_trips; ``` Example output: ``` +---------------+--------------+------------------+ | trip_distance | total_amount | tpep_pickup_time | +---------------+--------------+------------------+ | 1.2 | 9.80 | 2023-01-15 08:32 | | 3.4 | 18.50 | 2023-01-15 09:10 | +---------------+--------------+------------------+ Time: 0.012s. 2 rows. ``` Locally accelerated datasets provide significantly faster query times compared to remote sources. [Learn more about Data Accelerators](/docs/next/components/data-accelerators) for faster access. --- # Constraints Constraints enforce data integrity in a database. Spice supports constraints on locally accelerated tables to ensure data quality and configure behavior for data updates that violate constraints. Constraints are specified using [column references](#column-references) in the Spicepod via the `primary_key` field in the acceleration configuration. Additional unique constraints are specified via the [`indexes`](/docs/next/features/data-acceleration/indexes) field with the value `unique`. Data that violates these constraints will result in a [conflict](#handling-conflicts). If multiple rows in the incoming data violate any constraint, the entire incoming batch of data will be dropped. Example Spicepod: ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash # Define a primary key on the `hash` column indexes: '(number, timestamp)': unique # Add a unique index with a multicolumn key comprised of the `number` and `timestamp` columns ``` ## Column References[​](#column-references "Direct link to Column References") Column references can be used to specify which columns are part of the constraint. The column reference can be a single column name or a multicolumn key. The column reference must be enclosed in parentheses if it is a multicolumn key. Examples * `number`: Reference a constraint on the `number` column * `(hash, timestamp)`: Reference a constraint on the `hash` and `timestamp` columns ## Handling conflicts[​](#handling-conflicts "Direct link to Handling conflicts") The behavior of inserting data that violates the constraint can be configured via the `on_conflict` field to either `drop` the data that violates the constraint or `upsert` that data into the accelerated table (i.e. update all values other than the columns that are part of the constraint to match the incoming data). warning If there are multiple rows in the incoming data that violate any constraint, the entire incoming batch of data will be dropped. Example Spicepod: ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash # Define a primary key on the `hash` column indexes: '(number, timestamp)': unique # Add a unique index with a multicolumn key comprised of the `number` and `timestamp` columns on_conflict: # Upsert the incoming data when the primary key constraint on "hash" is violated, # alternatively "drop" can be used instead of "upsert" to drop the data update. hash: upsert ``` ### Advanced Upsert Options[​](#advanced-upsert-options "Direct link to Advanced Upsert Options") By default, even when `upsert` is configured, if there are constraint violations, such as duplicates within the same batch of ingested data, it will result in a constraint violation - as attempting to upsert data into the target acceleration engine results in an error if done in a single statement. (i.e. [PostgreSQL does not allow the same row to be proposed for insertion more than once](https://www.postgresql.org/docs/18/sql-insert.html)) Spice provides two `upsert` options to resolve duplicates within a single update: * `upsert_dedup`: Removes exact duplicates in the incoming batch if there is a constraint violation. (i.e. the equivalent of running `SELECT DISTINCT * FROM [batch]`) * `upsert_dedup_by_row_id`: Resolves conflicts by taking the row with the greatest row id. This is the behavior that would occur if the upsert were applied row-by-row. This guarantees that no constraint violations would result in an error, but it has the tradeoff of being effectively "random" if the incoming data is not ordered. The new behavior is only triggered when an incoming batch has a constraint violation, minimizing the effect of applying these computations to only when its necessary. However, they can have a performance impact and are not enabled by default. Full configuration example: ``` acceleration: enabled: true engine: duckdb mode: file primary_key: id on_conflict: id: upsert_dedup # upsert_dedup_by_row_id ``` Examples for advanced upsert behavior Take these two CSV files: `one.csv`: ``` foo,bar a,1 b,2 a,1 ``` Behavior on `one.csv` with a primary key on `foo` and `on_conflict` set to: * `upsert`: Will error with: `Constraint Violation: Incoming data violates uniqueness constraint on column(s): foo` * `upsert_dedup`: Will succeed in loading 2 rows, the `a,1` row is reduced to a single instance. * `upsert_dedup_by_row_id`: Same as `upsert_dedup` `two.csv`: ``` foo,bar a,1 b,2 a,10 ``` Behavior on `one.csv` with a primary key on `foo` and `on_conflict` set to: * `upsert`: Will error with: `Constraint Violation: Incoming data violates uniqueness constraint on column(s): foo` * `upsert_dedup`: Will error with: `Constraint Violation: Incoming data violates uniqueness constraint on column(s): foo` * `upsert_dedup_by_row_id`: Will succeed in loading 2 rows, `a,10` and `b,2`. The primary key violation is resolved to the row that occurred later. ## Limitations[​](#limitations "Direct link to Limitations") * **Single on\_conflict target supported**: Only a single `on_conflict` target can be specified, unless all `on_conflict` targets are specified with drop. * Examples for valid/invalid `on_conflict` targets The following Spicepod is invalid because it specifies multiple `on_conflict` targets with `upsert`: Invalid ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash indexes: '(number, timestamp)': unique on_conflict: hash: upsert '(number, timestamp)': upsert ``` The following Spicepod is valid because it specifies multiple `on_conflict` targets with `drop`, which is allowed: Valid ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash indexes: '(number, timestamp)': unique on_conflict: hash: drop '(number, timestamp)': drop ``` The following Spicepod is invalid because it specifies multiple `on_conflict` targets with `upsert` and `drop`: Invalid ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash indexes: '(number, timestamp)': unique on_conflict: hash: upsert '(number, timestamp)': drop ``` * **DuckDB Limitations:** * DuckDB does not support `upsert` for datasets with List or Map types. * Standard indexes unexpectedly act like unique indexes and block updates when `upsert` is configured. * Standard indexes blocking updates The following Spicepod specifies a standard index on the `number` column, which blocks updates when `upsert` is configured for the `hash` column: ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: duckdb primary_key: hash indexes: number: enabled on_conflict: hash: upsert ``` The following error is returned when attempting to upsert data into the `eth.recent_blocks` table: ``` ERROR runtime::accelerated_table::refresh: Error adding data for eth.recent_blocks: External error: Unable to insert into duckdb table: Binder Error: Can not assign to column 'number' because it has a UNIQUE/PRIMARY KEY constraint ``` This is a limitation of DuckDB. --- # Data Refresh Acceleration data can be refreshed (updated) by: * **API**: POST to `/v1/datasets/:name/acceleration/refresh`. See [Refresh Dataset HTTP API](/docs/next/api/HTTP/post-dataset-refresh). * **Interval**: Time-based refresh interval. See [Refresh Interval](#refresh-interval). * **Change Data Capture (CDC)**: CDC from a database using Debezium. See [Change Data Capture](/docs/next/features/cdc). * **Push**: Spice-to-Spice Push over Apache Arrow DoExchange. ![Spice.ai Open Source Acceleration Refresh](/assets/images/acceleration-refresh-28183446549fbcc17f21f07881cfe7d9.png). ## Refresh Modes[​](#refresh-modes "Direct link to Refresh Modes") Spice supports five modes to refresh/update local data from a connected data source. `full` is the default mode. | Mode | Description | Example | | ---------- | ---------------------------------------------------- | ---------------------------------------------------------------- | | `full` | Replace/overwrite the entire dataset on each refresh | A table of users | | `append` | Append/add data to the dataset on each refresh | Append-only, immutable datasets, such as time-series or log data | | `changes` | Apply incremental changes | Customer order lifecycle table | | `caching` | Read-through caching for SQL queries | API search results or dynamic content endpoints | | `snapshot` | Reload exclusively from the snapshot store | Read-only replicas bootstrapped from centralized snapshots | Learn more about each mode: * [Full Mode](/docs/next/features/data-acceleration/refresh-modes/full) * [Append Mode](/docs/next/features/data-acceleration/refresh-modes/append) * [Changes Mode](/docs/next/features/data-acceleration/refresh-modes/changes) * [Caching Mode](/docs/next/features/data-acceleration/refresh-modes/caching) * [Snapshot Mode](/docs/next/features/data-acceleration/refresh-modes/snapshot) Example: ``` datasets: - from: databricks:my_dataset name: accelerated_dataset acceleration: refresh_mode: full refresh_check_interval: 10m ``` ### Append[​](#append "Direct link to Append") Using `refresh_mode: append` requires the use of a [`time_column` dataset parameter](/docs/next/reference/spicepod/datasets#time_column), specifying a column to compare the local acceleration against the remote source. Data will be incrementally refreshed where the `time_column` value in the remote source is greater-than (gt) the `max(time_column)` value in the local acceleration. A date-typed `time_column` (`time_format: date`) is compared greater-than-or-equal (gte) against the start of that day instead, since every row of a day shares one value — see [Day-Granular Time Columns](/docs/next/features/data-acceleration/refresh-modes/append#day-granular-time-columns). E.g. ``` datasets: - from: databricks:my_dataset name: accelerated_dataset time_column: created_at acceleration: refresh_mode: append refresh_check_interval: 10m ``` Readiness with snapshots Append-mode accelerations that define a `time_column` wait to report ready until the first append refresh completes after [snapshot bootstrap](/docs/next/features/data-acceleration/snapshots). This keeps the dataset out of rotation until the freshest data is available while still benefiting from the snapshot-assisted startup. If late arriving data or clock-skew needs to be accounted for, an optional overlap can also be specified. See [`acceleration.refresh_append_overlap`](/docs/next/reference/spicepod/datasets#accelerationrefresh_append_overlap). #### `time_partition_column`[​](#time_partition_column "Direct link to time_partition_column") Datasets that are partitioned by a less-granular time-column (e.g. day, month, year) can also use the `time_partition_column` parameter in addition to the `time_column` parameter to specify the time-column to use for efficient partition pruning. Example: ``` datasets: - from: databricks:my_dataset name: accelerated_dataset time_column: created_at time_format: iso8601 time_partition_column: created_at_day time_partition_format: date ``` #### Append only modified files[​](#append-only-modified-files "Direct link to Append only modified files") Spice can automatically detect and append only newly created or updated files from object-store data sources. This is useful for append-only datasets where only new files are added to the source and existing files are not modified or deleted. Enable this feature by setting either `time_column` or `time_partition_column` to the special value `last_modified`. When configured this way with `refresh_mode: append`, Spice will use the file/object's metadata to determine which files are new or have been updated. This approach can drastically speed up incremental updates for large datasets, as Spice only needs to process the new files rather than scanning the entire dataset for changes to a column. This optimization is particularly valuable for datasets with many files or large file sizes. If `last_modified` already exists as a column in the parquet data, that column will take precedence over the metadata value from the file itself. Example using `time_column`: ``` datasets: - from: s3://my_bucket/my_dataset name: accelerated_dataset time_column: last_modified params: file_format: parquet acceleration: refresh_mode: append refresh_check_interval: 10m ``` Example using `time_partition_column`: ``` datasets: - from: s3://my_bucket/my_dataset name: accelerated_dataset time_column: created_at time_partition_column: last_modified params: file_format: parquet acceleration: refresh_mode: append refresh_check_interval: 10m ``` info Appending modified files is only supported for datasets that support setting the [file format parameter](/docs/next/reference/file_format), such as `s3://`, `abfs://`, `file://`, etc. ### Changes (CDC)[​](#changes-cdc "Direct link to Changes (CDC)") Datasets configured with acceleration `refresh_mode: changes` requires a [Change Data Capture (CDC)](/docs/next/features/cdc) supported data connector. Initial CDC support in Spice is supported by the [Debezium data connector](/docs/next/components/data-connectors/debezium). ### Caching[​](#caching "Direct link to Caching") The `caching` refresh mode is designed for HTTP-based datasets where request metadata acts as cache keys. This mode is particularly useful for API responses that return multiple rows for a single request, such as search results or dynamic content endpoints. See [Caching Mode](/docs/next/features/data-acceleration/refresh-modes/caching) for detailed documentation and examples. ### Snapshot[​](#snapshot "Direct link to Snapshot") The `snapshot` refresh mode creates a read-only acceleration that reloads exclusively from the [snapshot store](/docs/next/features/data-acceleration/snapshots). The federated data source is never queried for refreshes — instead, the runtime polls the snapshot store on a configurable interval and atomically swaps in newer snapshots when available. ``` snapshots: enabled: true location: s3://my-bucket/snapshots/ params: s3_auth: iam_role datasets: - from: postgres:public.my_table name: my_table acceleration: enabled: true engine: duckdb mode: file refresh_mode: snapshot refresh_check_interval: 30s # Poll interval; defaults to 1m snapshots: enabled params: duckdb_file: /nvme/my_table.db ``` **Requirements:** * `acceleration.snapshots` must be `enabled` or `bootstrap_only` * The acceleration engine must be a snapshot-capable file-based engine: **DuckDB**, **SQLite**, **Cayenne**, or **Turso** **Behavior:** * On startup, the runtime bootstraps from the most recent snapshot (same as other snapshot-enabled modes) * After bootstrap, the runtime polls the snapshot store at `refresh_check_interval` (default: 60 seconds) for newer snapshots * When a newer snapshot is found, its schema is validated against the current acceleration schema before downloading * The accelerator file is swapped atomically — queries continue to be served from the previous snapshot until the swap completes * `INSERT INTO` statements are rejected with an error since the acceleration is driven exclusively from snapshots tip Use `refresh_mode: snapshot` for read-only replicas that don't need direct access to the federated source — for example, edge nodes that receive snapshots from a centralized writer. ## Ready State[​](#ready-state "Direct link to Ready State") | | | | --------------------------- | --------- | | Supported in `refresh_mode` | Any | | Required | No | | Default Value | `on_load` | By default, Spice will return an error for queries against an accelerated dataset that is still loading its initial data. The endpoint [`/v1/ready`](/docs/next/api/HTTP/ready) is used in production deployments to control when queries are sent to the Spice runtime. The ready state for an accelerated dataset can be configured using the [`ready_state`](/docs/next/reference/spicepod/datasets#ready_state) parameter in the dataset configuration. * `ready_state: on_load`: Default. The dataset is considered ready after the initial load of the accelerated data. For file-based accelerated datasets that have existing data, this will be ready immediately. Queries against this dataset before the data is loaded will return an error. * `ready_state: on_registration`: The dataset is considered ready when the dataset is registered in Spice, even before the initial data is loaded. Queries against this dataset before the data is loaded will automatically fallback to the federated source. Once the data is loaded, queries will be served from the acceleration. * `ready_state: on_schema_resolved`: The dataset is considered ready once the federated source's schema has been resolved (which also verifies access to the source), without waiting for the initial data refresh. Queries fall back to the federated source until the initial load completes; subsequent refresh failures are still reported via dataset status and metrics. Example: ``` datasets: - from: s3://my_bucket/my_dataset name: my_dataset ready_state: on_load # or on_registration, on_schema_resolved acceleration: enabled: true ``` ## Fast Cold Starts with Snapshots[​](#fast-cold-starts-with-snapshots "Direct link to Fast Cold Starts with Snapshots") File-based acceleration engines (DuckDB, SQLite, Cayenne, or Turso) can rely on [acceleration snapshots](/docs/next/features/data-acceleration/snapshots) to download a pre-built database file on startup instead of waiting for the first refresh to finish. Configure a shared snapshot location under the top-level `snapshots` block and opt individual datasets in with `acceleration.snapshots: enabled`, `bootstrap_only`, or `create_only`. Snapshots are stored using Hive-style partitions (`month=YYYY-MM/day=YYYY-MM-DD/dataset=`) and are only supported when each dataset writes to its own acceleration file. ## Filtered Refresh[​](#filtered-refresh "Direct link to Filtered Refresh") Typically only a working subset of an entire dataset is used in an application or dashboard. Use these features to filter refresh data, creating a smaller subset for faster processing and to reduce the data transferred and stored locally. * [Refresh SQL](#refresh-sql) - Specify the filter as arbitrary SQL to be pushed down to the remote source. * [Refresh Data Window](#refresh-data-window) - Filters data from the remote source outside the specified time window. ### Refresh SQL[​](#refresh-sql "Direct link to Refresh SQL") | | | | --------------------------- | ----- | | Supported in `refresh_mode` | Any | | Required | No | | Default Value | Unset | Refresh SQL supports specifying filters for data accelerated from the connected source using arbitrary SQL. Filters will be pushed down to the remote source when possible, so only the requested data will be transferred over the network. Example: ``` datasets: - from: databricks:my_dataset name: accelerated_dataset acceleration: enabled: true refresh_mode: full refresh_check_interval: 10m refresh_sql: | SELECT * FROM accelerated_dataset WHERE city = 'Seattle' ``` The `refresh_sql` parameter can be updated at runtime on-demand using `PATCH /v1/datasets/:name/acceleration`. This change is temporary and will revert to the `spicepod.yml` definition at the next runtime restart. Columns can be selected in the query via the `SELECT` clause, but only column names are supported. Arbitrary expressions or aliases are not supported. Example: ``` curl -i -X PATCH \ -H "Content-Type: application/json" \ -d '{ "refresh_sql": "SELECT city, state FROM accelerated_dataset WHERE city = 'Bellevue'" }' \ 127.0.0.1:8090/v1/datasets/accelerated_dataset/acceleration ``` Queries that return zero results will fallback to the behavior specified by the [`on_zero_results` parameter](#behavior-on-zero-results), and will not have the `refresh_sql` applied to the results from the fallback. The `refresh_sql` only applies to acceleration refresh tasks. For the complete reference, view the `refresh_sql` section of [datasets](/docs/next/reference/spicepod/datasets#accelerationrefresh_sql). Limitations * When `refresh_mode: changes` is specified, Refresh SQL can only modify the selected columns and cannot apply filters. * Running queries while using refresh SQL will not fallback to the source if any query returns more than zero rows, even when querying against columns that are not explicitly filtered by the refresh SQL. This may result in queries returning partial data, depending on the filters applied in the refresh SQL. * Refresh SQL only supports filtering data from the current dataset - joining across other datasets is not supported. * Refresh SQL modifications made via API are temporary and will revert after a runtime restart. ### Refresh Data Window[​](#refresh-data-window "Direct link to Refresh Data Window") | | | | --------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default Value | Unset | The `refresh_data_window` parameter supports refreshing data that falls within the specified time window. The `refresh_data_window` is applied cumulatively to any filters specified by the [`refresh_sql`](#refresh-sql), and applies a time filter based on `now() - refresh_data_window`. For example, the following configuration: ``` time_column: column_time acceleration: refresh_sql: "SELECT * FROM my_dataset WHERE column_one = 'value'" refresh_data_window: 1d ``` In this example, `refresh_data_window` is converted into an effective Refresh SQL of `SELECT * FROM my_dataset WHERE column_one = 'value' AND column_time > (now() - interval '1' day)`. The `time_column` column can be specified in the `refresh_sql` in conjunction with the `refresh_data_window`, and both filters are combined with `AND`. This parameter relies on the `time_column` dataset parameter specifying a column that is a timestamp type. Optionally, the `time_format` can be specified to instruct the Spice runtime on how to interpret timestamps in the `time_column`. *Example with `refresh_sql`:* ``` datasets: - from: databricks:my_dataset name: accelerated_dataset time_column: created_at acceleration: enabled: true refresh_mode: full refresh_check_interval: 10m refresh_sql: | SELECT * FROM accelerated_dataset WHERE city = 'Seattle' refresh_data_window: 1d ``` This example will only accelerate data from the federated source that matches the filter `city = 'Seattle'` and is less than 1 day old. *Example with `on_zero_results`:* ``` datasets: - from: databricks:my_dataset name: accelerated_dataset time_column: created_at acceleration: enabled: true refresh_mode: full refresh_check_interval: 10m refresh_sql: | SELECT * FROM accelerated_dataset WHERE city = 'Seattle' refresh_data_window: 1d on_zero_results: use_source ``` This example will only accelerate data from the federated source that matches the filter `city = 'Seattle'` and is less than 1 day old. If a query against the accelerated data returns zero results, the query will fallback to the source and return the direct results without any filtering. If a query against the accelerated data returns some results, the query will not fall back. For example, attempting to query for the last 2 days of data would only return the last 1 day of data without falling back. ## Behavior on Zero Results[​](#behavior-on-zero-results "Direct link to Behavior on Zero Results") | | | | --------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default Value | `return_empty` | warning `on_zero_results` is ignored when `refresh_mode: caching` is set. Caching mode always queries the source on a cache miss, regardless of this setting. Remove `on_zero_results` from caching-mode dataset configurations to silence the runtime warning. By default, accelerated datasets only return locally materialized data. If this local data is a subset of the full dataset in the federated source—due to settings like `refresh_sql`, `refresh_data_window`, or retention policies—queries against the accelerated dataset may return zero results, even when the federated table would return results. To address this, `on_zero_results: use_source` can be configured in the acceleration configuration. Queries returning zero results will fall back to the federated source, returning results from querying the underlying data. `on_zero_results`: * `return_empty` (Default) - Return an empty result set when no data is found in the accelerated dataset. * `use_source` - Fall back to querying the federated table when no data is found in the accelerated dataset. Example: ``` datasets: - from: databricks:my_dataset name: accelerated_dataset acceleration: enabled: true refresh_sql: SELECT * FROM accelerated_dataset where city = 'Seattle' on_zero_results: use_source ``` In this example a query against `accelerated_dataset` within Spice like `SELECT * FROM accelerated_dataset WHERE city = 'Portland'` would initially query against the accelerated data, see that it returns zero results and then fallback to querying against the federated table in Databricks. warning * It is possible that even though an accelerated table returns some results, it may not contain all the data that would be returned by the federated table. `on_zero_results` only controls the behavior in the simple case where no data is returned by the acceleration for a given query. ## Refresh on Startup[​](#refresh-on-startup "Direct link to Refresh on Startup") | Parameter | Value | | --------------------------- | ------ | | Supported in `refresh_mode` | Any | | Required | No | | Default Value | `auto` | Controls the refresh behavior of an accelerated dataset across restarts. `refresh_on_startup` Options: * `auto` (Default) – Maintains refresh state across restarts: * With `refresh_check_interval`: Schedules next refresh based on last successful refresh time, triggering immediately if interval has already elapsed * Without `refresh_check_interval`: No refresh (on-demand only) * `always` – Forces a dataset refresh on every startup, regardless of the existing acceleration state. Setting `refresh_on_startup: always` ensures that accelerated data is always refreshed to match the source when the service restarts. This is useful in **development environments** or when **data consistency is critical** after deployment. Example Configuration: ``` datasets: - from: databricks:my_dataset name: accelerated_dataset acceleration: enabled: true refresh_on_startup: always ``` For the complete reference, view the `refresh_on_startup` section of [datasets](/docs/next/reference/spicepod/datasets#accelerationrefresh_on_startup). ## Refresh Interval[​](#refresh-interval "Direct link to Refresh Interval") | | | | --------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default Value | Unset | The [`refresh_check_interval`](/docs/next/reference/spicepod/datasets#accelerationrefresh_check_interval) parameter controls how often the accelerated dataset is refreshed. Example: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true refresh_mode: full refresh_check_interval: 10s ``` This configuration will refresh `taxi_trips` data every 10 seconds. ## Refresh On-Demand[​](#refresh-on-demand "Direct link to Refresh On-Demand") info Supported for accelerators with a `refresh_mode` of `full` or `append`. Accelerated datasets can be refreshed on-demand via the `refresh` CLI command or `POST /v1/datasets/:name/acceleration/refresh` API endpoint. CLI example: ``` spice refresh eth_recent_blocks ``` API example using cURL: ``` curl -i -XPOST 127.0.0.1:8090/v1/datasets/eth_recent_blocks/acceleration/refresh ``` with response: ``` HTTP/1.1 201 Created content-type: application/json content-length: 55 date: Thu, 11 Apr 2024 20:11:18 GMT {"message":"Dataset refresh triggered for eth_recent_blocks."} ``` Note On-demand refresh always initiates a new refresh, terminating any in-progress refresh for the dataset. ## Refresh Schedules[​](#refresh-schedules "Direct link to Refresh Schedules") | | | | --------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default Value | Unset | The [`refresh_cron`](/docs/next/reference/spicepod/datasets#accelerationrefresh_cron) parameter supports specifying a cron schedule which controls when datasets refresh. Example: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true refresh_mode: full refresh_cron: '0 12 * * 1-5' ``` This configuration will refresh `taxi_trips` data at midday every weekday. For more information about cron schedules, see the [cron schedule reference](/docs/next/reference/cron). The `refresh_cron` parameter cannot be specified in conjunction with a `refresh_check_interval` parameter. ## Refresh Retries[​](#refresh-retries "Direct link to Refresh Retries") | | | | ------------------------------------ | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default `refresh_retry_enabled` | `true` | | Default `refresh_retry_max_attempts` | Unset | By default, data refreshes for accelerated datasets are retried on transient errors (connectivity issues, compute warehouse goes idle, etc.) using a [Fibonacci](https://en.wikipedia.org/wiki/Fibonacci_sequence) backoff strategy. Retry behavior can be configured using the [`acceleration.refresh_retry_enabled`](/docs/next/reference/spicepod/datasets#accelerationrefresh_retry_enabled) and [`acceleration.refresh_retry_max_attempts`](/docs/next/reference/spicepod/datasets#accelerationrefresh_retry_max_attempts) parameters. Example: Disable retries ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: refresh_retry_enabled: false refresh_check_interval: 30s ``` Example: Limit retries to a maximum of 10 attempts ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: refresh_retry_max_attempts: 10 refresh_check_interval: 30s ``` ## Retention Policy[​](#retention-policy "Direct link to Retention Policy") | | | | ---------------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default `retention_check_enabled` | `false` | | Default `retention_period` | Unset | | Default `retention_sql` | Unset | | Default `retention_check_interval` | Unset | Accelerated datasets can be configured to automatically evict data using two different retention strategies: ### Time-based Retention[​](#time-based-retention "Direct link to Time-based Retention") Automatically evict time-series data exceeding a retention period by setting a retention policy based on the configured `time_column` and `acceleration.retention_period`. The policy is set using the [`acceleration.retention_check_enabled`](/docs/next/reference/spicepod/datasets#accelerationretention_check_enabled), [`acceleration.retention_period`](/docs/next/reference/spicepod/datasets#accelerationretention_period) and [`acceleration.retention_check_interval`](/docs/next/reference/spicepod/datasets#accelerationretention_check_interval) parameters, along with the [`time_column`](/docs/next/reference/spicepod/datasets#time_column) and [`time_format`](/docs/next/reference/spicepod/datasets#time_format) dataset parameters. When `retention_check_enabled` is set to `true`, `retention_check_interval` and `retention_period` are required parameters. Example: ``` datasets: - from: mysql:user_events name: user_events time_column: created_at acceleration: enabled: true refresh_mode: append retention_check_enabled: true retention_period: 30d retention_check_interval: 1h ``` ### Custom SQL-based Retention[​](#custom-sql-based-retention "Direct link to Custom SQL-based Retention") Evict data from an acceleration based on custom filter predicates using the [`acceleration.retention_sql`](/docs/next/reference/spicepod/datasets#accelerationretention_sql) parameter. This is useful for scenarios like soft-deleting rows in append datasets or removing data based on complex business logic. The `retention_sql` parameter takes the form of a `DELETE FROM
WHERE ` statement. Example - Soft delete retention: ``` datasets: - from: mysql:user_events name: user_events time_column: created_at acceleration: enabled: true refresh_mode: append primary_key: user_id on_conflict: user_id: upsert retention_check_enabled: true retention_check_interval: 5m retention_sql: DELETE FROM user_events WHERE status = 'archived' ``` note * Time-based retention (`retention_period`) and custom SQL retention (`retention_sql`) can be used independently or together. When both are configured, both retention policies will be applied during each retention check. ### End-to-End Incremental Ingestion Example[​](#end-to-end-incremental-ingestion-example "Direct link to End-to-End Incremental Ingestion Example") The following example combines the pieces above into a single configuration for keeping an accelerated dataset incrementally up-to-date from a source that supports soft deletes: * `refresh_mode: append` with a `time_column` for incremental queries * `refresh_check_interval` to poll for new/changed rows * `refresh_append_overlap` to tolerate clock skew and late-arriving rows without missing data * `primary_key` + `on_conflict: upsert` so rows updated in the source overwrite the accelerated copy instead of duplicating * `retention_period` to bound the working set by time * `retention_sql` to evict soft-deleted rows (`deleted_at IS NOT NULL`) ``` datasets: - from: postgres:public.orders name: orders time_column: updated_at acceleration: enabled: true engine: duckdb refresh_mode: append refresh_check_interval: 1m refresh_append_overlap: 5m primary_key: id on_conflict: id: upsert retention_check_enabled: true retention_check_interval: 10m retention_period: 90d retention_sql: DELETE FROM orders WHERE deleted_at IS NOT NULL ``` With this configuration Spice bootstraps from the source, then every minute fetches rows where `updated_at > max(updated_at) - 5m`, upserting on `id`. Rows older than 90 days — or rows the source has soft-deleted — are evicted on the retention check. ## Refresh Jitter[​](#refresh-jitter "Direct link to Refresh Jitter") | | | | -------------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default `refresh_jitter_enabled` | `false` | | Default `refresh_jitter_max` | Unset | Accelerated datasets can include a random jitter in their refresh interval to prevent the [Thundering herd problem](https://en.wikipedia.org/wiki/Thundering_herd_problem), where multiple datasets refresh simultaneously. The jitter is a random value between 0 and `refresh_jitter_max`, which is added to or subtracted from the base `refresh_check_interval`. If `refresh_jitter_max` is not specified, it defaults to 10% of `refresh_check_interval`. Refresh Jitter applies to the initial dataset load. If multiple similarly configured Spice instances are restarted at the same time, they will load with a jitter between 0 and `refresh_jitter_max`. Example: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: refresh_check_interval: 10s refresh_jitter_enabled: true refresh_jitter_max: 1s ``` In the configuration above: 1. The initial load will include a random delay between **0** and **1 second**. 2. Subsequent refresh intervals will vary randomly between **9 seconds** and **11 seconds**. Refresh jitter configuration: * [`refresh_jitter_enabled`](/docs/next/reference/spicepod/datasets#accelerationrefresh_jitter_enabled) * [`refresh_jitter_max`](/docs/next/reference/spicepod/datasets#accelerationrefresh_jitter_max) ## Configuration Examples[​](#configuration-examples "Direct link to Configuration Examples") ### Accelerating a full set of data that sometimes changes[​](#accelerating-a-full-set-of-data-that-sometimes-changes "Direct link to Accelerating a full set of data that sometimes changes") In this example, Spice connects with a dataset that changes infrequently and is not configured for CDC. For example, a list of product categories. ``` datasets: - from: mysql:product_categories name: product_categories acceleration: refresh_mode: full refresh_check_interval: 8h ``` In this scenario, Spice uses a simple acceleration configuration - full refreshes on an 8 hour schedule. No additional behaviors are enabled, so queries matching for new product codes will return no results until the next refresh cycle. ### Accelerating a subset of data that changes frequently[​](#accelerating-a-subset-of-data-that-changes-frequently "Direct link to Accelerating a subset of data that changes frequently") In this example, Spice connects with a dataset that has frequently changing data that is not configured for CDC. For example, user's posts on a social media platform. ``` datasets: - from: mysql:posts name: posts acceleration: refresh_mode: full refresh_check_interval: 10m refresh_sql: "SELECT * FROM posts WHERE updated_at > now() - interval '1' day" on_zero_results: use_source ``` With this configuration, Spice will refresh every 10 minutes accelerating posts that have been updated in the last day. When querying for posts by direct ID, if a post is not accelerated Spice will fallback to retrieving the post from the non-accelerated source due to the behavior of `on_zero_results: use_source`. However, if querying for a range of posts that includes some which have updated in the last day Spice will only return those results without falling back to the source. This could result in queries for a range of posts excluding posts that exist in the non-accelerated source because they have been filtered out due to their `updated_at` value. ### Accelerating application logs[​](#accelerating-application-logs "Direct link to Accelerating application logs") In this example, Spice connects to a data source that is immutable, receives new rows, and is not configured for CDC. For example, a database that contains some application logs. ``` datasets: - from: duckdb:logs name: logs time_column: created_at params: duckdb_open: logs.duckdb acceleration: refresh_mode: append refresh_check_interval: 10m refresh_sql: "SELECT * FROM logs WHERE asset = 'asset_id'" refresh_data_window: 1d on_zero_results: use_source retention_check_enabled: true retention_period: 7d retention_check_interval: 10m ``` This acceleration configuration applies a number of different behaviors: 1. A `refresh_data_window` was specified. When Spice starts, it will apply this `refresh_data_window` to the `refresh_sql`, and retrieve only the last day's worth of logs with an `asset = 'asset_id'`. 2. Because a `refresh_sql` is specified, every refresh (including initial load) will have the filter applied to the refresh query. 3. 10 minutes after loading, as specified by the `refresh_check_interval`, the first refresh will occur - retrieving new rows where `asset = 'asset_id'`. 4. Running a query to retrieve logs with an `asset` that is *not* `asset_id` will fall back to the source, because of the `on_zero_results: use_source` parameter. 5. Running a query to retrieve a log longer than 1 day ago will fall back to the source, because of the `on_zero_results: use_source` parameter. 6. Running a query to retrieve logs within a range of now to longer than 1 day ago will only return logs from the last day. This is due to the `refresh_data_window` only accelerating the last day's worth of logs, which will return some results. Because results are returned, Spice will not fall back to the source even though `on_zero_results: use_source` is specified. 7. Spice will retain newly appended log rows for 7 days before discarding them, as specified by the `retention_*` parameters. ## Cookbook[​](#cookbook "Direct link to Cookbook") * Configure accelerated dataset retention policy. [Accelerated Dataset Retention Policy](https://github.com/spiceai/cookbook/tree/trunk/retention#readme) * Dynamically refresh specific data at runtime by programmatically updating refresh\_sql and triggering data refreshes. [Advanced Data Refresh](https://github.com/spiceai/cookbook/tree/trunk/acceleration/data-refresh#readme) * Configure `refresh_data_window` to filter refreshed data to recent data [Refresh Data Window](https://github.com/spiceai/cookbook/tree/trunk/refresh-data-window#readme) --- # Hash Index for Arrow Acceleration Experimental Hash index is an experimental feature available in Spice v1.11.0-rc.2 and later. The hash index is an optional, high-performance indexing feature for Arrow-accelerated datasets. It provides O(1) point lookups on primary key and secondary index columns, dramatically improving query performance for equality predicates. ## Key Features[​](#key-features "Direct link to Key Features") * **O(1) Point Lookups**: Direct row access via primary key or secondary indexes without full table scans * **Secondary Indexes**: Optional indexes on non-primary-key columns for fast lookups * **256-Shard Design**: Minimizes lock contention for concurrent reads * **SIMD-Optimized Hashing**: Uses XXH3\_64 for fast, high-quality hashing * **Built-in Bloom Filter**: Fast negative lookups to skip unnecessary hash table probes ## Configuration[​](#configuration "Direct link to Configuration") Hash indexing activates automatically on Arrow-accelerated datasets when a `primary_key` or [secondary index](/docs/next/features/data-acceleration/indexes) is configured. No additional parameter is required. ``` datasets: - from: s3://bucket/orders.parquet name: orders acceleration: engine: arrow primary_key: order_id ``` The hash index activates whenever: * `engine` is `arrow` or `partitioned_arrow`, * `acceleration.enabled` is `true`, * and either `indexes` is set, or `primary_key` is set with a non-`caching` `refresh_mode`. ### Secondary Indexes[​](#secondary-indexes "Direct link to Secondary Indexes") Secondary indexes can be added on non-primary-key columns to accelerate equality lookups on those columns. Define them using the [`indexes`](/docs/next/features/data-acceleration/indexes) field in the acceleration configuration: ``` datasets: - from: s3://bucket/users.parquet name: users acceleration: engine: arrow primary_key: user_id indexes: email: unique status: enabled '(region, category)': unique ``` Index types: * **`unique`** — Enforces uniqueness and enables O(1) indexed lookups. * **`enabled`** — Permits duplicates. The index is built and maintained but does not currently accelerate queries (queries fall back to a full scan). Compound secondary indexes can be defined with a multicolumn key in parentheses, e.g. `'(col1, col2)': unique`, but are not yet used for query optimization. note Only single-column `unique` secondary indexes currently accelerate queries. Non-unique and compound secondary indexes are maintained for future use. ### Configuration Options[​](#configuration-options "Direct link to Configuration Options") | Parameter | Type | Required | Default | Description | | ------------- | -------------- | ----------------------------- | ------- | -------------------------------------------------------------------------------- | | `primary_key` | string or list | Yes (unless `indexes` is set) | None | Column(s) for the primary key index | | `indexes` | YAML map | No | None | Secondary indexes (see [indexes](/docs/next/features/data-acceleration/indexes)) | `hash_index` parameter is ignored The legacy `hash_index: enabled` parameter is accepted but no longer activates indexing on its own. When set, the runtime logs a warning and falls back to the automatic rules above. Remove `hash_index` from `params` to clear the warning. ## Supported Data Types[​](#supported-data-types "Direct link to Supported Data Types") The hash index supports the following primary key column types: ### Primitive Types[​](#primitive-types "Direct link to Primitive Types") * `Int8`, `Int16`, `Int32`, `Int64` * `UInt8`, `UInt16`, `UInt32`, `UInt64` ### String Types[​](#string-types "Direct link to String Types") * `Utf8`, `LargeUtf8` ### Binary Types[​](#binary-types "Direct link to Binary Types") * `Binary`, `LargeBinary` ## Query Optimization[​](#query-optimization "Direct link to Query Optimization") The hash index automatically accelerates queries with equality predicates on indexed columns. ### Optimized Queries[​](#optimized-queries "Direct link to Optimized Queries") ``` -- Primary key lookup (uses primary key index) SELECT * FROM my_dataset WHERE id = 123; -- Multiple key lookups (uses primary key index for each key) SELECT * FROM my_dataset WHERE id IN (1, 2, 3); -- Secondary index lookup (uses unique secondary index) SELECT * FROM my_dataset WHERE email = 'user@example.com'; -- Primary key lookup with additional filter (index + post-filter) SELECT * FROM my_dataset WHERE id = 123 AND status = 'active'; ``` When a primary key lookup is combined with additional filters (e.g. `WHERE id = 123 AND status = 'active'`), the index is used for the primary key lookup and the remaining filters are applied afterward by DataFusion. ### Non-Optimized Queries[​](#non-optimized-queries "Direct link to Non-Optimized Queries") ``` -- Range queries (full scan) SELECT * FROM my_dataset WHERE id > 100 AND id < 200; -- Pattern matching (full scan) SELECT * FROM my_dataset WHERE id LIKE 'A%'; -- Composite primary keys (full scan, not yet supported) SELECT * FROM my_dataset WHERE region = 'US' AND customer_id = 42; -- Non-unique secondary index (full scan, not yet optimized) SELECT * FROM my_dataset WHERE status = 'active'; ``` ## Index Size[​](#index-size "Direct link to Index Size") There is no minimum row count. Whenever the [activation rules](#configuration) are met, the index is built for the dataset regardless of size — including for an empty table, which is indexed at load and rebuilt as rows arrive. The row count is used only to pre-size the hash table and bloom filter, not to decide whether to index at all, so budget the [memory](#memory-usage) below for every indexed dataset. ## Performance[​](#performance "Direct link to Performance") ### Bloom Filter Performance[​](#bloom-filter-performance "Direct link to Bloom Filter Performance") The built-in bloom filter provides: * \~0.82% false positive rate (10 bits/item, 7 hash functions) * O(1) negative lookup confirmation * Reduced unnecessary hash table probes for non-existent keys ## Memory Usage[​](#memory-usage "Direct link to Memory Usage") | Component | Memory per Entry | | ------------ | ---------------------------------------- | | Hash slot | 16 bytes (8-byte hash + 8-byte location) | | Bloom filter | \~1.25 bytes | | **Total** | \~17.25 bytes per indexed row | ### Estimating Memory[​](#estimating-memory "Direct link to Estimating Memory") For a 10 million row dataset: ``` Index memory ≈ 10M × 17.25 bytes ≈ 165 MB ``` ## Architecture[​](#architecture "Direct link to Architecture") ### Sharded Hash Table[​](#sharded-hash-table "Direct link to Sharded Hash Table") The index uses 256 independent shards to minimize lock contention: ``` ┌────────────────────────────────────────────────┐ │ HashIndex │ ├────────────────────────────────────────────────┤ │ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ │ │ Shard 0 │ │ Shard 1 │ ... │Shard 255│ │ │ │ RwLock │ │ RwLock │ │ RwLock │ │ │ └────┬────┘ └────┬────┘ └────┬────┘ │ │ │ │ │ │ │ ▼ ▼ ▼ │ │ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ │ │ Hash │ │ Hash │ ... │ Hash │ │ │ │ Table │ │ Table │ │ Table │ │ │ └─────────┘ └─────────┘ └─────────┘ │ ├────────────────────────────────────────────────┤ │ ┌─────────────────────────────────────────┐ │ │ │ Optional Bloom Filter │ │ │ │ (Fast Negative Lookups) │ │ │ └─────────────────────────────────────────┘ │ └────────────────────────────────────────────────┘ ``` **Shard Selection**: Uses XOR-folded hash bits: `((hash >> 56) ^ (hash >> 48) ^ hash) & 0xFF` ### Row Location[​](#row-location "Direct link to Row Location") Each indexed key maps to a `RowLocation`: ``` RowLocation { partition: u32, // Partition index batch: u32, // Batch index within partition row: u32, // Row index within batch } ``` ### Hash Function[​](#hash-function "Direct link to Hash Function") Uses XXH3\_64 with a fixed seed (`0x5370_6963_6541_4920` = "SpiceAI ") for: * Deterministic hashing across instances * High-quality distribution (passes SMHasher) * SIMD acceleration on arm64/amd64 ## Limitations[​](#limitations "Direct link to Limitations") 1. **Arrow Engines Only**: Hash index is only available for `engine: arrow` and `engine: partitioned_arrow` acceleration 2. **Single-Column Primary Keys Only**: Composite primary keys are not yet supported for indexed lookups; only single-column primary keys use the index 3. **Experimental**: API and behavior may change in future releases 4. **No Persistence**: Index is rebuilt on restart (data persists, index is in-memory) 5. **Duplicate Keys**: Primary key columns must have unique values 6. **Secondary Index Limitations**: Only single-column `unique` secondary indexes accelerate queries. Non-unique and compound secondary indexes are built and maintained but do not yet optimize queries ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") ### "No index available for point lookup"[​](#no-index-available-for-point-lookup "Direct link to \"No index available for point lookup\"") **Cause**: A direct point lookup was issued against an indexed table that has no primary key index — `primary_key` is not set, so only the configured secondary indexes exist. **Solution**: Set `primary_key` on the acceleration to build the primary key index. Dataset size is not a factor: the index is built at any row count once the [activation rules](#configuration) are met. ### Warning: "The hash\_index acceleration parameter is ignored for Arrow acceleration"[​](#warning-the-hash_index-acceleration-parameter-is-ignored-for-arrow-acceleration "Direct link to Warning: \"The hash_index acceleration parameter is ignored for Arrow acceleration\"") **Cause**: `hash_index: enabled` is set in `params` but no longer activates indexing on its own. **Solution**: Remove `hash_index` from `params`. Hash indexing activates automatically when `primary_key` or `indexes` is configured on an Arrow-accelerated dataset (see [Configuration](#configuration)). ### Hash index not active despite `primary_key` being set[​](#hash-index-not-active-despite-primary_key-being-set "Direct link to hash-index-not-active-despite-primary_key-being-set") **Cause**: `refresh_mode: caching` disables hash indexing even when `primary_key` is set; the caching path uses its own lookup strategy. **Solution**: Use a non-caching `refresh_mode` (e.g. `full`, `append`, `changes`) for datasets that need point-lookup acceleration via the hash index. ### High Memory Usage[​](#high-memory-usage "Direct link to High Memory Usage") **Cause**: Index consumes \~17 bytes per row. **Solution**: * Remove `primary_key` for datasets where point lookups are rare (hash indexing stops being applied) * Consider using a different acceleration engine for very large datasets --- # Indexes Database indexes are essential for optimizing query performance. This document explains how to add indexes to tables created by Spice for local data acceleration. Example Spicepod: ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite indexes: number: enabled # Index the `number` column '(hash, timestamp)': unique # Add a unique index with a multicolumn key comprised of the `hash` and `timestamp` columns ``` ## Column References[​](#column-references "Direct link to Column References") Column references can be used to specify which columns to index. The column reference can be a single column name or a multicolumn key. The column reference must be enclosed in parentheses if it is a multicolumn key. Examples * `number`: Index the `number` column * `(hash, timestamp)`: Index the `hash` and `timestamp` columns ## Index Types[​](#index-types "Direct link to Index Types") There are two types of indexes that can be specified in a Spicepod: * `enabled`: Creates a standard index on the specified column(s). * Similar to specifying `CREATE INDEX my_index ON my_table (my_column)`. * `unique`: Creates a unique index on the specified column(s). See [Constraints](/docs/next/features/data-acceleration/constraints) for more information on working with unique constraints on locally accelerated tables. * Similar to specifying `CREATE UNIQUE INDEX my_index ON my_table (my_column)`. Limitations Traditional indexes are not supported for the in-memory Arrow or [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) acceleration engines. Use [DuckDB](/docs/next/components/data-accelerators/duckdb), [SQLite](/docs/next/components/data-accelerators/sqlite), [Turso](/docs/next/components/data-accelerators/turso), or [PostgreSQL](/docs/next/components/data-accelerators/postgres) as the acceleration engine to enable indexing. For Arrow acceleration, see [Hash Index](/docs/next/features/data-acceleration/hash-index) (experimental, v1.11.0-rc.2+) for O(1) point lookups on primary key columns. Spice Cayenne Point Lookup Performance While Spice Cayenne does not support traditional indexes, [Vortex](https://github.com/vortex-data/vortex) provides [100x faster random access reads](https://bench.vortex.dev) compared to Parquet through segment statistics (similar to zone-maps), fast random access encodings ([FSST](https://www.vldb.org/pvldb/vol13/p2649-boncz.pdf), [FastLanes](https://www.vldb.org/pvldb/vol16/p2132-afroozeh.pdf)), and compute push-down on compressed data. For many point lookup workloads, Spice Cayenne matches or exceeds indexed query performance without requiring explicit index configuration. See the [Spice Cayenne documentation](/docs/next/components/data-accelerators/cayenne#point-lookups-and-random-access) for details. --- # Partitioning Partitioning splits an accelerated dataset into multiple physical units (files or in-memory tables) keyed by an expression evaluated per row. Queries that filter on the partitioning expression — or on a column it references — only read the partitions that can match, dramatically reducing the data scanned. ``` datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet acceleration: enabled: true engine: cayenne mode: file partition_by: - bucket(50, PULocationID) ``` This config writes 50 separate partitions, each containing the rows whose `PULocationID` hashes to the same bucket. A query like `WHERE PULocationID = 132` only reads the single bucket partition that could contain `132`. ## How partitioning works[​](#how-partitioning-works "Direct link to How partitioning works") 1. **At refresh time**, Spice evaluates each `partition_by` expression for every row and routes the row to a partition keyed by the expression's value. 2. **At query time**, Spice rewrites filters that reference the partition column or expression into a partition selection, and only reads partitions that could satisfy the filter ([partition pruning](#partition-pruning)). 3. **Composite partitioning** (Arrow and Cayenne) layers multiple expressions hierarchically — e.g. `year` then `month` — to combine pruning across dimensions. Partitioning is most useful when: * Queries reliably filter on a small subset of values (`region = 'EU'`, `tenant_id IN (...)`, `created_at >= today() - INTERVAL '7 days'`). * A single column has high cardinality and partitioning by it directly would create too many tiny files — `bucket(N, col)` collapses to `N` partitions. * The dataset grows large enough that scanning the whole acceleration on every query is expensive. ## Configuration[​](#configuration "Direct link to Configuration") ### `partition_by`[​](#partition_by "Direct link to partition_by") Lives directly under `acceleration:` (not under `acceleration.params:`). It's a list of expressions; each entry is either a plain string or a single-entry `{ name: expression }` mapping: ``` acceleration: enabled: true engine: arrow partition_by: # Anonymous expression — auto-named "expr0", "expr1", … - "YEAR(created_at)" # Named expression - month: "MONTH(created_at)" ``` Multi-entry mappings (`- year: "…", month: "…"` on one list item) are rejected at load time. ### Supported engines[​](#supported-engines "Direct link to Supported engines") | Engine | Required `mode:` | Multi-expression | Layout | | --------- | ----------------- | ---------------- | ------------------------------------------------------------------ | | `arrow` | (memory; default) | Yes | One Arrow `MemTable` per partition value. | | `cayenne` | `file` | Yes | One Vortex table per partition; catalog in a SQLite metadata file. | `duckdb`, `sqlite`, `postgres`, and `turso` accelerators do **not** support `partition_by`; configuring it on those engines is rejected at load time. Use `arrow` or `cayenne` for partitioned acceleration. ## Partition transforms[​](#partition-transforms "Direct link to Partition transforms") `partition_by` accepts any DataFusion-compatible scalar SQL expression that returns a String, integer, Boolean, or Timestamp and references exactly one column from the dataset. The most common partition transforms are: ### `bucket(num_buckets, column)`[​](#bucketnum_buckets-column "Direct link to bucketnum_buckets-column") Hashes `column` into `num_buckets` deterministic buckets. * `num_buckets` must be a positive integer literal `≤ 1,000,000`. * The return type matches `num_buckets` — `bucket(50, …)` (Int64 literal) returns `Int64`; `bucket(50::int32, …)` returns `Int32`. * Same input always maps to the same bucket for a given `num_buckets` (uses ahash with a fixed seed). * Use for: high-cardinality columns where direct partitioning would create too many partitions (`user_id`, `account_id`, `device_id`). ``` partition_by: - bucket(100, user_id) ``` See the [`bucket` reference](/docs/next/reference/sql/scalar_functions#bucket) for the full SQL signature. ### `truncate(width, value)`[​](#truncatewidth-value "Direct link to truncatewidth-value") Truncates to the next-lower multiple of `width`. Iceberg's truncate transform. * `width` must be a positive `Int64` literal. * `value` may be any signed/unsigned integer, `Decimal128`/`Decimal256`, `Utf8` (string), or `Binary`. For strings/binary, returns the first `width` units. * Returns the same type as `value`. * Use for: floor-bucketing wide numeric ranges (`truncate(1000, amount)`), or grouping strings by prefix (`truncate(2, country_code)`). ``` partition_by: - truncate(1000, amount) ``` See the [`truncate` reference](/docs/next/reference/sql/scalar_functions#truncate) for examples. ### `date_part(unit, column)` and `date_trunc(unit, column)`[​](#date_partunit-column-and-date_truncunit-column "Direct link to date_partunit-column-and-date_truncunit-column") Built-in DataFusion datetime functions. Useful for time-based partitioning at year, month, day, or hour granularity. * `date_part('year', col)` returns the integer year (e.g. `2026`). * `date_trunc('day', col)` returns the timestamp truncated to the start of the day. ``` partition_by: - date_part('year', l_shipdate) ``` ``` partition_by: - day: "date_trunc('day', created_at)" ``` `YEAR(col)`, `MONTH(col)`, `DAY(col)`, etc. are aliases of `date_part(...)` and work identically. Filter pruning for `date_part()` is not yet implemented A `date_part('year', l_shipdate)` partition still produces correctly-distributed partitions, but a filter like `WHERE l_shipdate >= '2026-01-01'` does not currently translate to a partition pruning. Queries return correct results — they just scan more partitions than necessary. If your filter is on the bare partition expression (`WHERE date_part('year', l_shipdate) = 2026`), the equality form does prune. ### Modulo (`column % N`)[​](#modulo-column--n "Direct link to modulo-column--n") A plain modulo expression also produces stable, partition-prunable buckets: ``` partition_by: - "id % 16" ``` Range filters on the base column (`id BETWEEN 0 AND 1000`) are pruned for `%` partitions. ### Plain column reference[​](#plain-column-reference "Direct link to Plain column reference") A bare column name partitions one partition per distinct value — useful for low-cardinality columns: ``` partition_by: - region ``` Pruning supports equality, `IN`, `NOT IN`, and range filters on the column. ## Composite partitioning (Arrow + Cayenne)[​](#composite-partitioning-arrow--cayenne "Direct link to Composite partitioning (Arrow + Cayenne)") Arrow and Cayenne accelerations accept multiple `partition_by` expressions. Spice partitions hierarchically — first by the leftmost expression, then by the next, and so on. ``` acceleration: enabled: true engine: cayenne mode: file partition_by: - year: "date_part('year', created_at)" - month: "date_part('month', created_at)" - region: region ``` A query with `WHERE region = 'EU' AND date_part('year', created_at) = 2026` prunes on both axes. ## Partition pruning[​](#partition-pruning "Direct link to Partition pruning") Spice attempts to translate query-time filters into a partition selection. The matrix below summarizes which filter shapes prune which partition transforms: | Filter shape | Plain column / `truncate` / `date_trunc` / `% N` | `bucket(N, col)` | | ------------------------------------------------------ | ------------------------------------------------ | ------------------------------------------------------- | | `col = X` | ✓ | ✓ (filter substituted into bucket expression) | | `col != X` | ✓ | ✗ (no pruning) | | `col IN (a, b, …)` | ✓ | ✓ | | `col NOT IN (a, b, …)` | ✓ | ✗ (no pruning) | | `col < X` / `<= X` / `> X` / `>= X` | ✓ | ✓ only for **bounded** ranges (both lower and upper) | | `col BETWEEN a AND b` | ✓ | ✓ (when range expands to ≤ a few thousand Int32 values) | | `col = a OR col = b OR …` | ✓ | ✓ | | `partition_expr = X` (e.g. `bucket(50, c) = 7`) | ✓ | ✓ | | `expr1_filter AND expr2_filter` (composite partitions) | ✓ | ✓ | Pruning notes: * **Filter on the base column is substituted into the partition expression.** `WHERE user_id = 42` against a `bucket(100, user_id)` partition evaluates `bucket(100, 42)` once and reads only that partition. * **Bucket inequality pruning** enumerates candidate values within a bounded `Int32` range (capped at `MAX_BUCKET_ENUMERATION_I32` candidate values). Open-ended ranges (`col < X` with no lower bound) and non-`Int32` types fall back to no pruning. * **`date_part()` filter pruning is not yet implemented** for time-range filters. Equality on the partition expression still prunes; range filters on the base column do not. * **Filters that don't fully resolve at the partition layer** are passed through to the data layer — pruning is best-effort and never returns wrong rows. ## Engine-specific behavior[​](#engine-specific-behavior "Direct link to Engine-specific behavior") ### Arrow (`engine: arrow`)[​](#arrow-engine-arrow "Direct link to arrow-engine-arrow") In-memory MemTable per partition value. Partitions are rebuilt on every refresh. The simplest engine to start with for moderate datasets that fit in RAM. `hash_index` and `sort_columns` are propagated per partition. ### Cayenne (`engine: cayenne`)[​](#cayenne-engine-cayenne "Direct link to cayenne-engine-cayenne") Each partition is a separate Cayenne (Vortex) table; the partition catalog is tracked in a SQLite metadata file. Cayenne supports composite partitioning natively and is the right pick for very large datasets where Arrow would not fit in memory. Requires `mode: file`. ## Changing `partition_by` after refresh[​](#changing-partition_by-after-refresh "Direct link to changing-partition_by-after-refresh") Once partitions exist on disk, changing `partition_by` is rejected: > The `partition_by` expressions are different from the expressions used to create the existing partition files. Revert the `partition_by` expressions, delete the partition files, or change the location the partition files are stored to create new partitions. To re-partition: 1. Stop the runtime. 2. Delete the partitioned acceleration directory or point `accelerator_dir` at a fresh location. 3. Update `partition_by`. 4. Restart — Spice will rebuild from the source. There is no automatic in-place re-partitioning. ## Validation rules[​](#validation-rules "Direct link to Validation rules") A `partition_by` expression is rejected at startup if any of the following hold: * Result type is not `String`, an integer (signed or unsigned), `Boolean`, or `Timestamp`. * Expression references zero or multiple dataset columns. * Expression contains a subquery, `OUTER REFERENCE`, `UNNEST`, window function, aggregate function, `EXISTS`, `GROUPING SET`, or `PLACEHOLDER`. * Expression aliases the column (`col AS partition_key`). For the `cayenne` engine, `mode:` must be `file`. ## Examples[​](#examples "Direct link to Examples") ### High-cardinality hash partitioning[​](#high-cardinality-hash-partitioning "Direct link to High-cardinality hash partitioning") Partition events by user, collapsing millions of users into 100 stable buckets: ``` datasets: - from: s3://my-bucket/events/ name: events params: file_format: parquet acceleration: enabled: true engine: cayenne mode: file partition_by: - bucket(100, user_id) ``` ``` -- Reads the single bucket that hashes user_id 42: SELECT COUNT(*) FROM events WHERE user_id = 42; -- Reads at most 4 bucket partitions: SELECT * FROM events WHERE user_id IN (1, 2, 3, 4); ``` ### Time-based partitioning by year[​](#time-based-partitioning-by-year "Direct link to Time-based partitioning by year") ``` acceleration: enabled: true engine: cayenne mode: file partition_by: - date_part('year', l_shipdate) ``` ``` -- Equality on the partition expression prunes: SELECT * FROM lineitem WHERE date_part('year', l_shipdate) = 2026; ``` ### Composite year/month partitioning (Cayenne)[​](#composite-yearmonth-partitioning-cayenne "Direct link to Composite year/month partitioning (Cayenne)") ``` acceleration: enabled: true engine: cayenne mode: file partition_by: - year: "date_part('year', created_at)" - month: "date_part('month', created_at)" ``` ### Truncate-based numeric partitioning[​](#truncate-based-numeric-partitioning "Direct link to Truncate-based numeric partitioning") ``` acceleration: enabled: true engine: arrow partition_by: - "truncate(1000, order_total_cents)" ``` ``` -- Range pruning works because truncate is monotonic: SELECT * FROM orders WHERE order_total_cents BETWEEN 5000 AND 9999; ``` ### Plain column partitioning by region[​](#plain-column-partitioning-by-region "Direct link to Plain column partitioning by region") ``` acceleration: enabled: true engine: arrow partition_by: - region ``` ``` -- One partition read: SELECT * FROM events WHERE region = 'EU'; ``` ## Related[​](#related "Direct link to Related") * [`bucket` SQL reference](/docs/next/reference/sql/scalar_functions#bucket) * [`truncate` SQL reference](/docs/next/reference/sql/scalar_functions#truncate) * [`date_part` SQL reference](/docs/next/reference/sql/scalar_functions#date_part) * [`date_trunc` SQL reference](/docs/next/reference/sql/scalar_functions#date_trunc) * [Data refresh modes](/docs/next/features/data-acceleration/data-refresh) — `time_partition_column` for time-pruned refreshes * [Sharded deployment](/docs/next/deployment/architectures/sharded) — partitioning across multiple Spice instances rather than within one acceleration --- # Refresh Modes Spice supports five modes to refresh accelerated datasets. `full` is the default. | Mode | Description | Example | | -------------------------------------------------------------------------- | ---------------------------------------------------- | ---------------------------------------------------------------- | | [`full`](/docs/next/features/data-acceleration/refresh-modes/full) | Replace/overwrite the entire dataset on each refresh | A table of users | | [`append`](/docs/next/features/data-acceleration/refresh-modes/append) | Append/add data to the dataset on each refresh | Append-only, immutable datasets, such as time-series or log data | | [`changes`](/docs/next/features/data-acceleration/refresh-modes/changes) | Apply incremental inserts, updates, and deletes | Customer order lifecycle table | | [`caching`](/docs/next/features/data-acceleration/refresh-modes/caching) | Read-through caching for HTTP-based datasets | API search results or dynamic content endpoints | | [`snapshot`](/docs/next/features/data-acceleration/refresh-modes/snapshot) | Reload exclusively from the snapshot store | Read-only replicas bootstrapped from centralized snapshots | For cross-cutting refresh behavior — refresh intervals, on-demand refresh, retries, retention, and behavior on zero results — see [Data Refresh](/docs/next/features/data-acceleration/data-refresh). --- # Append Refresh Mode The `append` refresh mode incrementally adds new rows to the acceleration on each refresh. It is designed for append-only or immutable datasets such as time-series, event, and log data. Use `append` when: * New rows are continuously added to the source and existing rows are not modified or deleted. * A monotonic time or sequence column is available to identify new rows. * The full dataset is too large to refresh in `full` mode on each interval. ## Configuration[​](#configuration "Direct link to Configuration") `append` mode requires a [`time_column`](/docs/next/reference/spicepod/datasets#time_column) that identifies new rows by comparing the local maximum value to the source. Data is incrementally refreshed where `time_column` in the source is greater than `max(time_column)` in the acceleration — see [Day-Granular Time Columns](#day-granular-time-columns) for the one case where that comparison is inclusive instead. ``` datasets: - from: databricks:my_dataset name: accelerated_dataset time_column: created_at acceleration: enabled: true refresh_mode: append refresh_check_interval: 10m ``` ## Late-Arriving Data[​](#late-arriving-data "Direct link to Late-Arriving Data") To account for clock skew or late-arriving rows, configure an overlap window with [`acceleration.refresh_append_overlap`](/docs/next/reference/spicepod/datasets#accelerationrefresh_append_overlap). Rows within the overlap are re-read on each refresh. ## Day-Granular Time Columns[​](#day-granular-time-columns "Direct link to Day-Granular Time Columns") A date-typed `time_column` (Arrow `Date32`, configured as [`time_format: date`](/docs/next/reference/spicepod/datasets#time_format)) carries no time of day, so every row of a given day shares one value. A strictly-greater comparison against the accelerated maximum would exclude every row that arrives later for the day already loaded, on that refresh and on every one after it. For a day-granular time column the comparison is therefore **inclusive**, against the start of the high-water mark's day, and Spice's exact-row de-duplication drops the already-loaded rows that come back so nothing is appended twice. A `time_partition_column` of the same type is floored to that same day boundary, so the partition predicate does not exclude the day the new rows are in. The trade-off is that each refresh re-reads the boundary day from the source rather than nothing — the minimum needed to see that day's new rows at all. Sub-day time columns are unaffected and keep the strict comparison; if a single day is large, bound the re-read by using a sub-day `time_column`. `refresh_append_overlap` still subtracts its duration from the high-water mark, but on a day-granular column the resulting window is quantized to the day boundary, so an overlap shorter than a day re-reads the whole boundary day. ## Partition Pruning with `time_partition_column`[​](#partition-pruning-with-time_partition_column "Direct link to partition-pruning-with-time_partition_column") Datasets partitioned by a less-granular time column (day, month, year) can specify [`time_partition_column`](/docs/next/reference/spicepod/datasets#time_partition_column) in addition to `time_column` for efficient partition pruning at the source. ``` datasets: - from: databricks:my_dataset name: accelerated_dataset time_column: created_at time_format: iso8601 time_partition_column: created_at_day time_partition_format: date ``` ## Append Only Modified Files[​](#append-only-modified-files "Direct link to Append Only Modified Files") For object-store sources, set `time_column` or `time_partition_column` to the special value `last_modified` to append only newly created or updated files. Spice uses file metadata to determine which files are new, dramatically reducing scan time for large datasets. ``` datasets: - from: s3://my_bucket/my_dataset name: accelerated_dataset time_column: last_modified params: file_format: parquet acceleration: refresh_mode: append refresh_check_interval: 10m ``` If `last_modified` exists as a column in the data, the column value takes precedence over file metadata. This is supported for connectors that accept the [file format parameter](/docs/next/reference/file_format), such as `s3://`, `abfs://`, and `file://`. ## Readiness with Snapshots[​](#readiness-with-snapshots "Direct link to Readiness with Snapshots") Append-mode accelerations that define a `time_column` wait to report ready until the first append refresh completes after [snapshot bootstrap](/docs/next/features/data-acceleration/snapshots). This keeps the dataset out of rotation until the freshest data is available while still benefiting from snapshot-assisted startup. ## Combining with Upserts[​](#combining-with-upserts "Direct link to Combining with Upserts") Pair `refresh_mode: append` with a `primary_key` and `on_conflict: upsert` to handle source rows that are occasionally updated. See [End-to-End Incremental Ingestion Example](/docs/next/features/data-acceleration/data-refresh#end-to-end-incremental-ingestion-example). ## Related Topics[​](#related-topics "Direct link to Related Topics") * [Refresh Interval](/docs/next/features/data-acceleration/data-refresh#refresh-interval) * [Refresh on Startup](/docs/next/features/data-acceleration/data-refresh#refresh-on-startup) * [Refresh Retries](/docs/next/features/data-acceleration/data-refresh#refresh-retries) * [Retention Policy](/docs/next/features/data-acceleration/data-refresh#retention-policy) * [Refresh Data Window](/docs/next/features/data-acceleration/data-refresh#refresh-data-window) --- # Caching Refresh Mode The `caching` refresh mode provides intelligent caching for HTTP-based datasets where multiple result rows can share the same request metadata. This mode is specifically designed for scenarios like API responses where the same request parameters can return different content over time or multiple rows of data. The caching mode supports two key paradigms: 1. **Stale-While-Revalidate (SWR)** - Serves cached data immediately while refreshing in the background, optimizing for low latency and reduced API costs 2. **Cache Persistence** - Stores cached data to disk using file-based accelerators (DuckDB, SQLite, or Cayenne) for fast cold starts and durability ## Overview[​](#overview "Direct link to Overview") Unlike traditional refresh modes that treat datasets as single sources of truth, the `caching` mode treats HTTP request metadata (path, query parameters, and body) as cache keys. This approach is particularly useful for: * REST API responses that return multiple records for a single request * Search API results where the same query may return different results over time * Dynamic content APIs where responses change based on server state Future Enhancement While currently designed for HTTP-based datasets, future versions of Spice will extend the `caching` mode to support arbitrary queries from any data source, enabling flexible caching strategies across all connector types. ## How It Works[​](#how-it-works "Direct link to How It Works") The `caching` mode uses HTTP request filter values as cache keys rather than enforcing primary key constraints. When a refresh occurs: 1. **Cache Key Generation**: By default, the combination of `request_path`, `request_query`, and `request_body` acts as the cache key. If a `primary_key` is explicitly specified in the acceleration configuration, it will be used instead of the metadata fields. 2. **Row Replacement**: All existing rows matching the cache key are removed before inserting new data 3. **Multiple Results**: Multiple rows with identical request metadata can coexist, representing different content items from the same API response 4. **Timestamp Tracking**: Each row includes a `fetched_at` timestamp indicating when the data was retrieved ## Schema[​](#schema "Direct link to Schema") Datasets using `caching` mode include the following metadata fields in addition to the content data: | Field Name | Type | Description | | ------------------ | ------------------- | ------------------------------------------------------------------- | | `request_path` | String | The URL path used for the request | | `request_query` | String | The query parameters used for the request | | `request_body` | String | The request body (for POST requests) | | `content` | String | The response content | | `response_status` | UInt16 | The HTTP status code of the response (e.g., `200`, `404`, `500`) | | `response_headers` | Map(String, String) | The HTTP response headers as key-value pairs | | `fetched_at` | Timestamp | The timestamp when the data was fetched (based on HTTP Date header) | The `fetched_at` timestamp uses the HTTP `Date` response header when available, falling back to the current system time if not present. ## Configuration[​](#configuration "Direct link to Configuration") To use `caching` mode, configure an `HTTP`/`HTTPS` dataset with `refresh_mode: caching`: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_cache params: file_format: json allowed_request_paths: '/search/shows,/shows/*' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file refresh_check_interval: 30s params: caching_ttl: 10s # How long a cache entry is considered fresh caching_stale_while_revalidate_ttl: 30s # How long after the `caching_ttl` to serve stale data while refreshing in the background ``` ## Use Cases[​](#use-cases "Direct link to Use Cases") ### Caching TV Show Search Results[​](#caching-tv-show-search-results "Direct link to Caching TV Show Search Results") Cache TV show search API results where the same query may return different results over time: ``` datasets: - from: https://api.tvmaze.com name: tv_search_cache params: file_format: json allowed_request_paths: '/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 15s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 30s refresh_sql: | SELECT * FROM tv_search_cache WHERE request_path = '/search/shows' AND request_query = 'q=game+of+thrones' ``` This configuration: * Fetches search results for "game of thrones" every 30 seconds * Stores all result items with the same request metadata * Replaces all previous results for this query on each refresh * Preserves the timestamp of when results were fetched ### Caching TV Show Episodes[​](#caching-tv-show-episodes "Direct link to Caching TV Show Episodes") Cache responses from a TV show episodes API: ``` datasets: - from: https://api.tvmaze.com name: episodes_cache params: file_format: json allowed_request_paths: '/shows/*/episodes' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 10s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 20s refresh_sql: | SELECT * FROM episodes_cache WHERE request_path = '/shows/82/episodes' AND request_query = 'season=1' ``` ### Multi-Endpoint Caching[​](#multi-endpoint-caching "Direct link to Multi-Endpoint Caching") Cache responses from multiple TVMaze API endpoints: ``` datasets: - from: https://api.tvmaze.com name: multi_endpoint_cache params: file_format: json allowed_request_paths: '/shows/*,/search/shows,/people/*' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 10s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 30s refresh_sql: | SELECT * FROM multi_endpoint_cache WHERE (request_path = '/shows/82' OR request_path = '/shows/169') OR (request_path = '/search/shows' AND request_query = 'q=breaking+bad') ``` ## Querying Cached Data[​](#querying-cached-data "Direct link to Querying Cached Data") Query cached data using standard SQL, filtering by request metadata or content: ``` -- Get all cached search results for a specific TV show query SELECT content, fetched_at FROM tv_search_cache WHERE request_query = 'q=game+of+thrones' ORDER BY fetched_at DESC; -- Find the most recent cache entry for each unique request SELECT request_path, request_query, MAX(fetched_at) as last_fetched FROM tv_shows_cache GROUP BY request_path, request_query; -- Get cached results fetched within the last hour SELECT * FROM episodes_cache WHERE fetched_at > NOW() - INTERVAL '1 hour'; -- Parse JSON to extract show information SELECT json_get_str(content, 'name') as show_name, json_get_str(content, 'type') as show_type, fetched_at FROM tv_shows_cache WHERE request_path = '/shows/82'; ``` ## Stale-While-Revalidate Pattern[​](#stale-while-revalidate-pattern "Direct link to Stale-While-Revalidate Pattern") The `caching` mode supports the Stale-While-Revalidate (SWR) pattern for acceleration, providing optimal performance by serving cached data immediately while refreshing in the background. ### How SWR Works[​](#how-swr-works "Direct link to How SWR Works") When configured with background refresh, the caching mode: 1. **Serves stale data immediately** - Returns cached results without waiting for a refresh 2. **Triggers background refresh** - Initiates an asynchronous refresh of the cache 3. **Updates cache transparently** - Subsequent queries receive fresh data once the refresh completes This pattern reduces query latency by eliminating wait times for data fetches while keeping the cache reasonably fresh. ### Background Refresh Configuration[​](#background-refresh-configuration "Direct link to Background Refresh Configuration") Configure background refresh using `refresh_check_interval` to specify how frequently the cache should be updated: ``` datasets: - from: https://api.tvmaze.com name: shows_cache params: file_format: json allowed_request_paths: '/shows/*' acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file # Persist cache to disk params: caching_ttl: 10s # Data is fresh for 10 seconds caching_stale_while_revalidate_ttl: 10s # Serve stale data for 10 seconds while refreshing refresh_check_interval: 30s # Refresh every 30 seconds in background refresh_sql: | SELECT * FROM shows_cache WHERE request_path = '/shows/82' ``` ### SWR Benefits for API Caching[​](#swr-benefits-for-api-caching "Direct link to SWR Benefits for API Caching") The SWR pattern is particularly valuable for caching API responses: * **Reduced latency** - Queries return immediately from the cache without waiting for HTTP requests * **Lower API costs** - Fewer requests to external APIs reduce usage and costs * **Improved reliability** - Cached data remains available even if the API is temporarily unavailable * **Better user experience** - Consistent fast response times improve application performance ### Example: SWR with On-Demand Refresh[​](#example-swr-with-on-demand-refresh "Direct link to Example: SWR with On-Demand Refresh") Combine background refresh with on-demand refresh for maximum flexibility: ``` datasets: - from: https://api.tvmaze.com name: tv_search_swr params: file_format: json allowed_request_paths: '/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file # Persist cache to disk params: caching_ttl: 15s # Cache data is fresh for 15 seconds caching_stale_while_revalidate_ttl: 10s # Serve stale data for 10 seconds while refreshing refresh_check_interval: 30s # Background refresh every 30 seconds refresh_on_startup: always # Always refresh on startup refresh_sql: | SELECT * FROM tv_search_swr WHERE request_path = '/search/shows' AND request_query = 'q=breaking+bad' ``` With this configuration: * The cache refreshes every 30 seconds automatically * Queries are served immediately from the cache * Manual refresh is available via `/v1/datasets/tv_search_swr/acceleration/refresh` * Cache is guaranteed fresh on application startup ## Cache Persistence[​](#cache-persistence "Direct link to Cache Persistence") The `caching` mode supports persisting cached data to disk using file-based acceleration engines, enabling the cache to survive application restarts and reducing cold start times. ### File-Based Accelerators[​](#file-based-accelerators "Direct link to File-Based Accelerators") Three acceleration engines support file persistence for caching mode: * **DuckDB** - High-performance analytical database with excellent compression * **SQLite** - Lightweight, reliable database ideal for embedded scenarios * **Cayenne** - Spice's native accelerator optimized for analytical workloads ### Configuring File Persistence[​](#configuring-file-persistence "Direct link to Configuring File Persistence") Enable file persistence by setting `acceleration.mode: file` and specifying an acceleration engine: ``` datasets: - from: https://api.tvmaze.com name: shows_persistent_cache params: file_format: json allowed_request_paths: '/shows/*,/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb # or sqlite, cayenne mode: file # Enable file persistence params: caching_ttl: 10s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 30s refresh_sql: | SELECT * FROM shows_persistent_cache WHERE request_path IN ('/shows/82', '/shows/169') OR (request_path = '/search/shows' AND request_query = 'q=game+of+thrones') ``` ### DuckDB Persistence Example[​](#duckdb-persistence-example "Direct link to DuckDB Persistence Example") DuckDB provides excellent performance for analytical queries on cached data: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_duckdb params: file_format: json allowed_request_paths: '/shows/*,/shows/*/episodes' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 10s caching_stale_while_revalidate_ttl: 10s duckdb_file: tv_shows_cache.db # Specify custom file location refresh_check_interval: 30s refresh_sql: | SELECT * FROM tv_shows_duckdb WHERE request_path IN ('/shows/82', '/shows/169') OR (request_path = '/shows/82/episodes' AND request_query = 'season=1') ``` ### SQLite Persistence Example[​](#sqlite-persistence-example "Direct link to SQLite Persistence Example") SQLite is ideal for lightweight caching scenarios: ``` datasets: - from: https://api.tvmaze.com name: tv_search_sqlite params: file_format: json allowed_request_paths: '/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: sqlite mode: file params: caching_ttl: 15s caching_stale_while_revalidate_ttl: 10s sqlite_file: tv_search_cache.db refresh_check_interval: 30s refresh_sql: | SELECT * FROM tv_search_sqlite WHERE request_path = '/search/shows' AND request_query IN ('q=breaking+bad', 'q=game+of+thrones') ``` ### Cayenne Persistence Example[​](#cayenne-persistence-example "Direct link to Cayenne Persistence Example") Cayenne provides optimized performance for Spice workloads: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_cayenne params: file_format: json allowed_request_paths: '/shows/*' acceleration: enabled: true refresh_mode: caching engine: cayenne mode: file params: caching_ttl: 10s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 30s refresh_sql: | SELECT * FROM tv_shows_cayenne WHERE request_path IN ('/shows/82', '/shows/169', '/shows/73') ``` ### Benefits of File Persistence[​](#benefits-of-file-persistence "Direct link to Benefits of File Persistence") Persisting the cache to disk provides several advantages: * **Fast cold starts** - Cache is immediately available on application restart without fetching from APIs * **Reduced API load** - No need to refetch all data after restarts * **Cost savings** - Fewer API requests reduce metered API costs * **Offline capability** - Cached data remains queryable even when the API is unavailable * **Data durability** - Cache survives application crashes and restarts ### Memory vs. File Mode[​](#memory-vs-file-mode "Direct link to Memory vs. File Mode") Choose between in-memory and file-based caching based on your requirements: | Aspect | Memory Mode (`mode: memory`) | File Mode (`mode: file`) | | --------------- | --------------------------------- | ----------------------------- | | **Performance** | Fastest - all data in RAM | Fast with disk I/O | | **Persistence** | Lost on restart | Survives restarts | | **Capacity** | Limited by available memory | Limited by disk space | | **Cold start** | Slow - must refetch all data | Fast - loads from disk | | **Best for** | Small, frequently changing caches | Large, stable caches | | **Engines** | `arrow` (default) | `duckdb`, `sqlite`, `cayenne` | ### Combining SWR and Persistence[​](#combining-swr-and-persistence "Direct link to Combining SWR and Persistence") For optimal performance, combine the SWR pattern with file persistence: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_optimized params: file_format: json allowed_request_paths: '/shows/*,/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file # Persist to disk params: caching_ttl: 15s # Cache data is fresh for 15 seconds caching_stale_while_revalidate_ttl: 10s # Serve stale data for 10 seconds while refreshing refresh_check_interval: 30s # Background refresh (SWR) refresh_on_startup: auto # Use persisted cache on startup refresh_sql: | SELECT * FROM tv_shows_optimized WHERE request_path IN ('/shows/82', '/shows/169') OR (request_path = '/search/shows' AND request_query = 'q=game+of+thrones') ``` This configuration provides: * Immediate query response from persisted cache on startup * Background refresh every 30 seconds without blocking queries * Durable cache that survives application restarts * Reduced API requests and costs ## Behavior and Characteristics[​](#behavior-and-characteristics "Direct link to Behavior and Characteristics") ### Row-Level Replacement[​](#row-level-replacement "Direct link to Row-Level Replacement") The `caching` mode uses `InsertOp::Replace` to handle data updates. When new data is fetched for a given cache key (request metadata): 1. All existing rows matching that cache key are removed 2. All new rows are inserted 3. This operation is atomic, ensuring consistent cache state This behavior differs from other modes: * **`full` mode**: Replaces the entire dataset * **`append` mode**: Adds new rows without removing existing ones * **`changes` mode**: Applies CDC events * **`caching` mode**: Replaces rows matching the specific cache key ### Cache Key Behavior[​](#cache-key-behavior "Direct link to Cache Key Behavior") The `caching` mode determines cache keys based on the acceleration configuration: **Default (No Primary Key Specified)**: * Uses HTTP request metadata fields as the cache key: `request_path`, `request_query`, and `request_body` * Multiple result rows can share the same request metadata * The cache key serves as the logical grouping mechanism for row replacement * Content within a response may have duplicate values across different requests **With Primary Key Specified**: * Uses the explicitly configured `primary_key` columns as the cache key * Provides fine-grained control over cache key composition * Useful when caching requires uniqueness based on response content fields rather than request metadata Example with custom primary key: ``` datasets: - from: https://api.tvmaze.com name: tv_episodes_custom_key params: file_format: json allowed_request_paths: '/shows/*/episodes' acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file primary_key: [id, season, number] # Use episode fields as cache key params: caching_ttl: 15s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 30s refresh_sql: | SELECT * FROM tv_episodes_custom_key WHERE request_path = '/shows/82/episodes' ``` ### HTTP Date Header[​](#http-date-header "Direct link to HTTP Date Header") The `fetched_at` timestamp respects the HTTP `Date` response header when present. This provides: * Accurate server-side timestamps for cached responses * Consistency across distributed systems * Proper cache age calculation based on server time If the `Date` header is not present, the system falls back to using the current system time. ### Transient Error Handling[​](#transient-error-handling "Direct link to Transient Error Handling") When caching HTTP responses, transient server errors are automatically excluded from the cache to prevent temporary failures from polluting cached data. Specifically: * **5xx responses** (500–599) — Server errors indicating temporary issues (e.g., overload, outage) * **429 Too Many Requests** — Rate limiting responses These responses are still returned to the querying client, but they are **not written to the cache**. This ensures that subsequent cache reads return valid data rather than error responses from temporary failures. ## Refresh Configuration[​](#refresh-configuration "Direct link to Refresh Configuration") The `caching` mode supports standard refresh configuration options. See [Stale-While-Revalidate Pattern](#stale-while-revalidate-pattern) for background refresh details and [Cache Persistence](#cache-persistence) for file-based caching configuration. | Parameter | Description | Default | | ------------------------ | ---------------------------------------------------------------------------------------------------------------- | -------------- | | `refresh_check_interval` | How often to refresh cached data in the background | None | | `refresh_sql` | SQL query defining what data to cache | None | | `refresh_on_startup` | Whether to refresh on startup (`auto` or `always`) | `auto` | | `on_zero_results` | **Ignored in caching mode.** Caching mode always queries the source on a cache miss, regardless of this setting. | `return_empty` | | `engine` | Acceleration engine (`arrow`, `duckdb`, `sqlite`, `cayenne`) | `arrow` | | `mode` | Persistence mode (`memory` or `file`) | `memory` | ### Cache TTL (Time-to-Live)[​](#cache-ttl-time-to-live "Direct link to Cache TTL (Time-to-Live)") The caching mode provides parameters to control cache freshness and staleness behavior: | Parameter | Description | Default | | ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | | `caching_ttl` | Duration that cached data is considered fresh. After this period, data becomes stale and triggers a background refresh. | `30s` | | `caching_stale_while_revalidate_ttl` | Duration after `caching_ttl` expires during which stale data is served while refreshing in the background. After this period, queries wait for fresh data. | None | | `caching_stale_if_error` | When set to `enabled`, serves expired cached data if the upstream source returns an error. Valid values: `enabled`, `disabled`. | `disabled` | The `caching_ttl` parameter defines how long cached data is considered fresh before it becomes stale. Once cached data exceeds this age, the SWR pattern triggers background refresh to update the cache while continuing to serve the stale data during the `caching_stale_while_revalidate_ttl` window. If `caching_stale_while_revalidate_ttl` is omitted, cached data becomes rotten immediately after `caching_ttl` expires, and queries will wait for fresh data rather than returning stale results. When a value is specified, stale data is served during that window after `caching_ttl` expires while a background refresh occurs. Once the combined `caching_ttl + caching_stale_while_revalidate_ttl` period has passed, queries will wait for fresh data instead of returning stale results. **Configuring Cache TTL**: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_ttl params: file_format: json allowed_request_paths: '/shows/*' acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file # Persist cache to disk refresh_check_interval: 30s # Periodic check for stale data params: caching_ttl: 15s # Cache data is fresh for 15 seconds caching_stale_while_revalidate_ttl: 10s # Serve stale data for 10 seconds while refreshing ``` **How Cache TTL Works**: 1. When data is fetched, the `fetched_at` timestamp is recorded 2. On subsequent queries, the system checks `now - fetched_at > caching_ttl` 3. If data is within TTL, it is served immediately (fresh) 4. If data exceeds TTL, it becomes stale: * Stale data is served immediately (no query delay) if within `caching_stale_while_revalidate_ttl` * Background refresh is triggered to update the cache * Next query receives the refreshed data **TTL Format**: Duration strings support common units: * Seconds: `30s`, `90s` * Minutes: `5m`, `15m` * Hours: `1h`, `24h` * Mixed: `1h30m`, `2h15m30s` **Default Behavior**: When `caching_ttl` is not specified, it defaults to `30s` (30 seconds). This provides a reasonable balance between freshness and cache efficiency for most use cases. When `caching_stale_while_revalidate_ttl` is not specified, stale data is not served after the TTL expires, and queries will wait for fresh data. ### Stale-If-Error Behavior[​](#stale-if-error-behavior "Direct link to Stale-If-Error Behavior") The `caching_stale_if_error` parameter controls whether expired cached data is served when the upstream data source returns an error during a refresh attempt. This provides fault tolerance by returning stale data instead of failing the query when the upstream source is temporarily unavailable. ``` datasets: - from: https://api.tvmaze.com name: tv_shows_resilient acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 15s caching_stale_while_revalidate_ttl: 30s caching_stale_if_error: enabled # Serve stale data on upstream errors ``` When `caching_stale_if_error: enabled`: * If the upstream source returns an error during refresh, expired cached data is served instead of failing * Queries continue to return data even when the upstream API is unavailable * Useful for APIs with intermittent availability or rate limits When `caching_stale_if_error: disabled` (default): * Errors from the upstream source propagate to the query * Queries fail when fresh data cannot be fetched and cached data has expired Stale-While-Revalidate Configuration Conflict When using `refresh_mode: caching`, you cannot configure both the caching accelerator's `caching_stale_while_revalidate_ttl` and the [results cache](/docs/next/features/caching)'s `stale_while_revalidate_ttl` for the same dataset. These parameters control similar behavior at different layers, and having both enabled creates a conflict. Choose one approach: * **Caching accelerator SWR**: Use `acceleration.params.caching_stale_while_revalidate_ttl` for HTTP-based dataset caching * **Results cache SWR**: Use `runtime.caching.sql_results.stale_while_revalidate_ttl` for SQL query results caching **TTL Considerations**: * **Shorter TTL** (e.g., `10s`, `30s`): More frequent refresh, higher data freshness, more API requests * **Longer TTL** (e.g., `10m`, `1h`): Fewer refreshes, lower API costs, potentially stale data * **Matching patterns**: Set `caching_ttl` shorter than `refresh_check_interval` to define the staleness window ## Limitations[​](#limitations "Direct link to Limitations") * Currently only available for HTTP-based datasets using the [HTTPS connector](/docs/next/components/data-connectors/https). Future releases will extend support to arbitrary queries from any data source. * Requires `acceleration.enabled: true` * When no `primary_key` is specified, cache keys default to request metadata fields (`request_path`, `request_query`, `request_body`) * On-demand refresh via `/v1/datasets/:name/acceleration/refresh` API triggers a new refresh for all cache keys defined in `refresh_sql` ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") * [HTTPS Data Connector](/docs/next/components/data-connectors/https) - Detailed HTTP connector configuration * [Data Refresh](/docs/next/features/data-acceleration/data-refresh) - Overview of all refresh modes * [Refresh SQL](/docs/next/features/data-acceleration/data-refresh#refresh-sql) - Using SQL to control refresh behavior * [Special Metadata Fields](/docs/next/components/data-connectors/https#special-metadata-fields) - HTTP request metadata fields * [Data Accelerators](/docs/next/components/data-accelerators) - Acceleration engines for cache persistence * [DuckDB Accelerator](/docs/next/components/data-accelerators/duckdb) - DuckDB acceleration engine * [SQLite Accelerator](/docs/next/components/data-accelerators/sqlite) - SQLite acceleration engine --- # Changes Refresh Mode The `changes` refresh mode applies incremental inserts, updates, and deletes from a [Change Data Capture (CDC)](/docs/next/features/cdc) source. Unlike `append`, `changes` mode reflects modifications and deletions in the acceleration, keeping it consistent with sources where rows mutate over time. Use `changes` when: * The source supports CDC (e.g., a database with a transaction log). * Rows in the source are updated or deleted, not just inserted. * The acceleration must reflect the current state of the source row-for-row. ## Configuration[​](#configuration "Direct link to Configuration") `refresh_mode: changes` requires a CDC-capable data connector. Spice supports CDC via [PostgreSQL Logical Replication](/docs/next/features/cdc/postgres-replication), [MySQL Binlog Replication](/docs/next/features/cdc/mysql-replication), [MongoDB Change Streams](/docs/next/features/cdc/mongodb-streams), [DynamoDB Streams](/docs/next/features/cdc/dynamodb-streams), and [Debezium](/docs/next/features/cdc/debezium) (over Kafka). See [Supported Data Connectors](/docs/next/features/cdc#supported-data-connectors) for details. note [Apache Kafka](/docs/next/components/data-connectors/kafka) is a real-time streaming source but is append-only — it uses [`refresh_mode: append`](/docs/next/features/data-acceleration/refresh-modes/append), not `changes`. Any accelerator engine that supports writes can be a `changes` sink — `arrow`, `duckdb`, `sqlite`, and `cayenne`. [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) is recommended for large-scale CDC (incremental materialized views, in-memory CDC tier, and replication-lag/freshness SLOs). ``` datasets: - from: debezium:cdc.public.customer_orders name: customer_orders acceleration: enabled: true refresh_mode: changes engine: duckdb mode: file ``` The Debezium connector streams change events from a Kafka topic produced by Debezium. Each event is applied to the acceleration in order, preserving inserts, updates, and deletes from the source. ## Behavior[​](#behavior "Direct link to Behavior") * The acceleration is bootstrapped from the source snapshot, then continuously updated from the change stream. * `refresh_check_interval`, `refresh_cron`, on-demand refresh, `refresh_data_window`, and `retention_period` do not apply — updates are driven by the change stream rather than periodic polling. * [`refresh_sql`](/docs/next/features/data-acceleration/data-refresh#refresh-sql) can only modify selected columns in `changes` mode and cannot apply row filters. ## Related Topics[​](#related-topics "Direct link to Related Topics") * [Change Data Capture](/docs/next/features/cdc) * [Debezium Data Connector](/docs/next/components/data-connectors/debezium) --- # Full Refresh Mode The `full` refresh mode replaces the entire accelerated dataset on every refresh. It is the default refresh mode and the simplest way to keep an acceleration in sync with its source. Use `full` when: * The dataset is small enough to be re-read on each refresh. * Source rows can be inserted, updated, or deleted, and incremental tracking is not available. * Strong consistency with the source is preferred over minimizing source load. ## Configuration[​](#configuration "Direct link to Configuration") ``` datasets: - from: databricks:my_dataset name: accelerated_dataset acceleration: enabled: true refresh_mode: full refresh_check_interval: 10m ``` On each refresh, the runtime issues a single `SELECT` against the source, materializes the result into the acceleration engine, and atomically swaps the new data in. ## Behavior[​](#behavior "Direct link to Behavior") * Each refresh fully scans the source. Any [`refresh_sql`](/docs/next/features/data-acceleration/data-refresh#refresh-sql) and [`refresh_data_window`](/docs/next/features/data-acceleration/data-refresh#refresh-data-window) filters are pushed down to limit data transferred. * Queries continue to be served from the previous result set until the new refresh completes. * Supported with all data connectors and all acceleration engines. ## Related Topics[​](#related-topics "Direct link to Related Topics") For cross-cutting refresh behavior that applies to `full` mode, see: * [Refresh Interval](/docs/next/features/data-acceleration/data-refresh#refresh-interval) * [Refresh on Startup](/docs/next/features/data-acceleration/data-refresh#refresh-on-startup) * [Refresh Retries](/docs/next/features/data-acceleration/data-refresh#refresh-retries) * [Retention Policy](/docs/next/features/data-acceleration/data-refresh#retention-policy) * [Behavior on Zero Results](/docs/next/features/data-acceleration/data-refresh#behavior-on-zero-results) --- # Snapshot Refresh Mode The `snapshot` refresh mode creates a read-only acceleration that reloads exclusively from the [snapshot store](/docs/next/features/data-acceleration/snapshots). The federated source is never queried for refreshes — instead, the runtime polls the snapshot store on a configurable interval and atomically swaps in newer snapshots when available. Use `snapshot` when: * A separate writer publishes acceleration snapshots to object storage. * Read replicas need fast, source-independent startup and refresh. * The federated source should not be queried by the replica (e.g., edge nodes, security boundaries, or to reduce source load). ## Configuration[​](#configuration "Direct link to Configuration") ``` snapshots: enabled: true location: s3://my-bucket/snapshots/ params: s3_auth: iam_role datasets: - from: postgres:public.my_table name: my_table acceleration: enabled: true engine: duckdb mode: file refresh_mode: snapshot refresh_check_interval: 30s # Poll interval; defaults to 1m snapshots: enabled params: duckdb_file: /nvme/my_table.db ``` ## Requirements[​](#requirements "Direct link to Requirements") * `acceleration.snapshots` must be `enabled` or `bootstrap_only`. * The acceleration engine must be a snapshot-capable file-based engine: **DuckDB**, **SQLite**, **Cayenne**, or **Turso**. ## Behavior[​](#behavior "Direct link to Behavior") * On startup, the runtime bootstraps from the most recent snapshot, identical to other snapshot-enabled modes. * After bootstrap, the runtime polls the snapshot store at `refresh_check_interval` (default: 60s) for newer snapshots. * When a newer snapshot is found, its schema is validated against the current acceleration schema before downloading. * The accelerator file is swapped atomically — queries continue to be served from the previous snapshot until the swap completes. * `INSERT INTO` statements are rejected with an error since the acceleration is driven exclusively from snapshots. tip Use `refresh_mode: snapshot` for read-only replicas that should not access the federated source — for example, edge nodes that receive snapshots from a centralized writer. ## Related Topics[​](#related-topics "Direct link to Related Topics") * [Acceleration Snapshots](/docs/next/features/data-acceleration/snapshots) * [Refresh Interval](/docs/next/features/data-acceleration/data-refresh#refresh-interval) --- # Snapshots Enterprise edition Acceleration Snapshots are available in the Spice [Enterprise edition](https://docs.spice.ai/docs/enterprise/getting-started/distributions). ## Spicepod Example[​](#spicepod-example "Direct link to Spicepod Example") ``` snapshots: enabled: true location: s3://some_bucket/some_folder/ bootstrap_on_failure_behavior: warn params: s3_auth: iam_role datasets: - name: some_table acceleration: engine: duckdb mode: file snapshots: enabled snapshots_trigger: refresh_complete params: duckdb_file: /nvme/some_table.db ``` ## Overview[​](#overview "Direct link to Overview") Acceleration snapshots let Spice reuse a pre-built acceleration file on startup instead of waiting for a full refresh. When a dataset uses a file-mode acceleration engine (DuckDB, SQLite, Cayenne, or Turso) and the local file is missing (for example on first boot or when using ephemeral NVMe storage), Spice downloads the most recent snapshot from object storage and moves the dataset straight to a ready state. ## How it works[​](#how-it-works "Direct link to How it works") * On startup, Spice checks whether the file supplied in `acceleration.params` (for example `duckdb_file`) exists. * If the file is missing and snapshots are enabled, Spice looks under the configured snapshot location and downloads the newest snapshot for that dataset. * If no snapshot is available, the acceleration boots empty and refreshes from the source. * Spice creates new snapshots based on the configured `snapshots_trigger` mode. Snapshots are organized with Hive-style partitioning so they are easy to retain and prune. For a dataset named `my_dataset`, Spice writes files such as: ``` s3://some_bucket/some_folder/month=2025-09/day=2025-09-30/dataset=my_dataset/my_dataset_20250919T134522Z.db ``` The timestamp is recorded in UTC using ISO 8601 without punctuation. Dedicated files only Every accelerated dataset must write to its own file (for example, `/nvme/my_dataset.db`). Sharing a single file across multiple datasets is not supported. ## Configure snapshot storage[​](#configure-snapshot-storage "Direct link to Configure snapshot storage") Snapshots are controlled with a top-level `snapshots` block in the Spicepod. The location can point to S3, Azure ADLS Gen2, Google Cloud Storage, or the local filesystem. ``` snapshots: enabled: true location: s3://some_bucket/some_folder/ # Folder where snapshots are written bootstrap_on_failure_behavior: warn # retry | fallback | warn params: s3_auth: iam_role # Defaults to iam_role for snapshots ``` ### Supported storage backends[​](#supported-storage-backends "Direct link to Supported storage backends") | Backend | URL scheme | Environment variables | | -------------------- | ------------------------- | --------------------------------------------------------------------------------------------------------------------------- | | Amazon S3 | `s3://` | Standard AWS credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, etc.) | | Azure ADLS Gen2 | `abfss://`, `abfs://` | `AZURE_STORAGE_ACCOUNT_NAME`, `AZURE_STORAGE_ACCOUNT_KEY`, `AZURE_CLIENT_ID`/`AZURE_TENANT_ID`/`AZURE_FEDERATED_TOKEN_FILE` | | Google Cloud Storage | `gs://` | `GOOGLE_APPLICATION_CREDENTIALS`, Workload Identity | | Local filesystem | Absolute or relative path | N/A | When the location is an S3 bucket, the configuration accepts any [S3 dataset parameters](/docs/next/components/data-connectors/s3) under `params`. Azure and GCS locations also accept their respective connector parameters under `params` for explicit credential overrides. When no explicit credentials are supplied, Spice reads standard environment variables for each cloud provider. ### Failure behavior[​](#failure-behavior "Direct link to Failure behavior") `bootstrap_on_failure_behavior` controls what Spice does when it cannot load the most recent snapshot. * `retry` – keep retrying the newest snapshot until it succeeds. * `fallback` – try older snapshot files until one loads successfully. * `warn` – log a warning and continue with an empty acceleration. (Default.) ## Enable snapshots per dataset[​](#enable-snapshots-per-dataset "Direct link to Enable snapshots per dataset") Each dataset opts into snapshotting through the `acceleration.snapshots` field. Four modes are available: * `enabled` – download snapshots on startup and write a new snapshot after each refresh. * `bootstrap_only` – only download snapshots; never write new ones. * `create_only` – write new snapshots after refreshes, but never download them on startup. * `disabled` – disable snapshot usage for this dataset. (Default.) Complete configuration: ``` acceleration: snapshots: enabled | disabled # default: disabled snapshots_trigger: # see trigger modes below snapshots_trigger_threshold: # threshold for time_interval or stream_batches snapshots_compaction: enabled | disabled # default: disabled (DuckDB only) snapshots_reset_expiry_on_load: enabled | disabled # default: disabled (DuckDB only with Caching refresh mode) ``` ### Snapshot triggers[​](#snapshot-triggers "Direct link to Snapshot triggers") The `snapshots_trigger` setting controls when Spice creates new snapshots. The available triggers depend on the dataset's refresh mode. #### Batch-based datasets[​](#batch-based-datasets "Direct link to Batch-based datasets") Datasets using `refresh_mode: full`, `refresh_mode: caching`, or `refresh_mode: append` with a `time_column` support the following triggers: | Trigger | Description | | ------------------ | --------------------------------------------------------------- | | `refresh_complete` | Create a snapshot after each data refresh completes. (Default.) | | `time_interval` | Create snapshots at a fixed time interval. | Example with default trigger: ``` datasets: - from: s3://some_bucket/some_table/ name: some_table acceleration: enabled: true engine: duckdb mode: file snapshots: enabled # snapshots_trigger defaults to refresh_complete params: duckdb_file: /nvme/some_table.db ``` Example with time-based trigger: ``` datasets: - from: s3://some_bucket/some_table/ name: some_table acceleration: enabled: true engine: duckdb mode: file snapshots: enabled snapshots_trigger: time_interval snapshots_trigger_threshold: 30m params: duckdb_file: /nvme/some_table.db ``` #### Stream-based datasets[​](#stream-based-datasets "Direct link to Stream-based datasets") Datasets using `refresh_mode: changes`, or `refresh_mode: append` without a `time_column`, support the following triggers: | Trigger | Description | | ---------------- | -------------------------------------------------------------------- | | `time_interval` | Create snapshots at a fixed time interval. (Default: 10m.) | | `stream_batches` | Create a snapshot after a specified number of batches are processed. | Example with time-based trigger (default): ``` datasets: - from: debezium:cdc_source name: cdc_table acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes snapshots: enabled # snapshots_trigger defaults to time_interval # snapshots_trigger_threshold defaults to 10m params: duckdb_file: /nvme/cdc_table.db ``` Example with batch-based trigger: ``` datasets: - from: debezium:cdc_source name: cdc_table acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes snapshots: enabled snapshots_trigger: stream_batches snapshots_trigger_threshold: 300 params: duckdb_file: /nvme/cdc_table.db ``` ### Snapshot compaction[​](#snapshot-compaction "Direct link to Snapshot compaction") For DuckDB-based accelerations, enable `snapshots_compaction` to compact the database before uploading. This uses DuckDB's internal mechanism (`COPY DATABASE`) to reduce file size and improve read performance. ``` acceleration: enabled: true engine: duckdb mode: file snapshots: enabled snapshots_compaction: enabled params: duckdb_file: /nvme/some_table.db ``` info Compaction is only available for the DuckDB acceleration engine. ### Snapshot Resetting Expiry on Load[​](#snapshot-resetting-expiry-on-load "Direct link to Snapshot Resetting Expiry on Load") When using `Caching` refresh mode with DuckDB-based acceleration, you can enable `snapshots_reset_expiry_on_load` to extend the data's expiry to `now() + TTL` each time a snapshot is loaded. ``` acceleration: enabled: true engine: duckdb mode: file refresh_mode: caching snapshots: enabled snapshots_reset_expiry_on_load: enabled params: caching_ttl: 1m caching_stale_while_revalidate_ttl: 1m ``` ## Complete example[​](#complete-example "Direct link to Complete example") ``` snapshots: enabled: true location: s3://some_bucket/some_folder/ bootstrap_on_failure_behavior: warn params: s3_auth: iam_role datasets: # Batch dataset with refresh-triggered snapshots - from: s3://some_bucket/batch_table/ name: batch_table params: file_format: parquet s3_auth: iam_role acceleration: enabled: true engine: duckdb mode: file snapshots: enabled snapshots_trigger: refresh_complete snapshots_compaction: enabled params: duckdb_file: /nvme/batch_table.db # Stream dataset with time-interval snapshots - from: debezium:cdc_source name: stream_table acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes snapshots: enabled snapshots_trigger: time_interval snapshots_trigger_threshold: 5m params: duckdb_file: /nvme/stream_table.db ``` Readiness with append refreshes Append-mode accelerations that define a `time_column` wait to report ready until the first append refresh completes after snapshot bootstrap. This keeps the dataset out of rotation until the freshest data is available while still benefiting from the snapshot-assisted startup. See [Fast Cold Starts](/docs/next/features/data-acceleration/data-refresh#fast-cold-starts-with-snapshots) for additional context. ## Best practices[​](#best-practices "Direct link to Best practices") * **Pair with ephemeral storage:** Deployments commonly place the acceleration file on fast ephemeral disks (such as NVMe instance storage) while relying on snapshots for persistence across restarts. * **Enable compaction for large datasets:** Use `snapshots_compaction: enabled` for DuckDB accelerations to reduce snapshot size and improve bootstrap performance. * **Tune trigger thresholds for stream datasets:** For high-throughput streaming datasets, balance snapshot frequency against I/O overhead by adjusting `snapshots_trigger_threshold`. * **Align retention policies:** Apply an object storage lifecycle rule that mirrors the desired snapshot retention policy. * **Monitor bootstraps:** Track warning logs emitted when Spice falls back to an empty acceleration so operators can respond quickly if snapshot loading fails. For the full reference, see [`snapshots` in the Spicepod specification](/docs/next/reference/spicepod#snapshots) and [`acceleration.snapshots`](/docs/next/reference/spicepod/datasets#accelerationsnapshots). For the production deployment pattern that uses snapshots to separate ingest from read workloads, see [Read/Write Separation](/docs/next/deployment/read-write-separation). Limitations * Only datasets are supported for snapshots. Views are not supported. --- # Data Ingestion Data can be ingested into the Spice runtime using the following methods: 1. **Acceleration Refresh Modes** – Pull data from a source connector into a local accelerator using one of the standard [refresh modes](/docs/next/features/data-acceleration/refresh-modes) (`full`, `append`, `changes`, `snapshot`, `caching`). This is the most common ingestion path for keeping a local accelerator in sync with an upstream system. 2. **SQL Statements** – Write data directly to [write-capable connectors](/docs/next/tags/write) using standard SQL `INSERT` (and, where supported, `UPDATE`/`DELETE`) syntax. 3. **OpenTelemetry (OTEL) Ingestion** – Stream OTEL metrics for real-time processing and acceleration. Data ingestion is useful for scenarios such as keeping a local accelerator continuously in sync with an upstream database, collecting metrics from edge devices, writing application events for later analysis, or populating datasets from external sources. ## Ingestion via Acceleration Refresh Modes[​](#ingestion-via-acceleration-refresh-modes "Direct link to Ingestion via Acceleration Refresh Modes") When a dataset is configured with `acceleration.enabled: true`, Spice ingests rows from the source connector into a local engine (Arrow, DuckDB, SQLite, PostgreSQL, or Cayenne). The `refresh_mode` controls how that ingestion happens. | Refresh Mode | What it ingests | Typical source | | -------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- | | [`full`](/docs/next/features/data-acceleration/refresh-modes/full) | Replaces the accelerator's contents with a fresh read of the source on every refresh. | Slowly-changing reference tables; small lookup datasets. | | [`append`](/docs/next/features/data-acceleration/refresh-modes/append) | Inserts only rows newer than the highest seen `time_column` value on each refresh. | Time-series, event/log data, append-only tables. | | [`changes`](/docs/next/features/data-acceleration/refresh-modes/changes) | Streams row-level inserts, updates, and deletes from a source [CDC feed](/docs/next/features/cdc) (PostgreSQL logical replication, DynamoDB Streams, MongoDB Change Streams, Debezium, Kafka, etc.). | Operational databases where you need near real-time mirror of the source. | | [`snapshot`](/docs/next/features/data-acceleration/refresh-modes/snapshot) | Loads exclusively from an external snapshot store; no source reads. | Read-only replicas bootstrapped from a centralized snapshot, e.g. for fan-out reader fleets. | | [`caching`](/docs/next/features/data-acceleration/refresh-modes/caching) | Read-through caches per-request HTTP/HTTPS responses with a TTL. | API search results or other request-keyed content fetched lazily. | For cross-cutting refresh behavior — refresh intervals, on-demand refresh, retries, retention, and zero-results handling — see [Data Refresh](/docs/next/features/data-acceleration/data-refresh). ### Example: continuous CDC ingestion into an accelerator[​](#example-continuous-cdc-ingestion-into-an-accelerator "Direct link to Example: continuous CDC ingestion into an accelerator") ``` datasets: - from: postgres:public.users name: users params: pg_host: pg.internal pg_port: '5432' pg_user: spice pg_pass: ${secrets:pg_pass} pg_db: myapp acceleration: enabled: true engine: duckdb mode: file # Persistence so resume across restarts is cheap refresh_mode: changes ``` This uses [PostgreSQL Logical Replication](/docs/next/features/cdc/postgres-replication) to ingest every `INSERT`, `UPDATE`, and `DELETE` from `public.users` into a local DuckDB accelerator with low latency. ## SQL Statements[​](#sql-statements "Direct link to SQL Statements") Spice supports writing data to **compatible data connectors** using standard SQL `INSERT INTO` syntax. ### Write-Capable Connectors[​](#write-capable-connectors "Direct link to Write-Capable Connectors") Data connectors that support write operations are tagged as [write](/docs/next/tags/write): * **[Apache Iceberg](/docs/next/components/data-connectors/iceberg)** - Write to Iceberg tables via data connector or [catalog connector](/docs/next/components/catalogs/iceberg) * **[AWS Glue](/docs/next/components/data-connectors/glue)** - Write to Glue Data Catalog tables via data connector or [catalog connector](/docs/next/components/catalogs/glue) ### Configuration for Write Operations[​](#configuration-for-write-operations "Direct link to Configuration for Write Operations") To enable write operations, configure your dataset or catalog with [read\_write access](/docs/next/reference/spicepod/datasets#access): ``` datasets: - from: glue:my_catalog.my_schema.my_table name: my_table access: read_write params: # ... connector-specific parameters ``` ### Example SQL[​](#example-sql "Direct link to Example SQL") ``` INSERT INTO my_table (column1, column2) VALUES ('value1', 'value2'); INSERT INTO my_table (column1, column2) SELECT source_column1, source_column2 FROM source_table WHERE condition = 'filter'; ``` For more details on the `INSERT` statement syntax, see the [SQL INSERT documentation](/docs/next/reference/sql/dml#insert). ## OpenTelemetry Data Ingestion[​](#opentelemetry-data-ingestion "Direct link to OpenTelemetry Data Ingestion") By default, the runtime exposes an [OpenTelemetry](https://opentelemetry.io) (OTEL) endpoint at grpc://127.0.0.1:50051 for the OTEL data ingestion. OTEL metrics will be inserted into datasets with matching names (metric name = dataset name) and optionally replicated to the dataset source. ### Supported metric types[​](#supported-metric-types "Direct link to Supported metric types") | OTLP metric type | Supported | Notes | | ---------------------- | --------- | ---------------------------------------------------------------------------------------------------------- | | `Gauge` | Yes | Ingested as number data points. | | `Sum` | Yes | Ingested as number data points. | | `Histogram` | Yes | Ingested with explicit bucket bounds and counts. | | `ExponentialHistogram` | No | Not written. Its data points are counted as rejected, and an unsupported metric data type error is logged. | | `Summary` | No | Not written. Its data points are counted as rejected, and an unsupported metric data type error is logged. | ### Export outcomes[​](#export-outcomes "Direct link to Export outcomes") Data points that cannot be written — a metric with no matching writable dataset, an unsupported metric type, or a batch that fails to build — are counted and reported back to the exporter: | Export | Response | | ---------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- | | No data points rejected | Success. | | Some data points rejected | Success with `ExportMetricsPartialSuccess`, carrying `rejected_data_points` and the message `Some data points were rejected`. | | Every data point rejected | Fails with `INVALID_ARGUMENT` and the message `All data points were rejected`. | | No data points at all (e.g. only metrics carrying no data) | Success — nothing was rejected. | A metric whose existing table has more than one column of the same name (Arrow permits duplicate field names, so an attribute sharing a name with one of the metric's value columns can produce one) can never be matched by a new batch. Rather than warning on every export, the ingest fails the export with an error naming the metric, the duplicated column, and the remedy: drop and recreate that metric's dataset. ### Ingested schema[​](#ingested-schema "Direct link to Ingested schema") `Gauge` and `Sum` metrics produce the following columns: | Column | Type | Description | | ---------------------- | -------------------- | --------------------------------------------------------------------------------------------------------- | | `value` | `Int64` or `Float64` | The data point value. The type is fixed by the first data point (or the existing table's `value` column). | | `time_unix_nano` | `UInt64` | Data point timestamp, in nanoseconds since the Unix epoch. | | `start_time_unix_nano` | `UInt64` | Start of the data point's aggregation interval, in nanoseconds since the Unix epoch. | `Histogram` metrics have a fixed set of value columns instead of `value`: | Column | Type | Description | | ----------------- | --------------- | ------------------------------------------------------------------------- | | `count` | `UInt64` | Number of values in the population. | | `sum` | `Float64` | Sum of the values. `NULL` when the exporter does not record it. | | `min` / `max` | `Float64` | Extrema over the interval. `NULL` when the exporter does not record them. | | `bucket_counts` | `List` | Per-bucket counts. | | `explicit_bounds` | `List` | The explicit bucket boundaries — one fewer element than `bucket_counts`. | Histograms carry the same `time_unix_nano` and `start_time_unix_nano` columns as number data points. Data point attributes become additional columns named after the attribute key, typed by the attribute value (`Utf8`, `Boolean`, `Int64`, `Float64`, or `Binary`). When a metric starts reporting a new attribute, Spice evolves an accelerated table's schema in place before writing, subject to the dataset's [`on_schema_change`](/docs/next/reference/spicepod/datasets#on_schema_change) policy. Because OTLP timestamps are nanoseconds, set [`time_format: unix_nanos`](/docs/next/reference/spicepod/datasets#time_format) when using `time_unix_nano` as a dataset's `time_column`: ``` datasets: - from: spice.ai/coolorg/metrics/datasets/http_server_duration name: http_server_duration # must match the OTEL metric name access: read_write time_column: time_unix_nano time_format: unix_nanos acceleration: enabled: true ``` ## Benefits[​](#benefits "Direct link to Benefits") Spice.ai OSS includes built-in data ingestion support for collecting the latest data from edge nodes for use in subsequent queries. This feature eliminates the need for additional ETL pipelines and improves the speed of the feedback loop. For example, consider CPU usage anomaly detection. When CPU metrics are sent to the Spice OpenTelemetry endpoint, they are immediately queryable alongside accelerated data, so an application can compare the most recent observations against historical baselines and act on the edge node. This process occurs quickly on the edge itself, within milliseconds, and without generating network traffic. Additionally, Spice will periodically replicate the data to the data connector for further use. ## Considerations[​](#considerations "Direct link to Considerations") Data Quality: Use Spice SQL capabilities to transform and cleanse ingested edge data, ensuring high-quality inputs. Data Security: Evaluate data sensitivity and secure network connections between the edge and data connector when replicating data for further use. Implement encryption, access controls, and secure protocols. ## Example[​](#example "Direct link to Example") ### [Disk SMART](https://en.wikipedia.org/wiki/Self-Monitoring,_Analysis_and_Reporting_Technology)[​](#disk-smart "Direct link to disk-smart") Start Spice with the following dataset: ``` datasets: - from: spice.ai/coolorg/smart/datasets/drive_stats name: smart_attribute_raw_value access: read_write replication: enabled: true acceleration: enabled: true ``` Start telegraf with the following config: ``` [[inputs.smart]] attributes = true [[outputs.opentelemetry]] service_address = "localhost:50051" [agent] interval = "1s" flush_interval = "1s" ``` SMART data will be available in the `smart_attribute_raw_value` dataset in Spice.ai OSS and replicated to the `coolorg.smart.drive_stats` dataset in Spice.ai Cloud. ## Limitations[​](#limitations "Direct link to Limitations") Current Limitations * Write Support: Only selected [write-capable connectors and catalogs](/docs/next/tags/write) support write operations. * Only Spice.ai replication is supported for OpenTelemetry ingestion --- # Distributed Query Learn how to configure and run Spice in distributed mode to handle larger scale queries across multiple nodes. Preview Multi-node distributed query execution based on Apache Ballista is available as a preview feature in Spice `v1.9.0`. ## Overview[​](#overview "Direct link to Overview") Spice integrates [Apache Ballista](https://github.com/apache/datafusion-ballista) to schedule and coordinate distributed queries across multiple executor nodes. This integration is useful when querying large, partitioned datasets in data lake formats such as Parquet, Delta Lake, or Iceberg. For smaller workloads or non-partitioned data, a single Spice instance is typically sufficient. ## Architecture[​](#architecture "Direct link to Architecture") A distributed Spice cluster consists of two components: * **Scheduler** – Plans distributed queries and manages the work queue for the executor fleet. Also manages async query jobs when `scheduler.state_location` is configured. * **Executors** – One or more nodes responsible for executing physical query plans. The scheduler holds the cluster-wide configuration for a Spicepod, while executors connect to the scheduler to receive work. A cluster can run with a single scheduler for simplicity, or multiple schedulers for [high availability](#high-availability). ### Network Ports[​](#network-ports "Direct link to Network Ports") Spice separates public and internal cluster traffic across different ports: | Port | Service | Description | | ----- | --------------- | --------------------------------------------------------------------- | | 50051 | Flight SQL | Public query endpoint | | 8090 | HTTP API | Public REST API | | 9090 | Prometheus | Metrics endpoint | | 50052 | Cluster Service | Internal scheduler/executor communication (mTLS enforced, by default) | Internal cluster services are isolated on port 50052 with mTLS enforced by default. ## Secure Cluster Communication (mTLS)[​](#secure-cluster-communication-mtls "Direct link to Secure Cluster Communication (mTLS)") Distributed query cluster mode uses mutual TLS (mTLS) for secure communication between schedulers and executors. Internal cluster communication includes highly privileged RPC calls like fetching Spicepod configuration and expanding secrets. mTLS ensures only authenticated nodes can join the cluster and access sensitive data. ### Certificate Requirements[​](#certificate-requirements "Direct link to Certificate Requirements") Each node in the cluster requires: * A CA certificate (`ca.crt`) trusted by all nodes * A node certificate with the node's advertise address in the Subject Alternative Names (SANs) * A private key for the node certificate Production deployments should use the organization's PKI infrastructure to generate certificates with proper SANs for each node. ### Development Certificates[​](#development-certificates "Direct link to Development Certificates") For local development and testing, the Spice CLI provides commands to generate self-signed certificates: ``` # Initialize CA and generate CA certificate spice cluster tls init # Generate certificate for the scheduler node spice cluster tls add scheduler1 # Generate certificate for an executor node spice cluster tls add executor1 ``` Certificates are stored in `~/.spice/pki/` by default. warning CLI-generated certificates are not intended for production use. Production deployments should use certificates issued by the organization's PKI or a trusted certificate authority. ### Insecure Mode[​](#insecure-mode "Direct link to Insecure Mode") For local development and testing, mTLS can be disabled using the `--allow-insecure-connections` flag: ``` spiced --role scheduler --allow-insecure-connections ``` warning Do not use `--allow-insecure-connections` in production environments. This flag disables authentication and encryption for internal cluster communication. ## Getting Started[​](#getting-started "Direct link to Getting Started") Cluster deployment typically starts with a scheduler instance, followed by one or more executors that register with the scheduler. The following examples use CLI-generated development certificates. For production, substitute certificates from your organization's PKI. ### Generate Development Certificates[​](#generate-development-certificates "Direct link to Generate Development Certificates") ``` spice cluster tls init spice cluster tls add scheduler1 spice cluster tls add executor1 ``` ### Start the Scheduler[​](#start-the-scheduler "Direct link to Start the Scheduler") The scheduler is the only `spiced` process that needs to be configured (i.e. have a `spicepod.yaml` in the current dir). Override the Flight bind address when it must be reachable outside of `localhost`: ``` spiced --role scheduler \ --flight 0.0.0.0:50051 \ --node-mtls-ca-certificate-file ~/.spice/pki/ca.crt \ --node-mtls-certificate-file ~/.spice/pki/scheduler1.crt \ --node-mtls-key-file ~/.spice/pki/scheduler1.key ``` ### Start Executors[​](#start-executors "Direct link to Start Executors") Executors connect to the scheduler's internal cluster port (50052) to register and pull work. Executors do not require a `spicepod.yaml`; they fetch the configuration from the scheduler. Each executor automatically selects a free port if the default is unavailable: ``` spiced --role executor \ --scheduler-address https://scheduler1:50052 \ --node-mtls-ca-certificate-file ~/.spice/pki/ca.crt \ --node-mtls-certificate-file ~/.spice/pki/executor1.crt \ --node-mtls-key-file ~/.spice/pki/executor1.key ``` Specifying `--scheduler-address` implies `--role executor`. ## Query Execution[​](#query-execution "Direct link to Query Execution") Queries run against the scheduler endpoint. The `EXPLAIN` output confirms that distributed planning is active—Spice includes a `distributed_plan` section showing how stages are split across executors: ``` EXPLAIN SELECT count(id) FROM my_dataset; ``` Limitations * In open source, distributed query targets partitioned data lake sources (e.g. Parquet, Delta Lake, Iceberg). Distributing **accelerated** datasets across executors is a Spice.ai Enterprise feature (see below). * As a preview feature, clusters may encounter stability or performance issues. Spice.ai Enterprise [**Distributed Accelerations**](https://docs.spice.ai/docs/enterprise/features/distributed-accelerations) partition-shard accelerated datasets across executors — each executor materializes and serves only the partitions it owns, with partition-aware read routing and write-through semantics. This is available in [Spice.ai Enterprise](https://spice.ai) and requires a [`SpicepodCluster`](https://docs.spice.ai/docs/enterprise/kubernetes-operator/spicepodcluster). ## Async Queries API[​](#async-queries-api "Direct link to Async Queries API") For long-running queries, the async queries API enables submitting queries for background execution, polling for status, and retrieving paginated results when ready. warning The async queries API is experimental and requires `scheduler.state_location` to be configured. ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Spice runtime running in cluster mode with `--role scheduler` * `scheduler.state_location` configured in the Spicepod (see [High Availability > Configuration](#configuration)) * At least one executor node connected to the scheduler ### Enabling Async Queries[​](#enabling-async-queries "Direct link to Enabling Async Queries") Configure `runtime.scheduler.state_location` in your `spicepod.yaml` to enable the async queries API: ``` runtime: scheduler: state_location: s3://my-bucket/spice-state params: region: us-east-1 ``` The state location is a shared object store (S3, GCS, Azure Blob, or local filesystem via `file://`) used to persist async query job state and result chunks. For local development: ``` runtime: scheduler: state_location: "file://.data/scheduler-state" ``` ### Query Ownership[​](#query-ownership "Direct link to Query Ownership") A query is owned by the principal that submitted it, and **only that principal can list, poll, fetch results from, or cancel it**. This applies to the HTTP, Arrow Flight, and CLI surfaces alike. Ownership follows the same principal boundary as [Per-Principal Cache Isolation](/docs/next/features/caching#per-principal-cache-isolation), and is not configurable: | Caller | Owner scope | | ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ | | Unauthenticated (or [authentication](/docs/next/api/auth) not enabled) | A single shared `public` scope — every caller sees every query, matching pre-isolation behavior. | | Authenticated principal (e.g., API key) | The principal's own isolated scope. | | Internal / background runtime tasks | An internal `system` scope that can never reach a user-submitted query. | There is no administrator or organization-wide scope. Every authenticated principal sees only its own queries; no principal can list, poll, or cancel another's, and there is no role or setting that grants a wider view. For a query it does not own, a caller always receives **404 Not Found** — never 403, and never 410. Ownership is checked before result expiry, so the response does not reveal that the query exists: | Caller | Query is running or complete | Query's results have expired | | ------------------- | ---------------------------- | ---------------------------- | | The owner | `200 OK` | `410 Gone` | | Any other principal | `404 Not Found` | `404 Not Found` | Ownership tracking was introduced in **v2.2.0**. A job written by an earlier runtime carries no owner and is treated as belonging to the `public` scope. ### HTTP REST API[​](#http-rest-api "Direct link to HTTP REST API") Base path: `/v1/queries` #### Endpoints[​](#endpoints "Direct link to Endpoints") | Method | Path | Description | | ------ | ----------------------------------------------------- | --------------------------------------- | | `POST` | `/v1/queries` | Submit a query for async execution | | `GET` | `/v1/queries` | List the caller's queries | | `GET` | `/v1/queries/{query_id}` | Get query status and first result chunk | | `GET` | `/v1/queries/{query_id}/status` | Get query status only | | `GET` | `/v1/queries/{query_id}/results` | Get results (with pagination) | | `GET` | `/v1/queries/{query_id}/results/chunks/{chunk_index}` | Get a specific result chunk | | `POST` | `/v1/queries/{query_id}/cancel` | Cancel a running query | #### Submit Query[​](#submit-query "Direct link to Submit Query") `POST /v1/queries` Submits a SQL query for asynchronous execution and returns immediately with a job ID. **Request Body** (`application/json`): | Field | Type | Required | Description | | ----------------- | ------- | -------- | -------------------------------------------------------------------------------- | | `sql` | string | Yes | SQL statement to execute | | `parameters` | array | No | Bind variables for parameterized queries (`$1`, `$2`, ...) | | `timeout_seconds` | integer | No | Maximum execution time in seconds. The query is cancelled and failed on timeout. | | `maximum_size` | integer | No | Maximum result size in bytes. The query is failed if results exceed this limit. | **Request Example**: ``` { "sql": "SELECT * FROM large_table WHERE status = $1 AND created_at > $2", "parameters": ["active", "2025-01-01"], "timeout_seconds": 300, "maximum_size": 104857600 } ``` **Response** (HTTP 202 Accepted): ``` { "query_id": "01ABC-DEF-456-7890AB", "status": "PENDING", "error": null, "status_url": "/v1/queries/01ABC-DEF-456-7890AB/status", "results_url": "/v1/queries/01ABC-DEF-456-7890AB/results" } ``` #### Get Query[​](#get-query "Direct link to Get Query") `GET /v1/queries/{query_id}` Returns the full query status, result manifest, and the first result chunk (if completed successfully). **Response** (HTTP 200): ``` { "query_id": "01ABC-DEF-456-7890AB", "status": "SUCCEEDED", "error": null, "manifest": { "format": "ARROW_IPC", "schema": { "column_count": 3, "columns": [ { "name": "id", "type_name": "Int64", "nullable": false, "position": 0 }, { "name": "status", "type_name": "Utf8", "nullable": true, "position": 1 }, { "name": "created_at", "type_name": "Timestamp(Microsecond, Some(\"UTC\"))", "nullable": true, "position": 2 } ] }, "total_row_count": 25000, "total_chunk_count": 3 }, "result": { "chunk_index": 0, "row_offset": 0, "row_count": 10000, "next_chunk_index": 1, "next_chunk_url": "/v1/queries/01ABC-DEF-456-7890AB/results/chunks/1", "data_array": [ { "id": 1, "status": "active", "created_at": "2025-06-15T10:30:00Z" } ] }, "created_at": "2026-03-02T12:00:00+00:00", "started_at": "2026-03-02T12:00:00.050+00:00", "completed_at": "2026-03-02T12:00:05.200+00:00", "expires_at": "2026-03-03T00:00:05.200+00:00" } ``` #### Get Status[​](#get-status "Direct link to Get Status") `GET /v1/queries/{query_id}/status` Returns the current status of a query without result data. Use this for lightweight polling. **Response** (HTTP 200): ``` { "status": "RUNNING", "error": null } ``` When the query has failed: ``` { "status": "FAILED", "error": { "error_code": "EXECUTION_FAILED", "message": "Table 'missing_table' not found", "sql_state": null } } ``` #### Get Results[​](#get-results "Direct link to Get Results") `GET /v1/queries/{query_id}/results` Returns result data for a completed query. Use the `partition` query parameter to paginate through chunks. **Query Parameters**: | Parameter | Type | Default | Description | | ----------- | ------- | ------- | --------------------------------- | | `partition` | integer | `0` | Chunk index to retrieve (0-based) | **Response** (HTTP 200): ``` { "chunk_index": 0, "row_offset": 0, "row_count": 10000, "next_chunk_index": 1, "next_chunk_url": "/v1/queries/01ABC-DEF-456-7890AB/results/chunks/1", "data_array": [ { "id": 1, "status": "active" } ] } ``` When the last chunk is reached, `next_chunk_index` and `next_chunk_url` are `null`. #### Get Chunk[​](#get-chunk "Direct link to Get Chunk") `GET /v1/queries/{query_id}/results/chunks/{chunk_index}` Returns a specific result chunk by index. Same response format as **Get Results**. #### Cancel Query[​](#cancel-query "Direct link to Cancel Query") `POST /v1/queries/{query_id}/cancel` Cancels a running query. Also cancels the underlying distributed query on the Ballista scheduler. Only the principal that submitted the query can cancel it; any other caller receives 404 Not Found and the query keeps running — see [Query Ownership](#query-ownership). **Response** (HTTP 200): ``` { "query_id": "01ABC-DEF-456-7890AB", "status": "CANCELLED", "error": null, "manifest": null, "result": null, "created_at": "2026-03-02T12:00:00+00:00", "started_at": "2026-03-02T12:00:00.050+00:00", "completed_at": "2026-03-02T12:00:02.100+00:00", "expires_at": null } ``` #### List Queries[​](#list-queries "Direct link to List Queries") `GET /v1/queries` Lists the queries submitted by the calling principal, optionally filtered by status. Queries submitted by other principals are never listed — see [Query Ownership](#query-ownership). **Query Parameters**: | Parameter | Type | Default | Description | | --------- | ------- | ------- | --------------------------------------------------------------------------------------------------------- | | `status` | string | *all* | Filter by status: `queued`/`pending`, `running`, `completed`/`succeeded`, `failed`, `cancelled`, `closed` | | `limit` | integer | `100` | Maximum number of results | **Response** (HTTP 200): ``` { "queries": [ { "query_id": "01ABC-DEF-456-7890AB", "status": "RUNNING", "sql_preview": "SELECT * FROM large_table WHERE status = ...", "created_at": "2026-03-02T12:00:00+00:00" } ], "total_count": 1 } ``` The `sql_preview` field contains the first 100 characters of the SQL statement. #### HTTP Error Responses[​](#http-error-responses "Direct link to HTTP Error Responses") | HTTP Status | Condition | | ------------------------- | ----------------------------------------------------------------------------------------- | | 202 Accepted | Query successfully submitted | | 200 OK | Status/results retrieved successfully | | 404 Not Found | Query ID, chunk, or result not found, or the query was submitted by a different principal | | 409 Conflict | Query not yet complete (when fetching results by chunk) | | 410 Gone | Query results have expired | | 425 Too Early | Query still running (results endpoint) | | 500 Internal Server Error | Execution or serialization failure | | 503 Service Unavailable | Not running in scheduler cluster mode, or executor not yet initialized | ### Arrow Flight API[​](#arrow-flight-api "Direct link to Arrow Flight API") The async query API is also available via Apache Arrow Flight `DoAction` requests. This is more efficient for programmatic access since results are returned in Arrow IPC binary format instead of JSON. Every action below is scoped to the principal that submitted the query, exactly as the HTTP endpoints are — see [Query Ownership](#query-ownership). | Action Type | Request Body (JSON) | Response | | --------------------- | --------------------------------------- | --------------------------------------------------------------------- | | `SubmitAsyncQuery` | `{"sql": "...", "parameters": [...]}` | JSON: `{"query_id": "...", "status": "PENDING"}` | | `GetAsyncQueryStatus` | `{"query_id": "..."}` | JSON: query status with error/result metadata | | `GetAsyncQueryResult` | `{"query_id": "...", "chunk_index": 0}` | Binary: Arrow IPC stream | | `CancelAsyncQuery` | `{"query_id": "..."}` | JSON: `{"query_id": "...", "cancelled": true, "status": "CANCELLED"}` | #### SubmitAsyncQuery[​](#submitasyncquery "Direct link to SubmitAsyncQuery") **Request**: ``` { "sql": "SELECT * FROM large_table", "parameters": [] } ``` **Response** (JSON): ``` { "query_id": "01ABC-DEF-456-7890AB", "status": "PENDING" } ``` #### GetAsyncQueryStatus[​](#getasyncquerystatus "Direct link to GetAsyncQueryStatus") **Request**: ``` { "query_id": "01ABC-DEF-456-7890AB" } ``` **Response** (JSON): ``` { "query_id": "01ABC-DEF-456-7890AB", "status": "SUCCEEDED", "error": null, "result": { "total_row_count": 25000, "total_chunk_count": 3 } } ``` #### GetAsyncQueryResult[​](#getasyncqueryresult "Direct link to GetAsyncQueryResult") **Request**: ``` { "query_id": "01ABC-DEF-456-7890AB", "chunk_index": 0 } ``` **Response**: Arrow IPC binary stream containing the `RecordBatch` data for the requested chunk. #### CancelAsyncQuery[​](#cancelasyncquery "Direct link to CancelAsyncQuery") **Request**: ``` { "query_id": "01ABC-DEF-456-7890AB" } ``` **Response** (JSON): ``` { "query_id": "01ABC-DEF-456-7890AB", "cancelled": true, "status": "CANCELLED" } ``` ### CLI[​](#cli "Direct link to CLI") The `spice query` command provides a CLI and interactive REPL for the async queries API. #### Submit and Wait[​](#submit-and-wait "Direct link to Submit and Wait") ``` spice query "SELECT * FROM orders WHERE total > 100 LIMIT 50;" ``` The CLI auto-polls with a spinner and displays results when ready. Press `Ctrl+C` to stop waiting — the query continues running in the background. #### Submit Without Waiting[​](#submit-without-waiting "Direct link to Submit Without Waiting") ``` spice query "SELECT * FROM large_table;" --no-wait ``` #### Options[​](#options "Direct link to Options") | Option | Default | Description | | ----------------------- | ------- | ------------------------------------------------------------------------------------------------- | | `--no-wait` | `false` | Submit the query and return immediately without waiting for results | | `--timeout ` | *none* | Maximum client-side wait time (e.g., `30s`, `5m`). The query itself continues running on timeout. | | `-o, --output ` | `table` | Output format: `table` or `json` | #### Subcommands[​](#subcommands "Direct link to Subcommands") ``` spice query list [--status X] [--limit N] # List queries spice query status # Check query status spice query results # Fetch results of completed query spice query cancel # Cancel a running query ``` These subcommands reach only the queries submitted with the same credentials — see [Query Ownership](#query-ownership). #### Interactive REPL[​](#interactive-repl "Direct link to Interactive REPL") When invoked without arguments, `spice query` starts an interactive REPL: ``` query> SELECT COUNT(*) > FROM large_table > WHERE status = 'active'; Submitted query: 01ABC-DEF-456-7890AB (PENDING) Press Ctrl+C to stop waiting (query continues in background) ⠹ RUNNING (2.3s)... ✓ SUCCEEDED (5.1s) +----------+ | count(*) | +----------+ | 42000 | +----------+ Time: 5.10000000 seconds. 1 rows. ``` **REPL Commands**: | Command | Description | | ---------------------- | ---------------------------------------------- | | `.list` | List all queries tracked in this REPL session | | `.status ` | Show detailed status of a query | | `.results ` | Fetch and display results of a completed query | | `.wait ` | Resume waiting for a query to complete | | `.cancel ` | Cancel a running query | | `.clear` | Clear the local tracked queries list | | `.clear history` | Clear command history | | `.help` | Show available commands | | `.exit`, `.quit`, `.q` | Exit the REPL | Query IDs can be abbreviated if they uniquely identify a query within the tracked session. ### Job Lifecycle[​](#job-lifecycle "Direct link to Job Lifecycle") ``` PENDING → RUNNING → SUCCEEDED → CLOSED (after 12h TTL) → FAILED → CANCELLED ``` | Status | Description | | ----------- | ---------------------------------------------------- | | `PENDING` | Job is queued but not yet executing | | `RUNNING` | Job is actively executing on the distributed cluster | | `SUCCEEDED` | Job completed successfully, results are available | | `FAILED` | Job execution failed (see `error` field for details) | | `CANCELLED` | Job was cancelled by the user | | `CLOSED` | Job results have expired and been cleaned up | ### Error Codes[​](#error-codes "Direct link to Error Codes") When a query fails, the `error` object contains an `error_code` field: | Error Code | Description | | -------------------------- | ------------------------------------------------------- | | `SCHEDULER_UNAVAILABLE` | The Ballista scheduler is not reachable | | `SUBMISSION_FAILED` | Failed to submit the query to the distributed scheduler | | `EXECUTION_FAILED` | The query failed during execution | | `FETCHING_RESULTS_FAILED` | Failed to retrieve results from executor nodes | | `CANCELLED` | The query was explicitly cancelled | | `PARAMETER_BINDING_FAILED` | Failed to bind the provided query parameters | | `NOT_FOUND` | The referenced query or job was not found | | `INTERNAL` | An unexpected internal error occurred | | `TIMEOUT` | The query exceeded the configured `timeout_seconds` | ### Storage Layout[​](#storage-layout "Direct link to Storage Layout") Job state and result chunks are stored in the shared object store configured via `scheduler.state_location`: ``` {base_prefix}/ ├── jobs/ │ ├── {job_id}.json # Job state (JSON) │ └── {job_id}/ │ ├── chunk_0.arrow # Result chunk 0 (Arrow IPC) │ ├── chunk_1.arrow # Result chunk 1 │ └── ... ``` ### Defaults and Limitations[​](#defaults-and-limitations "Direct link to Defaults and Limitations") | Setting | Default | | ---------- | ----------- | | Chunk size | 10,000 rows | | Result TTL | 12 hours | | List limit | 100 queries | * Only available in cluster mode with `--role scheduler` * Requires `scheduler.state_location` to be configured * The `format` query parameter on the results endpoint is declared but not yet implemented (results are always JSON over HTTP, Arrow IPC over Flight) * Result TTL is not yet configurable per-query (fixed at 12 hours) * Chunk size is not yet configurable per-query (fixed at 10,000 rows) ## High Availability[​](#high-availability "Direct link to High Availability") For production deployments, Spice supports running multiple active schedulers in an active/active configuration. This eliminates the scheduler as a single point of failure and enables graceful handling of node failures. ### HA Architecture[​](#ha-architecture "Direct link to HA Architecture") In an HA cluster: * Multiple schedulers run simultaneously, each capable of accepting queries * Schedulers share state via an S3-compatible object store * Executors discover all schedulers automatically * A load balancer distributes client queries across schedulers ``` ┌─────────────────────┐ │ Load Balancer │ └─────────────────────┘ │ ┌────────────────┼────────────────┐ ▼ ▼ ▼ ┌────────────┐ ┌────────────┐ ┌────────────┐ │ Scheduler │ │ Scheduler │ │ Scheduler │◄──► Object Store │ │ │ │ │ │ (S3) └────────────┘ └────────────┘ └────────────┘ ▲ ▲ ▲ │ (executor-initiated) │ ┌────────────┐ ┌────────────┐ ┌────────────┐ │ Executor │ │ Executor │ │ Executor │ └────────────┘ └────────────┘ └────────────┘ ``` ### Configuration[​](#configuration "Direct link to Configuration") Enable HA by configuring `runtime.scheduler.state_location` in the Spicepod to point to an S3-compatible object store: ``` runtime: scheduler: state_location: s3://my-bucket/spice-cluster params: region: us-east-1 ``` The object store is used for scheduler registration and discovery, and to persist [async query](#async-queries-api) job state (the execution graph plus its status) so that schedulers are effectively stateless for async queries. ### Scheduler Failover[​](#scheduler-failover "Direct link to Scheduler Failover") When `runtime.scheduler.state_location` is configured, each async query's execution graph and status are persisted to the shared object store. If the scheduler driving an async query becomes unavailable, another scheduler detects the orphaned job and resumes it to completion from the persisted execution graph — the query is re-driven rather than replanned, and consumers and executors do not need to know which scheduler is running it. Takeover is single-winner: ownership transfers via a compare-and-set on the job's metadata, and a scheduler never reclaims its own in-flight jobs. This failover applies to async queries, which require `scheduler.state_location`. Synchronous queries in flight on a scheduler that becomes unavailable are not resumed automatically; the client should retry them against another scheduler. Without `scheduler.state_location`, job state is held in memory and a single-scheduler cluster behaves as before (no failover). ### S3 Configuration[​](#s3-configuration "Direct link to S3 Configuration") The `runtime.scheduler.params` section supports the following S3 parameters: | Parameter | Description | Default | | ------------------ | ----------------------------------------------- | ---------- | | `s3_region` | AWS region for the S3 bucket | - | | `s3_endpoint` | Custom S3-compatible endpoint URL | - | | `s3_auth` | Authentication method: `iam_role` or `key` | `iam_role` | | `s3_key` | AWS access key ID (when `auth: key`) | - | | `s3_secret` | AWS secret access key (when `auth: key`) | - | | `s3_session_token` | AWS session token for temporary credentials | - | | `client_timeout` | S3 client timeout | - | | `allow_http` | Allow HTTP (non-TLS) connections to S3 endpoint | `false` | Example with explicit credentials: ``` runtime: scheduler: state_location: s3://my-bucket/spice-cluster params: region: us-east-1 auth: key key: ${secrets:aws_access_key} secret: ${secrets:aws_secret_key} ``` ### Starting an HA Cluster[​](#starting-an-ha-cluster "Direct link to Starting an HA Cluster") 1. **Configure shared state** in `spicepod.yaml`: ``` runtime: scheduler: state_location: s3://my-bucket/spice-cluster params: region: us-east-1 ``` 2. **Start multiple schedulers**, each with unique certificates: ``` # Scheduler 1 spiced --role scheduler \ --flight 0.0.0.0:50051 \ --node-mtls-ca-certificate-file ~/.spice/pki/ca.crt \ --node-mtls-certificate-file ~/.spice/pki/scheduler1.crt \ --node-mtls-key-file ~/.spice/pki/scheduler1.key # Scheduler 2 (on a different node) spiced --role scheduler \ --flight 0.0.0.0:50051 \ --node-mtls-ca-certificate-file ~/.spice/pki/ca.crt \ --node-mtls-certificate-file ~/.spice/pki/scheduler2.crt \ --node-mtls-key-file ~/.spice/pki/scheduler2.key ``` 3. **Start executors** (they discover all schedulers automatically): ``` spiced --role executor \ --scheduler-address https://scheduler1:50052 \ --node-mtls-ca-certificate-file ~/.spice/pki/ca.crt \ --node-mtls-certificate-file ~/.spice/pki/executor1.crt \ --node-mtls-key-file ~/.spice/pki/executor1.key ``` 4. **Configure a load balancer** to distribute queries across scheduler Flight SQL endpoints (port 50051). ### HA Considerations[​](#ha-considerations "Direct link to HA Considerations") * **Object store latency** – The object store is accessed during scheduler coordination. Use a low-latency object store (e.g., S3 Express One Zone) for best performance. Object Store Requirements The object store must support conditional writes (S3 ETags). Most S3-compatible stores support this, including AWS S3, MinIO, and Google Cloud Storage (with S3 compatibility mode). --- # Embedding Datasets Learn how to define and augment datasets with embedding columns for advanced search capabilities. ## Overview[​](#overview "Direct link to Overview") Spice provides three distinct methods for handling embedding columns in datasets: 1. **[Just-in-Time (JIT) Embeddings](/docs/next/components/embeddings#jit-embeddings)**: Dynamically computes embeddings, on-demand, during query execution, without precomputing data. 2. **[Accelerated Embeddings](/docs/next/components/embeddings#accelerated-embeddings)**: Precomputes embeddings by transforming and augmenting the source dataset for faster query and search performance. 3. **[Passthrough Embeddings](/docs/next/components/embeddings#passthrough-embeddings)**: Uses pre-existing embeddings directly from the underlying source datasets, bypassing any additional computation. ## Configuring Embedding Models[​](#configuring-embedding-models "Direct link to Configuring Embedding Models") Before configuring dataset embeddings, define the embedding models in the `spicepod.yaml`. For example: ``` embeddings: - name: local_embedding_model from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 - from: openai name: remote_service params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } ``` See [Embedding components](/docs/next/components/embeddings) for more information on embedding models. ## Vector Searches[​](#vector-searches "Direct link to Vector Searches") Spice supports complex searches by utilizing embeddings. Both local and remote embedding models can be used for vector searches. To run a vector search, embeddings must be defined for the relevant columns in your dataset. Once configured, similarity searches can be performed using the defined embeddings. For detailed instructions and examples on running vector searches, refer to the [Vector-Based Search documentation](/docs/next/features/search/vector-search). ## Generating Embeddings in Queries[​](#generating-embeddings-in-queries "Direct link to Generating Embeddings in Queries") The [`embed()` scalar function](/docs/next/reference/sql/scalar_functions#ai-and-embed) generates embeddings directly within SQL queries. This function can process both single text strings and arrays of text, making it useful for ad-hoc embedding generation and comparison operations. --- # Functions Functions extend Spice's SQL engine with custom logic declared in your Spicepod. Functions can be **scalar** (one value per row) or **table** (returning multiple rows and columns). Each function can be: * **Called directly in SQL** like any built-in function (`SELECT my_fn(col) FROM ...`). * **Surfaced to LLMs as tools** for tool-calling workflows. * **Listed via SQL** with the `list_udfs()` UDTF and via the HTTP API at `GET /v1/functions`. Functions are declared in the top-level `functions:` block of `spicepod.yaml`. The full YAML reference is on the [Functions Spicepod reference](/docs/next/reference/spicepod/functions) page. ## Quickstart[​](#quickstart "Direct link to Quickstart") Enable functions and declare a SQL function: ``` runtime: functions: enabled: true # Required — registration is off by default functions: - name: double_it from: sql description: Double a 64-bit integer. volatility: immutable signature: args: - { name: x, type: int64 } returns: int64 body: 'x * 2' ``` Call it from SQL: ``` SELECT double_it(21); -- 42 ``` The function is automatically registered both as a SQL function and as a callable LLM tool (set `as_tool: false` to keep it SQL-only). ## Execution Tiers[​](#execution-tiers "Direct link to Execution Tiers") Spice supports two tiers for functions, selected by the `from:` field: | Tier | `from:` scheme | Where it runs | When to use | | ------ | --------------------- | ------------------------------------ | ----------------------------------------------------------------------------- | | SQL | `sql` | In-process, in the DataFusion engine | Pure expressions, math, string transforms, business logic over column values. | | Remote | `http://`, `https://` | A remote HTTP + JSON service | Custom logic in another language, ML inference, calls to internal APIs. | ### SQL tier (`from: sql`)[​](#sql-tier-from-sql "Direct link to sql-tier-from-sql") The function `body` is a single SQL expression evaluated against the function's arguments. It can call any DataFusion built-in function (math, string, datetime, JSON, regex, etc.). ``` functions: - name: haversine_km from: sql description: Haversine great-circle distance in kilometres. volatility: immutable signature: args: - { name: lat1, type: float64 } - { name: lon1, type: float64 } - { name: lat2, type: float64 } - { name: lon2, type: float64 } returns: float64 body: | 6371 * acos( cos(radians(lat1)) * cos(radians(lat2)) * cos(radians(lon2) - radians(lon1)) + sin(radians(lat1)) * sin(radians(lat2)) ) ``` ``` SELECT haversine_km(37.7749, -122.4194, 40.7128, -74.0060) AS km; -- 4129.085647... ``` #### External SQL files with `body_ref`[​](#external-sql-files-with-body_ref "Direct link to external-sql-files-with-body_ref") For non-trivial SQL, keep the body in its own file with proper editor support: ``` functions: - name: shipping_class from: sql signature: args: - { name: weight_g, type: int64 } - { name: country, type: utf8 } returns: utf8 body_ref: ./functions/shipping_class.sql ``` `body_ref` is read from the local filesystem at registration time. For portable spicepods loaded from object storage, use inline `body:` instead. ### Remote tier (`from: http://...`)[​](#remote-tier-from-http "Direct link to remote-tier-from-http") The runtime POSTs row batches to the configured endpoint and reads the resulting values. Use this tier to delegate logic to a service in another language, an ML model server, or an internal API. ``` functions: - name: classify_intent from: http://classifier.internal/v1/classify description: Classify a user prompt via a remote service. volatility: volatile signature: args: - { name: prompt, type: utf8 } returns: utf8 params: timeout: 5s batch_size: 256 batch_concurrency: 4 auth_bearer: ${secrets:CLASSIFIER_TOKEN} ``` ``` SELECT user_id, classify_intent(latest_message) AS intent FROM conversations WHERE date = today(); ``` #### Wire contract[​](#wire-contract "Direct link to Wire contract") The runtime sends a single HTTP `POST` per batch with `Content-Type: application/json`: **Request body:** ``` { "rows": [ { "prompt": "How do I cancel my subscription?" }, { "prompt": "What's the weather in Paris?" } ] } ``` **Response body** (HTTP `200`): ``` { "values": ["billing", "smalltalk"] } ``` `values.len()` must equal `rows.len()` — a mismatch is treated as an error. Each row contains every declared argument under its argument `name`. Output values are decoded into the declared `returns` Arrow type using Arrow's JSON reader. #### Remote `params:` knobs[​](#remote-params-knobs "Direct link to remote-params-knobs") | Parameter | Default | Description | | ------------------- | ------- | ------------------------------------------------------------------------------------------------- | | `timeout` | `30s` | Per-call timeout. Accepts plain integer seconds or `Ns` / `Nms` suffix strings. | | `batch_size` | `1024` | Maximum rows per HTTP request. Capped at `100 000`. | | `batch_concurrency` | `4` | Maximum in-flight HTTP batches per function invocation. Capped at `64`. | | `auth_bearer` | unset | When set, the runtime adds `Authorization: Bearer ` to each request. Use `${secrets:...}`. | Calls to remote functions require the runtime to be configured with [`runtime.auth.api-key`](/docs/next/reference/spicepod/runtime#runtimeauth) — they execute under the read-write API key context. ## Table Functions[​](#table-functions "Direct link to Table Functions") Table functions (`kind: table`) return multiple rows and columns instead of a single scalar value. They are called using standard SQL table-function syntax: `SELECT ... FROM my_table_fn(args)`. ### Declaring a table function[​](#declaring-a-table-function "Direct link to Declaring a table function") Set `kind: table` and provide `returns` as a list of output columns (instead of a single type string): ``` functions: - name: emit_pair from: sql kind: table description: Emit an input row and its successor. volatility: immutable signature: args: [{ name: x, type: int64 }] returns: - { name: value, type: int64 } - { name: doubled, type: int64 } body: | SELECT x AS value, x * 2 AS doubled FROM args UNION ALL SELECT x + 1 AS value, (x + 1) * 2 AS doubled FROM args ``` ``` SELECT value, doubled FROM emit_pair(4) ORDER BY value; -- value | doubled -- ------+-------- -- 4 | 8 -- 5 | 10 ``` ### Key differences from scalar functions[​](#key-differences-from-scalar-functions "Direct link to Key differences from scalar functions") | Aspect | Scalar | Table | | -------------------- | ---------------------------------------- | ------------------------------------- | | `kind:` | `scalar` (default) | `table` | | `signature.returns:` | Single Arrow type string (e.g., `int64`) | List of `{name, type}` output columns | | `body:` (SQL tier) | Single SQL expression | Full SQL `SELECT` query | | LLM tool exposure | `as_tool: true` (default) | Always SQL-only | ### SQL body for table functions[​](#sql-body-for-table-functions "Direct link to SQL body for table functions") The `body` of a SQL table function is a complete `SELECT` query (not an expression). Scalar arguments are exposed through a virtual one-row table named `args`: ``` body: | SELECT x AS value, x * 2 AS doubled FROM args ``` The query's output columns must match the declared `returns` schema. ### Filter pushdown with args[​](#filter-pushdown-with-args "Direct link to Filter pushdown with args") Scalar args are inlined as literals before logical planning, so they participate in filter pushdown. This makes it possible to parameterize a table function against connectors that require concrete filter values — for example, the HTTP connector's `request_path`: ``` functions: - name: hn_user from: sql kind: table signature: args: - { name: username, type: utf8 } returns: - { name: username, type: utf8 } - { name: karma, type: utf8 } body: | SELECT json_get_str(content, 'username') AS username, json_get_str(content, 'karma') AS karma FROM raw_users WHERE request_path = concat('/users/', (SELECT username FROM args)) ``` Calling `SELECT * FROM hn_user('pg')` resolves the `(SELECT username FROM args)` subquery to the literal `'pg'` before planning, so the resulting `WHERE request_path = '/users/pg'` predicate is pushed down to the HTTP connector. ### Wrapping search functions[​](#wrapping-search-functions "Direct link to Wrapping search functions") A SQL table function can wrap the [search functions](/docs/reference/sql/search) — [`vector_search()`](/docs/reference/sql/search#vector-search-vector_search), [`text_search()`](/docs/reference/sql/search#full-text-search-text_search), and either of them nested inside [`rrf()`](/docs/reference/sql/search#reciprocal-rank-fusion-rrf) — passing its own argument through as the search query. This packages a hybrid-search query as one named, callable function: ``` functions: - name: search_cookbook from: sql kind: table description: Hybrid search over the cookbook. signature: args: - { name: q, type: utf8 } returns: - { name: path, type: utf8 } - { name: content, type: utf8 } body: | SELECT path, content FROM rrf( vector_search(cookbook_files, q), text_search(cookbook_files, q, content) ) ``` ``` SELECT * FROM search_cookbook('how does hybrid search work') LIMIT 10; ``` Search functions need their query as a string literal while the plan is being built, which is earlier than the `args` table can supply a value. The runtime handles this by substituting the literal into the search function's query position before the body is planned. What that supports, and what it does not: * **The argument must be a declared `utf8` argument, referenced by name.** `vector_search(docs, q)` resolves; an expression built from one — `vector_search(docs, concat(q, ' docs'))` — is not rewritten, and planning still rejects it with `Second argument must be a query string`. Compose the query text in the caller instead. * **Both positional and `query =>` forms work.** `text_search(docs, query => q, column => body)` is accepted; the named query argument is moved into the second positional slot, which is where every search function reads it from. * **Only the query position is rewritten.** Every other reference to a function argument in the body — a `WHERE` predicate, a `LIMIT`, a join key — still goes through the `args` table as described above. ## Volatility[​](#volatility "Direct link to Volatility") Volatility tells the optimizer how the function behaves across calls. Pick the strongest level that's actually true — the default (`volatile`) is the safest but disables constant folding, query-level caching, and pushdown. | Volatility | Meaning | Optimizer behavior | | ----------- | ------------------------------------------------------------------------- | ------------------------------------------------------------------- | | `immutable` | Same inputs always yield the same output. E.g. `abs`, `upper`. | May be constant-folded at plan time and cached aggressively. | | `stable` | Stable within a single query but may change across queries. E.g. `now()`. | Cached per query, not constant-folded. | | `volatile` | (default) Unpredictable on every call. E.g. `random()`. | Never cached, never constant-folded, never pushed across executors. | Set volatility explicitly on every function — it strongly affects performance: ``` functions: - name: shout from: sql volatility: immutable # opts into constant folding & caching signature: args: [{ name: s, type: utf8 }] returns: utf8 body: 'upper(s)' ``` ## Types[​](#types "Direct link to Types") Argument and return types use Arrow logical types. Both Spicepod aliases and Arrow display forms are accepted. **Scalar aliases:** | Spicepod alias | Arrow type | | ------------------------------------------------ | ----------------------- | | `int8` / `int16` / `int32` (or `int`) / `int64` | `Int8` … `Int64` | | `uint8` / `uint16` / `uint32` / `uint64` | `UInt8` … `UInt64` | | `float32` (or `float`) / `float64` (or `double`) | `Float32`, `Float64` | | `utf8` (or `string`) / `large_utf8` | `Utf8`, `LargeUtf8` | | `boolean` (or `bool`) | `Boolean` | | `binary` / `large_binary` | `Binary`, `LargeBinary` | | `date32` / `date64` | `Date32`, `Date64` | **Complex types:** | Spicepod alias | Arrow type | | ------------------------------ | ------------------------------------- | | `list` | `List(Int64)` | | `large_list` | `LargeList(Utf8)` | | `struct` | `Struct(name: Utf8, age: Int32)` | | `decimal(38, 10)` | `Decimal128(38, 10)` | | `decimal256(76, 20)` | `Decimal256(76, 20)` | | `timestamp(us, utc)` | `Timestamp(Microsecond, Some("UTC"))` | The corresponding Arrow display forms (e.g. `Int64`, `List(Int64)`, `Decimal128(38, 10)`) are also accepted. ## Discovering registered functions[​](#discovering-registered-functions "Direct link to Discovering registered functions") ### From SQL[​](#from-sql "Direct link to From SQL") ``` SELECT * FROM list_udfs() WHERE source = 'user'; ``` The `list_udfs()` UDTF returns every function registered in the runtime, including built-ins. Filter by `source = 'user'` to see only declared functions: | Column | Description | | ------------- | ------------------------------------------------------------------- | | `name` | Function identifier. | | `source` | `user` for declared functions, `builtin` for Spice/DataFusion ones. | | `kind` | `scalar` or `table` for user functions, `NULL` for built-ins. | | `volatility` | `immutable` / `stable` / `volatile`. | | `from` | `sql`, `http://...`, or `https://...`. | | `description` | The declared description, if any. | ### From the HTTP API[​](#from-the-http-api "Direct link to From the HTTP API") ``` curl http://localhost:8090/v1/functions ``` Returns a JSON array of user functions only (built-ins are excluded). Each entry includes `name`, `kind`, `volatility`, `from`, and `description`. The endpoint returns an empty array when `runtime.functions.enabled` is `false`. ## Functions as LLM tools[​](#functions-as-llm-tools "Direct link to Functions as LLM tools") Every declared scalar function is automatically callable from LLMs as a tool with the same name and description. This lets a model reason in natural language and then invoke `haversine_km(...)` or `classify_intent(...)` directly. Table functions (`kind: table`) are always SQL-only and are not surfaced as LLM tools. Many functions? Use the Tool Registry A Spicepod with many functions can quickly cross the threshold where injecting every function definition into every chat turn becomes expensive. The [Tool Registry](/docs/next/features/tool-registry) replaces individual tool definitions with searchable `tool_search` / `tool_invoke` meta-tools, typically saving \~10× the per-turn tool-definition tokens. Set `tools: auto` on the model and the registry kicks in automatically once the function count crosses the threshold. To keep a function SQL-only: ``` functions: - name: internal_hash from: sql as_tool: false # Hidden from the tool registry; still callable from SQL signature: args: [{ name: x, type: int64 }] returns: int64 body: 'x * 2654435761' ``` The reverse — a `tools:` entry that's also callable from SQL — is supported via `as_sql: true` on a tool. See [Tools Spicepod reference](/docs/next/reference/spicepod/tools). ## Examples[​](#examples "Direct link to Examples") ### String normalization (SQL tier, immutable)[​](#string-normalization-sql-tier-immutable "Direct link to String normalization (SQL tier, immutable)") ``` functions: - name: shout from: sql volatility: immutable signature: args: [{ name: s, type: utf8 }] returns: utf8 body: 'upper(s)' ``` ``` SELECT shout('hello world'); -- "HELLO WORLD" ``` ### Geospatial helper (SQL tier, immutable)[​](#geospatial-helper-sql-tier-immutable "Direct link to Geospatial helper (SQL tier, immutable)") See `haversine_km` above. Because it is `immutable`, repeated calls with the same arguments are constant-folded. ### Remote ML classifier (Remote tier, volatile)[​](#remote-ml-classifier-remote-tier-volatile "Direct link to Remote ML classifier (Remote tier, volatile)") ``` functions: - name: classify_intent from: https://classifier.internal/v1/classify volatility: volatile signature: args: [{ name: prompt, type: utf8 }] returns: utf8 params: timeout: 5s batch_size: 256 auth_bearer: ${secrets:CLASSIFIER_TOKEN} ``` ``` SELECT id, classify_intent(message) AS intent FROM support_tickets WHERE created_at >= now() - INTERVAL '1' DAY; ``` The runtime batches up to 256 rows per HTTP call and issues up to four batches in parallel. ### Returning a struct (SQL tier)[​](#returning-a-struct-sql-tier "Direct link to Returning a struct (SQL tier)") ``` functions: - name: split_full_name from: sql volatility: immutable signature: args: [{ name: full, type: utf8 }] returns: struct body: | named_struct( 'first', split_part(full, ' ', 1), 'last', split_part(full, ' ', 2) ) ``` ### List-typed arguments (SQL tier)[​](#list-typed-arguments-sql-tier "Direct link to List-typed arguments (SQL tier)") ``` functions: - name: total_quantity from: sql volatility: immutable signature: args: [{ name: quantities, type: list }] returns: int64 body: 'array_sum(quantities)' ``` ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ | ------------------------------------------------------------------------------- | | Function appears defined but `SELECT my_fn(...)` errors with "function not found" | `runtime.functions.enabled` is not set to `true`. | Add `runtime.functions.enabled: true` to the spicepod. | | `the from scheme '...' is unsupported` | `from:` is not `sql`, `http://`, or `https://`. | Use one of the supported schemes. | | `body: and body_ref: are mutually exclusive` | Both fields are set on a SQL function. | Provide exactly one. | | `failed to parse function body as a SQL expression` | The body is not a single valid DataFusion SQL expression. | The body must be one expression (no statements), referencing arguments by name. | | `body expression evaluates to type ... not coercible to declared return type ...` | The body's computed type doesn't match `returns:`. | Adjust the body, declare a wider numeric `returns:`, or cast inside the body. | | Remote function returns `expected N values, got M` | The HTTP service returned the wrong number of `values`. | The service must return exactly one value per input row, in input order. | | Remote function calls fail with auth errors | `runtime.auth.api-key` is not configured, or `auth_bearer` is missing/invalid. | Configure runtime auth and supply a valid bearer token via `${secrets:...}`. | | Function registered but not surfaced as a tool | `as_tool: false` is set, or the function is `enabled: false`. | Remove `as_tool: false`; ensure `enabled: true` (default). | --- # Large Language Models Spice provides a high-performance, OpenAI API-compatible AI Gateway optimized for managing and scaling large language models (LLMs). It offers tools for Enterprise Retrieval-Augmented Generation (RAG), such as SQL query across federated datasets and an advanced search feature (see [Search](/docs/next/features/search)). ![ai-gateway](https://github.com/user-attachments/assets/4a45cd62-ebfc-4a73-956d-661f1ab44cd8) Spice supports **full OpenTelemetry observability**, helping with detailed tracking of model tool use, recursion, data flows and requests for full transparency and easier debugging. ## Quickstart[​](#quickstart "Direct link to Quickstart") Add a language model to your `spicepod.yaml` to start using AI capabilities: ``` models: - from: openai:gpt-4o-mini name: my_model params: openai_api_key: ${ env:OPENAI_API_KEY } tools: auto # Gives the model access to datasets for data-grounded responses ``` Start the runtime and use the chat REPL: ``` spice run # In another terminal: spice chat chat> What tables are available? ``` Or call the OpenAI-compatible HTTP API directly: ``` curl -X POST http://localhost:8090/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "my_model", "messages": [{"role": "user", "content": "What tables are available?"}] }' ``` ## Configuring Language Models[​](#configuring-language-models "Direct link to Configuring Language Models") Spice supports a variety of LLMs (see [Model Providers](/docs/next/components/models)). ### Core Features[​](#core-features "Direct link to Core Features") * **SQL Integration**: Invoke LLMs directly within SQL queries using the `ai()` function for text generation tasks. See [SQL Reference: ai function](/docs/next/reference/sql/scalar_functions#ai-and-embed). * **Custom Tools**: Provide models with tools to interact with the Spice runtime. See [Tools](/docs/next/features/large-language-models/tools). * **System Prompts**: Customize system prompts and override defaults for [`v1/chat/completion`](/docs/next/api/HTTP/post-chat-completions). See [Parameter Overrides](/docs/next/features/large-language-models/parameter_overrides). * **Memory**: Provide LLMs with memory persistence tools to store and retrieve information across conversations. See [Memory](/docs/next/features/large-language-models/memory). * **Vector Search**: Perform advanced vector-based searches using embeddings. See [Vector Search](/docs/next/features/search/vector-search). * **Local Models**: Load and serve models locally from various sources, including local filesystems and Hugging Face. See [Local Models](/docs/next/features/large-language-models/serving). For API usage, refer to the [API Documentation](/docs/next/api). ## [📄️Tools](/docs/next/features/large-language-models/tools) [Learn how LLMs interact with the Spice runtime.](/docs/next/features/large-language-models/tools) ## [📄️MCP](/docs/next/features/large-language-models/mcp) [Learn how to use the Model Context Protocol (MCP) with Spice.](/docs/next/features/large-language-models/mcp) ## [📄️Memory](/docs/next/features/large-language-models/memory) [Learn how to provide LLMs with memory](/docs/next/features/large-language-models/memory) ## [📄️Evals (Deprecated)](/docs/next/features/large-language-models/evals) [Language model evals are no longer supported in Spice.](/docs/next/features/large-language-models/evals) ## [📄️Parameter Overrides](/docs/next/features/large-language-models/parameter_overrides) [Learn how to override default LLM hyperparameters in Spice.](/docs/next/features/large-language-models/parameter_overrides) ## [📄️Local Models](/docs/next/features/large-language-models/serving) [Learn how to load and serve large learning models.](/docs/next/features/large-language-models/serving) ## [📄️Parameterized Prompts](/docs/next/features/large-language-models/parameterized_prompts) [Learn how to update system prompts for each request with Jinja-styled templating.](/docs/next/features/large-language-models/parameterized_prompts) --- # Evaluating Language Models (Deprecated) Deprecated Language model evals — the `evals` Spicepod component and the `/v1/evals` API endpoint — are no longer supported and were removed in Spice v2.0 ([spiceai/spiceai#9420](https://github.com/spiceai/spiceai/pull/9420)). For documentation on evals in the last version that supported them, see the [v1.11.x Evals documentation](https://docs.spiceai.org/docs/1.11.x/features/large-language-models/evals). --- # Model Context Protocol (MCP) The Model Context Protocol (MCP) helps integrate external tools and services into the Spice runtime. MCP tools can be run internally or connected over HTTP using the [Streamable HTTP](https://modelcontextprotocol.io/specification/2025-03-26/basic/transports#streamable-http) transport. ![Spice.ai Open Source Model-Context-Protocol (MCP) support](/assets/images/mcp-ea1443c9eb03219835e6945023b5b036.png) ## Overview[​](#overview "Direct link to Overview") MCP enables Spice to: 1. Run stdio-based MCP servers internally. 2. Connect to external MCP servers over Streamable HTTP. This flexibility helps extend the capabilities of language models by providing access to external tools and services. ## Configuring MCP Tools[​](#configuring-mcp-tools "Direct link to Configuring MCP Tools") To configure MCP tools, define them in the `tools` section of your `spicepod.yaml` file. The `from` field specifies the transport mechanism (e.g., `mcp:npx` for stdio or an HTTP URL for Streamable HTTP). ### Example: Adding an MCP Tool[​](#example-adding-an-mcp-tool "Direct link to Example: Adding an MCP Tool") ``` tools: - name: google_maps from: mcp:npx params: mcp_args: -y @modelcontextprotocol/server-google-maps ``` ### Example: Connecting to an External MCP Server[​](#example-connecting-to-an-external-mcp-server "Direct link to Example: Connecting to an External MCP Server") ``` tools: - name: external_mcp_server from: mcp:http://example.com/v1/mcp ``` ## Using MCP Tools with Models[​](#using-mcp-tools-with-models "Direct link to Using MCP Tools with Models") Once configured, MCP tools can be assigned to models via the `tools` parameter. ``` models: - name: model_with_mcp from: openai:gpt-4o params: tools: google_maps ``` ## Spice as an MCP Server[​](#spice-as-an-mcp-server "Direct link to Spice as an MCP Server") Spice can also act as an MCP server, exposing its tools over Streamable HTTP. This enables other Spice instances or external systems to connect and use the tools. ### Authentication[​](#authentication "Direct link to Authentication") The `/v1/mcp` endpoint requires [`runtime.auth`](/docs/next/reference/spicepod/runtime#runtimeauth) to be configured. Without it, every request is rejected with `401 Unauthorized`, including requests from local clients: ``` runtime: auth: api_key: enabled: true keys: - ${secrets:SPICE_API_KEY} ``` Clients then authenticate with an `X-API-KEY` header. ### Example: Connecting to another Spice instance via MCP[​](#example-connecting-to-another-spice-instance-via-mcp "Direct link to Example: Connecting to another Spice instance via MCP") Because the remote endpoint requires authentication, pass credentials with `mcp_headers` — or `mcp_auth_token` to send `Authorization: Bearer`. See [Connecting to an Auth-Enabled MCP Server](/docs/next/components/tools/mcp#example-connecting-to-an-auth-enabled-mcp-server-streamable-http). ``` tools: - name: spice_instance from: mcp:http://localhost:8090/v1/mcp params: mcp_headers: 'X-API-KEY: ${secrets:SPICE_API_KEY}' ``` ### Allowed Hosts[​](#allowed-hosts "Direct link to Allowed Hosts") By default the `/v1/mcp` endpoint only accepts requests with a `Host` header matching `localhost`, `127.0.0.1`, or `::1` to prevent DNS rebinding attacks. To allow additional hosts, configure [`runtime.mcp.allowed_hosts`](/docs/next/reference/spicepod/runtime#runtimemcp): ``` runtime: mcp: allowed_hosts: - localhost - my-host.internal:8090 ``` Set `allowed_hosts: ["*"]` to disable host checking entirely. ## Additional Configuration Options[​](#additional-configuration-options "Direct link to Additional Configuration Options") ### `from`[​](#from "Direct link to from") The `from` field specifies the transport mechanism for the MCP tool: * **Streamable HTTP**: Use an HTTP URL pointing to the MCP endpoint (e.g., `http://localhost:8090/v1/mcp`). * **Stdio**: Use commands like `mcp:npx` or `mcp:docker`. Additional arguments can be passed via `params.mcp_args`. ### `params`[​](#params "Direct link to params") The `params` field provides additional configuration for MCP tools. For stdio-based tools, use `mcp_args` to specify command-line arguments. ``` tools: - name: custom_tool from: mcp:npx params: mcp_args: -y @custom/tool ``` ### `env`[​](#env "Direct link to env") For stdio-based MCP tools, environment variables can be set using the `env` field. ``` tools: - name: tool_with_env from: mcp:docker env: API_KEY: your_api_key ``` For more details, see the [MCP Tools Reference](/docs/next/components/tools/mcp). ## Example: Spice as an MCP gateway for GitHub and Jira[​](#example-spice-as-an-mcp-gateway-for-github-and-jira "Direct link to Example: Spice as an MCP gateway for GitHub and Jira") Spice can act as a unified MCP gateway, proxying multiple external MCP servers alongside its own built-in tools. Clients connect to a single `/v1/mcp` endpoint and see one combined tool catalog. ``` version: v1 kind: Spicepod name: mcp-gateway # Required — /v1/mcp returns 401 to every request without runtime.auth runtime: auth: api_key: enabled: true keys: - ${secrets:SPICE_API_KEY} # GitHub issues and PRs accelerated in-memory for fast SQL queries datasets: - from: github:github.com/your-org/your-repo/issues name: github_issues description: GitHub issues — filterable by state, label, assignee, or milestone params: github_token: ${secrets:GITHUB_PERSONAL_ACCESS_TOKEN} github_query_mode: search time_column: updated_at acceleration: enabled: true refresh_check_interval: 5m - from: github:github.com/your-org/your-repo/pulls name: github_pulls description: GitHub pull requests — open, merged, or closed params: github_token: ${secrets:GITHUB_PERSONAL_ACCESS_TOKEN} github_query_mode: search time_column: updated_at acceleration: enabled: true refresh_check_interval: 5m # MCP servers proxied through Spice's /v1/mcp endpoint tools: - name: github from: mcp:npx description: GitHub tools — create issues, review PRs, search code params: mcp_args: -y @modelcontextprotocol/server-github env: GITHUB_PERSONAL_ACCESS_TOKEN: ${secrets:GITHUB_PERSONAL_ACCESS_TOKEN} - name: jira from: mcp:uvx description: Jira and Confluence tools — query tickets, update status, search projects params: mcp_args: mcp-atlassian env: JIRA_URL: ${secrets:JIRA_URL} JIRA_USERNAME: ${secrets:JIRA_USERNAME} JIRA_API_TOKEN: ${secrets:JIRA_API_TOKEN} ``` With this configuration, clients connecting to `/v1/mcp` see a single catalog that includes: * Spice built-in tools: `sql`, `list_datasets`, `table_schema`, `search`, and more * GitHub MCP tools: `github__create_issue`, `github__list_pull_requests`, `github__search_code`, and more * Jira MCP tools: `jira__jira_get_issue`, `jira__jira_search`, `jira__jira_create_issue`, and more ### Proxied tool names[​](#proxied-tool-names "Direct link to Proxied tool names") A proxied tool is exposed as `__`, joined by a **double underscore**. The `tool-name` is the `name` given in the `tools` section, so renaming the tool renames every tool it proxies. The separator is not a `/`, because MCP clients such as Claude and Cursor reject any tool name that does not match `^[a-zA-Z0-9_-]{1,64}$`. Where an upstream server already namespaces its own tools, the prefix appears doubled — `mcp-atlassian` exposes `jira_get_issue`, so a tool named `jira` yields `jira__jira_get_issue`. This is expected. ### Jira credentials[​](#jira-credentials "Direct link to Jira credentials") The `jira` tool contributes **no tools at all** unless all three `JIRA_*` values resolve. `mcp-atlassian` starts successfully without them, registers nothing, and Spice logs no tool-loading error — the runtime reports healthy with Jira silently missing from the catalog. Check the startup log for `runtime_secrets` errors if `jira__*` tools do not appear. `JIRA_URL` depends on which kind of API token is used, and a mismatch is the most common cause of an empty Jira catalog: | Token type | `JIRA_URL` | | ---------------------------- | ---------------------------------------------- | | Classic (unscoped) API token | `https://your-org.atlassian.net` | | Scoped API token | `https://api.atlassian.com/ex/jira/` | Scoped tokens authenticate only against the `api.atlassian.com` gateway and return `401` against the site URL. Retrieve the cloud ID for a site from `https://your-org.atlassian.net/_edge/tenant_info`, which requires no authentication. Both token types use `JIRA_USERNAME` (the account email address) with HTTP basic auth. ### Connecting a client[​](#connecting-a-client "Direct link to Connecting a client") Connect Claude Code to this gateway with a single command, using a key from `runtime.auth`: ``` claude mcp add --transport http spice http://localhost:8090/v1/mcp --header "X-API-KEY: " ``` Large tool catalogs The two servers above expose close to 100 tools between them, which degrades tool selection in most MCP clients. Restrict each server to the tools actually needed — for `mcp-atlassian`, set `TOOLSETS` in its `env` block. See the [mcp-server cookbook](https://github.com/spiceai/cookbook/tree/trunk/mcp-server) for a complete runnable example. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") ### A tool is missing from the catalog[​](#a-tool-is-missing-from-the-catalog "Direct link to A tool is missing from the catalog") A failing MCP server does not stop the runtime. Spice logs a warning, retries the connection in the background, and reports healthy — the tool is simply absent from `tools/list`. Verify the runtime log at startup: ``` grep "Unable to load tool" ``` A server that starts but authenticates with empty or invalid credentials is harder to spot: it registers zero tools without producing that warning. Check for `runtime_secrets` errors, which indicate a `${secrets:...}` reference that did not resolve. ### Every request returns `401`[​](#every-request-returns-401 "Direct link to every-request-returns-401") `/v1/mcp` requires [`runtime.auth`](#authentication). The response body names the missing configuration: ``` { "message": "MCP endpoint (/v1/mcp) requires `runtime.auth` to be configured. ..." } ``` ### Every request returns `403`[​](#every-request-returns-403 "Direct link to every-request-returns-403") The `Host` header is not in [`runtime.mcp.allowed_hosts`](#allowed-hosts). This is the DNS-rebinding guard, and it applies to any host other than `localhost`, `127.0.0.1`, or `::1`. ### Listing the catalog directly[​](#listing-the-catalog-directly "Direct link to Listing the catalog directly") To confirm what a client actually sees, call `tools/list` over the Streamable HTTP transport. The `Accept` header must offer both content types, and the session ID returned by `initialize` is required on subsequent requests: ``` curl -sS -D headers.txt -X POST http://localhost:8090/v1/mcp \ -H 'Content-Type: application/json' \ -H 'Accept: application/json, text/event-stream' \ -H 'X-API-KEY: ' \ -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"curl","version":"1"}}}' SID=$(grep -i mcp-session-id headers.txt | tr -d '\r' | awk '{print $2}') curl -sS -X POST http://localhost:8090/v1/mcp \ -H 'Content-Type: application/json' \ -H 'Accept: application/json, text/event-stream' \ -H 'X-API-KEY: ' \ -H "Mcp-Session-Id: $SID" \ -d '{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}}' ``` Responses are returned as SSE frames prefixed with `data:`. --- # Language Model Memory Spice provides memory persistence tools that help language models store and retrieve information across conversations. These tools are available through the `memory` tool group. Memory tools are useful for applications where context from previous interactions should influence future responses, such as chatbots, assistants, or multi-turn workflows. ## Enabling Memory Tools[​](#enabling-memory-tools "Direct link to Enabling Memory Tools") To enable memory tools for Spice models, define a `store` [memory](/docs/next/components/data-connectors/memory) dataset and specify `memory` in the model's `tools` parameter. ### Example: Enabling Memory Tools[​](#example-enabling-memory-tools "Direct link to Example: Enabling Memory Tools") ``` datasets: - from: memory:store name: llm_memory access: read_write models: - name: memory-enabled-model from: openai:gpt-4o params: tools: memory, sql # Can be combined with other tool groups ``` For more information on tools, see [Tool components](/docs/next/components/tools). --- # Language Model Overrides ### Chat Completion Parameter Overrides[​](#chat-completion-parameter-overrides "Direct link to Chat Completion Parameter Overrides") The [`v1/chat/completion`](/docs/next/api/HTTP/post-chat-completions) endpoint is compatible with OpenAI's API. It supports a subset of request body parameters defined in the [OpenAI reference documentation](https://platform.openai.com/docs/api-reference/chat/create). Spice helps configure different defaults for these request parameters. Supported parameters: * [`frequency_penalty`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-frequency_penalty) * [`logit_bias`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-logit_bias) * [`logprobs`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-logprobs) * [`max_completion_tokens`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-max_completion_tokens) * [`metadata`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-metadata) * [`n`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-n) * [`parallel_tool_calls`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-parallel_tool_calls) * [`presence_penalty`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-presence_penalty) * [`response_format`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format) * [`seed`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-seed) * [`stop`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-stop) * [`store`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-store) * [`stream`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-stream) * [`stream_options`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-stream_options) * [`temperature`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-temperature) * [`tool_choice`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-tool_choice) * [`tools`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-tools) * [`top_logprobs`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-top_logprobs) * [`top_p`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-top_p) * [`user`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-user) ### Example: Setting Default Overrides[​](#example-setting-default-overrides "Direct link to Example: Setting Default Overrides") Deprecated Default Overrides Parameters The `openai_` prefix is deprecated for non-OpenAI model providers. Use the [model provider prefix](/docs/next/components/models#model-provider-prefix) instead. To specify a default override for a parameter, use the [model provider prefix](/docs/next/components/models#model-provider-prefix) followed by the parameter name. For example, to set the `temperature` parameter to `0.1` for all requests with this model for Hugging Face model, use `hf_temperature: 0.1`. A `temperature` parameter in the request body will still override the default. ``` models: - name: pirate-haikus from: openai:gpt-4o params: openai_temperature: 0.1 openai_response_format: { 'type': 'json_object' } ``` When sending this payload to spice `/v1/chat/completions`: ``` { "model": "pirate-haikus", "messages": [ { "role": "user", "content": "What is the capital of France?" } ], "temperature": 0.5 } ``` Will be passed to the OpenAI API as: ``` { "model": "gpt-4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ], "temperature": 0.5, // temperature overriden by value in request body "response_format": { "type": "json_object" } // default response format from model configuration } ``` ### System Prompt[​](#system-prompt "Direct link to System Prompt") In addition to any system prompts provided in message dialogue, or added by model providers, Spice can configure an additional system prompt. ``` models: - name: pirate-haikus from: openai:gpt-4o params: system_prompt: | Write everything in Haiku like a pirate ``` Any request to [HTTP `v1/chat/completion`](/docs/next/api/HTTP/post-chat-completions) or `v1/responses` will include the configured system prompt. For `v1/responses`, the system prompt is set as the `instructions` field. If client-provided `instructions` are also included in the request, both are combined: the configured system prompt appears first, followed by the client instructions. ### Example: Enforcing default structured output and using system prompt[​](#example-enforcing-default-structured-output-and-using-system-prompt "Direct link to Example: Enforcing default structured output and using system prompt") This example demonstrates how to create a specialized math tutoring model by combining system prompts with structured JSON output. The configuration ensures consistent, step-by-step mathematical solutions in a machine-readable format. ``` models: - name: math-tutor from: openai:gpt-4o params: system_prompt: | You are a helpful math tutor. Guide the user through the solution step by step. openai_response_format: type: json_schema json_schema: name: math_reasoning schema: type: object properties: steps: type: array items: type: object properties: explanation: type: string output: type: string required: - explanation - output additionalProperties: false final_answer: type: string required: - steps - final_answer additionalProperties: false strict: true ``` To use the configured math tutor, send a simple request to the chat completions endpoint: ``` curl -s -XPOST http://localhost:8090/v1/chat/completions -H "Content-Type: application/json" -d \ '{ "model": "math-tutor", "messages": [{ "role": "user", "content" :"how can I solve 8x + 7 = -23" }] }' \ | jq '.choices[0].message.content | fromjson' ``` Example response: ``` { "final_answer": "x = -3.75", "steps": [ { "explanation": "We start with the given equation that we need to solve.", "output": "8x + 7 = -23" }, { "explanation": "Our goal is to solve for x. We can start by isolating the term with x on one side of the equation. To do this, we need to eliminate the constant term (7) on the left side. We subtract 7 from both sides of the equation in order to keep it balanced.", "output": "8x + 7 - 7 = -23 - 7" }, { "explanation": "Subtracting 7 from both sides simplifies the equation. On the left side, the +7 and -7 cancel out, leaving just the term with the variable.", "output": "8x = -30" }, { "explanation": "Now, we have 8 times x equals -30. To solve for x, we divide both sides of the equation by the coefficient of x, which is 8.", "output": "8x / 8 = -30 / 8" }, { "explanation": "Dividing both sides results in x on the left side and simplifies the fraction on the right side. The fraction -30/8 can be simplified further by dividing both the numerator and the denominator by their greatest common divisor, which is 2.", "output": "x = -3.75" }, { "explanation": "The solution has been simplified completely, giving us the value of x.", "output": "x = -3.75" } ] } ``` Visit [OpenAI Structured Outputs](https://platform.openai.com/docs/guides/structured-outputs) for more information on how to use structured output formats. ### Prompt Caching[​](#prompt-caching "Direct link to Prompt Caching") Spice supports provider-aware prompt caching to reduce latency and cost for repeated prompts. Set `prompt_cache_key` on a model to enable the provider's native caching mechanism. ``` models: - from: openai:gpt-4o name: my_model params: prompt_cache_key: "schema-context" ``` When `prompt_cache_key` is set as a model default, it is injected into every Chat API and Responses API request to that model (unless the request itself provides one). The key is mapped into the appropriate provider-native mechanism — for example, Anthropic's `cache_control`, xAI's `x-grok-conv-id` header, or Bedrock's `CachePoint` block. See the [full provider mapping](/docs/next/reference/spicepod/models#paramsprompt_cache_key). For the OpenAI Responses API, `prompt_cache_retention` can also be set to request a retention duration (e.g. `"24h"`). The `prompt_cache_key` can also be passed per-request in the [`/v1/nsql` API](/docs/next/api/HTTP/post-nsql) body to enable caching for text-to-SQL queries. For local models using mistral-rs, paged-attention scheduling is enabled automatically on supported backends (CUDA + Unix) for KV-cache prefix reuse — no configuration is needed. --- # System Prompt parameterization Spice supports defining system prompts for Large Language Models (LLM)s in the [spicepod](/docs/next/features/large-language-models/parameter_overrides#system-prompt). **Example**: ``` models: - name: advice from: openai:gpt-4o params: system_prompt: | Write everything in Haiku like a pirate from Australia ``` More than this, system prompts can use Jinja syntax so that system prompts can be altered on each [v1/chat/completion](/docs/next/api/HTTP/post-chat-completions) request. This involves three steps: 1. Add `parameterized_prompt: enabled` to the model. 2. Use Jinja syntax in the `system_prompt` parameter for the model in the spicepods. ``` models: - name: advice from: openai:gpt-4o params: parameterized_prompt: enabled system_prompt: | Write everything in {{ form }} like a {{ user.character }} from {{ user.country }} ``` 3. Provide the required variables in [v1/chat/completion](/docs/next/api/HTTP/post-chat-completions) via the `.metadata` field. ``` curl -X POST http://localhost:8090/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "advice", "messages": [ {"role": "user", "content": "Where should I visit in San Francisco?"} ], "metadata": { "form": "haiku", "user": { "character": "pirate", "country": "australia" } } }' ``` --- # Load and Serve Models Locally Spice supports loading and serving LLMs from various sources for embeddings and inference, including local filesystems and Hugging Face. ### Example: Loading a LLM from Hugging Face[​](#example-loading-a-llm-from-hugging-face "Direct link to Example: Loading a LLM from Hugging Face") ``` models: - name: llama_3.2_1B from: huggingface:huggingface.co/meta-llama/Llama-3.2-1B params: hf_token: ${ secrets:HF_TOKEN } ``` ## Filesystem[​](#filesystem "Direct link to Filesystem") Models can be hosted on a local filesystem and referenced directly in the configuration. For more details, see the [Filesystem Model Component](/docs/next/components/models/filesystem). ## Hugging Face[​](#hugging-face "Direct link to Hugging Face") Spice integrates with Hugging Face, enabling you to use a wide range of pre-trained models. For more information, see the [Hugging Face Model Component](/docs/next/components/models/huggingface). --- # Language Models Tools Spice provides tools that help LLMs interact with the runtime. To specify these tools for a Spice model, include them in its `params.tools`. For a list of available tools, or how to define additional tools, see [Tool Components](/docs/next/components/tools). ### Tool Modes[​](#tool-modes "Direct link to Tool Modes") The `tools` parameter on a model controls how tools are provided to the LLM: | Value | Description | | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `auto` | Automatically choose between direct tools and searchable registry discovery. When the number of available tools exceeds 20 and an embedding model is available, `auto` switches to registry-based discovery; otherwise it uses direct tools. | | `all` | Provide all built-in and Spicepod-configured tools directly to the LLM. | | `search_registry` | Use searchable registry discovery. The LLM receives `tool_search` and `tool_invoke` meta-tools instead of individual tool definitions. Requires an embedding model (see `tool_embedding_model`). | | `nsql` | Provide only the built-in tools relevant to text-to-SQL: `table_schema`, `sql`, `list_datasets`, `get_current_datetime`, `random_sample`, `sample_distinct_columns`, and `top_n_sample`. | | `disabled` | Provide no tools to the LLM. | | `, , ...` | Provide only the named tools directly. | ### Example: Specifying Tools for a Model[​](#example-specifying-tools-for-a-model "Direct link to Example: Specifying Tools for a Model") ``` models: - name: sql-model from: openai:gpt-4o params: tools: list_datasets, sql, table_schema ``` ### Example: Using all default tools directly[​](#example-using-all-default-tools-directly "Direct link to Example: Using all default tools directly") ``` models: - name: full-runtime from: openai:gpt-4o params: tools: all ``` ### Example: Using auto mode (default)[​](#example-using-auto-mode-default "Direct link to Example: Using auto mode (default)") ``` models: - name: full-runtime from: openai:gpt-4o params: tools: auto ``` For details on tool groups, see [Tool Components](/docs/next/components/tools#tool-groups). ### Tool Registry[​](#tool-registry "Direct link to Tool Registry") When the runtime exposes many tools (multiple MCP servers, lots of dataset-bound tools, or many [Functions](/docs/next/features/functions) declared with `as_tool: true`), passing every tool definition into every chat turn consumes a large portion of the context window and degrades selection accuracy. The **Tool Registry** replaces individual tool definitions with two meta-tools — `tool_search` and `tool_invoke` — backed by a hybrid search index over the runtime's tool catalog. This typically saves **\~10×** the per-turn tool-definition tokens for tool-heavy Spicepods. `tools: auto` enables the registry automatically when there are more than 20 tools and an embedding model is configured; `tools: search_registry` requires it. See **[Tool Registry](/docs/next/features/tool-registry)** for the full reference, including the hybrid-search algorithm, `tool_search` / `tool_invoke` parameters and response shapes, and guidance on when to prefer direct tools. ### Example: Specifying tools and tool groups[​](#example-specifying-tools-and-tool-groups "Direct link to Example: Specifying tools and tool groups") ``` models: - name: full-runtime from: openai:gpt-4o params: tools: memory, sql ``` ### Tool Recursion Limit[​](#tool-recursion-limit "Direct link to Tool Recursion Limit") When a model requests to call a runtime tool, Spice runs the tool internally and feeds the result back to the model. The model may then request another tool call based on the result, creating a chain of tool invocations. The `tool_recursion_limit` parameter limits the depth of this internal recursion. By default, this limit is set to 10. Lowering the limit can help prevent runaway tool chains in cases where a model repeatedly invokes tools without converging on a final answer. ``` models: - name: my-model from: openai params: tool_recursion_limit: 3 ``` --- # Machine Learning Models > **Deprecated in vNext**: Support for loading and serving traditional machine learning (ONNX) models for inference was removed in vNext, along with the `/v1/predict` and `/v1/models/{name}/predict` prediction endpoints. See the [v2.1 docs](https://docs.spiceai.org/docs/2.1.x/features/machine-learning-models) for documentation of this feature. Spice no longer loads or serves traditional machine learning (ONNX) models. The [Models](/docs/components/models) component now serves large language models (LLMs) only. For LLM configuration and usage, see [Large Language Models](/docs/next/features/large-language-models). --- # Observability & Monitoring Spice provides monitoring and observability through three mechanisms: * **Prometheus-compatible metrics endpoint**: Exposes metrics in the [Prometheus exposition format](https://prometheus.io/docs/instrumenting/exposition_formats/#basic-info) for scraping by monitoring systems like [Datadog](https://www.datadoghq.com/), [New Relic](https://newrelic.com/), and [Chronosphere](https://chronosphere.io/). * **OpenTelemetry metrics export**: Pushes metrics to an [OpenTelemetry](https://opentelemetry.io/) collector using gRPC. * **Distributed tracing**: Integrates with [Zipkin](https://zipkin.io/) and compatible tracing systems for request tracing. ![observability](https://github.com/user-attachments/assets/2468e3e7-4fb4-4a74-8b26-45eeeee90310) ### Monitoring Integrations[​](#monitoring-integrations "Direct link to Monitoring Integrations") * [Datadog](/docs/next/monitoring/datadog) * [Grafana & Prometheus](/docs/next/monitoring/grafana) * [New Relic](/docs/next/monitoring/new-relic) * [Zipkin](/docs/next/monitoring/zipkin) ## Prometheus Metrics Endpoint[​](#prometheus-metrics-endpoint "Direct link to Prometheus Metrics Endpoint") Spice exposes a Prometheus-compatible metrics endpoint that monitoring systems can scrape. The endpoint serves metrics in the [Prometheus exposition format](https://prometheus.io/docs/instrumenting/exposition_formats/), which is supported by most enterprise monitoring platforms including Datadog, New Relic, Chronosphere, Grafana Cloud, and others. ### Default Configuration[​](#default-configuration "Direct link to Default Configuration") The metrics endpoint listens on port `9090` by default. The endpoint address is logged at startup: ``` 2024-11-28T19:48:10.942003Z INFO runtime::metrics_server: Spice Runtime Metrics listening on 127.0.0.1:9090 ``` ### Custom Port Binding[​](#custom-port-binding "Direct link to Custom Port Binding") Use the `--metrics` flag to bind to a specific address and port: ``` spiced --metrics 0.0.0.0:9091 ``` For Docker deployments: ``` FROM spiceai/spiceai:latest CMD ["--metrics", "0.0.0.0:9090"] EXPOSE 9090 ``` ### Verifying the Endpoint[​](#verifying-the-endpoint "Direct link to Verifying the Endpoint") Verify the metrics endpoint is working with a GET request: ``` curl http://localhost:9090/metrics # HELP runtime_flight_server_started Indicates the runtime Flight server has started. # TYPE runtime_flight_server_started counter runtime_flight_server_started 1 # HELP runtime_http_server_started Indicates the runtime HTTP server has started. # TYPE runtime_http_server_started counter runtime_http_server_started 1 # HELP dataset_load_state Status of the dataset. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. # TYPE dataset_load_state gauge dataset_load_state{dataset="taxi_trips"} 2 dataset_load_state{dataset="taxi_trips_accelerated"} 2 # HELP dataset_active_count Number of currently loaded datasets. # TYPE dataset_active_count gauge dataset_active_count{engine="None"} 1 dataset_active_count{engine="duckdb"} 1 ... ``` ## OpenTelemetry Metrics Exporter[​](#opentelemetry-metrics-exporter "Direct link to OpenTelemetry Metrics Exporter") Spice can push metrics to an [OpenTelemetry](https://opentelemetry.io/) collector, enabling integration with platforms such as [Jaeger](https://www.jaegertracing.io/), [New Relic](https://newrelic.com/), [Honeycomb](https://www.honeycomb.io/), and other OpenTelemetry-compatible backends. ### Configuration[​](#configuration "Direct link to Configuration") Configure the OpenTelemetry exporter in `spicepod.yaml` under `runtime.telemetry.otel_exporter`: | Parameter | Required | Default | Description | | --------------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `enabled` | No | `true` | Whether the OpenTelemetry exporter is enabled. | | `endpoint` | Yes | - | The OpenTelemetry collector endpoint. Protocol (gRPC or HTTP) is inferred from the format. | | `push_interval` | No | `60s` | How frequently metrics are pushed to the collector. | | `metrics` | No | `[]` | List of metric names to export. When empty, all metrics are exported. | | `headers` | No | `{}` | Map of headers to send with each export request. For HTTP: sent as HTTP headers. For gRPC: sent as metadata entries (keys must be lowercase ASCII). Values support the `${secrets:...}` replacement syntax. | ### Protocol[​](#protocol "Direct link to Protocol") Spice infers the OTLP protocol from the `endpoint` format: * **gRPC** — bare host :port with no scheme (e.g. `localhost:4317`). Default port: `4317`. * **HTTP** — includes the `http://` or `https://` scheme and ends in `/v1/metrics` (e.g. `http://localhost:4318/v1/metrics`, `https://otlp.us3.datadoghq.com/v1/metrics`). Default port: `4318`. ### Authentication[​](#authentication "Direct link to Authentication") For collectors that require authentication (Datadog, Grafana Cloud, New Relic, Honeycomb, etc.), set the `headers` map. Secret values should be loaded from a [supported secret store](/docs/next/components/secret-stores) using the `${secrets:...}` [replacement syntax](/docs/next/components/secret-stores#using-secrets) rather than committed to source: ``` runtime: telemetry: otel_exporter: endpoint: 'https://otlp.example.com/v1/metrics' headers: Authorization: 'Bearer ${secrets:otlp_token}' ``` gRPC metadata keys must be lowercase When exporting over gRPC, header keys are sent as gRPC metadata and **must be lowercase ASCII** — use `authorization`, not `Authorization`. The runtime fails fast at startup if any gRPC metadata key is invalid. HTTP exports preserve the casing you provide. ### Examples[​](#examples "Direct link to Examples") #### Local gRPC collector[​](#local-grpc-collector "Direct link to Local gRPC collector") ``` runtime: telemetry: enabled: true otel_exporter: endpoint: 'localhost:4317' push_interval: '30s' ``` #### Local HTTP collector[​](#local-http-collector "Direct link to Local HTTP collector") ``` runtime: telemetry: enabled: true otel_exporter: endpoint: 'http://localhost:4318/v1/metrics' push_interval: '30s' ``` #### Datadog (OTLP/HTTP)[​](#datadog-otlphttp "Direct link to Datadog (OTLP/HTTP)") Replace `us3` with your Datadog site (`us3`, `us5`, `eu`, `ap1`, etc.) and store the API key in a secret store: ``` runtime: telemetry: enabled: true otel_exporter: endpoint: 'https://otlp.us3.datadoghq.com/v1/metrics' push_interval: '30s' headers: DD-API-KEY: ${secrets:dd_api_key} ``` Equivalent standard OTLP environment-variable form (for cross-reference): ``` export OTEL_EXPORTER_OTLP_ENDPOINT="https://otlp.us3.datadoghq.com" export OTEL_EXPORTER_OTLP_HEADERS="DD-API-KEY=${DD_API_KEY}" ``` For a complete Datadog setup including metric prefixing and custom tags via OTLP resource attributes, see the [Datadog monitoring guide](/docs/next/monitoring/datadog#opentelemetry-otlp-export). #### Grafana Cloud (OTLP/HTTP)[​](#grafana-cloud-otlphttp "Direct link to Grafana Cloud (OTLP/HTTP)") Grafana Cloud's OTLP gateway expects HTTP Basic authentication. Obtain the base64-encoded `instanceID:accessPolicyToken` credential from the Grafana Cloud "OpenTelemetry" connection page and store it in a secret: ``` runtime: telemetry: enabled: true otel_exporter: endpoint: 'https://otlp-gateway-us-central2.grafana.net/otlp/v1/metrics' push_interval: '30s' headers: Authorization: 'Basic ${secrets:grafana_cloud_auth}' ``` Equivalent standard OTLP environment-variable form (for cross-reference): ``` export OTEL_EXPORTER_OTLP_ENDPOINT="https://otlp-gateway-us-central2.grafana.net/otlp" export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Basic ${GRAFANA_CLOUD_AUTH}" ``` Match the region in the URL to your Grafana Cloud stack (`us-central2`, `eu-west-2`, `prod-ap-south-0`, etc.). #### gRPC collector with auth metadata[​](#grpc-collector-with-auth-metadata "Direct link to gRPC collector with auth metadata") ``` runtime: telemetry: enabled: true otel_exporter: endpoint: 'otel-collector.internal:4317' push_interval: '30s' headers: # Keys MUST be lowercase for gRPC api-key: ${secrets:collector_api_key} ``` ## Metric Naming and Custom Tags[​](#metric-naming-and-custom-tags "Direct link to Metric Naming and Custom Tags") Two runtime fields control how exported metrics are named and labeled across **all** readers (Prometheus scrape, cluster OTLP reader, and the `otel_exporter` push exporter): * [`runtime.telemetry.metric_prefix`](/docs/next/reference/spicepod/runtime#runtimetelemetrymetric_prefix) — prepends a string to every metric name (e.g. `spiceai.query_duration_ms`). Useful for namespacing in shared backends. * [`runtime.telemetry.properties`](/docs/next/reference/spicepod/runtime#runtimetelemetryproperties) — attaches custom key/value attributes as OpenTelemetry resource attributes, which most backends surface as dimensions or tags. ``` runtime: telemetry: metric_prefix: 'spiceai.' properties: environment: prod region: us-west-2 team: data-platform ``` Both fields apply to every exporter the runtime has enabled. See the [Datadog monitoring guide](/docs/next/monitoring/datadog#opentelemetry-otlp-export) for backend-specific notes (Datadog requires `dd-otel-metric-config` to map resource attributes to tags). ### Metric Filtering[​](#metric-filtering "Direct link to Metric Filtering") To export only specific metrics, use the `metrics` parameter: ``` runtime: telemetry: enabled: true otel_exporter: endpoint: 'localhost:4317' metrics: - query_duration_ms - query_executions - dataset_load_state ``` When `metrics` is empty or omitted, all available metrics are exported. Filtering happens after `metric_prefix` is applied The whitelist is matched against the **final** metric name, after `runtime.telemetry.metric_prefix` has been prepended. If you set `metric_prefix: 'spiceai.'`, the entries under `metrics:` must include the prefix (e.g. `spiceai.query_duration_ms`), otherwise nothing will match and no metrics will be exported. For full configuration details, see the [runtime.telemetry reference](/docs/next/reference/spicepod/runtime#runtimetelemetry). ## Available Metrics[​](#available-metrics "Direct link to Available Metrics") Spice exposes the following metrics. The **Dimensions** column lists labels available for filtering and aggregation; `—` indicates the metric is emitted without dimensions. Dimensions annotated *(request context)* expand to: `protocol`, `client`, `client_version`, `client_system`, `user_agent`, `runtime`, `runtime_version`, `runtime_system` (individual labels are only emitted when the corresponding request attribute is present). The cache counters (`results_cache_*`, `search_results_cache_*`, `embeddings_cache_*`) are published at zero when the runtime starts, so a counter that has not yet fired appears in a scrape as a zero series rather than being absent — an absent series indicates a scrape or exporter problem, not an idle cache. | Metric | Type | Dimensions | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------- | ------------------------------------------------------------------------------------------------------------ | | `accelerated_ready_state_federated_fallback`

Number of times the federated table was queried due to the accelerated table loading the initial data. | *count* | `dataset_name` | | `accelerated_zero_results_federated_fallback`

Number of times the federated table was queried due to the accelerated table returning zero results. | *count* | `dataset_name` | | `ai_inferences_with_spice_count`

AI Inferences with Spice count. | *count* | `tools_used` | | `catalog_acceleration_accelerated_tables`

Accelerated tables in a CDC-accelerated catalog, labeled by catalog and acceleration kind (primary\_key, unique\_index, full). | *gauge* | `catalog`, `kind` | | `catalog_acceleration_tables`

Relations resolved by a CDC-accelerated catalog, labeled by catalog and category (accelerated, skipped, excluded, views\_not\_replicated). | *gauge* | `catalog`, `category` | | `catalog_load_errors`

Number of errors loading the catalog provider. | *count* | — | | `catalog_load_state`

Status of the catalog provider. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | *gauge* | `catalog` | | `component_metric_registered_count`

Number of currently registered component metrics. | *gauge* | — | | `dataset_acceleration_ingestion_lag_ms`

Lag between the current wall-clock time and the maximum time\_column value after the refresh operation, in milliseconds. [Disabled by default](/docs/next/reference/spicepod/runtime#runtimemetrics) | *gauge* | `dataset`, `mode` | | `dataset_acceleration_last_refresh_unix_time_ms`

Unix timestamp in milliseconds when the last refresh completed. [Disabled by default](/docs/next/reference/spicepod/runtime#runtimemetrics) | *gauge* | `dataset` | | `dataset_acceleration_max_timestamp_after_refresh_ms`

Maximum value of the dataset's time\_column after the refresh operation, in milliseconds. [Disabled by default](/docs/next/reference/spicepod/runtime#runtimemetrics) | *gauge* | `dataset`, `mode` | | `dataset_acceleration_max_timestamp_before_refresh_ms`

Maximum value of the dataset's time\_column before the refresh operation, in milliseconds. [Disabled by default](/docs/next/reference/spicepod/runtime#runtimemetrics) | *gauge* | `dataset`, `mode` | | `dataset_acceleration_refresh_data_fetches_skipped`

Number of refresh data fetches skipped due to unchanged file metadata. | *count* | `dataset`, `mode` | | `dataset_acceleration_refresh_duration_ms`

Duration in milliseconds to load a full or appended refresh data. | *histogram* | `dataset`, `mode` | | `dataset_acceleration_refresh_errors`

Number of errors refreshing the dataset. | *count* | `dataset`, `mode` | | `dataset_acceleration_refresh_lag_ms`

Difference between the maximum time\_column value after and before the refresh operation, in milliseconds. | *gauge* | `dataset`, `mode` | | `dataset_acceleration_refresh_rows_written`

Cumulative number of rows read from the federated source and written into the accelerated table. | *count* | `dataset` | | `dataset_acceleration_refresh_bytes_written`

Cumulative number of bytes (Arrow in-memory size) read from the federated source and written into the accelerated table. | *count* | `dataset` | | `dataset_acceleration_refresh_worker_panics`

Number of times a refresh worker panicked while refreshing a dataset. | *count* | `dataset` | | `dataset_acceleration_size_bytes`

Size of the accelerated table storage in bytes. | *gauge* | `dataset` | | `dataset_acceleration_snapshot_bootstrap_bytes`

Number of bytes downloaded when bootstrapping the acceleration from a snapshot. | *gauge* | `dataset` | | `dataset_acceleration_snapshot_bootstrap_checksum`

Checksum of the snapshot downloaded during bootstrap (emitted with `checksum` attribute). | *gauge* | `dataset`, `checksum` | | `dataset_acceleration_snapshot_bootstrap_duration_ms`

Time in milliseconds taken to download the snapshot used to bootstrap acceleration. | *count* | `dataset` | | `dataset_acceleration_snapshot_failure_count`

Number of failures encountered while writing snapshots. | *count* | `dataset` | | `dataset_acceleration_snapshot_write_bytes`

Number of bytes written for the most recent snapshot. | *gauge* | `dataset` | | `dataset_acceleration_snapshot_write_checksum`

Checksum of the most recent snapshot write (emitted with `checksum` attribute). | *gauge* | `dataset`, `checksum` | | `dataset_acceleration_snapshot_write_duration_ms`

Time in milliseconds taken to write the latest snapshot to object storage. | *histogram* | `dataset` | | `dataset_acceleration_snapshot_write_timestamp`

Unix timestamp (seconds) when the most recent snapshot write completed. | *gauge* | `dataset` | | `dataset_active_count`

Number of currently loaded datasets. | *gauge* | `engine` | | `dataset_load_errors`

Number of errors loading the dataset. | *count* | — | | `dataset_load_state`

Status of the dataset. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | *gauge* | `dataset` | | `dataset_unavailable_time_ms`

Time dataset went offline in milliseconds. | *gauge* | `dataset` | | `embeddings_active_count`

Number of currently loaded embeddings. | *gauge* | `embeddings`, `source` | | `embeddings_cache_evictions`

Number of entries removed from the cache, split by `reason`: `size` (over `max_size`), `expired` (past `item_ttl`), or `invalidated` (a refresh or DML write dropped the entries referencing a table). | *count* | `reason` | | `embeddings_cache_hit_ratio`

Cache hit ratio (hits / total requests). | *gauge* | — | | `embeddings_cache_hits`

Cache hit count. | *count* | — | | `embeddings_cache_items_count`

Number of items currently in the cache. | *gauge* | — | | `embeddings_cache_max_size_bytes`

Maximum allowed size of the cache in bytes. | *gauge* | — | | `embeddings_cache_misses`

Cache miss count. | *count* | — | | `embeddings_cache_requests`

Number of requests to get a key from the cache. | *count* | — | | `embeddings_cache_size_bytes`

Size of the cache in bytes. | *gauge* | — | | `embeddings_cache_stale_swr_count`

Number of stale-while-revalidate background refreshes skipped due to existing in-flight revalidation. | *count* | — | | `embeddings_cache_swr_background_query_count`

Number of background queries triggered for stale-while-revalidate cache refreshes. | *count* | — | | `embeddings_failures`

Number of embedding failures. | *count* | `model`, `encoding_format`, `user`, `dimensions` | | `embeddings_internal_request_duration_ms`

The duration of running an embedding(s) internally. | *histogram* | `model`, `encoding_format`, `user`, `dimensions` | | `embeddings_load_errors`

Number of errors loading the embedding. | *count* | — | | `embeddings_load_state`

Status of the embedding. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | *gauge* | `model` | | `embeddings_requests`

Number of embedding requests. | *count* | `model`, `encoding_format`, `user`, `dimensions` | | `executor_assigned_partitions_count`

Number of acceleration partitions currently assigned to this executor in distributed query mode. | *gauge* | `node_id`, `dataset` | | `executor_scheduler_active_connections`

Active control-stream connections from the executor to each scheduler (0 or 1 per scheduler). | *gauge* | `node_id`, `scheduler` | | `executor_scheduler_connection_retries`

Reconnections initiated by the executor to a scheduler. | *count* | `node_id`, `scheduler` | | `flight_do_exchange_data_updates_sent`

Number of data updates sent via DoExchange. | *count* | — | | `flight_do_put_bytes_written`

Cumulative number of bytes (Arrow in-memory size) received and written via Flight DoPut. | *count* | `dataset` | | `flight_do_put_rows_written`

Cumulative number of rows received and written via Flight DoPut. | *count* | `dataset` | | `flight_request_duration_ms`

Measures the duration of Flight requests in milliseconds. | *histogram* | `method`, `command`, *(request context)* | | `flight_requests`

Total number of Flight requests. | *count* | `method`, `command`, *(request context)* | | `http_requests`

Number of HTTP requests, counted when the response head is produced. A streaming response that fails after its head was sent is still counted here with the `200` status it started with — use `http_responses` to see the outcome. | *count* | `method`, `path`, `status`, *(request context)* | | `http_requests_duration_ms`

Measures the duration of HTTP requests in milliseconds, up to the response head. For streaming responses (such as the default `/v1/sql` JSON format) the body continues after this is recorded — use `http_responses_duration_ms` for end-to-end duration. | *histogram* | `method`, `path`, `status`, *(request context)* | | `http_responses`

Number of HTTP responses, counted once the response body terminates. The `outcome` label reports how it terminated, so a streaming response that fails after its `200` head is not counted as a success. | *count* | `method`, `path`, `status`, `outcome` (one of `complete`, `error`, `incomplete`), *(request context)* | | `http_responses_duration_ms`

End-to-end HTTP response duration in milliseconds, from the request arriving to the response body terminating. Unlike `http_requests_duration_ms`, this includes the time spent streaming the body. | *histogram* | `method`, `path`, `status`, `outcome` (one of `complete`, `error`, `incomplete`), *(request context)* | | `llm_failures`

Number of LLM failures. | *count* | `model`, `stream`, `request_level_tools`, `tool_choice`, `user`, `metadata`, `responses_api`, `instructions` | | `llm_internal_request_duration_ms`

The duration of running an LLM request internally. | *histogram* | `model`, `stream`, `request_level_tools`, `tool_choice`, `user`, `metadata`, `responses_api`, `instructions` | | `llm_load_state`

Status of the LLM model. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | *gauge* | `model` | | `llm_requests`

Number of LLM requests. | *count* | `model`, `stream`, `request_level_tools`, `tool_choice`, `user`, `metadata`, `responses_api`, `instructions` | | `model_active_count`

Number of currently loaded models. | *gauge* | `model`, `source` | | `model_load_duration_ms`

Duration in milliseconds to load the model. | *histogram* | — | | `model_load_errors`

Number of errors loading the model. | *count* | — | | `model_load_state`

Status of the model. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | *gauge* | `model` | | `process_resident_memory_bytes`

Resident set size (RSS) of the `spiced` process, in bytes — the number the kernel's OOM decision is made on. Budgets and pool gauges describe intent; this describes fact, and the gap between them is off-pool memory no budget covers. Sampled every 2 seconds, and only emitted when [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) acceleration is configured. | *gauge* | — | | `query_active_count`

Number of concurrent top-level queries actively being processed in the runtime. | *histogram* | `protocol` (one of `http`, `flight`, `flightsql`, `internal`) | | `query_duration_ms`

The total amount of time spent planning and executing queries in milliseconds. | *histogram* | `tags`, `datasets`, *(request context)* | | `query_execution_duration_ms`

The total amount of time spent only executing queries (0 for cached queries). | *histogram* | `tags`, `datasets`, *(request context)* | | `query_executions`

Number of query executions. | *count* | `tags`, `datasets`, *(request context)* | | `query_executor_count`

Number of executors selected per query during partition-aware planning (distributed query). | *histogram* | `node_id` | | `query_failures`

Number of query failures. | *count* | `tags`, `datasets`, `err_code`, *(request context)* | | `query_memory_pool_used_bytes`

Live bytes reserved in the query memory pool (`runtime.query.memory_limit`), excluding the in-memory CDC tier's mirror account. Sampled every 2 seconds, and only emitted when [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) acceleration is configured. | *gauge* | — | | `query_planning_failures`

Queries that failed during partition-aware planning before execution. Indicates missing partitions or unavailable executors. | *count* | `node_id`, `error_type` (`missing_partitions`, `no_executors`) | | `query_processed_bytes`

Number of bytes processed by the runtime. | *count* | *(request context)* | | `query_produced_spills`

Number of spills produced by the query. | *count* | *(request context)* | | `query_returned_bytes`

Number of bytes returned to query clients. | *count* | *(request context)* | | `query_returned_rows`

Number of rows returned to query clients. | *histogram* | *(request context)* | | `query_spilled_bytes`

Number of spilled bytes produced by the query. | *count* | *(request context)* | | `query_spilled_rows`

Number of spilled rows produced by the query. | *count* | *(request context)* | | `results_cache_evictions`

Number of entries removed from the cache, split by `reason`: `size` (over `max_size`), `expired` (past `item_ttl`), or `invalidated` (a refresh or DML write dropped the entries referencing a table). | *count* | `reason` | | `results_cache_hit_ratio`

Cache hit ratio (hits / total requests). | *gauge* | — | | `results_cache_hits`

Cache hit count. | *count* | — | | `results_cache_items_count`

Number of items currently in the cache. | *gauge* | — | | `results_cache_max_size_bytes`

Maximum allowed size of the cache in bytes. | *gauge* | — | | `results_cache_misses`

Cache miss count. | *count* | — | | `results_cache_requests`

Number of requests to get a key from the cache. | *count* | — | | `results_cache_size_bytes`

Size of the cache in bytes. | *gauge* | — | | `results_cache_stale_rejections`

Number of lookups that found an entry but refused to serve it because a table the result read had since been invalidated. These are also counted as misses. | *count* | — | | `results_cache_stale_swr_count`

Number of stale-while-revalidate background refreshes skipped due to existing in-flight revalidation. | *count* | — | | `results_cache_swr_background_query_count`

Number of background queries triggered for stale-while-revalidate cache refreshes. | *count* | — | | `runtime_flight_server_started`

Indicates the runtime Flight server has started. | *count* | — | | `runtime_http_server_started`

Indicates the runtime HTTP server has started. | *count* | — | | `runtime_tls_reload_total`

Number of TLS certificate hot-reload attempts. | *count* | `scope` (`public`, `cluster`), `result` (`ok`, `io_error`, `parse_error`) | | `scheduler_active_executors_count`

Number of executors currently connected to the scheduler node. | *gauge* | `node_id` | | `scheduler_executor_active_connections`

Active control-stream connections from the scheduler to each executor (0 or 1 per executor). | *gauge* | `node_id`, `executor` | | `scheduler_executor_connection_retries`

Reconnections observed by the scheduler for an executor. | *count* | `node_id`, `executor` | | `scheduler_partition_assignments`

Acceleration-partition assignment operations executed by the scheduler. | *count* | `node_id`, `executor`, `status` (`committed`, `failed`) | | `scheduler_partition_discovery_duration_ms`

Duration of acceleration-partition discovery against the upstream source. | *histogram* | `node_id`, `dataset` | | `scheduler_partition_state_operations`

Partition status update operations on the scheduler. | *count* | `node_id`, `status` (`added`, `removed`, `reassigned`) | | `scheduler_partitioned_write_forwards`

Partitioned writes forwarded by the scheduler to executors. | *count* | `node_id`, `executor`, `status` (`completed`, `failed`) | | `scheduler_partitions_count`

Number of acceleration partitions known to the scheduler, broken down by assignment status. | *gauge* | `node_id`, `dataset`, `status` (`assigned`, `unassigned`) | | `search_results_cache_evictions`

Number of entries removed from the cache, split by `reason`: `size` (over `max_size`), `expired` (past `item_ttl`), or `invalidated` (a refresh or DML write dropped the entries referencing a table). | *count* | `reason` | | `search_results_cache_hit_ratio`

Cache hit ratio (hits / total requests). | *gauge* | — | | `search_results_cache_hits`

Search cache hit count. | *count* | — | | `search_results_cache_items_count`

Number of items currently in the search cache. | *gauge* | — | | `search_results_cache_max_size_bytes`

Maximum allowed size of the search cache in bytes. | *gauge* | — | | `search_results_cache_misses`

Cache miss count. | *count* | — | | `search_results_cache_requests`

Number of requests to get a key from the search cache. | *count* | — | | `search_results_cache_size_bytes`

Size of the search cache in bytes. | *gauge* | — | | `search_results_cache_stale_swr_count`

Number of stale-while-revalidate background refreshes skipped due to existing in-flight revalidation. | *count* | — | | `search_results_cache_swr_background_query_count`

Number of background queries triggered for stale-while-revalidate cache refreshes. | *count* | — | | `secrets_store_load_duration_ms`

Duration in milliseconds to load the secret stores. | *histogram* | — | | `spiced_cpu_budget_cores`

CPU cores the runtime sizes itself for, in whole cores. See [`runtime.cpu`](/docs/next/reference/spicepod/runtime#runtimecpu). | *gauge* | `source` (`configured`, `cgroup_quota`, `request_burst`, `all_cores`, `affinity`, `fallback`) | | `spiced_cpu_budget_millicores`

CPU millicores the runtime sizes itself for — the same entitlement, exact. | *gauge* | `source` (`configured`, `cgroup_quota`, `request_burst`, `all_cores`, `affinity`, `fallback`) | | `spiced_cpu_limit_millicores`

CPU limit read from the cgroup CPU quota (Kubernetes `limits.cpu`). Not emitted when the cgroup expresses no limit, so it is never confused with a real zero. | *gauge* | — | | `spiced_cpu_request_millicores`

The pod's own CPU request (Kubernetes `requests.cpu`), in millicores. Sizing derives from it only when nothing outranks it — no CPU limit, and no explicit `runtime.cpu.cores`; the `source` label above reports whether it did. Not emitted when no request was declared. The runtime cannot read the request itself: the pod spec must pass it in through `SPICE_CPU_REQUEST_MILLICORES` (see below). | *gauge* | — | | `tokio_runtime_alive_tasks`

Alive (spawned, not-yet-completed) tasks per tokio runtime. | *gauge* | `runtime` (`main`, and `cpu`, `refresh`, `cdc_apply`, `compaction` when active) | | `tokio_runtime_global_queue_depth`

Tasks waiting in the runtime's global (injection) queue per tokio runtime. | *gauge* | `runtime` | | `tokio_runtime_worker_busy_seconds`

Cumulative worker-busy time (summed across workers) per tokio runtime; `rate()` / `tokio_runtime_workers` gives the busy ratio. Only emitted in builds compiled with the `tokio_unstable` cfg. | *gauge* | `runtime` | | `tokio_runtime_worker_park_count`

Cumulative worker park count (summed across workers) per tokio runtime. Only emitted in builds compiled with the `tokio_unstable` cfg. | *gauge* | `runtime` | | `tokio_runtime_worker_steal_count`

Cumulative task-steal count (summed across workers) per tokio runtime. Only emitted in builds compiled with the `tokio_unstable` cfg. | *gauge* | `runtime` | | `tokio_runtime_workers`

Worker threads per tokio runtime. | *gauge* | `runtime` | | `tool_active_count`

Number of currently loaded LLM tools. | *gauge* | `tool` or `tool_catalog` | | `tool_load_errors`

Number of errors loading the LLM tool. | *count* | — | | `tool_load_state`

Status of the LLM tools. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | *gauge* | `tool` or `tool_catalog` | | `view_load_errors`

Number of errors loading the view. | *count* | — | | `view_load_state`

Status of the views. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | *gauge* | `view` | | `worker_active_count`

Number of currently loaded workers. | *gauge* | `worker` | | `workers_load_duration_ms`

Duration in milliseconds to load the worker. | *histogram* | — | #### CPU entitlement metrics[​](#cpu-entitlement-metrics "Direct link to CPU entitlement metrics") `spiced_cpu_budget_cores` carries a `source` label naming the rung of the [detection ladder](/docs/next/reference/spicepod/runtime#detection) the entitlement came from. It is the authority on *why* a pod is sized the way it is, which makes a fleet greppable for pods that resolved somewhere unexpected: | `source` | Meaning | | --------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `configured` | An explicit `runtime.cpu.cores`, `SPICE_CPU_CORES`, or `--cpu-cores` quantity. Used exactly as written. | | `cgroup_quota` | A cgroup CPU quota — Kubernetes `resources.limits.cpu`, `docker run --cpus`. Capped by the CPU affinity mask. | | `request_burst` | Derived from the pod's declared CPU request as `min(max(2 cores, request × 2), available CPUs)`. The 2× exceeds the request deliberately, so a burstable pod can still burst above its scheduling floor. | | `all_cores` | `runtime.cpu.cores: all` — every CPU the process may use, ignoring any CPU request. A CPU limit still caps it. | | `affinity` | No quota, no declared request, nothing configured: every CPU the process may run on. Bare metal, `docker run` with no CPU flags, and every benchmark land here. | | `fallback` | Nothing could be determined; one core. | `request_burst` requires the pod spec to pass the request in — the runtime cannot read `resources.requests.cpu` on its own. The Kubernetes downward API supplies it: ``` env: - name: SPICE_CPU_REQUEST_MILLICORES valueFrom: resourceFieldRef: containerName: spiceai resource: requests.cpu divisor: 1m # required: makes the value millicores ``` The Spice Helm chart and the Spice Kubernetes Operator both emit this by default whenever the pod sets a CPU request, so a deployment using either gets `request_burst` with no configuration. A hand-written pod spec that omits it reports `affinity` instead — sized for the machine rather than the request — and the runtime warns at startup naming the variable to set. Component Metrics In addition to these core metrics, individual components can expose their own metrics. For example, the MySQL data connector exposes [connection pool metrics](/docs/next/components/data-connectors/mysql#metrics). See [Component Metrics](/docs/next/features/observability/component_metrics) for more information. --- # Component Metrics Component metrics provide detailed insights into the internal state and performance of individual components in Spice. Each component can expose its own set of metrics that can be enabled selectively to monitor specific aspects of its operation. ## Enabling Component Metrics[​](#enabling-component-metrics "Direct link to Enabling Component Metrics") Most component metrics are disabled by default and can be enabled by adding a `metrics` section to the component configuration. Each metric can be enabled individually by specifying its name in the metrics list. Some metrics are **auto-registered**: they export without any `metrics` configuration, and an entry with `enabled: false` is what turns one off. Each component's own documentation states which of its metrics are auto-registered. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - from: some_component:my_resource name: my_resource metrics: - name: metric_one enabled: true - name: metric_two enabled: true - name: metric_three enabled: false params: param_one: value_one param_two: value_two ``` ## Available Metrics[​](#available-metrics "Direct link to Available Metrics") Each component defines its own set of available metrics. These metrics are exposed in the Prometheus format with the following naming convention: ``` {component_type}_{component_name}_{metric_name} ``` For example, a MySQL dataset component's metrics would be prefixed with `dataset_mysql_`. ## Monitoring Component Metrics[​](#monitoring-component-metrics "Direct link to Monitoring Component Metrics") Component metrics are exposed through the same Prometheus-compatible metrics endpoint as other Spice metrics. These metrics can be accessed using standard Prometheus tools or any monitoring system that supports Prometheus metrics. To view the metrics, make a GET request to the metrics endpoint: ``` curl http://localhost:9090/metrics ``` The response will include all enabled component metrics in Prometheus format, with proper HELP and TYPE annotations. ## Component-Specific Metrics[​](#component-specific-metrics "Direct link to Component-Specific Metrics") For detailed information about metrics available for specific components, view all [components that expose metrics](/docs/next/tags/component-metrics). --- # Query Federation Spice provides a high-performance SQL query engine built on Apache DataFusion, supporting query federation across multiple data sources including databases (PostgreSQL, MySQL), data warehouses (Databricks, Snowflake, BigQuery), and data lakes (S3, MinIO). ![Spice.ai Open Source Query Federation](/assets/images/query-federation-d1077a7b9c335e975aeec36b9dffd8ae.png) For a full list of supported sources, see [Data Connectors](/docs/next/components/data-connectors). ## When to Use Query Federation[​](#when-to-use-query-federation "Direct link to When to Use Query Federation") Query federation is useful when: * Data lives in multiple systems (e.g., PostgreSQL + S3 + Snowflake) and needs to be joined without ETL pipelines. * Applications need a single SQL interface to query across databases, data lakes, and warehouses. * SQL queries should be pushed down to source databases to minimize data transfer. ## Minimal Example[​](#minimal-example "Direct link to Minimal Example") Query data from PostgreSQL and S3 through a single SQL interface: ``` version: v1 kind: Spicepod name: federation_example datasets: - from: postgres:public.customers name: customers params: pg_host: localhost pg_db: mydb pg_user: reader pg_pass: ${secrets:PG_PASS} - from: s3://analytics-bucket/orders/ name: orders params: file_format: parquet s3_auth: iam_role ``` ``` -- Join across PostgreSQL and S3 in a single query SELECT c.name, COUNT(o.order_id) as order_count FROM customers c JOIN orders o ON c.id = o.customer_id GROUP BY c.name ORDER BY order_count DESC; ``` ## Query Methods[​](#query-methods "Direct link to Query Methods") Spice supports multiple ways to execute queries: * **SQL Queries**: Execute standard SQL queries against datasets using the HTTP API, Arrow Flight SQL, JDBC, ODBC, or ADBC. * **Parameterized Queries**: Execute prepared statements with parameter binding for improved security and performance. * **Federated Queries**: Join and query data across multiple sources in a single SQL statement. ## API Endpoints[​](#api-endpoints "Direct link to API Endpoints") | Protocol | Endpoint | Description | | ---------------- | ------------------------ | -------------------------------------- | | HTTP | `/v1/sql` | Execute SQL queries over HTTP | | Arrow Flight SQL | `grpc://localhost:50051` | High-performance Arrow-native queries | | JDBC/ODBC | Flight SQL compatible | Connect from BI tools and applications | | ADBC | Flight SQL driver | Arrow Database Connectivity | ### HTTP API[​](#http-api "Direct link to HTTP API") Execute a query using the HTTP API: ``` curl -X POST http://localhost:8090/v1/sql \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT * FROM my_table LIMIT 10"}' ``` ### Arrow Flight SQL[​](#arrow-flight-sql "Direct link to Arrow Flight SQL") Connect using Arrow Flight SQL for high-performance data transfer: ``` import adbc_driver_flightsql.dbapi conn = adbc_driver_flightsql.dbapi.connect('grpc://localhost:50051') cursor = conn.cursor() cursor.execute("SELECT * FROM my_table LIMIT 10") result = cursor.fetch_arrow_table() ``` ### SQL REPL[​](#sql-repl "Direct link to SQL REPL") Use the Spice CLI for interactive queries: ``` spice sql ``` ``` SELECT * FROM my_table LIMIT 10; ``` ## Query Features[​](#query-features "Direct link to Query Features") ## [📄️Parameterized Queries](/docs/next/features/query-federation/parameterized-queries) [Learn how to use prepared statements and parameterized queries in Spice for improved security and performance.](/docs/next/features/query-federation/parameterized-queries) ## [📄️URL Tables](/docs/next/features/query-federation/url-tables) [Query object store files directly using URLs without pre-registering datasets](/docs/next/features/query-federation/url-tables) ## Federated Query Example[​](#federated-query-example "Direct link to Federated Query Example") To start using federated queries in Spice, follow these steps: **Step 1.** Install Spice by following the [installation instructions](/docs/next/getting-started). **Step 2.** Clone the Spice Cookbook repository and navigate to the `federation` directory. ``` git clone https://github.com/spiceai/cookbook.git cd cookbook/federation ``` **Step 3.** Login to the demo Dremio. ``` spice login dremio -u demo -p demo1234 ``` **Step 4.** Create a new Spice app called `demo`. ``` # Create Spice app "demo" spice init demo # Change to demo directory. cd demo ``` **Step 5.** Add the `spiceai/fed-demo` Spicepod. ``` # Change to demo directory. cd demo spice add spiceai/fed-demo ``` Note in the Spice runtime output several datasets are loaded. **Step 6.** Start the Spice runtime. ``` spice run ``` **Step 7.** Show available tables and query them, regardless of source. ``` # Start the Spice SQL REPL. spice sql ``` Show the available tables: ``` show tables; ``` Execute the queries: ``` -- Query S3 (Parquet) SELECT * FROM s3_source LIMIT 10; -- Query S3 (Parquet) accelerated SELECT * FROM s3_source_accelerated LIMIT 10; -- Query Dremio SELECT * FROM dremio_source LIMIT 10; -- Query Dremio accelerated SELECT * FROM dremio_source_accelerated LIMIT 10; ``` **Step 8.** Join tables across remote sources and locally accelerated source ``` -- Query across S3 and Dremio WITH all_sales AS ( SELECT sales FROM s3_source UNION ALL select fare_amount+tip_amount as sales from dremio_source ) SELECT SUM(sales) as total_sales, COUNT(*) AS total_transactions, MAX(sales) AS max_sale, AVG(sales) AS avg_sale FROM all_sales; +--------------------+--------------------+----------+--------------------+ | total_sales | total_transactions | max_sale | avg_sale | +--------------------+--------------------+----------+--------------------+ | 11501140.079999998 | 102823 | 14082.8 | 111.85376890384445 | +--------------------+--------------------+----------+--------------------+ Time: 1.079320792 seconds. 1 rows. ``` **Step 9.** Join tables across locally accelerated sources and query ``` -- Query across S3 accelerated and Dremio accelerated WITH all_sales AS ( SELECT sales FROM s3_source_accelerated UNION ALL select fare_amount+tip_amount as sales from dremio_source_accelerated ) SELECT SUM(sales) as total_sales, COUNT(*) AS total_transactions, MAX(sales) AS max_sale, AVG(sales) AS avg_sale FROM all_sales; +-------------+--------------------+----------+--------------------+ | total_sales | total_transactions | max_sale | avg_sale | +-------------+--------------------+----------+--------------------+ | 11501140.08 | 102823 | 14082.8 | 111.85376890384447 | +-------------+--------------------+----------+--------------------+ Time: 0.011524375 seconds. 1 rows. ``` ### Acceleration[​](#acceleration "Direct link to Acceleration") The query in step 8 returns results from federated remote data sources, but performance is affected by network latency and data transfer overhead. Step 9 demonstrates the same query executed against locally materialized datasets using [Data Accelerators](/docs/next/components/data-accelerators). By storing data locally, queries avoid network round-trips and achieve significantly faster response times. Limitations * **Query Performance:** Without acceleration, federated queries will be slower than local queries due to network latency and data transfer. * **Query Capabilities:** Not all SQL features and data types are supported across all data sources. More complex data type queries may not work as expected. ## Related Topics[​](#related-topics "Direct link to Related Topics") * [Distributed Query](/docs/next/features/distributed-query) - Scale queries across multiple nodes * [Results Caching](/docs/next/features/caching) - Cache query results for improved performance * [Arrow Flight SQL API](/docs/next/api/arrow-flight-sql) - High-performance query protocol * [ADBC](/docs/next/api/adbc) - Arrow Database Connectivity --- # Parameterized Queries Parameterized queries separate SQL logic from data values, providing protection against SQL injection attacks and improving query performance through prepared statement caching. ## Overview[​](#overview "Direct link to Overview") Instead of embedding values directly in SQL strings, parameterized queries use placeholders (e.g., `$1`, `$2`) that are bound to values at execution time. This approach: * **Prevents SQL injection**: Values are never interpreted as SQL code * **Improves performance**: Query plans can be cached and reused * **Enhances code clarity**: Query structure is separate from data ## Placeholder Syntax[​](#placeholder-syntax "Direct link to Placeholder Syntax") Spice uses positional placeholders following the PostgreSQL convention: ``` SELECT * FROM users WHERE id = $1 AND status = $2; ``` Parameters are bound in order: `$1` receives the first parameter, `$2` the second, and so on. ## Using Parameterized Queries[​](#using-parameterized-queries "Direct link to Using Parameterized Queries") * Python (ADBC) * Go * Java * Rust * Dotnet * HTTP API ADBC provides native support for parameterized queries through the FlightSQL driver: ``` from adbc_driver_flightsql.dbapi import connect with connect("grpc://127.0.0.1:50051") as conn: with conn.cursor() as cur: # Single parameter cur.execute("SELECT * FROM users WHERE id = $1", parameters=(42,)) table = cur.fetch_arrow_table() print(table) # Multiple parameters cur.execute( "SELECT * FROM orders WHERE customer_id = $1 AND total > $2", parameters=(100, 50.0) ) table = cur.fetch_arrow_table() print(table) ``` The [gospice](https://github.com/spiceai/gospice) SDK (v8+) provides `SqlWithParams()` for parameterized queries: ``` import "github.com/spiceai/gospice/v8" spice := gospice.NewSpiceClient() defer spice.Close() if err := spice.Init(); err != nil { panic(err) } // Query with parameters reader, err := spice.SqlWithParams( context.Background(), "SELECT * FROM customers WHERE c_custkey > $1 LIMIT 10", 100, ) if err != nil { panic(err) } defer reader.Release() for reader.Next() { record := reader.RecordBatch() defer record.Release() fmt.Println(record) } ``` Multiple parameters with different types: ``` reader, err := spice.SqlWithParams( context.Background(), "SELECT * FROM orders WHERE customer_id = $1 AND order_date > $2 AND total > $3", 42, // int "2024-01-01", // string 100.50, // float64 ) ``` The [spice-java](https://github.com/spiceai/spice-java) SDK (v0.5.0+) provides `queryWithParams()` for parameterized queries: ``` import org.apache.arrow.vector.VectorSchemaRoot; import org.apache.arrow.vector.ipc.ArrowReader; import ai.spice.SpiceClient; public class Example { public static void main(String[] args) { try (SpiceClient client = SpiceClient.builder().build()) { // Query with automatic type inference ArrowReader reader = client.queryWithParams( "SELECT * FROM taxi_trips WHERE trip_distance > $1 LIMIT 10", 5.0); // Double is inferred as Float64 while (reader.loadNextBatch()) { VectorSchemaRoot root = reader.getVectorSchemaRoot(); System.out.println(root.contentToTSVString()); } reader.close(); } catch (Exception e) { System.err.println("Error: " + e.getMessage()); } } } ``` Multiple parameters: ``` ArrowReader reader = client.queryWithParams( "SELECT * FROM taxi_trips WHERE trip_distance > $1 AND fare_amount > $2 LIMIT 10", 5.0, 20.0); ``` The [spice-rs](https://github.com/spiceai/spice-rs) SDK (v3.0.0+) provides `query_with_params()`: ``` use spice_rs::Client; use arrow::record_batch::RecordBatch; let client = Client::new("http://localhost:50051").await?; // Create parameter batch let params = RecordBatch::try_new( Arc::new(Schema::new(vec![ Field::new("$1", DataType::Float64, false), ])), vec![Arc::new(Float64Array::from(vec![5.0]))], )?; let batches = client.query_with_params( "SELECT * FROM taxi_trips WHERE trip_distance > $1 LIMIT 10", params ).await?; ``` The [spice-dotnet](https://github.com/spiceai/spice-dotnet) SDK (v1.1.0+) provides `Query()` with dictionary parameters: ``` using Spice; var client = new SpiceClientBuilder().Build(); var parameters = new Dictionary { { "min_distance", 5.0 }, { "min_fare", 20.0 } }; var data = await client.Query( "SELECT * FROM taxi_trips WHERE trip_distance > :min_distance AND fare_amount > :min_fare", parameters ); ``` The HTTP API supports parameterized queries through the `/v1/sql` endpoint: ``` curl -X POST http://localhost:8090/v1/sql \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT * FROM users WHERE id = $1 AND status = $2", "parameters": [42, "active"] }' ``` ## Supported Parameter Types[​](#supported-parameter-types "Direct link to Supported Parameter Types") Spice supports a wide range of parameter types with automatic type inference: | Type Category | Examples | Arrow Type | | ----------------- | ----------------------------------------- | -------------------- | | Signed integers | `int`, `int8`, `int16`, `int32`, `int64` | Int8/16/32/64 | | Unsigned integers | `uint`, `uint8`, `uint16`, `uint32` | Uint8/16/32/64 | | Floating point | `float`, `double` | Float32/Float64 | | Text | `string` | Utf8 | | Boolean | `bool`, `boolean` | Boolean | | Binary | `byte[]`, `[]byte` | Binary | | Date/Time | `LocalDate`, `LocalDateTime`, `time.Time` | Date32/64, Timestamp | | Decimal | `BigDecimal`, `decimal` | Decimal128/256 | | Null | `null`, `nil` | Null | ## Explicit Type Control[​](#explicit-type-control "Direct link to Explicit Type Control") For precise control over Arrow types, SDKs provide typed parameter constructors: * Go * Java * Rust ``` import "github.com/spiceai/gospice/v8" reader, err := spice.SqlWithParams( ctx, "SELECT * FROM financial WHERE amount >= $1 AND timestamp > $2", gospice.Decimal128Param(amountBytes, 19, 4), // Decimal with precision & scale gospice.TimestampParam(ts, arrow.Microsecond, "UTC"), // Timestamp with unit & timezone ) ``` Available constructors: `Int8Param`, `Int16Param`, `Int32Param`, `Int64Param`, `Uint8Param`, `Uint16Param`, `Uint32Param`, `Uint64Param`, `Float16Param`, `Float32Param`, `Float64Param`, `StringParam`, `LargeStringParam`, `BinaryParam`, `LargeBinaryParam`, `FixedSizeBinaryParam`, `BoolParam`, `Date32Param`, `Date64Param`, `Time32Param`, `Time64Param`, `TimestampParam`, `DurationParam`, `MonthIntervalParam`, `DayTimeIntervalParam`, `MonthDayNanoIntervalParam`, `Decimal128Param`, `Decimal256Param`, `NullParam` ``` import ai.spice.Param; ArrowReader reader = client.queryWithParams( "SELECT * FROM orders WHERE order_id = $1 AND amount >= $2", Param.int64(12345), Param.decimal128(new BigDecimal("99.99"), 10, 2)); ``` Available constructors: `int8`, `int16`, `int32`, `int64`, `uint8`, `uint16`, `uint32`, `uint64`, `float16`, `float32`, `float64`, `string`, `largeString`, `binary`, `largeBinary`, `fixedSizeBinary`, `bool`, `date32`, `date64`, `time32`, `time64`, `timestamp`, `duration`, `decimal128`, `decimal256`, `nullValue` You can also use generic constructors: `Param.of(value)` for automatic type inference, or `Param.of(value, arrowType)` for explicit Arrow type. Rust uses Arrow `RecordBatch` for parameters, giving full control over the Arrow schema: ``` use arrow::datatypes::{DataType, Field, Schema, TimeUnit}; use arrow::array::{TimestampMicrosecondArray, Decimal128Array}; let params = RecordBatch::try_new( Arc::new(Schema::new(vec![ Field::new("$1", DataType::Decimal128(19, 4), false), Field::new("$2", DataType::Timestamp(TimeUnit::Microsecond, Some("UTC".into())), false), ])), vec![ Arc::new(Decimal128Array::from(vec![Some(9999)])), Arc::new(TimestampMicrosecondArray::from(vec![Some(1704067200000000)])), ], )?; ``` ## Security Benefits[​](#security-benefits "Direct link to Security Benefits") Parameterized queries protect against SQL injection by ensuring user input is never interpreted as SQL code. * Go * Java * Python (ADBC) ``` // ❌ DANGEROUS: User input directly in SQL string userId := getUserInput() // Could be: "1 OR 1=1" sql := fmt.Sprintf("SELECT * FROM users WHERE id = %s", userId) reader, err := spice.Sql(ctx, sql) // ✅ SAFE: User input passed as parameter userId := getUserInput() reader, err := spice.SqlWithParams(ctx, "SELECT * FROM users WHERE id = $1", userId) ``` ``` // ❌ DANGEROUS: User input directly in SQL string String userId = getUserInput(); // Could be: "1 OR 1=1" String sql = "SELECT * FROM users WHERE id = " + userId; FlightStream stream = client.query(sql); // ✅ SAFE: User input passed as parameter ArrowReader reader = client.queryWithParams( "SELECT * FROM users WHERE id = $1", userId); ``` ``` # ❌ DANGEROUS: User input directly in SQL string user_id = get_user_input() # Could be: "1 OR 1=1" sql = f"SELECT * FROM users WHERE id = {user_id}" cur.execute(sql) # ✅ SAFE: User input passed as parameter cur.execute("SELECT * FROM users WHERE id = $1", parameters=(user_id,)) ``` With parameterized queries, even if `userId` contains malicious SQL like `1 OR 1=1`, it is treated as a literal string value, not as SQL code. ## Performance Considerations[​](#performance-considerations "Direct link to Performance Considerations") Parameterized queries can improve performance in several ways: 1. **Query Plan Caching**: The database can cache and reuse execution plans for parameterized queries 2. **Reduced Parsing**: Parameter binding avoids repeated SQL parsing 3. **Batch Operations**: Multiple executions with different parameters use the same prepared statement ## SDK Support[​](#sdk-support "Direct link to SDK Support") | SDK | Version | Method | Status | | ---------------------------------------------------------------- | ------- | ------------------------------------------- | ------- | | [gospice](https://github.com/spiceai/gospice) (Go) | v8.0.0+ | `SqlWithParams()` with typed constructors | ✅ Full | | [spice-rs](https://github.com/spiceai/spice-rs) (Rust) | v3.0.0+ | `query_with_params()` with `RecordBatch` | ✅ Full | | [spice-dotnet](https://github.com/spiceai/spice-dotnet) (Dotnet) | v1.1.0+ | `Query()` with `Dictionary` | ✅ Full | | [spice-java](https://github.com/spiceai/spice-java) (Java) | v0.5.0+ | `queryWithParams()` with `Param` class | ✅ Full | | [spice.js](https://github.com/spiceai/spice.js) (JavaScript) | v3.1.0+ | `sql()` with positional or named parameters | ✅ Full | | [spicepy](https://github.com/spiceai/spicepy) (Python) | v3.1.0+ | `query_with_params()` with `Param` class | ✅ Full | | ADBC (Python) | - | `cursor.execute()` with `parameters` | ✅ Full | | JDBC | - | `PreparedStatement` | ✅ Full | | ODBC | - | Parameterized queries | ✅ Full | ## Examples[​](#examples "Direct link to Examples") ### Filtering with Multiple Conditions[​](#filtering-with-multiple-conditions "Direct link to Filtering with Multiple Conditions") ``` # Python with ADBC cur.execute(""" SELECT order_id, customer_name, total FROM orders WHERE status = $1 AND order_date BETWEEN $2 AND $3 AND total > $4 ORDER BY order_date DESC LIMIT $5 """, parameters=("completed", "2024-01-01", "2024-12-31", 100.0, 50)) ``` ### Batch Processing Pattern[​](#batch-processing-pattern "Direct link to Batch Processing Pattern") ``` // Go: Process multiple customers with same query customerIDs := []int{100, 200, 300, 400, 500} for _, id := range customerIDs { reader, err := spice.SqlWithParams( ctx, "SELECT * FROM orders WHERE customer_id = $1", id, ) if err != nil { log.Printf("Error querying customer %d: %v", id, err) continue } processOrders(reader) reader.Release() } ``` ### Using with Aggregations[​](#using-with-aggregations "Direct link to Using with Aggregations") ``` # Calculate statistics for a specific time range cur.execute(""" SELECT DATE_TRUNC('day', order_date) as day, COUNT(*) as order_count, SUM(total) as daily_total FROM orders WHERE order_date >= $1 AND order_date < $2 GROUP BY DATE_TRUNC('day', order_date) ORDER BY day """, parameters=("2024-01-01", "2024-02-01")) ``` ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") ### Parameter Count Mismatch[​](#parameter-count-mismatch "Direct link to Parameter Count Mismatch") Ensure the number of parameters matches the number of placeholders: ``` # Error: 3 placeholders but only 2 parameters cur.execute("SELECT * FROM t WHERE a = $1 AND b = $2 AND c = $3", parameters=(1, 2)) # Correct: 3 placeholders and 3 parameters cur.execute("SELECT * FROM t WHERE a = $1 AND b = $2 AND c = $3", parameters=(1, 2, 3)) ``` ### Type Mismatch[​](#type-mismatch "Direct link to Type Mismatch") If you encounter type errors, use explicit type constructors (Go) or ensure Python types match expected column types: ``` // If int is inferred as Int64 but column expects Int32 reader, err := spice.SqlWithParams(ctx, "SELECT $1", gospice.Int32Param(42)) ``` ### Connection Issues[​](#connection-issues "Direct link to Connection Issues") For ADBC connections, verify the Spice runtime is running and accessible: ``` # Test connection try: conn = connect("grpc://localhost:50051") cursor = conn.cursor() cursor.execute("SELECT 1") print("Connection successful") except Exception as e: print(f"Connection failed: {e}") ``` ## Related Topics[​](#related-topics "Direct link to Related Topics") * [ADBC API](/docs/next/api/adbc) - Arrow Database Connectivity documentation * [Arrow Flight SQL API](/docs/next/api/arrow-flight-sql) - Flight SQL protocol details * [Go SDK](/docs/next/sdks/golang) - gospice SDK documentation * [Python SDK](/docs/next/sdks/python) - spicepy SDK documentation * [Results Caching](/docs/next/features/caching) - Query result caching --- # URL Tables URL tables enable querying files in object stores directly using their URLs, without pre-registering datasets in a Spicepod. This provides an ad-hoc query capability for exploring data stored in S3, Azure Blob Storage, or HTTP endpoints. ## Enabling URL Tables[​](#enabling-url-tables "Direct link to Enabling URL Tables") URL tables are disabled by default and must be explicitly enabled in the Spicepod configuration: ``` runtime: params: url_tables: enabled ``` ## Supported URL Schemes[​](#supported-url-schemes "Direct link to Supported URL Schemes") | Scheme | Description | Example | | ---------- | ---------------------------- | ------------------------------------------------------ | | `s3://` | Amazon S3 | `s3://bucket/path/file.parquet` | | `abfs://` | Azure Blob Storage | `abfs://container@account/path/file.parquet` | | `abfss://` | Azure Data Lake Storage Gen2 | `abfss://container@account.dfs.core.windows.net/path/` | | `https://` | HTTPS endpoints | `https://example.com/data.parquet` | | `http://` | HTTP endpoints | `http://localhost:8080/data.csv` | ## Query Patterns[​](#query-patterns "Direct link to Query Patterns") ### Single File[​](#single-file "Direct link to Single File") Query a single file by specifying its full URL: ``` SELECT * FROM 's3://my-bucket/data/sales.parquet' LIMIT 10; ``` ### Directory or Prefix[​](#directory-or-prefix "Direct link to Directory or Prefix") Query all files under a directory or prefix by including a trailing slash: ``` -- All files in a directory SELECT * FROM 's3://my-bucket/data/'; -- All files in a bucket SELECT * FROM 's3://my-bucket/'; ``` ### Glob Patterns[​](#glob-patterns "Direct link to Glob Patterns") Use glob patterns to match specific files: ``` -- All parquet files in a directory SELECT * FROM 's3://my-bucket/data/*.parquet'; -- Files matching a pattern across subdirectories SELECT * FROM 's3://my-bucket/year=2024/month=*/data.parquet'; ``` ### Hive-Style Partitions[​](#hive-style-partitions "Direct link to Hive-Style Partitions") Hive-style partitions are automatically inferred from the path structure, enabling partition pruning: ``` -- If data is stored at s3://bucket/data/year=2024/month=01/file.parquet -- the year and month columns are available for filtering SELECT * FROM 's3://my-bucket/data/' WHERE year = '2024' AND month = '01'; ``` ## Authentication[​](#authentication "Direct link to Authentication") URL tables use the same authentication mechanisms as the corresponding data connectors. Credentials are loaded automatically from environment variables or cloud provider defaults. ### S3[​](#s3 "Direct link to S3") For S3, credentials are loaded from: 1. Environment variables: `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN` 2. Shared AWS credentials file (`~/.aws/credentials`) 3. IAM instance profiles or roles For public buckets, no authentication is required. ### Azure Blob Storage[​](#azure-blob-storage "Direct link to Azure Blob Storage") For Azure, set the storage account name via environment variable: ``` export AZURE_STORAGE_ACCOUNT=mystorageaccount ``` Alternatively, include the account name in the URL: ``` SELECT * FROM 'abfss://container@mystorageaccount.dfs.core.windows.net/path/file.parquet'; ``` Additional authentication options: * Environment variable: `AZURE_STORAGE_KEY` for access key authentication * Azure Managed Identity (automatic when running on Azure) * Azure CLI credentials ## Examples[​](#examples "Direct link to Examples") ### S3 Query[​](#s3-query "Direct link to S3 Query") ``` runtime: params: url_tables: enabled ``` ``` -- Query a public S3 dataset SELECT VendorID, passenger_count, trip_distance FROM 's3://spiceai-public-datasets/taxi_small_samples/taxi_sample.parquet' LIMIT 5; ``` ### Azure Blob Storage Query[​](#azure-blob-storage-query "Direct link to Azure Blob Storage Query") ``` runtime: params: url_tables: enabled ``` Set the account via environment variable: ``` export AZURE_STORAGE_ACCOUNT=mystorageaccount export AZURE_STORAGE_KEY=${your_access_key} ``` Or include the account in the URL: ``` SELECT * FROM 'abfss://mycontainer@mystorageaccount.dfs.core.windows.net/data/' LIMIT 10; ``` ### Cross-Source Query[​](#cross-source-query "Direct link to Cross-Source Query") URL tables can be combined with registered datasets in federated queries: ``` runtime: params: url_tables: enabled datasets: - from: postgres:orders name: orders params: pg_host: localhost pg_db: mydb ``` ``` -- Join a registered dataset with an ad-hoc S3 query SELECT o.order_id, o.customer_id, s.product_name FROM orders o JOIN 's3://my-bucket/products.parquet' s ON o.product_id = s.id; ``` ## Considerations[​](#considerations "Direct link to Considerations") * **Schema Inference**: The schema is inferred from the files at query time. For best performance with large datasets, consider registering datasets in the Spicepod. * **File Format Detection**: File formats are automatically inferred from file extensions. Supported formats include Parquet, CSV, and JSON. * **Performance**: URL tables query data directly from the object store without local acceleration. For frequently accessed data or performance-critical queries, register datasets with [data acceleration](/docs/next/components/data-accelerators). * **Authentication Scope**: URL table queries use environment-level credentials. For queries requiring different credentials per source, register datasets with explicit authentication parameters. ## Related Topics[​](#related-topics "Direct link to Related Topics") * [S3 Data Connector](/docs/next/components/data-connectors/s3) - Register S3 datasets with full configuration options * [Azure BlobFS Data Connector](/docs/next/components/data-connectors/abfs) - Register Azure datasets with full configuration options * [Query Federation](/docs/next/features/query-federation/) - Learn about federated queries across multiple sources * [Data Acceleration](/docs/next/components/data-accelerators) - Accelerate query performance with local caching --- # Search Functionality > 🎓 For a practical walkthrough, see the: [Amazon S3 Vectors with Spice](https://spice.ai/blog/amazon-s3-vectors-with-spice) engineering blog post. Spice provides comprehensive search capabilities enabling developers to query datasets beyond traditional SQL, including semantic (vector-based) search, full-text keyword search, and hybrid search methods. ## Search Methods Overview[​](#search-methods-overview "Direct link to Search Methods Overview") Spice supports multiple search methods: * **Vector Search**: Semantic search using embeddings to retrieve data by meaning and similarity. * **Multi-Vector Search**: Search over columns of vectors, including ColBERT-style late-interaction queries. * **Full-Text Search**: Keyword-driven search optimized for text data retrieval. * **Hybrid Search**: Combine multiple search methods using Reciprocal Rank Fusion (RRF) for improved relevance. * **Reranking**: Reorder search results using dedicated reranker models or LLM-as-reranker for improved relevance. * **SQL Search**: Traditional SQL queries for precise and structured searches. ### Vector Search[​](#vector-search "Direct link to Vector Search") Vector search uses embeddings—numerical representations of data—to identify similar or related content based on semantic meaning. **Requirements:** * Configured data connectors or accelerators * Defined embeddings for datasets **Getting Started:** * [Configure Embeddings](/docs/next/components/embeddings) * [Performing Vector Search](/docs/next/features/search/vector-search) **Example SQL Vector Search:** ``` SELECT id, extra_column, score FROM vector_search(my_table, 'search query') WHERE date_published > '2021-01-01' ORDER BY score DESC LIMIT 5; ``` For complete SQL UDTF specifications, see [Vector-Based Search SQL UDTF](/docs/next/features/search/vector-search#sql-udtf). ### Multi-Vector Search[​](#multi-vector-search "Direct link to Multi-Vector Search") Multi-vector search operates on columns that store many vectors per row, such as per-tag or per-section embeddings. It also supports ColBERT-style late-interaction queries where the query itself is an array of strings. **Requirements:** * A list-typed source column (`List`) embedded with a multi-vector aggregation **Getting Started:** * [Multi-Vector Search Docs](/docs/next/features/search/multi-vector) **Example SQL Multi-Vector Search:** ``` SELECT product_id, name, score FROM vector_search(products, ['hiking', 'waterproof'], tags) ORDER BY score DESC LIMIT 10; ``` ### Full-Text Search[​](#full-text-search "Direct link to Full-Text Search") Full-text search efficiently retrieves records matching specific keywords. **Requirements:** * Indexed columns within datasets **Getting Started:** * [Full-Text Search Docs](/docs/next/features/search/full-text) **Example SQL Full-Text Search:** ``` SELECT id, extra_column, score FROM text_search(my_table, 'search terms') WHERE date_published > '2021-01-01' ORDER BY score DESC LIMIT 5; ``` For detailed SQL UDTF instructions, see [Full-Text Search SQL UDTF](/docs/next/features/search/full-text#searching-with-sql). ### Hybrid Search with RRF[​](#hybrid-search-with-rrf "Direct link to Hybrid Search with RRF") Reciprocal Rank Fusion (RRF) combines results by merging rankings from multiple search methods to improve relevance. This is useful when neither vector search nor full-text search alone provides optimal results. **Requirements:** * Multiple search methods configured (vector, full-text, etc.) **When to use hybrid search:** * The query contains both semantic concepts and specific keywords. * Results from a single method are missing relevant documents. * Improved ranking is needed across diverse content types. **Example SQL Hybrid Search:** ``` SELECT id, title, content, fused_score FROM rrf( vector_search(documents, 'machine learning algorithms'), text_search(documents, 'neural networks deep learning', content), join_key => 'id' -- join key for optimal performance ) ORDER BY fused_score DESC LIMIT 5; ``` For complete RRF syntax and parameters, see [Search SQL Reference](/docs/next/reference/sql/search#reciprocal-rank-fusion-rrf). ### Reranking[​](#reranking "Direct link to Reranking") Reranking reorders search results using a dedicated reranker model (Cohere, Voyage, Jina, or a custom HTTP endpoint) or any registered chat model as an LLM-as-reranker. This two-stage retrieve-then-rerank pattern improves relevance beyond initial retrieval scores. **Requirements:** * A registered reranker (in the `rerankers:` spicepod section) or a registered chat model **Getting Started:** * [Reranking Docs](/docs/next/features/search/rerank) **Example SQL Rerank:** ``` SELECT * FROM rerank( rrf( vector_search(docs, 'delta lake time travel', limit => 50), text_search(docs, 'delta lake time travel', limit => 50) ), document => 'content', model => 'cohere_rr', limit => 10 ); ``` For complete rerank syntax and parameters, see [Search SQL Reference](/docs/next/reference/sql/search#reranking-rerank). ## Dataset Readiness[​](#dataset-readiness "Direct link to Dataset Readiness") An accelerated dataset cannot be searched until its initial load completes. With the default [`ready_state: on_load`](/docs/next/reference/spicepod/datasets#ready_state), searching a dataset that is still loading fails with the same error a SQL query returns: ``` Acceleration not ready; loading initial data for ``` This applies to every search method — vector search (whether served just-in-time or from a vector index) and full-text search — so a partially-loaded index never answers with the subset that happens to have loaded so far. When a search request does not name any datasets and instead sweeps every searchable dataset, a dataset that is not ready is excluded from the results rather than failing the whole request. Naming that dataset explicitly returns the error above. Datasets configured `ready_state: on_registration` or `ready_state: on_schema_resolved` are unaffected: they remain searchable while the accelerator loads, served from the federated source. Datasets accelerated with `refresh_mode: caching` are also exempt, as they serve queries without an initial load. ## [📄️Vector Search](/docs/next/features/search/vector-search) [Learn how Spice can perform searches using vector-based methods.](/docs/next/features/search/vector-search) ## [📄️Full-text Search](/docs/next/features/search/full-text) [Learn how Spice can perform full text search](/docs/next/features/search/full-text) ## [📄️Multi-Vector Search](/docs/next/features/search/multi-vector) [Embed list-of-strings columns as a column of vectors and use ColBERT-style late-interaction search in Spice.](/docs/next/features/search/multi-vector) ## [📄️Reranking](/docs/next/features/search/rerank) [Rerank search results using dedicated reranker models or LLM-as-reranker for improved relevance.](/docs/next/features/search/rerank) --- # Full-Text Search Spice provides full-text search functionality with BM25 scoring. This search method is optimized for keyword-based queries and is useful when: * Users search for specific terms or phrases * Exact keyword matching is important * Searching structured text fields like titles, tags, or names Datasets can be augmented with a full-text search index that enables efficient search. Dataset columns are included in the full-text index based on the column configuration. ## Engines[​](#engines "Direct link to Engines") Spice supports two full-text search engines: | Engine | Description | | --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Tantivy** (default) | Built-in, in-process BM25 engine. No external dependencies. | | **Elasticsearch** | Indexes into an external Elasticsearch cluster, fronted by a local Tantivy warm tier that serves searches. Useful when Elasticsearch is already part of the infrastructure or when its operational characteristics (sharding, replication, snapshots) are preferred. | When no engine is specified, Tantivy is used automatically. ### Text Analysis[​](#text-analysis "Direct link to Text Analysis") The built-in Tantivy engine indexes text with Tantivy's `en_stem` tokenizer: terms are lowercased and reduced to their English (Snowball) stem, with token positions retained so phrase queries keep working. Query terms are analyzed the same way, so a search for `running` also matches documents containing `run` and `runs`. Stemming is always on for the built-in engine and has no configuration parameter. It is English-only — text in other languages is still tokenized and lowercased, but not stemmed. The local warm index used with `engine: elasticsearch` is a Tantivy index and analyzes text the same way. Searches [served directly by Elasticsearch](#warm-tier) — multi-column datasets, a warm index that could not be built, or an `on_zero_results: use_source` fallback — are analyzed by Elasticsearch's own analyzer for that index instead. ## Enabling Full-Text Search[​](#enabling-full-text-search "Direct link to Enabling Full-Text Search") To enable full-text search, configure your dataset columns within your dataset definition as follows: ``` datasets: - from: github:github.com/spiceai/docs/pulls name: doc.pulls params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true columns: - name: title full_text_search: enabled: true row_id: - id - name: body full_text_search: enabled: true ``` In this example, full-text search indexing is enabled on both the `title` and `body` columns using the default Tantivy engine. The `row_id` specifies a unique identifier for referencing search results and retrieving additional data. ### Index Storage[​](#index-storage "Direct link to Index Storage") By default the built-in Tantivy index is held in memory and rebuilt on every start. Set `index_store: file` on a column to persist it to disk instead, so a restart reopens the existing index rather than re-indexing the dataset: ``` columns: - name: body full_text_search: enabled: true index_store: file # Optional. Defaults to `.spice/data/fts///
/` index_directory: ./my-index ``` | Field | Default | Description | | ----------------- | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `index_store` | `memory` | Where the index lives: `memory` (rebuilt on every start) or `file` (persisted on disk). If any indexed column of a dataset sets `file`, the dataset's index is persisted. | | `index_directory` | `.spice/data/fts///
/` | Directory for a persisted index. Only applies with `index_store: file` — combining it with `index_store: memory` logs a warning and keeps the in-memory store. | #### Changing the configuration of a persisted index[​](#changing-the-configuration-of-a-persisted-index "Direct link to Changing the configuration of a persisted index") A persisted index records the schema it was built with, and that schema is what serves queries. On start, Spice compares the persisted schema against the current configuration: * **The index fails to open** — with an error naming the column, how its indexing changed, and the directory to delete so the index is rebuilt — when a configured column is absent from the persisted index, or when a column's value type, indexed flag, or tokenized/untokenized indexing differs. Searches, filters, and primary-key deletes over such a column cannot behave as configured, so this is reported rather than served. Adding a search column to a dataset with a persisted index therefore requires deleting the index directory. * **A warning is logged** when only the text analysis differs (for example an index built before [stemming](#text-analysis) became the default). Queries stay consistent with what the index actually holds; delete the directory to rebuild with the configured analysis. `index_store: memory` is never affected — it rebuilds from scratch on every start and so always matches the configuration. ### Filtering on other columns[​](#filtering-on-other-columns "Direct link to Filtering on other columns") Beyond the indexed text, the built-in index also carries the dataset's primary key (or `row_id`) and every column declared with [`metadata.vectors`](/docs/next/reference/spicepod/datasets#columnsmetadatavectors). Those columns are returned by searches and can be filtered on, and a `WHERE` predicate over them is applied **inside** the index scan — before the search's row limit, so a filtered search returns the best matches among the rows that pass the filter rather than whatever survives filtering the unfiltered top matches: ``` columns: - name: body full_text_search: enabled: true row_id: - id - name: state metadata: vectors: filterable ``` ``` SELECT id, score FROM text_search(doc.pulls, 'search keywords', body, 5) WHERE state = 'open'; ``` A column whose type the index cannot represent — a date or timestamp — is skipped when the index is built, with a warning naming the column; predicates on it are applied above the scan as before. See [Filter Pushdown](/docs/next/reference/sql/search#full-text-filter-pushdown) for the full set of predicates the index applies. ## Using Elasticsearch as the FTS Engine[​](#using-elasticsearch-as-the-fts-engine "Direct link to Using Elasticsearch as the FTS Engine") To use Elasticsearch instead of the built-in Tantivy engine, add a dataset-level `full_text_search` block with `engine: elasticsearch` and the connection parameters: ``` datasets: - from: file:./articles.parquet name: articles acceleration: enabled: true engine: arrow full_text_search: engine: elasticsearch params: elasticsearch_endpoint: http://localhost:9200 elasticsearch_user: ${secrets:ES_USER} elasticsearch_pass: ${secrets:ES_PASS} elasticsearch_index: articles-fts columns: - name: title full_text_search: enabled: true row_id: - id - name: body full_text_search: enabled: true row_id: - id ``` The dataset-level `full_text_search` block selects the engine and provides connection parameters. Column-level `full_text_search.enabled` controls which columns are indexed. #### Warm tier[​](#warm-tier "Direct link to Warm tier") With `engine: elasticsearch`, Spice maintains a **local Tantivy warm index in front of the Elasticsearch index**: * **Writes fan out to both tiers** — every indexed row is written to the local Tantivy index and to Elasticsearch. * **Searches are served from the local warm index.** Elasticsearch is queried only as a fallback, and only when the dataset sets [`acceleration.on_zero_results: use_source`](/docs/next/reference/spicepod/datasets#accelerationon_zero_results); with the default `return_empty`, searches are served from the warm tier alone. Two cases fall back to querying Elasticsearch directly, each logged at startup: * **More than one full-text column on the dataset.** The warm index is single-column, so multi-column datasets keep the Elasticsearch-only behavior. * **The warm index cannot be built or paired with the Elasticsearch index.** Warm-tier construction never fails dataset load — a warning is logged and Elasticsearch is registered alone. A primary key whose type Elasticsearch normalizes (for example `LargeUtf8` → `Utf8`) is a common cause, because the two tiers must agree on the search column and primary-key fields. Enterprise edition The Elasticsearch full-text search engine is available in the Spice [Enterprise edition](https://docs.spice.ai/docs/enterprise/getting-started/distributions). ### Elasticsearch FTS Parameters[​](#elasticsearch-fts-parameters "Direct link to Elasticsearch FTS Parameters") | Parameter | Description | Example | | ------------------------ | ------------------------------------------------------------------------ | ----------------------- | | `elasticsearch_endpoint` | Required. Elasticsearch cluster URL. | `http://localhost:9200` | | `elasticsearch_user` | Optional. Username for HTTP basic authentication. | `${secrets:ES_USER}` | | `elasticsearch_pass` | Optional. Password for HTTP basic authentication. | `${secrets:ES_PASS}` | | `elasticsearch_index` | Optional. ES index name for FTS documents. Defaults to the dataset name. | `articles-fts` | | `client_timeout` | Optional. Total HTTP request timeout. Default: `30s`. | `30s` | | `connect_timeout` | Optional. HTTP connect timeout. Default: `10s`. | `10s` | ### Elasticsearch Ingestion Tuning[​](#elasticsearch-ingestion-tuning "Direct link to Elasticsearch Ingestion Tuning") Optional parameters to control Elasticsearch index creation and write behavior: | Parameter | Description | Default | | ---------------------------- | ----------------------------------------------------------------------------------------------- | ------------------------------ | | `number_of_shards` | ES `number_of_shards` index setting (applied at index creation). | ES default | | `number_of_replicas` | ES `number_of_replicas` index setting (applied at index creation). | ES default | | `refresh_interval` | ES `refresh_interval` index setting (applied at index creation). | ES default | | `bulk_load_refresh_interval` | Temporary `refresh_interval` during bulk writes. Set to `-1` to disable refresh during loading. | Not set | | `force_merge_after_write` | Run `_forcemerge` after full/append writes. | `false` | | `force_merge_segments` | Max segments for `_forcemerge`. Setting this also enables force merge. | `1` (when force merge enabled) | | `batch_write_rows` | Max rows per `_bulk` request. | `1000` | | `index_settings` | JSON object passed as ES index settings at creation. | Not set | ### YAML Anchor Reuse[​](#yaml-anchor-reuse "Direct link to YAML Anchor Reuse") When multiple datasets or columns share the same Elasticsearch connection, use YAML anchors to avoid repeating config: ``` x-elasticsearch-fts: &elasticsearch_fts enabled: true engine: elasticsearch params: elasticsearch_endpoint: http://localhost:9200 elasticsearch_user: ${secrets:ES_USER} elasticsearch_pass: ${secrets:ES_PASS} datasets: - from: file:./articles.parquet name: articles acceleration: enabled: true full_text_search: <<: *elasticsearch_fts params: elasticsearch_endpoint: http://localhost:9200 elasticsearch_index: articles-fts columns: - name: title full_text_search: enabled: true row_id: - id ``` ### Combining with the Elasticsearch Vector Engine[​](#combining-with-the-elasticsearch-vector-engine "Direct link to Combining with the Elasticsearch Vector Engine") Elasticsearch can serve as both the vector engine and the FTS engine for the same dataset. Configure `vectors` and `full_text_search` independently: ``` datasets: - from: file:./articles.parquet name: articles acceleration: enabled: true vectors: enabled: true engine: elasticsearch params: elasticsearch_endpoint: http://localhost:9200 elasticsearch_index: articles-vectors full_text_search: engine: elasticsearch params: elasticsearch_endpoint: http://localhost:9200 elasticsearch_index: articles-fts columns: - name: body embeddings: - from: my_embedding_model row_id: - id full_text_search: enabled: true row_id: - id ``` Use [`rrf()`](/docs/next/reference/sql/search#reciprocal-rank-fusion-rrf) to combine vector and full-text results with hybrid search. ## Searching with the HTTP API[​](#searching-with-the-http-api "Direct link to Searching with the HTTP API") After enabling indexing, you can perform searches using the HTTP API endpoint `/v1/search`. Results will be ranked based on the relevance to your keyword query across indexed columns (`title` and `body` in this example). For details on using this endpoint, see the [API reference for `/v1/search`](/docs/next/api/HTTP/post-search). ## Searching with SQL[​](#searching-with-sql "Direct link to Searching with SQL") Spice also provides full-text search through SQL using a user-defined table function (UDTF), `text_search()`. ### Example SQL Query[​](#example-sql-query "Direct link to Example SQL Query") Here's how you can query using SQL: ``` SELECT id, title, score FROM text_search(doc.pulls, 'search keywords', body) ORDER BY score DESC LIMIT 5; ``` This returns the top 5 results from the `doc.pulls` dataset that best match your search keywords within the `body` column. ### Function Signature[​](#function-signature "Direct link to Function Signature") The `text_search()` function has the following signature: ``` text_search( table IDENTIFIER, -- Dataset name (required, unquoted) query STRING, -- Keyword or phrase to search (required) col IDENTIFIER, -- Column name to search (required if dataset has multiple indexed columns, unquoted) limit INTEGER, -- Maximum results returned (optional, defaults to 1000) include_score BOOLEAN -- Include relevance scores in results (optional, defaults to TRUE) ) RETURNS TABLE -- Original table columns plus an optional FLOAT column `score` ``` By default, `text_search` retrieves up to 1000 results. To adjust this, specify the `limit` parameter in the function call. Use this function to integrate full-text search directly into your data workflows. --- # Multi-Vector Search A multi-vector column stores many embedding vectors per row rather than a single vector. Spice produces a multi-vector column by embedding each element of a `List` source column independently, yielding a `List>` embedding column. Multi-vector embeddings are useful when a single row has several distinct pieces of text — for example, a product with many tags, a paper with multiple titles and section headings, or a user with a set of historical queries. Each element is embedded and scored separately, and per-row results are produced by aggregating the per-element similarities. ## How Multi-Vector Differs from Chunking[​](#how-multi-vector-differs-from-chunking "Direct link to How Multi-Vector Differs from Chunking") Chunking splits one long string (such as a document body) into pieces and embeds each piece. Multi-vector starts from a column that is already a list of independent strings and embeds each list element as-is. | Source column type | Embedding mode | Produced embedding type | | ------------------ | ---------------------- | --------------------------------- | | `Utf8` | Scalar (default) | `FixedSizeList` | | `Utf8` + chunking | Chunked | `List>` | | `List` | Multi-vector (default) | `List>` | Multi-vector and chunked columns share the same Arrow type, but the per-element offsets column (`_offsets`) is only produced for chunked columns. ## Configuring a Multi-Vector Column[​](#configuring-a-multi-vector-column "Direct link to Configuring a Multi-Vector Column") Define an embedding on a `List` column the same way as a scalar string column. Spice detects the list type and embeds each element independently. ``` datasets: - from: file:products.parquet name: products acceleration: enabled: true columns: - name: tags # List embeddings: - from: local_embedding_model aggregation: max max_elements_per_row: 64 embeddings: - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: local_embedding_model ``` ### Aggregation Strategies[​](#aggregation-strategies "Direct link to Aggregation Strategies") When a multi-vector column is queried with a single query string, each element's similarity to the query is computed, and the per-row score is the aggregate of those similarities. | `aggregation` | Description | | ------------- | ---------------------------------------------------------------------------------- | | `max` | ColBERT-style `MaxSim`. Row scores as high as its best-matching element (default). | | `mean` | Average similarity across elements. Favors rows where most elements are relevant. | | `sum` | Sum of similarities. Biases toward rows with many matching elements. | ### Element Caps[​](#element-caps "Direct link to Element Caps") Multi-vector columns default to embedding the first 32 elements per row. Raise the cap with `max_elements_per_row` (hard-capped at `1024`). Excess elements are dropped with a warning log so that rows with unbounded tag counts do not blow up embedding cost. ## Querying with `vector_search`[​](#querying-with-vector_search "Direct link to querying-with-vector_search") A multi-vector column is queried with the standard `vector_search` UDTF. The configured `aggregation` is applied automatically. ``` SELECT product_id, name, score FROM vector_search(products, 'travel accessories', tags) ORDER BY score DESC LIMIT 10; ``` ## Late-Interaction (Multi-Query) Search[​](#late-interaction-multi-query-search "Direct link to Late-Interaction (Multi-Query) Search") Multi-vector columns also support ColBERT-style late-interaction search, where the query itself is an array of strings. Each query is embedded independently, the best-matching element is selected for each query (`MaxSim`), and the per-row score is the sum across queries: ``` score(d) = Σ_{q ∈ Q} max_{e ∈ d} cos(q, e) ``` ``` SELECT product_id, name, score FROM vector_search( products, ['hiking', 'waterproof', 'lightweight'], tags ) ORDER BY score DESC LIMIT 10; ``` Late-interaction search is only supported on multi-vector columns; passing an array of queries to a scalar or chunked column returns an error. A maximum of 32 query strings are accepted per call. ## Passthrough Multi-Vector Columns[​](#passthrough-multi-vector-columns "Direct link to Passthrough Multi-Vector Columns") Datasets that already contain multi-vector columns can be used directly when their schema matches the conventions in [Vector-Based Search](/docs/next/features/search/vector-search#using-existing-embeddings): * Column name: `_embedding` * Type: `List>` * No offsets column (that is only required for chunked scalar columns) Declare the underlying column's embedding in `spicepod.yaml` so that Spice knows which embedding model the existing vectors came from. ## Limitations[​](#limitations "Direct link to Limitations") * Multi-vector embeddings require the source column to be `List` or `LargeList`. * Late-interaction search accepts at most 32 query strings per call. * Multi-vector columns cannot currently be stored in an external vector engine; use a [data accelerator](/docs/next/components/data-accelerators) with `acceleration.enabled: true` to cache embeddings. --- # Reranking Reranking reorders a set of candidate results — from `vector_search`, `text_search`, `rrf`, or a plain table — using a dedicated reranker model. This produces more relevant rankings than the initial retrieval scores alone. ## How It Works[​](#how-it-works "Direct link to How It Works") 1. An initial search (vector, full-text, or hybrid) retrieves candidate documents. 2. The `rerank()` UDTF sends each candidate's text to a reranker model alongside the query. 3. The reranker scores each document for relevance and returns them in order. This two-stage pattern (retrieve then rerank) is standard in modern search and RAG pipelines. ## Configuration[​](#configuration "Direct link to Configuration") ### Reranker Models[​](#reranker-models "Direct link to Reranker Models") Define dedicated reranker models in the `rerankers:` section of the spicepod. Supported providers: | Provider prefix | Description | | ---------------------- | ----------------------------------------------------------------- | | `cohere:` | Cohere Rerank API (e.g., `cohere:rerank-v3.5`) | | `voyage:` | Voyage Rerank API (e.g., `voyage:rerank-2`) | | `jina:` | Jina Rerank API (e.g., `jina:jina-reranker-v2-base-multilingual`) | | `http://` / `https://` | Any HTTP endpoint implementing the standard rerank API schema | ``` rerankers: - from: cohere:rerank-v3.5 name: cohere_rr params: api_key: ${secrets:COHERE_API_KEY} - from: voyage:rerank-2 name: voyage_rr params: api_key: ${secrets:VOYAGE_API_KEY} - from: jina:jina-reranker-v2-base-multilingual name: jina_rr params: api_key: ${secrets:JINA_API_KEY} - from: https://rerank.internal/v1/rerank name: byo_rr params: api_key: ${secrets:INTERNAL_RR_KEY} # optional ``` ### LLM-as-Reranker[​](#llm-as-reranker "Direct link to LLM-as-Reranker") Any registered chat model can also be used as a reranker without additional configuration. When a model name resolves to a chat model instead of a dedicated reranker, Spice wraps it in an LLM-based reranking adapter. ``` models: - from: openai:gpt-4o-mini name: gpt_mini ``` ``` SELECT * FROM rerank( vector_search(kb, 'onboarding checklist', limit => 40), document => 'content', model => 'gpt_mini', -- resolves to the registered chat model strategy => 'pointwise', -- or 'listwise' (default) limit => 10 ); ``` The `strategy` parameter controls how the LLM scores documents: * **`listwise`** (default): sends all candidates in a single prompt and asks the LLM to rank them. * **`pointwise`**: sends each candidate individually and asks the LLM to score it. ## SQL Usage[​](#sql-usage "Direct link to SQL Usage") ``` SELECT * FROM rerank( , document => '', model => '', limit => ); ``` For the complete parameter reference, see [Reranking SQL Reference](/docs/next/reference/sql/search#reranking-rerank). ## Examples[​](#examples "Direct link to Examples") ### Hybrid Recall + Rerank[​](#hybrid-recall--rerank "Direct link to Hybrid Recall + Rerank") ``` SELECT * FROM rerank( rrf( vector_search(docs, 'delta lake time travel', limit => 50), text_search(docs, 'delta lake time travel', limit => 50) ), document => 'content', model => 'cohere_rr', limit => 10 ); ``` The query is automatically extracted from the nested search UDTFs — no need to specify `query` explicitly. ### Bare-Table Rerank[​](#bare-table-rerank "Direct link to Bare-Table Rerank") When reranking a plain table (not a search UDTF), provide the `query` explicitly: ``` SELECT * FROM rerank( tickets, query => 'auth failures', document => 'body', model => 'voyage_rr', limit => 5 ); ``` ### Custom LLM Prompt[​](#custom-llm-prompt "Direct link to Custom LLM Prompt") ``` SELECT * FROM rerank( vector_search(kb, 'onboarding checklist', limit => 40), document => 'content', model => 'gpt_mini', strategy => 'pointwise', prompt_template => 'Rate 0-1: is this useful for a new hire?\nQuery: {query}\nDoc: {document}', limit => 10 ); ``` --- # Vector-Based Search > 🎓 Learn how it works with the [Amazon S3 Vectors with Spice](https://spice.ai/blog/amazon-s3-vectors-with-spice) engineering blog post. Vector search uses embeddings (numerical representations of text or data) to find semantically similar content. Unlike keyword search, vector search understands meaning and context, making it useful for: * Finding documents with similar meaning but different wording * Semantic similarity matching * Retrieval-augmented generation (RAG) applications * Recommendation systems For embedding columns that contain many vectors per row (for example, one vector per tag or per section), see [Multi-Vector Search](/docs/next/features/search/multi-vector). ## Embedding Models[​](#embedding-models "Direct link to Embedding Models") Spice supports two types of embedding providers: * **Local embedding models** e.g., [sentence-transformers/all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2). * **Remote embedding services** e.g., [OpenAI Embeddings API](https://platform.openai.com/docs/api-reference/embeddings/create). Embedding models are defined in the `spicepod.yaml` file as top-level components. ``` embeddings: - name: openai_embeddings from: openai params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } - name: local_embedding_model from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 ``` ## Configuring Datasets for Embeddings[​](#configuring-datasets-for-embeddings "Direct link to Configuring Datasets for Embeddings") To enable vector search, specify embeddings for the dataset columns in `spicepod.yaml`: ``` datasets: - from: github:github.com/spiceai/spiceai/issues name: spiceai.issues params: github_token: ${ secrets:GITHUB_TOKEN } acceleration: enabled: true columns: - name: body embeddings: - from: local_embedding_model ``` This configuration instructs Spice to create embeddings from the `body` column, enabling similarity searches on body content. ## Performing a Vector Search[​](#performing-a-vector-search "Direct link to Performing a Vector Search") Execute similarity searches using Spice's HTTP API: ``` curl -X POST http://localhost:8090/v1/search \ -H 'Content-Type: application/json' \ -d '{ "datasets": ["spiceai.issues"], "text": "cutting edge AI", "where": "author=\"jeadie\"", "additional_columns": ["title", "state"], "limit": 2 }' ``` For detailed API documentation, see [Search API Reference](/docs/next/api/HTTP/post-search). ## Retrieving Full Documents[​](#retrieving-full-documents "Direct link to Retrieving Full Documents") If the dataset uses chunking, Spice returns relevant chunks. To retrieve entire documents, include the embedding column in `additional_columns`: ``` curl -X POST http://localhost:8090/v1/search \ -H 'Content-Type: application/json' \ -d '{ "datasets": ["spiceai.issues"], "text": "cutting edge AI", "where": "array_has(assignees, \"jeadie\")", "additional_columns": ["title", "state", "body"], "limit": 2 }' ``` Response: ```` { "matches": [ { "value": "implements a scalar UDF `array_distance`:\n```\narray_distance(FixedSizeList[Float32], FixedSizeList[Float32])", "dataset": "spiceai.issues", "metadata": { "title": "Improve scalar UDF array_distance", "state": "Closed", "body": "## Overview\n- Previous PR https://github.com/spiceai/spiceai/pull/1601 implements a scalar UDF `array_distance`:\n```\narray_distance(FixedSizeList[Float32], FixedSizeList[Float32])\narray_distance(FixedSizeList[Float32], List[Float64])\n```\n\n### Changes\n - Improve using Native arrow function, e.g. `arrow_cast`, [`sub_checked`](https://arrow.apache.org/rust/arrow/array/trait.ArrowNativeTypeOp.html#tymethod.sub_checked)\n - Support a greater range of array types and numeric types\n - Possibly create a sub operator and UDF, e.g.\n\t- `FixedSizeList[Float32] - FixedSizeList[Float32]`\n\t- `Norm(FixedSizeList[Float32])`" }, "score": 0.66, }, { "value": "est external tools being returned for toolusing models", "dataset": "spiceai.issues", "metadata": { "title": "Automatic NSQL retries in /v1/nsql ", "state": "Open", "body": "To mimic our ability for LLMs to repeatedly retry tools based on errors, the `/v1/nsql`, which does not use this same paradigm, should retry internally.\n\nIf possible, improve the structured output to increase the likelihood of valid SQL in the response. Currently we just inforce JSON like this\n```json\n{\n "sql": "SELECT ..."\n}\n```" }, "score": 0.52, } ], "duration_ms": 45 } ```` ## SQL UDTF[​](#sql-udtf "Direct link to SQL UDTF") The embedding index can also be used to perform search in SQL, via a user-defined table function (UDTF). ``` SELECT id, title, score FROM vector_search('sales', 'cutting edge AI') ORDER BY score DESC LIMIT 5; ``` **SQL Function Signature of `vector_search`:** ``` vector_search( table STRING, -- Dataset name (required) query STRING, -- Search text (required) col STRING, -- Column name (optional if single embedding column) limit INTEGER, -- Results limit (default: 1000) include_score BOOLEAN -- Include relevance scores (default: TRUE) ) RETURNS TABLE -- The original table and: -- - A FLOAT column `score` (if `include_score`). ``` By default, `vector_search` retrieves up to 1000 results. To adjust this limit, specify the `limit` parameter in the function call. When using a specific vector engine, such as `s3_vectors` the limit defaults to that of the vector engine. ``` SELECT id, title, score FROM vector_search('sales', 'cutting edge AI', 1500) ORDER BY score DESC; ``` `WHERE` predicates on base table columns are pushed down as pre-filters — only matching rows are scored and ranked. See [Search in SQL](/docs/next/reference/sql/search#vector-search-vector_search) for details. Limitations * `vector_search` UDTF does not yet support chunked embedding columns. Chunking support is on the roadmap. ## Using Existing Embeddings[​](#using-existing-embeddings "Direct link to Using Existing Embeddings") Spice supports vector searches on datasets with pre-existing embeddings. Ensure the dataset meets these requirements: 1. **Column Naming**: The embedding column name must be `_embedding`. 2. **Data Types**: Embedding columns must use Arrow types: * Non-chunked: `FixedSizeList[Float32|Float64, N]` * Chunked: `List[FixedSizeList[Float32|Float64, N]]` 3. **Offset Columns**: For chunked embeddings, an additional offset column (`_offsets`) is required: * Type: `List[FixedSizeList[Int32, 2]]`, indicating chunk boundaries. Example dataset structure (`sales` table): Non-chunked: ``` sql> describe sales; +-------------------+-----------------------------------------+-------------+ | column_name | data_type | is_nullable | +-------------------+-----------------------------------------+-------------+ | order_number | Int64 | YES | | quantity_ordered | Int64 | YES | | price_each | Float64 | YES | | order_line_number | Int64 | YES | | address | Utf8 | YES | | address_embedding | FixedSizeList( | NO | | | Field { | | | | name: "item", | | | | data_type: Float32, | | | | nullable: false, | | | | dict_id: 0, | | | | dict_is_ordered: false, | | | | metadata: {} | | | | }, | | | | 384 | | +-------------------+-----------------------------------------+-------------+ ``` Chunked: ``` sql> describe sales; +-------------------+-----------------------------------------+-------------+ | column_name | data_type | is_nullable | +-------------------+-----------------------------------------+-------------+ | order_number | Int64 | YES | | quantity_ordered | Int64 | YES | | price_each | Float64 | YES | | order_line_number | Int64 | YES | | address | Utf8 | YES | | address_embedding | List(Field { | NO | | | name: "item", | | | | data_type: FixedSizeList( | | | | Field { | | | | name: "item", | | | | data_type: Float32, | | | | }, | | | | 384 | | | | ), | | | | }) | | +-------------------+-----------------------------------------+-------------+ | address_offset | List(Field { | NO | | | name: "item", | | | | data_type: FixedSizeList( | | | | Field { | | | | name: "item", | | | | data_type: Int32, | | | | }, | | | | 2 | | | | ), | | | | }) | | +-------------------+-----------------------------------------+-------------+ ``` ### Constraints[​](#constraints "Direct link to Constraints") 1. **Underlying Column Presence:** * The underlying column must exist in the table, and be of `string` [Arrow data type](/docs/next/reference/datatypes/accelerators) . 2. **Embeddings Column Naming Convention:** * For each underlying column, the corresponding embeddings column must be named as `_embedding`. For example, a `customer_reviews` table with a `review` column must have a `review_embedding` column. 3. **Embeddings Column Data Type:** * The embeddings column must have the following [Arrow data type](/docs/next/reference/datatypes/accelerators) when loaded into Spice: 1. `FixedSizeList[Float32 or Float64, N]`, where `N` is the dimension (size) of the embedding vector. `FixedSizeList` is used for efficient storage and processing of fixed-size vectors. 2. If the column is [**chunked**](/docs/next/components/embeddings#chunking), use `List[FixedSizeList[Float32 or Float64, N]]`. 4. **Offset Column for Chunked Data:** * If the underlying column is chunked, there must be an additional offset column named `_offsets` with the following Arrow data type: 1. `List[FixedSizeList[Int32, 2]]`, where each element is a pair of integers `[start, end]` representing the start and end indices of the chunk in the underlying text column. This offset column maps each chunk in the embeddings back to the corresponding segment in the underlying text column. * *For instance, `[[0, 100], [101, 200]]` indicates two chunks covering indices 0–100 and 101–200, respectively.* By following these guidelines, you can ensure that your dataset with pre-existing embeddings is fully compatible with the vector search and other embedding functionalities provided by Spice. ### Example[​](#example "Direct link to Example") A table `sales` with an `address` column and corresponding embedding column(s). --- # Semantic Model The semantic model is the layer that gives every dataset, view, and column a human-readable description plus structured metadata, and exposes that context uniformly to SQL queries, LLM tools, and language models. Every description ends up in two places at once: * on the dataset's Arrow schema as `description` metadata — readable from SQL via [`obj_description`](/docs/next/reference/sql/scalar_functions#obj_description) and [`col_description`](/docs/next/reference/sql/scalar_functions#col_description); and * in the LLM tool context — surfaced by the built-in [`list_datasets`](/docs/next/components/tools#available-tools) and [`table_schema`](/docs/next/components/tools#available-tools) tools, so models see your descriptions automatically. ## Defining a Semantic Model[​](#defining-a-semantic-model "Direct link to Defining a Semantic Model") Semantic data models are defined within the `spicepod.yaml` file under the `datasets` section. Each dataset supports `description`, `metadata`, and a `columns` field where individual columns are described with metadata and features for utility and clarity. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - name: taxi_trips description: NYC taxi trip rides metadata: instructions: Always provide citations with reference URLs. reference_url_template: https://d37ci6vzurychx.cloudfront.net/trip-data/yellow_tripdata_.parquet columns: - name: tpep_pickup_time description: 'The time the passenger was picked up by the taxi' - name: notes description: 'Optional notes about the trip' embeddings: - from: hf_minilm # A defined Spice Model chunking: enabled: true target_chunk_size: 512 overlap_size: 128 trim_whitespace: true ``` ## Dataset Metadata[​](#dataset-metadata "Direct link to Dataset Metadata") Datasets can be defined with the following metadata: * `instructions`: Optional. Instructions to provide to a language model when using this dataset. * `reference_url_template`: Optional. A URL template for citation links. Arbitrary additional keys may be added under `metadata:` and are passed through to the Arrow schema unchanged, so anything an LLM tool or downstream consumer expects to find there (e.g. governance labels, owning team, source-of-truth links) can be attached without code changes. For detailed `metadata` configuration, see the [Dataset Reference](/docs/next/reference/spicepod/datasets#metadata). ## Column Definitions[​](#column-definitions "Direct link to Column Definitions") Each column can be defined with the following attributes: * `description`: Optional. A description of the column's contents and purpose. * `type` (alias `data_type`): Optional. Declared column type (Postgres-style or Arrow display form). * `nullable`: Optional. Override column nullability. * `embeddings`: Optional. Vector embeddings configuration for this column. * `full_text_search`: Optional. Full-text-search configuration. * `metadata`: Optional. Arbitrary key/value metadata attached to the column. For detailed `columns` configuration, see the [Dataset Reference](/docs/next/reference/spicepod/datasets#columns). ## Source-Side Comments[​](#source-side-comments "Direct link to Source-Side Comments") When a dataset is loaded from a source that exposes table or column comments, Spice automatically imports those comments into the Arrow schema metadata under the `description` key. This means database-native `COMMENT ON TABLE` and `COMMENT ON COLUMN` annotations show up alongside Spicepod-defined `description` values, giving the semantic model the same context that already lives in the source database. Source-side comments are imported automatically from: | Source | What is imported | | ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | **PostgreSQL** | Table and column comments via `obj_description` / `col_description` in `pg_catalog`. | | **MySQL** | `information_schema.tables.table_comment` and `information_schema.columns.column_comment`. | | **Snowflake** | `information_schema.tables.comment` and `information_schema.columns.comment`. | | **Databricks SQL Warehouse** | Table and column comments returned by the SQL Warehouse driver's metadata API. | | **BigQuery** (via ADBC catalog) | `INFORMATION_SCHEMA.TABLE_OPTIONS` (`description`) for the table and `INFORMATION_SCHEMA.COLUMN_FIELD_PATHS.description` for top-level columns. | Connectors not listed above (object stores, file formats, HTTP, etc.) do not have a comment concept on the source side; for those datasets, attach descriptions in the Spicepod. ### Precedence[​](#precedence "Direct link to Precedence") When both a Spicepod `description` and a source-side comment exist for the same table or column, **the Spicepod value wins**. This lets you override unhelpful or out-of-date source comments without round-tripping a DDL change through the upstream database. If the Spicepod does not set a `description`, the source-side comment passes through unchanged. ## Reading descriptions from SQL[​](#reading-descriptions-from-sql "Direct link to Reading descriptions from SQL") Two PostgreSQL-compatible scalar UDFs read the `description` metadata at query time: ``` -- Table-level description SELECT obj_description('public.taxi_trips'); -- Column-level description (by name or 1-based ordinal) SELECT col_description('public.taxi_trips', 'fare_amount'); SELECT col_description('public.taxi_trips', 3); ``` Full signatures, including PostgreSQL-shaped `(oid, 'pg_class')` variants and four-part `(catalog, schema, table, column)` forms, are documented in [`obj_description`](/docs/next/reference/sql/scalar_functions#obj_description) and [`col_description`](/docs/next/reference/sql/scalar_functions#col_description). Because these UDFs read the same metadata that source-side comment extraction populates, queries against PostgreSQL or Snowflake datasets that were authored with `COMMENT ON …` will return those comments without any additional configuration. ## How LLMs see the semantic model[​](#how-llms-see-the-semantic-model "Direct link to How LLMs see the semantic model") The built-in [LLM tools](/docs/next/components/tools) surface descriptions and metadata into the model's context automatically. ### `list_datasets`[​](#list_datasets "Direct link to list_datasets") Returns every dataset, view, and catalog table visible to the runtime as JSON, with each entry's fully-qualified `table` name, `description`, `metadata`, columns, and capability flags (e.g. `can_search_documents`). Descriptions come from the Spicepod, falling back to source-side comments when no Spicepod description is set. ### `table_schema`[​](#table_schema "Direct link to table_schema") Returns one or more tables' columns as a markdown table. When called with `output: full` (the default), the rendered schema includes a `Metadata:` block at the top with the table-level `description` and any other dataset metadata keys, and a per-column `Metadata` cell containing the column's `description` and any column metadata. With `output: minimal`, only column names, types, and nullability are returned. ``` **Table: spice.public.taxi_trips** Metadata: description: NYC taxi trip rides instructions: Always provide citations with reference URLs. | Column | Sql Type | Arrow Type | Nullable | Metadata | | ---------------- | --------- | ---------------------- | -------- | --------------------------------------------------- | | tpep_pickup_time | TIMESTAMP | Timestamp(Microsecond) | true | description: The time the passenger was picked up | | notes | VARCHAR | Utf8 | true | description: Optional notes about the trip | ``` Because both `list_datasets` and `table_schema` are part of the [`auto` and `all` tool groups](/docs/next/components/tools#tool-groups), no extra configuration is required — every model declared with `tools: auto` (or any explicit list that includes `table_schema` / `list_datasets`) gets the semantic model in its context. When the [Tool Registry](/docs/next/features/tool-registry) is active, dataset descriptions also feed the hybrid search index used by `tool_search` to match user questions to relevant dataset-bound tools, so well-described datasets are easier for the model to discover at scale. ## End-to-end flow[​](#end-to-end-flow "Direct link to End-to-end flow") ``` spicepod.yaml description source COMMENT ON TABLE / COLUMN | | | (Spicepod wins on overlap) | +---------------+----------------+ v Arrow schema metadata (`description` key) | +---------------+-----------------+ v v v obj_description table_schema list_datasets col_description (LLM context) (LLM context) (SQL) ``` --- # Tool Registry The Tool Registry is a runtime-level capability that replaces large lists of individual tool definitions with two **meta-tools** — `tool_search` and `tool_invoke` — backed by a hybrid search index over the runtime's tool catalog. It's used to keep per-turn token cost bounded as the catalog grows, while preserving the model's ability to discover and call any tool on demand. The registry indexes every tool that's callable from an LLM, regardless of where it came from: * Built-in Spice runtime tools (`sql`, `list_datasets`, `table_schema`, `search`, `random_sample`, …) * [MCP tools](/docs/next/features/large-language-models/mcp) (servers connected over `sse` or `stdio`) * [Functions](/docs/next/features/functions) declared in the Spicepod with `as_tool: true` (the default) * [`tools:`](/docs/next/reference/spicepod/tools) entries with `as_sql: true` (callable from both SQL and the LLM) If it can be called from a chat completion, it goes through the registry. ## Why the Tool Registry?[​](#why-the-tool-registry "Direct link to Why the Tool Registry?") Each tool exposed to a model carries a name, a description, and a JSON Schema for its parameters. A typical tool is **200–500 tokens** of schema; a Spicepod with rich [MCP integrations](/docs/next/features/large-language-models/mcp), several datasets exposed via `sql` / `table_schema` / `search`, and custom [user-defined functions](/docs/next/features/functions) can quickly cross **50 tools** and **10,000+ tokens** of tool definitions injected into every chat turn. That cost is paid on every request: * **Tokens**: tool definitions are part of the system context, billed on every prompt. * **Latency**: more input tokens means slower first-token times. * **Accuracy**: research and practice both show LLM tool-selection accuracy degrades when the model is faced with dozens of similarly-named tools. * **Context window**: tool definitions compete with conversation history, retrieved documents, and reasoning scratch space. The Tool Registry replaces every individual tool definition with just two meta-tools: * **`tool_search(query, ...)`** — Searches the registry for tools relevant to a natural-language query. Returns the top N tools with their full schemas. * **`tool_invoke(tool_id, arguments)`** — Invokes a tool returned by `tool_search`. For a workload with 50 tools, this is roughly a **10× reduction** in tool-definition tokens injected per turn — the model now only sees the schemas of the tools it actively asks for. `list_datasets` is always exposed directly alongside the meta-tools so the model can orient itself ("what tables exist?") in a single call without first asking the registry. ## When to Use the Registry[​](#when-to-use-the-registry "Direct link to When to Use the Registry") The registry is the right default for any model that has access to **a substantial number of tools** — particularly when those tools include: * Multiple [MCP servers](/docs/next/features/large-language-models/mcp) each contributing several tools. * A Spicepod with many [Functions](/docs/next/features/functions) declared as tools. * Multiple datasets, each contributing dataset-specific tools (e.g. via the [`tools:`](/docs/next/reference/spicepod/tools) section). It's **less useful** when: * The Spicepod has a small, focused tool set (under \~20 tools). * The model needs to chain tools without round-tripping through `tool_search` (saves one tool call per turn). * Deterministic tool exposure is required for evaluation or compliance reasons. For everything else — especially Spicepods that compose multiple tool sources — `tools: auto` is the recommended default. ## Enabling the Registry[​](#enabling-the-registry "Direct link to Enabling the Registry") The registry is controlled via the `tools` parameter on a model. Set it to `search_registry` to require registry-based discovery, or `auto` to let Spice decide: ``` embeddings: - name: tool_embeddings from: openai:text-embedding-3-small models: - name: my-model from: openai:gpt-4o params: tools: search_registry tool_embedding_model: tool_embeddings ``` `tools: auto` switches to the registry **only when both** of these are true: * The number of available tools exceeds **20** (`AUTO_SEARCH_TOOL_THRESHOLD`). * An embedding model is available. Otherwise `auto` falls back to providing tools directly — keeping small Spicepods ergonomic while large ones automatically benefit. See the [Tool Modes table](/docs/next/features/large-language-models/tools#tool-modes) for the full set of values. ### Configuring `tool_embedding_model`[​](#configuring-tool_embedding_model "Direct link to configuring-tool_embedding_model") The registry's vector channel uses a configured embedding model: * **One embedding configured** → used automatically. * **Multiple embeddings configured** → `tool_embedding_model` is required and must name one of them. * **No embedding configured** → `tools: search_registry` is rejected; `tools: auto` falls back to direct tools with a warning log. ``` embeddings: - name: openai_embed from: openai:text-embedding-3-small - name: local_embed from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 models: - name: my-model from: openai:gpt-4o params: tools: search_registry tool_embedding_model: openai_embed # disambiguate ``` ## How `tool_search` Ranks Results[​](#how-tool_search-ranks-results "Direct link to how-tool_search-ranks-results") `tool_search` runs a **hybrid search** over four channels and fuses the results with [Reciprocal Rank Fusion (RRF)](/docs/next/reference/sql/search#reciprocal-rank-fusion-rrf): | Channel | Signal | | ----------- | -------------------------------------------------------------------------------------------------------------- | | `full_text` | TF-IDF over tokenized tool name (×3 weight), description (×2), and parameters (×1). | | `keyword` | Exact-phrase and token matches against name / description / parameter text. Weighted by where the match lands. | | `schema` | Matches against the **parameter keys** in the tool's JSON Schema (e.g. `dataset`, `query`). | | `vector` | Cosine similarity between the query embedding and per-tool document embeddings. | Each channel produces a ranked list; RRF combines the ranks (not the scores) so a tool that places top-3 in two channels usually outranks one that places top-1 in a single channel. The final `score` is normalized to `0.0–1.0` against the highest-scoring tool in the result set. Per-tool embeddings are computed lazily on first search and cached for the lifetime of the registry instance — a model keeps its instance, and therefore its embeddings, until the model is reloaded. The `/v1/tools` HTTP endpoints build their instances through a separate bounded cache (up to 64 entries) keyed on `(runtime, embedding model, tools hash)`, so repeated calls reuse the same embeddings. That cache is not an LRU: once it is full, an arbitrary existing entry is evicted to make room. ## `tool_search` Reference[​](#tool_search-reference "Direct link to tool_search-reference") The model calls `tool_search` with a JSON object: | Parameter | Type | Description | | ----------- | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `query` | `string` (required) | Natural-language description of the capability the model needs. | | `keywords` | `string[]` | Optional exact-match phrases — useful for column or table names. Their tokens are added to the `full_text` and `schema` channels, and they *replace* `query` as the phrase the `keyword` channel matches on (with `keywords` omitted, the whole `query` string is that phrase). | | `limit` | `integer` | Maximum results to return. Defaults to **5**, capped at **20**. | | `min_score` | `number` | Optional minimum score, clamped to 0.0–1.0. Scores are normalized against the best match, so the top result always scores `1.0` — a cutoff trims the tail and can never empty the list. If no tool matched any channel at all, every score is `0.0` and the cutoff is skipped entirely: the first `limit` tools are returned in `tool_id` order. | Example call (issued by the model): ``` { "query": "count distinct values in a column", "keywords": ["distinct", "count"], "limit": 3 } ``` ### `tool_search` Response[​](#tool_search-response "Direct link to tool_search-response") ``` { "query": "count distinct values in a column", "keywords": ["distinct", "count"], "search_mode": "hybrid_rrf", "tools": [ { "tool_id": "sql", "description": "Execute SQL queries on the runtime.", "parameters": { "type": "object", "properties": { "query": { "type": "string" } } }, "score": 1.0, "matched_terms": ["count", "distinct", "sql"], "match_sources": [ { "source": "full_text", "rank": 1, "score": 4.231 }, { "source": "keyword", "rank": 1, "score": 9.0 }, { "source": "vector", "rank": 1, "score": 0.812 } ] } ] } ``` `match_sources` is intentionally surfaced — it lets the model (or a debugger) reason about *why* a tool was returned. A tool that only matched on `vector` but not `full_text` may be a semantic match for an unfamiliar phrasing; one that matched all four is a high-confidence hit. ## `tool_invoke` Reference[​](#tool_invoke-reference "Direct link to tool_invoke-reference") | Parameter | Type | Description | | ----------- | -------- | ---------------------------------------------------------------------------- | | `tool_id` | `string` | Tool name returned by `tool_search`. | | `arguments` | `object` | JSON object matching the selected tool's parameter schema. Defaults to `{}`. | Example: ``` { "tool_id": "sql", "arguments": { "query": "SELECT COUNT(DISTINCT customer_id) FROM orders" } } ``` ### `tool_invoke` Response[​](#tool_invoke-response "Direct link to tool_invoke-response") ``` { "tool_id": "sql", "result": [{ "count": 1247 }] } ``` Errors propagate the underlying tool's error message, prefixed with the `tool_id` so the model can decide whether to retry, ask for a different tool, or surface the failure to the user. ## Functions in the Registry[​](#functions-in-the-registry "Direct link to Functions in the Registry") Every [function](/docs/next/features/functions) declared with `as_tool: true` (the default) is registered both as a SQL UDF and as a tool, and therefore participates in the registry. This means a Spicepod with many domain-specific UDFs benefits from the registry exactly the same way as one with many MCP tools — the model only sees the function definitions for the few it actually asks about. ``` runtime: functions: enabled: true embeddings: - name: tool_embeddings from: openai:text-embedding-3-small functions: - name: haversine_km from: sql description: Haversine great-circle distance in kilometres. volatility: immutable signature: args: - { name: lat1, type: float64 } - { name: lon1, type: float64 } - { name: lat2, type: float64 } - { name: lon2, type: float64 } returns: float64 body: | 6371 * acos( cos(radians(lat1)) * cos(radians(lat2)) * cos(radians(lon2) - radians(lon1)) + sin(radians(lat1)) * sin(radians(lat2)) ) # ...many more models: - name: my-model from: openai:gpt-4o params: tools: auto # registry kicks in automatically once the function count crosses the threshold tool_embedding_model: tool_embeddings ``` To keep a function out of the registry (and out of the LLM tool surface entirely) while still callable from SQL, set `as_tool: false`: ``` functions: - name: internal_hash from: sql as_tool: false signature: args: [{ name: x, type: int64 }] returns: int64 body: 'x * 2654435761' ``` User-defined table functions (UDTFs) are SQL-only and are not currently registered as LLM tools, so they don't appear in the registry. ## Reserved Tool Names and Conflicts[​](#reserved-tool-names-and-conflicts "Direct link to Reserved Tool Names and Conflicts") `tool_search` and `tool_invoke` are reserved names. If a user-defined tool, function, or MCP tool registers under either name: * **`tools: search_registry`** → fails at startup with a clear error. * **`tools: auto`** → logs a warning and falls back to direct tools. Rename the offending tool, or set `as_tool: false` to keep it SQL-only. ## Discovering What's in the Registry[​](#discovering-whats-in-the-registry "Direct link to Discovering What's in the Registry") Two ways to inspect the catalog from outside the model: * **From SQL** — `SELECT * FROM list_udfs() WHERE source = 'user';` lists every user-declared function, regardless of whether it's currently in the registry. * **From the HTTP API** — `GET /v1/tools` lists every tool the runtime has registered, which is what the registry indexes. Tool routes are refused with a `401` unless [`runtime.auth`](/docs/next/reference/spicepod/runtime#runtimeauth) is configured. `GET /v1/functions` is a *function* listing, not a tool one: it returns every enabled entry in the Spicepod's `functions:` section that registered successfully — scalar and table functions alike, whether or not they are exposed as tools — and returns an empty list when `runtime.functions.enabled` is false. Use it to check that a function loaded; use `GET /v1/tools` to check what the model can reach. For tools (built-in plus MCP plus function-derived), the model can call `tool_search` with an open-ended query (e.g. `query: "*"`) — though in practice, asking for the tools relevant to the current step is what the model actually wants. --- # Views Views in Spice are virtual tables defined by SQL queries. They help simplify complex queries and promote reuse across different applications by encapsulating query logic in a single, reusable entity. ## Defining a View[​](#defining-a-view "Direct link to Defining a View") To define a view in the `spicepod.yaml` configuration file, specify the `views` section. Each view definition must include a `name` and a `sql` field. ### Example[​](#example "Direct link to Example") The following example demonstrates how to define a view named `rankings` that lists the top five products based on the total count of orders: ``` views: - name: rankings sql: | WITH a AS ( SELECT products.id, SUM(count) AS count FROM orders INNER JOIN products ON orders.product_id = products.id GROUP BY products.id ) SELECT name, count FROM products LEFT JOIN a ON products.id = a.id ORDER BY count DESC LIMIT 5 ``` ### Fields[​](#fields "Direct link to Fields") * `name`: The view's identifier, used for referencing in queries. * `sql`: The SQL query defining the view, supporting joins, subqueries, and aggregations. * `acceleration`: Views can be [locally accelerated](/docs/next/features/data-acceleration). ## Limitations and Considerations[​](#limitations-and-considerations "Direct link to Limitations and Considerations") * Views are read-only; insert, update, and delete operations are not supported. * Performance depends on SQL complexity and underlying data. * Ensure queries are optimized to prevent slow execution. ## Schema Inference and Evolution[​](#schema-inference-and-evolution "Direct link to Schema Inference and Evolution") Views derive their schema from the SQL query that defines them. When a view is accelerated, Spice materializes this schema into the acceleration engine at startup. If the underlying dataset schemas change while the runtime is running, the accelerated view will fail to refresh because the materialized schema no longer matches the source data. Restart the runtime to re-derive the view schema from the updated datasets. For more detail on schema inference and runtime schema changes, see the [Data Connectors schema inference](/docs/next/components/data-connectors#schema-inference) documentation. --- # Web Search Spice provides web search functionality through OpenAI's hosted tools, enabling access to recent information and relevant context. note The `websearch` tool (backed by Perplexity) is no longer supported and was deprecated in [spiceai/spiceai#9910](https://github.com/spiceai/spiceai/pull/9910). Use OpenAI's hosted web search tool as described below. ## Web Search Through OpenAI Hosted Tools[​](#web-search-through-openai-hosted-tools "Direct link to Web Search Through OpenAI Hosted Tools") Spice supports web search using [OpenAI's hosted web search tool](https://platform.openai.com/docs/guides/tools-web-search?api-mode=responses) when using [OpenAI's Responses API](https://platform.openai.com/docs/api-reference/responses). Sample Spicepod configuration: ``` models: - from: openai:gpt-4o-mini # Or any other model supported by OpenAI's Responses API name: openai_model params: openai_api_key: ${secrets:OPENAI_API_KEY} tools: auto responses_api: enabled # Required for using web search openai_responses_tools: web_search # Allowlist the web search tool via OpenAI's Responses API ``` **Sample Usage (with the above configuration)** Start Spice with `spice run`. Then, execute the following command, which makes a request to the `/v1/responses` endpoint. ``` curl -s -H "Content-Type: application/json" -X POST "http://localhost:8090/v1/responses" -d '{"model": "openai_model", "input": "what is the latest news today? use the web search tool"}' | jq -r '.output[] | select(.type=="message") | .content[] | select(.type=="output_text") | .text' ``` To invoke this model, use [`spice chat --responses`](/docs/next/cli/reference/chat) for an interactive REPL or the OpenAI-compatible `/v1/responses` HTTP endpoint in the runtime. To learn more about configuring models provided by OpenAI, view [the reference](/docs/next/components/models/openai). ## References[​](#references "Direct link to References") * [OpenAI Web Search Tool Documentation](https://platform.openai.com/docs/guides/tools-web-search?api-mode=responses) * [OpenAI Responses API Documentation](https://platform.openai.com/docs/api-reference/responses) --- # Workers Workers in the Spice runtime coordinate and manage interactions between models and tools — providing load-balancing strategies such as round-robin and fallback across multiple LLM providers. Each worker is defined in the `workers` section of the `spicepod.yaml` file, specifying its behavior and interaction logic. ## Configuration[​](#configuration "Direct link to Configuration") Workers are configured in the `workers` section of the `spicepod.yaml` file. Each worker definition includes a name, description, and a list of models or tools it encapsulates. **Example `spicepod.yaml` configuration:** ``` workers: - name: round-robin description: | Distributes requests between 'llama3_2' and 'gpt4_1' models in a round-robin fashion. load_balance: routing: - from: llama3_2 - from: gpt4_1 - name: fallback description: | Attempts 'gpt4_1' first, then 'llama3_2', then 'anth_haiku' if previous models fail. load_balance: routing: - from: llama3_2 order: 2 - from: gpt4_1 order: 1 - from: anth_haiku order: 3 - name: weighted description: | Routes 80% of traffic to 'llama3_2'. load_balance: routing: - from: llama3_2 weight: 4 - from: gpt4_1 weight: 1 ``` ## Use-Cases[​](#use-cases "Direct link to Use-Cases") Workers currently help implement: * Model fallback and error handling * Load balancing across multiple models ## Usage[​](#usage "Direct link to Usage") Workers can be invoked using the same API endpoints as individual models. For example, to call a worker named `fallback` using the OpenAI-compatible HTTP API: ``` curl http://localhost:8090/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "fallback", "messages": [{ "role": "user", "content": "Tell me a joke"}] }' ``` ## Roadmap[​](#roadmap "Direct link to Roadmap") The vision for workers includes support for dynamic serverless compute, enabling execution of user-defined functions within the Spice runtime. This direction aims to help developers define custom logic and orchestration patterns directly in the worker configuration, supporting more advanced workflows and automation. Further details and implementation timelines will be provided in future updates. For ongoing progress, refer to the project repository and documentation. ## Further Reading[​](#further-reading "Direct link to Further Reading") For a complete specification of worker configuration, routing rules, and available options, refer to the [Spicepod Workers Reference](/docs/next/reference/spicepod/workers). --- # Getting Started with Spice.ai OSS ### Follow these steps to get started with Spice[​](#follow-these-steps-to-get-started-with-spice "Direct link to Follow these steps to get started with Spice") Download the latest version of Spice, connect to a dataset in S3, and ask questions about the data using AI, in less than 5 minutes. **Step 1.** Install the Spice CLI: * macOS, Linux, and WSL * Windows ### Install Script[​](#install-script "Direct link to Install Script") ``` curl https://install.spiceai.org | /bin/bash ``` ### Homebrew[​](#homebrew "Direct link to Homebrew") ``` brew install spiceai/spiceai/spice ``` ### PowerShell Install Script[​](#powershell-install-script "Direct link to PowerShell Install Script") ``` iex ((New-Object System.Net.WebClient).DownloadString("https://install.spiceai.org/Install.ps1")) ``` **Step 2.** Initialize a new Spice app with the `spice init` command: ``` spice init spice_qs ``` A `spicepod.yaml` file is created in the `spice_qs` directory. Change to that directory: ``` cd spice_qs ``` **Step 3.** Add the `spiceai/quickstart` Spicepod. A Spicepod is a package of configuration defining datasets and ML models. ``` spice add spiceai/quickstart ``` The `spicepod.yaml` file will be updated with the `spiceai/quickstart` dependency. This Spicepod includes a `taxi_trips` dataset sourced from S3. ``` version: v1 kind: Spicepod name: spice_qs dependencies: - spiceai/quickstart ``` **Step 4.** Add an [OpenAI](/docs/components/models/openai) model to the Spicepod with tools enabled, so the model can query the `taxi_trips` dataset directly: ``` models: - from: openai:gpt-4o-mini name: openai_model params: openai_api_key: ${ env:OPENAI_API_KEY } tools: auto ``` Add your OpenAI API key to a `.env` file that Spice automatically loads at startup: ``` echo "OPENAI_API_KEY=sk-..." > .env ``` Start the Spice runtime: ``` spice run ``` The runtime starts and loads the `taxi_trips` dataset. Wait for the dataset to finish loading before querying: ``` Spice.ai runtime starting... 2024-08-05T13:02:40.247484Z INFO runtime::flight: Spice Runtime Flight listening on 127.0.0.1:50051 2024-08-05T13:02:40.247949Z INFO runtime: Initialized results cache; max size: 128.00 MiB, item ttl: 1s 2024-08-05T13:02:40.248611Z INFO runtime::http: Spice Runtime HTTP listening on 127.0.0.1:8090 ``` **Step 5.** Start the Spice Chat REPL and ask a question: ``` $ spice chat Using model: openai_model chat> How many taxi trips were taken? A total of 2,964,624 trips were taken according to the dataset. ``` The OpenAI model was automatically provided with the `taxi_trips` dataset and was able to answer the question. **Step 6.** Start the Spice SQL REPL: ``` spice sql ``` The SQL REPL interface will be shown: ``` Welcome to the Spice.ai SQL REPL! Type 'help' for help. show tables; -- list available tables sql> ``` Enter `show tables;` to display the available tables for query: ``` sql> show tables +---------------+--------------+---------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------+---------------+------------+ | spice | public | taxi_trips | BASE TABLE | | spice | runtime | task_history | BASE TABLE | | spice | runtime | metrics | BASE TABLE | +---------------+--------------+---------------+------------+ Time: 0.022671708 seconds. 3 rows. ``` Enter a query to display the longest taxi trips: ``` sql> SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10; ``` Output: ``` +---------------+--------------+ | trip_distance | total_amount | +---------------+--------------+ | 312722.3 | 22.15 | | 97793.92 | 36.31 | | 82015.45 | 21.56 | | 72975.97 | 20.04 | | 71752.26 | 49.57 | | 59282.45 | 33.52 | | 59076.43 | 23.17 | | 58298.51 | 18.63 | | 51619.36 | 24.2 | | 44018.64 | 52.43 | +---------------+--------------+ Time: 0.045150667 seconds. 10 rows. ``` ## Next Steps[​](#next-steps "Direct link to Next Steps") Now that Spice is running, explore the key capabilities: ## [📄️Spicepods](/docs/getting-started/spicepods) [Understand the Spicepod configuration format that defines datasets, models, and secrets.](/docs/getting-started/spicepods) ## [📄️Data Connectors](/docs/components/data-connectors) [Connect to PostgreSQL, MySQL, S3, Snowflake, Databricks, and more.](/docs/components/data-connectors) ## [📄️Data Acceleration](/docs/features/data-acceleration) [Materialize and cache datasets locally for sub-second query performance.](/docs/features/data-acceleration) ## [📄️Large Language Models](/docs/features/large-language-models) [Configure AI models with OpenAI-compatible APIs, local serving, and tool use.](/docs/features/large-language-models) ## [📄️Search](/docs/features/search) [Vector, full-text, and hybrid search across structured and unstructured data.](/docs/features/search) ## [🔗Cookbook](https://github.com/spiceai/cookbook) [Over 65 recipes demonstrating federated queries, RAG, text-to-SQL, and more.](https://github.com/spiceai/cookbook) --- # Community Data The [Spice.ai Cloud Platform](https://docs.spice.ai) includes a comprehensive set of free, ready-to-query [sample datasets](https://spicerack.org/). The Spice runtime can query these datasets using the [Spice.ai Data Connector](/docs/next/components/data-connectors/spiceai). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Spice CLI installed ([Installation guide](/docs/next/installation)) * A free Spice.ai Cloud Platform account at [spice.ai](https://spice.ai/) ## Quickstart[​](#quickstart "Direct link to Quickstart") To access these community datasets, navigate to [spice.ai](https://spice.ai/), and create a new account by clicking 'Start for Free'. After logging in, create an app in order to get an API key. ![create\_app-1](https://github.com/spiceai/spiceai/assets/112157037/d2446406-1f06-40fb-8373-1b6d692cb5f7) This quickstart will use the `taxi_trips` dataset from Spice.ai app. **Step 1.** Initialize a new project: ``` # Initialize a new Spice app spice init spice_app # Change to app directory cd spice_app ``` **Step 2.** Log in to the Spice Cloud Platform from the command line using the `spice login` command. A pop up browser window will prompt you to authenticate: ``` spice login ``` Logging in will create or update a `.env` file in the project directory with the API key. **Step 3.** Start the runtime: ``` # Start the runtime spice run ``` **Step 4.** Configure the dataset: In a new terminal window, configure a new dataset using the `spice dataset configure` command: ``` spice dataset configure ``` Enter a dataset name that will be used to reference the dataset in queries. This name does not need to match the name in the dataset source. ``` dataset name: (spice_app) taxi_trips ``` Enter the description of the dataset: ``` description: Taxi trips dataset ``` Enter the location of the dataset: ``` from: spice.ai/spiceai/quickstart/datasets/taxi_trips ``` Select `y` when prompted whether to accelerate the data: ``` Locally accelerate (y/n)? y ``` You should see the following output from your runtime terminal: ``` 2024-12-16T05:12:45.803694Z INFO runtime::init::dataset: Dataset taxi_trips registered (spice.ai/spiceai/quickstart/datasets/taxi_trips), acceleration (arrow, 10s refresh), results cache enabled. 2024-12-16T05:12:45.805494Z INFO runtime::accelerated_table::refresh_task: Loading data for dataset taxi_trips 2024-12-16T05:13:24.218345Z INFO runtime::accelerated_table::refresh_task: Loaded 2,964,624 rows (8.41 GiB) for dataset taxi_trips in 38s 412ms. ``` **Step 5.** In a new terminal window, use the Spice SQL REPL to query the dataset ``` spice sql ``` ``` SELECT tpep_pickup_datetime, passenger_count, trip_distance from taxi_trips LIMIT 10; ``` The output displays the results of the query along with the query execution time: ``` +----------------------+-----------------+---------------+ | tpep_pickup_datetime | passenger_count | trip_distance | +----------------------+-----------------+---------------+ | 2024-01-11T12:55:12 | 1 | 0.0 | | 2024-01-11T12:55:12 | 1 | 0.0 | | 2024-01-11T12:04:56 | 1 | 0.63 | | 2024-01-11T12:18:31 | 1 | 1.38 | | 2024-01-11T12:39:26 | 1 | 1.01 | | 2024-01-11T12:18:58 | 1 | 5.13 | | 2024-01-11T12:43:13 | 1 | 2.9 | | 2024-01-11T12:05:41 | 1 | 1.36 | | 2024-01-11T12:20:41 | 1 | 1.11 | | 2024-01-11T12:37:25 | 1 | 2.04 | +----------------------+-----------------+---------------+ Time: 0.00538925 seconds. 10 rows. ``` You can experiment with the time it takes to generate queries when using non-accelerated datasets. You can change the acceleration setting from `true` to `false` in the datasets.yaml file. ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` -- Query to display the average trip distance SELECT AVG(trip_distance) FROM taxi_trips; ``` The output displays the average gas used: ``` +-------------------------------+ | avg(taxi_trips.trip_distance) | +-------------------------------+ | 3.652169178958276 | +-------------------------------+ Time: 0.031145625 seconds. 1 rows. ``` --- # Spicepods ## Overview[​](#overview "Direct link to Overview") A **Spicepod** is a configuration package that defines application-specific datasets, catalogs, machine learning (ML) models, and secrets. It functions similarly to a code packaging system (such as npm or pip), but is designed for data and AI components rather than code libraries. Spicepods are defined in a YAML manifest file, typically named `spicepod.yaml`, and can be shared, versioned, and reused across projects. To create a new Spicepod, run: ``` spice init my_app ``` This generates a `spicepod.yaml` file in the `my_app` directory with the minimum required fields: ``` version: v1 kind: Spicepod name: my_app ``` ## Structure[​](#structure "Direct link to Structure") A Spicepod is described by a YAML manifest file, typically named `spicepod.yaml`, which includes the following key sections: * **Metadata:** Basic information about the Spicepod, such as its name and version. * **Datasets:** Definitions of datasets that are used or produced within the Spicepod. * **Catalogs:** Definitions of catalogs that are used within the Spicepod. * **Models:** Definitions of language models that the Spicepod manages, including their sources and associated datasets. * **Secrets:** Configuration for any secret stores used within the Spicepod. ## Example Manifest[​](#example-manifest "Direct link to Example Manifest") ``` version: v1 kind: Spicepod name: my_spicepod datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true models: - from: openai:gpt-4o-mini name: openai_model params: openai_api_key: ${ env:OPENAI_API_KEY } tools: auto secrets: - from: env name: env ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` version: v1 kind: Spicepod name: another_spicepod datasets: - from: databricks:spiceai_demo.public.dataset name: sample_ds params: mode: delta_lake databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_token: ${secrets:my_token} databricks_aws_access_key_id: ${secrets:aws_access_key_id} databricks_aws_secret_access_key: ${secrets:aws_secret_access_key} acceleration: enabled: true refresh_mode: full models: - from: huggingface.co/microsoft/Phi-3.5-mini-instruct name: phi secrets: - from: env name: env ``` ## Key Components[​](#key-components "Direct link to Key Components") ### Datasets[​](#datasets "Direct link to Datasets") Datasets in a Spicepod define the tables available for SQL queries. Each dataset specifies a source (using the `from` field) and optionally an acceleration engine for local materialization. Sources include local files, databases (PostgreSQL, MySQL), cloud warehouses (Snowflake, Databricks), object storage (S3), and more. ``` datasets: - from: postgres:public.orders name: orders params: pg_host: localhost pg_port: "5432" pg_db: mydb pg_user: reader pg_pass: ${secrets:PG_PASSWORD} acceleration: enabled: true engine: duckdb refresh_check_interval: 30s ``` Learn more at [Datasets](/docs/next/reference/spicepod/datasets). ### Catalogs[​](#catalogs "Direct link to Catalogs") Catalogs in a Spicepod can contain multiple schemas. Each schema, in turn, contains multiple tables where the actual data is stored. Learn more at [Catalogs](/docs/next/reference/spicepod/catalogs). ### Models[​](#models "Direct link to Models") ML and language models are configured in the Spicepod similarly to datasets. Models can reference hosted services (OpenAI, Anthropic) or local files (Hugging Face models). When tools are enabled, models can query datasets and run SQL during inference. ``` models: - from: openai:gpt-4o-mini name: assistant params: openai_api_key: ${ env:OPENAI_API_KEY } tools: auto # Gives the model access to dataset schemas and SQL ``` Learn more at [Models](/docs/next/reference/spicepod/models). ### Secrets[​](#secrets "Direct link to Secrets") Spice supports various secret stores to manage sensitive information such as API keys or database credentials. Supported secret store types include environment variables, files, AWS Secrets Manager, Kubernetes secrets, and keyrings. Reference secrets in dataset or model params using the `${secrets:KEY_NAME}` syntax. The `env` secret store (enabled by default) reads from environment variables and `.env` files: ``` secrets: - from: env name: env datasets: - from: postgres:users name: users params: pg_pass: ${secrets:DB_PASSWORD} # Reads DB_PASSWORD from environment or .env file ``` Learn more at [Secret Stores](/docs/next/components/secret-stores) --- # Telemetry Spice collects anonymous telemetry data to help improve the product. Usage telemetry is anonymous and aggregated. ## Data Collected[​](#data-collected "Direct link to Data Collected") The following anonymous information is collected: * The version of Spice being used (i.e. `v1.0.0`) * An anonymous identifier for the Spice instance, computed as `sha256(hostname + spicepod.name)`. * An anonymous identifier for the Spicepod, computed as `sha256(spicepod.name)`. * The code to calculate these identifiers is here: * Various metrics related to usage of features of the runtime, see the full list here: Data collected is sent to `https://telemetry.spiceai.org` once every hour. ## Disabling Telemetry[​](#disabling-telemetry "Direct link to Disabling Telemetry") Open Source builds In Spice.ai Open Source builds that include the `anonymous_telemetry` feature (the default), setting `runtime.telemetry.enabled: false` in a Spicepod or passing `--telemetry-enabled=false` **does not** disable anonymous usage telemetry. The runtime will log a warning when these settings are detected but are not applied. To fully remove anonymous telemetry from an Open Source build, compile from source without the `anonymous_telemetry` feature (option 3 below), or use Spice.ai Enterprise. Telemetry can be disabled in the following ways: 1. **Spice.ai Enterprise**: Anonymous telemetry respects the `runtime.telemetry.enabled` and `--telemetry-enabled` settings. 2. **Compile without the `anonymous_telemetry` feature** (Open Source): ``` cargo build --release --no-default-features --features "" ``` i.e. ``` cargo build --release --no-default-features --features "duckdb,postgres,sqlite,mysql,flightsql,delta_lake,databricks,dremio,clickhouse,spark,snowflake,ftp,debezium" ``` ### Configuration Settings (Enterprise only)[​](#configuration-settings-enterprise-only "Direct link to Configuration Settings (Enterprise only)") The following settings disable telemetry in Spice.ai Enterprise builds: Running the Spice runtime with the CLI flag `--telemetry-enabled false`: ``` spice run -- --telemetry-enabled false ``` or ``` spiced --telemetry-enabled false ``` Adding the following configuration to the Spicepod configuration file (`spicepod.yaml`): ``` runtime: telemetry: enabled: false ``` --- # Install Spice.ai OSS Installation options for Spice.ai OSS **Prerequisites:** macOS (Apple Silicon), Linux (x86\_64 or aarch64), or Windows 10+/WSL. No other dependencies are required for the pre-built binary. For deployment options, such as to Kubernetes, see [`Deployment`](/docs/next/deployment). * macOS, Linux, and WSL * Windows ### Install Script[​](#install-script "Direct link to Install Script") ``` curl https://install.spiceai.org | /bin/bash ``` ### Homebrew[​](#homebrew "Direct link to Homebrew") ``` brew install spiceai/spiceai/spice ``` ### PowerShell Install Script[​](#powershell-install-script "Direct link to PowerShell Install Script") ``` iex ((New-Object System.Net.WebClient).DownloadString("https://install.spiceai.org/Install.ps1")) ``` ## Direct Download[​](#direct-download "Direct link to Direct Download") Binaries for Linux, Windows, and macOS are available for download from GitHub at [github.com/spiceai/spiceai/releases](https://github.com/spiceai/spiceai/releases). **Verify the installation:** After installing, verify Spice is installed correctly: ``` spice version ``` Expected output: ``` CLI version: 1.x.x Runtime version: 1.x.x ``` If the command is not found, ensure the Spice binary directory is in your `PATH`. ## What's Next?[​](#whats-next "Direct link to What's Next?") After installing, follow the [Getting Started guide](/docs/next/getting-started) to initialize a Spice app, connect to a dataset, and run your first query in under 5 minutes. ## Building from Source[​](#building-from-source "Direct link to Building from Source") To build Spice from source, including CUDA and Metal hardware acceleration options, see [CONTRIBUTING.md](https://github.com/spiceai/spiceai/blob/trunk/CONTRIBUTING.md) in the Spice.ai GitHub repository. --- # Building Intelligent AI Applications with Spice.ai ## Building Data-Driven AI Applications with Spice.ai[​](#building-data-driven-ai-applications-with-spiceai "Direct link to Building Data-Driven AI Applications with Spice.ai") Spice.ai represents a paradigm shift in how intelligent applications are developed, deployed, and managed. As outlined in the blog post [Making Apps That Learn and Adapt](https://blog.spiceai.org/posts/2021/11/05/making-apps-that-learn-and-adapt/), the goal of Spice.ai is to reduce the technical complexity that often hampers developers when building AI-powered solutions. By colocating federated data and machine learning models with applications, Spice.ai provides lightweight, high-performance, and highly scalable AI copilot sidecars. These sidecars streamline application workflows, improving both speed and efficiency. At its core, Spice.ai addresses the fragmented nature of traditional AI infrastructure. Data often resides in multiple systems: modern cloud-based warehouses, legacy databases, or unstructured formats like files on FTP servers. Integrating these disparate sources into a unified application pipeline typically requires extensive engineering effort. Spice.ai simplifies this process by federating data across all these sources, materializing it locally for low-latency access, and offering a unified SQL API. This eliminates the need for complex and costly ETL pipelines or federated query engines that operate with high latency. Spice.ai also colocates machine learning models with the application runtime. This approach reduces the data transfer overhead that occurs when sending data to external inference services. By performing inference locally, applications can respond faster and operate more reliably, even in environments with intermittent network connectivity. The result is an infrastructure that enables developers to focus on building value-driven features rather than wrestling with data and deployment complexities. *** ## The Intelligent Application Workflow[​](#the-intelligent-application-workflow "Direct link to The Intelligent Application Workflow") The workflow for creating intelligent applications with Spice.ai is designed to provide developers with a straightforward, efficient path from data to decision-making. It begins with the creation of `Spicepods`, self-contained packages that define datasets and machine learning models. These packages can be distributed through [Spicerack.org](https://spicerack.org), a registry where developers can publish, share, and reuse datasets and models for various applications. Once deployed, federated datasets are materialized locally within the Spice runtime. Materialization involves prefetching and precomputing data, storing it in high-performance local stores like DuckDB or SQLite. This approach ensures that queries are executed with minimal latency, offering high concurrency and predictable performance. Accelerated access is made possible through advanced caching and query optimization techniques, enabling applications to perform even complex operations without relying on remote databases. Applications interact with the Spice runtime through high-performance APIs, calling machine learning models for inference tasks such as predictions, recommendations, or anomaly detection. These models are colocated with the runtime, giving them direct access to the same locally materialized datasets. For example, an e-commerce application could use this infrastructure to provide real-time product recommendations based on user behavior, or a manufacturing system could detect equipment failures before they happen by analyzing time-series sensor data. As the application runs, contextual and environmental data—such as user actions or external sensor readings are ingested into the runtime. This data is replicated back to centralized compute clusters where machine learning models are retrained and fine-tuned to improve accuracy and performance. The updated models are automatically versioned and deployed to the runtime, where they can be A/B tested in real time. This continuous feedback loop ensures that applications evolve and improve without manual intervention, reducing time to value while maintaining model relevance. ![Spice.ai Intelligent Application Workflow](https://github.com/spiceai/docs/assets/80174/22b02c5e-5fcb-4856-b79d-911ac5d084c6) *** ## Why Spice.ai Is the Future of Intelligent Applications[​](#why-spiceai-is-the-future-of-intelligent-applications "Direct link to Why Spice.ai Is the Future of Intelligent Applications") Spice.ai introduces an entirely new way of thinking about application infrastructure by making intelligent applications accessible to developers of all skill levels. It replaces the need for custom integrations and fragmented tools with a unified runtime optimized for AI and data-driven applications. Unlike traditional architectures that rely heavily on centralized databases or cloud-based inference engines, Spice.ai focuses on tier-optimized deployments. This means that data and computation are colocated wherever the application runs—whether in the cloud, on-prem, or at the edge. Federation and materialization are at the heart of Spice.ai’s architecture. Instead of querying remote data sources directly, Spice.ai materializes working datasets locally. For example, a logistics application might materialize only the last seven days of shipment data from a cloud data lake, ensuring that 99% of queries are served locally while retaining the ability to fall back to the full dataset as needed. This reduces both latency and costs while improving the user experience. Machine learning models benefit from the same localized efficiency. Because the models are colocated with the runtime, inference happens in milliseconds rather than seconds, even for complex operations. This is critical for use cases like fraud detection, where split-second decisions can save businesses millions, or real-time personalization, where user engagement depends on instant feedback. Spice.ai also integrates with diverse infrastructures. It supports modern cloud-native systems like Snowflake and Databricks, legacy databases like SQL Server, and even unstructured sources like files stored on FTP servers. With support for industry-standard APIs like JDBC, ODBC, and Arrow Flight, it fits into existing applications without requiring extensive refactoring. In addition to data acceleration and model inference, Spice.ai provides comprehensive observability and monitoring. Every query, inference, and data flow can be tracked and audited, ensuring that applications meet enterprise standards for security, compliance, and reliability. This makes Spice.ai particularly well-suited for industries such as healthcare, finance, and manufacturing, where data privacy and traceability are paramount. *** ## Getting Started with Spice.ai[​](#getting-started-with-spiceai "Direct link to Getting Started with Spice.ai") Developers can start building intelligent applications with Spice.ai by installing the open-source runtime from [GitHub](https://github.com/spiceai/spiceai). The installation process is simple, and the runtime can be deployed across cloud, on-premises, or edge environments. Once installed, developers can create and manage `Spicepods` to define datasets and machine learning models. These `Spicepods` serve as the building blocks for their applications, streamlining data and model integration. For those looking to accelerate development, [Spicerack.org](https://spicerack.org) provides a curated library of reusable datasets and models. By using these pre-built components, developers can reduce time to deployment while focusing on the unique features of their applications. The Spice.ai community is an essential resource for new and experienced developers alike. Through forums, documentation, and hands-on support, the community helps developers unlock the full potential of intelligent applications. Whether you’re building real-time analytics systems, AI-enhanced enterprise tools, or edge-based IoT applications, Spice.ai provides the infrastructure you need to succeed. In a world where intelligent applications are increasingly becoming the norm, Spice.ai stands out as the definitive platform for building fast, scalable, and secure AI-driven solutions. Its unified approach to data, computation, and machine learning sets a new standard for how applications are developed and deployed. --- # Monitoring ![Spice.ai Monitoring](/assets/images/observability-9095d33215e242807146ca3255f60915.png) Spice provides integration with monitoring systems for production deployments: * [Datadog](/docs/next/monitoring/datadog) - Enterprise monitoring and analytics * [Grafana & Prometheus](/docs/next/monitoring/grafana) - Open source metrics and visualization * [New Relic](/docs/next/monitoring/new-relic) - Observability platform with OTLP intake * [Spice Cloud Platform](/docs/next/monitoring/spice-cloud) - Centralize task history from self-hosted runtimes * [Zipkin](/docs/next/monitoring/zipkin) - Distributed tracing --- # Datadog Spice can be monitored with [Datadog](https://www.datadoghq.com/) using the [Spice Metrics Endpoint](/docs/next/features/observability) and pre-built dashboards available in the [Spice repository](https://github.com/spiceai/spiceai/tree/trunk/monitoring). ## Datadog Agent Configuration[​](#datadog-agent-configuration "Direct link to Datadog Agent Configuration") Prerequisite: [Datadog Agent version 6.5.0 or later is installed](https://docs.datadoghq.com/getting_started/agent/). Configure the Datadog Agent to scrape the Spice metrics endpoint: 1. Edit the `openmetrics.d/conf.yaml` file in the `conf.d/` folder at the root of your [Agent’s configuration directory](https://docs.datadoghq.com/agent/guide/agent-configuration-files/#agent-configuration-directory): ``` init_config: instances: - openmetrics_endpoint: http://localhost:9090/metrics # Spice metrics endpoint namespace: spiceai metrics: - .* collect_histogram_buckets: true # Required for percentile queries, see below histogram_buckets_as_distributions: true # Required for percentile queries, see below max_returned_metrics: 3000 # Default is 2000, which Spice can exceed tags: - service_instance_id: # e.g. hostname; required by the Spice dashboard filter ``` 1. [Restart the Agent](https://docs.datadoghq.com/agent/guide/agent-commands/#start-stop-and-restart-the-agent) to start collecting Spice metrics. 2. Refer to [Prometheus and OpenMetrics metrics collection from a host](https://docs.datadoghq.com/integrations/guide/prometheus-host-collection/) for all available configuration options and supported parameters. 3. Open Datadog Metrics Explorer and type `spiceai` to confirm Spice telemetry information is successfully collected. ![](/img/datadog/spice_datadog_metrics_explorer.png) The `namespace` sets the metric prefix Every scraped metric is prefixed with the configured `namespace`, so `query_duration_ms` is stored as `spiceai.query_duration_ms`. Keep this value consistent with the prefix used by the queries in the Spice dashboard JSON, and with [`runtime.telemetry.metric_prefix`](#namespace-spice-metrics-with-a-prefix) if the same deployment also pushes metrics over OTLP. ### Histogram Percentiles (Distributions)[​](#histogram-percentiles-distributions "Direct link to Histogram Percentiles (Distributions)") Spice exports latency and size metrics — `query_duration_ms`, `flight_request_duration_ms`, `http_requests_duration_ms`, `dataset_acceleration_refresh_duration_ms`, and others — as Prometheus histograms. Configuring the Agent to submit those buckets as [Datadog distributions](https://docs.datadoghq.com/integrations/guide/prometheus-metrics/?tab=latestversion#histogram) makes percentiles queryable over the selected time window: ``` p99:spiceai.flight_request_duration_ms{method:do_get} by {command} ``` | Option | Default | Why Spice needs it | | ------------------------------------ | ------- | --------------------------------------------------------------------------------------------------------------------------------------- | | `collect_histogram_buckets` | `true` | Sends the `_bucket` series that percentiles are computed from. | | `histogram_buckets_as_distributions` | `false` | Submits those buckets as a Datadog distribution rather than a set of counters. Required for `p50:`, `p90:`, `p95:`, and `p99:` queries. | Percentile aggregations must be enabled on the metric Submitting a distribution is not sufficient on its own. Datadog computes `p50`/`p90`/`p95`/`p99` for a distribution metric only once **percentile aggregations** are enabled for that metric on the Metrics Summary page — see [Enabling advanced query functionality](https://docs.datadoghq.com/metrics/distributions/#enabling-advanced-query-functionality). Until then a `p99:` query returns no data and the widget renders empty rather than reporting an error. Why not the `_summary` metrics The metrics endpoint also exposes a `_summary` family carrying `quantile` tags, derived from each histogram at scrape time. The underlying histogram is cumulative, so those quantiles describe the **entire lifetime of the process** rather than the queried time window: they flatten the longer a pod stays up and will not surface a latency regression. Prefer distributions for any percentile that needs to track a time window. ## Kubernetes (Operator / Autodiscovery)[​](#kubernetes-operator--autodiscovery "Direct link to Kubernetes (Operator / Autodiscovery)") With the [Datadog Agent](https://docs.datadoghq.com/containers/kubernetes/installation/) in the cluster, annotate the Spice container (`spiceai` in the Helm chart) so OpenMetrics scrapes include the `service_instance_id` tag used by the dashboard instance filter: ``` ad.datadoghq.com/spiceai.checks: | { "openmetrics": { "instances": [ { "openmetrics_endpoint": "http://%%host%%:9090/metrics", "namespace": "spiceai", "metrics": [".*"], "collect_histogram_buckets": true, "histogram_buckets_as_distributions": true, "max_returned_metrics": 3000, "tags": ["service_instance_id:%%kube_pod_name%%"] } ] } } ``` `collect_histogram_buckets` and `histogram_buckets_as_distributions` serve the same purpose here as in the host configuration — see [Histogram Percentiles (Distributions)](#histogram-percentiles-distributions). `max_returned_metrics` raises the check's default limit of 2000, which Spice can exceed. ## Instance Identity[​](#instance-identity "Direct link to Instance Identity") The Spice dashboard identifies each running instance by a `service_instance_id` tag: the `instance` filter selects on it, and every per-instance panel groups by it. The Agent supplies no instance-level tag of its own — its [unified service tags](https://docs.datadoghq.com/getting_started/tagging/unified_service_tagging/) are `env`, `service`, and `version` — so set it explicitly, as in both examples above: | Deployment | Value | | ------------------- | -------------------------------------------------- | | Kubernetes | `%%kube_pod_name%%` | | Host, VM, or Docker | Hostname, or another stable per-process identifier | Panels collapse silently without this tag Datadog does not error on a missing tag. Without `service_instance_id`, each panel renders a single `N/A` series summing every instance — a 12-replica deployment reports 12 times its real dataset count rather than showing nothing. Panels from the Kubernetes integration (CPU, memory, and PVC utilization) group by `pod_name` instead, as those metrics never carry `service_instance_id`. ## Import the Spice Datadog Dashboard[​](#import-the-spice-datadog-dashboard "Direct link to Import the Spice Datadog Dashboard") 1. Create [New Datadog Dashboard](https://docs.datadoghq.com/dashboards/#get-started) ![](/img/datadog/spice_datadog_dashboard_new.png) 2. Click **Import dashboard JSON** and drag and drop [monitoring/datadog-dashboard.json](https://raw.githubusercontent.com/spiceai/spiceai/trunk/monitoring/datadog-dashboard.json) file ![](/img/datadog/spice_datadog_dashboard_import.png) 3. Dashboard is now configured to display Spice.ai OSS key performance metrics ![](/img/datadog/spice_datadog_dashboard.png) ## OpenTelemetry OTLP Export[​](#opentelemetry-otlp-export "Direct link to OpenTelemetry OTLP Export") As an alternative to scraping the Prometheus endpoint with the Datadog Agent, Spice can push metrics directly to Datadog's [OTLP Metrics Intake Endpoint](https://docs.datadoghq.com/opentelemetry/setup/intake_endpoint/) over HTTP. This is the recommended approach for agentless deployments (e.g. serverless, ephemeral containers) and for environments where the Datadog API key is managed through Spice's [secret stores](/docs/next/components/secret-stores). ### Minimal Configuration[​](#minimal-configuration "Direct link to Minimal Configuration") Replace `us3` with the Datadog site for the target account (`us3`, `us5`, `eu`, `ap1`, etc.) and store the Datadog API key in a secret: ``` runtime: telemetry: otel_exporter: endpoint: https://otlp.us3.datadoghq.com/v1/metrics headers: DD-API-KEY: ${secrets:DD_API_KEY} ``` Metrics begin appearing in the Datadog Metrics Explorer within a minute or two. ### Namespace Spice Metrics with a Prefix[​](#namespace-spice-metrics-with-a-prefix "Direct link to Namespace Spice Metrics with a Prefix") Use [`runtime.telemetry.metric_prefix`](/docs/next/reference/spicepod/runtime#runtimetelemetrymetric_prefix) to prepend a string to every exported metric name. This avoids collisions with metrics from other services in the same Datadog account: ``` runtime: telemetry: metric_prefix: 'spiceai.' ``` The runtime metric `query_duration_ms` is then exported as `spiceai.query_duration_ms`. Combining `metric_prefix` with metric filtering If you also set [`runtime.telemetry.otel_exporter.metrics`](/docs/next/reference/spicepod/runtime#runtimetelemetryotel_exporter) to whitelist specific metrics, the entries must include the prefix. The filter runs after the prefix is applied, so e.g. `query_duration_ms` will not match when `metric_prefix: 'spiceai.'` is set — use `spiceai.query_duration_ms` instead. ### Add Custom Tags via Resource Attributes[​](#add-custom-tags-via-resource-attributes "Direct link to Add Custom Tags via Resource Attributes") Attach custom key/value pairs to every metric using [`runtime.telemetry.properties`](/docs/next/reference/spicepod/runtime#runtimetelemetryproperties). Spice sends these as OpenTelemetry resource attributes: ``` runtime: telemetry: properties: environment: prod region: us-west-2 team: data-platform ``` For these resource attributes to surface as **tags** in Datadog, the Datadog OTLP intake also requires the `dd-otel-metric-config` header with `resource_attributes_as_tags` enabled (see [Datadog OTLP Metrics Intake Endpoint](https://docs.datadoghq.com/opentelemetry/setup/intake_endpoint/)): ``` runtime: telemetry: otel_exporter: endpoint: https://otlp.us3.datadoghq.com/v1/metrics headers: DD-API-KEY: ${secrets:DD_API_KEY} dd-otel-metric-config: '{"resource_attributes_as_tags": true}' ``` Tags can lag behind metrics Datadog typically ingests OTLP metrics within seconds, but the associated tags (from resource attributes) can take noticeably longer to appear in the UI — sometimes several minutes after the first datapoints. The metrics and tags do eventually converge. Manage tag cardinality in Datadog Datadog [bills on custom metric cardinality](https://docs.datadoghq.com/account_management/billing/custom_metrics/), driven by the number of unique tag-value combinations per metric. The custom tags added via `runtime.telemetry.properties` are typically low-cardinality (`environment`, `region`, `team`), but Spice metrics also carry a number of automatically populated dimensions — for example `dataset`, `protocol`, `client`, `client_version`, `client_system`, `user_agent`, `runtime`, `runtime_version`, `runtime_system` (see [Available Metrics](/docs/next/features/observability#available-metrics)) — some of which can grow with the size of the deployment. Datadog's [Metrics without Limits™](https://docs.datadoghq.com/metrics/metrics-without-limits/) decouples ingestion from indexing for exactly this case. With Metrics without Limits™, every tag Spice emits is still ingested, but each metric is configured with one of: * an **allowlist** that keeps only the tags actually used in dashboards, monitors, and queries (e.g. keep `dataset` and `environment`, drop the rest), or * a **blocklist** that drops specific auto-populated tags that are not useful for a given metric (e.g. exclude `user_agent` or `client_version`). Only the indexed (queryable) tag combinations count toward custom metric billing. Configuration is done per metric in the Metrics Summary page or via the Metrics API, and the in-app UI surfaces an estimated indexed-metric volume before saving and can pre-populate an allowlist from tags actively queried in dashboards, monitors, and notebooks. ### Full Example[​](#full-example "Direct link to Full Example") A complete `runtime.telemetry` block combining metric prefixing, custom tags, and Datadog OTLP export: ``` runtime: telemetry: metric_prefix: 'spiceai.' properties: environment: prod region: us-west-2 team: data-platform otel_exporter: endpoint: https://otlp.us3.datadoghq.com/v1/metrics headers: DD-API-KEY: ${secrets:DD_API_KEY} dd-otel-metric-config: '{"resource_attributes_as_tags": true}' ``` With this configuration, every Spice metric (e.g. `spiceai.query_duration_ms`, `spiceai.query_executions`) arrives in Datadog tagged with `environment:prod`, `region:us-west-2`, and `team:data-platform`. For general OTLP exporter options (push interval, metric filtering, gRPC vs HTTP), see [OpenTelemetry Metrics Exporter](/docs/next/features/observability#opentelemetry-metrics-exporter). --- # Grafana & Prometheus Spice can be monitored with [Grafana](https://grafana.com/grafana/) using the [Spice Metrics Endpoint](/docs/next/features/observability) and pre-built dashboards available in the [Spice repository](https://github.com/spiceai/spiceai/tree/trunk/monitoring). ## Import Grafana Dashboard[​](#import-grafana-dashboard "Direct link to Import Grafana Dashboard") Navigate to the Dashboards section in Grafana and click "New" > "Import". ![](/img/grafana/import-dashboard-button.png) Copy the dashboard JSON from [monitoring/grafana-dashboard.json](https://github.com/spiceai/spiceai/blob/trunk/monitoring/grafana-dashboard.json) into the Grafana import box. ![](/img/grafana/import-dashboard.png) Click "Load". ## Kubernetes[​](#kubernetes "Direct link to Kubernetes") View the [Kubernetes](/docs/next/deployment/kubernetes/helm) deployment guide for configuring the Prometheus Operator to scrape metrics from Spice pods (`monitoring.podMonitor.enabled=true`). The dashboard **Kubernetes Resource Utilization** panels need cluster metrics in Prometheus (`k8s_pod_cpu_usage`, `k8s_pod_memory_working_set`, `k8s_volume_capacity`, `k8s_volume_available` with label `k8s_pod_name`). A common source is an OpenTelemetry Collector [kubeletstats](https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/receiver/kubeletstatsreceiver) receiver exporting those names (without unit suffixes). The scrape `instance` label should match the pod name so the Instances filter applies to those panels. ## Prometheus[​](#prometheus "Direct link to Prometheus") Configure a Prometheus instance to scrape metrics from the Spice runtimes. ``` global: scrape_interval: 1s scrape_configs: - job_name: spiceai static_configs: - targets: ['127.0.0.1:9090'] # Change to your Spice runtime endpoint + port ``` ## Local Quickstart[​](#local-quickstart "Direct link to Local Quickstart") This tutorial creates and configures Grafana and Prometheus locally to scrape and display metrics from several Spice instances. It assumes: * Two Spice runtimes, `spiced-main` and `spiced-edge`, are running on `127.0.0.1:9091` and `127.0.0.1:9092` respectively. 1. Create a `compose.yaml`: ``` version: '3' services: prometheus: image: prom/prometheus:latest volumes: - ./prometheus.yaml:/etc/prometheus/prometheus.yml ports: - 9090:9090 network_mode: 'host' grafana: image: grafana/grafana:latest volumes: - ./.grafana/provisioning:/etc/grafana/provisioning ports: - 3000:3000 network_mode: 'host' ``` 2. Create a `prometheus.yaml` to ``` global: scrape_interval: 1s scrape_configs: - job_name: spiced-main static_configs: - targets: ['127.0.0.1:9091'] - job_name: spiced-edge static_configs: - targets: ['127.0.0.1:9092'] ``` 3. Add a prometheus as a source to grafana. Create a `.grafana/provisioning/datasources/prometheus.yml` ``` apiVersion: 1 datasources: - name: Prometheus type: prometheus access: proxy url: http://localhost:9090 isDefault: true ``` 4. Run the Docker Compose ``` docker-compose up ``` 5. Go to `http://localhost:3000/dashboard/import` and add the JSON from [monitoring/grafana-dashboard.json](https://github.com/spiceai/spiceai/blob/trunk/monitoring/grafana-dashboard.json). 6. The dashboard will have data from the Spice runtimes. ![](/img/grafana/screenshot.png) ## Query Spice as a Grafana Data Source[​](#query-spice-as-a-grafana-data-source "Direct link to Query Spice as a Grafana Data Source") In addition to monitoring Spice with Grafana, you can query datasets served by Spice and visualize the results in Grafana panels using the [Infinity data source](https://grafana.com/grafana/plugins/yesoreyeram-infinity-datasource/), which can query Spice's [HTTP SQL API](/docs/next/api/HTTP/post-sql). 1. Install the [Infinity data source](https://grafana.com/grafana/plugins/yesoreyeram-infinity-datasource/) plugin from the Grafana plugin catalog. 2. Add a new Infinity data source. No base URL or authentication is required at the data source level when targeting a Spice runtime that does not require an API key — credentials can be configured per query if needed (see [API Auth](/docs/next/api/auth)). 3. Create a panel backed by the Infinity data source and configure the query as an HTTP request against the Spice SQL endpoint: * **Type**: `JSON` * **Method**: `POST` * **URL**: `http://localhost:8090/v1/sql` * **Headers**: `Content-Type: application/json` * **Body** (raw): ``` { "sql": "SELECT passenger_count, AVG(total_amount) FROM taxi_trips GROUP BY passenger_count ORDER BY passenger_count" } ``` The endpoint returns a JSON array of row objects (the default `application/json` response format), which Infinity parses directly into table rows for visualization. note The legacy `spiceai-spicexyz-datasource` Grafana plugin is no longer maintained and targets the earlier Spice.ai product, not the current runtime. Use the Infinity data source against the HTTP SQL API as shown above. --- # New Relic Spice can be monitored with [New Relic](https://newrelic.com/) using either the [Spice Metrics Endpoint](/docs/next/features/observability) (Prometheus scrape via the New Relic infrastructure agent) or the [OpenTelemetry Metrics Exporter](/docs/next/features/observability#opentelemetry-metrics-exporter) (push directly to New Relic's OTLP intake). For agent-based collection, see the New Relic [Prometheus integrations overview](https://docs.newrelic.com/docs/infrastructure/prometheus-integrations/get-started/send-prometheus-metric-data-new-relic/). The walkthrough below covers the agentless OTLP path, which is recommended for serverless and ephemeral deployments and for environments where the New Relic license key is managed through Spice's [secret stores](/docs/next/components/secret-stores). ## OpenTelemetry OTLP Export[​](#opentelemetry-otlp-export "Direct link to OpenTelemetry OTLP Export") New Relic accepts OpenTelemetry metrics on a hosted OTLP endpoint. Spice pushes directly to it without requiring an OpenTelemetry collector or the New Relic agent. ### Minimal Configuration[​](#minimal-configuration "Direct link to Minimal Configuration") Pick the endpoint that matches your account region (see [New Relic OTLP endpoint configuration](https://docs.newrelic.com/docs/opentelemetry/best-practices/opentelemetry-otlp/)) and store the New Relic license key in a secret: ``` runtime: telemetry: otel_exporter: endpoint: https://otlp.nr-data.net/v1/metrics headers: api-key: ${secrets:new_relic_license_key} ``` | Region | OTLP/HTTP endpoint | | ------------ | ------------------------------------------ | | US (default) | `https://otlp.nr-data.net/v1/metrics` | | EU | `https://otlp.eu01.nr-data.net/v1/metrics` | | FedRAMP | `https://gov-otlp.nr-data.net/v1/metrics` | The header name is `api-key` (lowercase). Use a New Relic [license key](https://docs.newrelic.com/docs/apis/intro-apis/new-relic-api-keys/#license-key) — either the account's ingest license key or an ingest-specific key. Metrics begin appearing in New Relic's [Metrics Explorer](https://docs.newrelic.com/docs/data-apis/understand-data/metric-data/query-metric-data-type/) within a minute or two. ### Namespace Spice Metrics with a Prefix[​](#namespace-spice-metrics-with-a-prefix "Direct link to Namespace Spice Metrics with a Prefix") Use [`runtime.telemetry.metric_prefix`](/docs/next/reference/spicepod/runtime#runtimetelemetrymetric_prefix) to prepend a string to every exported metric name. This avoids collisions with metrics from other services in the same New Relic account: ``` runtime: telemetry: metric_prefix: 'spiceai.' ``` The runtime metric `query_duration_ms` is then exported as `spiceai.query_duration_ms`. Combining `metric_prefix` with metric filtering If you also set [`runtime.telemetry.otel_exporter.metrics`](/docs/next/reference/spicepod/runtime#runtimetelemetryotel_exporter) to whitelist specific metrics, the entries must include the prefix. The filter runs after the prefix is applied, so e.g. `query_duration_ms` will not match when `metric_prefix: 'spiceai.'` is set — use `spiceai.query_duration_ms` instead. ### Add Custom Attributes via Resource Attributes[​](#add-custom-attributes-via-resource-attributes "Direct link to Add Custom Attributes via Resource Attributes") Attach custom key/value pairs to every metric using [`runtime.telemetry.properties`](/docs/next/reference/spicepod/runtime#runtimetelemetryproperties). Spice sends these as OpenTelemetry resource attributes, which New Relic surfaces as queryable dimensions on each metric: ``` runtime: telemetry: properties: environment: prod region: us-west-2 team: data-platform ``` These attributes are available in NRQL via the `WHERE` and `FACET` clauses, e.g.: ``` SELECT average(spiceai.query_duration_ms) FROM Metric WHERE environment = 'prod' FACET region SINCE 1 hour ago ``` ### Full Example[​](#full-example "Direct link to Full Example") A complete `runtime.telemetry` block combining metric prefixing, custom attributes, and New Relic OTLP export: ``` runtime: telemetry: metric_prefix: 'spiceai.' properties: environment: prod region: us-west-2 team: data-platform otel_exporter: endpoint: https://otlp.nr-data.net/v1/metrics push_interval: '30s' headers: api-key: ${secrets:new_relic_license_key} ``` With this configuration, every Spice metric (e.g. `spiceai.query_duration_ms`, `spiceai.query_executions`) arrives in New Relic with `environment`, `region`, and `team` available as dimensions for use in NRQL queries, dashboards, and alerts. For general OTLP exporter options (push interval, metric filtering, gRPC vs HTTP), see [OpenTelemetry Metrics Exporter](/docs/next/features/observability#opentelemetry-metrics-exporter). --- # Spice Cloud Platform A self-hosted Spice runtime — running in Kubernetes, Docker, on a VM, or on a laptop — can connect to the [Spice Cloud Platform](https://spice.ai) to centralize [task history](/docs/next/reference/task_history) across one or many runtimes. Once connected, query logs, refresh activity, AI tool invocations, and runtime errors from every connected instance are visible in a single SCP app, queryable as a standard `runtime.task_history` table, and retained beyond each runtime's local in-memory window. The connection is configured declaratively in `spicepod.yaml` with a top-level `management:` block. No CLI command, sidecar, or extra agent is required — the runtime itself streams task history to SCP over Arrow Flight. ## What gets sent[​](#what-gets-sent "Direct link to What gets sent") The `management:` block enables a single export path: rows from the local [`runtime.task_history`](/docs/next/reference/task_history) accelerated table are appended every 5 seconds to the `runtime.task_history` dataset of the target Spice Cloud app. Each row is a span representing one unit of execution — a SQL query, an AI chat completion, a dataset refresh, a tool call — including its inputs, outputs, duration, and error status. The export covers a **rolling 3-day window** from the runtime's current time. Records older than 3 days are not backfilled when a runtime first connects. The following are **not** sent through this path: * Prometheus metrics — use the [Prometheus endpoint](/docs/next/features/observability) or [OTLP exporter](/docs/next/features/observability) for [Datadog](/docs/next/monitoring/datadog), [New Relic](/docs/next/monitoring/new-relic), Grafana, or any OTLP-compatible backend. * Distributed traces — export to [Zipkin](/docs/next/monitoring/zipkin) or an OTLP tracing backend. * Application logs — Spice writes logs to stdout/stderr; collect them with the platform's log shipper. For a full observability stack on self-hosted deployments, combine `management:` with the metrics and tracing exporters above. `management:` covers task history; the others cover everything else. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A [Spice Cloud Platform](https://spice.ai/login) account. * A Spice Cloud app to receive task history. Apps are created from the Spice Cloud Console or via the CLI. * An API key for that app. * A self-hosted Spice runtime at v1.10.0 or later, with [task history enabled](/docs/next/reference/task_history) (the default). ### Create an app and API key[​](#create-an-app-and-api-key "Direct link to Create an app and API key") From the Spice Cloud Console, create a new app (any name and visibility). Copy the API key shown on the app's settings page — it is only displayed at creation time and on demand from the **API Keys** tab. To create the app from the CLI: ``` spice cloud login spice cloud create app my-observability --visibility private ``` The output includes the app's primary API key. Treat the key as a secret — anyone with the key can write task history to (and read it from) the app. For an existing app, retrieve the current key with: ``` spice cloud api-keys --app / ``` ## Configuration[​](#configuration "Direct link to Configuration") Add a `management:` block to `spicepod.yaml` on each self-hosted runtime. The API key should be sourced from a [secret store](/docs/next/components/secret-stores) rather than hard-coded. ### Minimal example[​](#minimal-example "Direct link to Minimal example") ``` version: v2 kind: Spicepod name: my-runtime management: api_key: ${secrets:SPICEAI_API_KEY} ``` With `SPICEAI_API_KEY` set in the environment (or `.env` file picked up by the [env secret store](/docs/next/components/secret-stores/env)), the runtime connects to the default Spice Cloud region on startup. On success, the runtime logs: ``` INFO runtime::management: Connected to Spice Cloud for management and monitoring ``` ### Selecting a region[​](#selecting-a-region "Direct link to Selecting a region") By default, the runtime targets the Spice Cloud default region. To target a specific region, set `region` in `params`: ``` management: api_key: ${secrets:SPICEAI_API_KEY} params: region: us-east-1 ``` The region must match the region the API key was issued in. Available regions are listed in the Spice Cloud Console. ### Custom endpoint[​](#custom-endpoint "Direct link to Custom endpoint") For VPC-peered or otherwise-routed Cloud endpoints, override the Flight endpoint explicitly with `data_endpoint`: ``` management: api_key: ${secrets:SPICEAI_API_KEY} params: data_endpoint: https://us-east-1-prod-aws-flight.spiceai.io ``` `data_endpoint` takes precedence over `region`. The scheme must be `https://` for production; `http://` is accepted for local testing only. ### Disabling without removing config[​](#disabling-without-removing-config "Direct link to Disabling without removing config") Set `enabled: false` to turn off the export without removing the block (useful for staging vs. production overrides): ``` management: enabled: false api_key: ${secrets:SPICEAI_API_KEY} ``` ## Reference[​](#reference "Direct link to Reference") The `management:` block lives at the top level of `spicepod.yaml`, alongside `runtime:`, `datasets:`, and `models:`. | Field | Type | Required | Default | Description | | ---------------------- | ------- | -------- | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | | `enabled` | boolean | No | `true` | Whether the management export is active. When `false`, the block is parsed but no connection is established. | | `api_key` | string | Yes | — | API key for the target Spice Cloud app. Resolved through any [secret store](/docs/next/components/secret-stores) via `${secrets:KEY}`. | | `params.region` | string | No | — | Spice Cloud region (e.g. `us-east-1`). Used to build the Flight endpoint when `data_endpoint` is not set. | | `params.data_endpoint` | string | No | Built from `region` | Flight endpoint URL override. Must be `https://` or `http://`. Takes precedence over `region`. | ### Behavior[​](#behavior "Direct link to Behavior") * **Transport:** Apache Arrow Flight over gRPC, authenticated with the API key via HTTP Basic on the Flight handshake. * **Export interval:** Every 5 seconds. Records pending at runtime shutdown are flushed once more before the process exits. * **Retention window:** Each export sends task history rows with an `end_time` within the last 3 days. Older rows are not backfilled. * **Retry policy:** Failed exports use Fibonacci backoff with up to 10 retries before the export attempt is dropped (the next 5-second tick retries from scratch). * **Task history dependency:** When `runtime.task_history.enabled` is `false`, the management export stays initialized but never sends data. Task history is enabled by default; see the [task history reference](/docs/next/reference/task_history) for retention and capture options. ## Verification[​](#verification "Direct link to Verification") After restarting the runtime, confirm the connection is established: 1. **Check the runtime log** for the connection line: ``` INFO runtime::management: Connected to Spice Cloud for management and monitoring ``` For deeper logging during setup, run with `RUST_LOG=runtime::management=debug`. 2. **Run a query** against the local runtime to generate task history: ``` spice sql ``` ``` SELECT 1; ``` 3. **Wait 5–10 seconds**, then query the target Spice Cloud app's task history. From any client logged in to Spice Cloud: ``` spice sql --cloud / ``` ``` SELECT start_time, input, error_message FROM runtime.task_history ORDER BY start_time DESC LIMIT 10; ``` The `SELECT 1` query issued on the self-hosted runtime appears in the result set. ## Connecting multiple runtimes[​](#connecting-multiple-runtimes "Direct link to Connecting multiple runtimes") A single Spice Cloud app can receive task history from any number of self-hosted runtimes — share the same `api_key` across them. To distinguish runtimes when querying the consolidated table, tag each runtime in `spicepod.yaml`: ``` name: edge-eu-west-1 runtime: params: deployment_label: edge-eu-west-1 ``` The `name` field is recorded with every task history span and is the simplest way to filter: ``` SELECT spicepod_name, count(*) AS task_count FROM runtime.task_history WHERE start_time > now() - INTERVAL '1 hour' GROUP BY spicepod_name ORDER BY task_count DESC; ``` ## Limitations[​](#limitations "Direct link to Limitations") * **Task history only.** Metrics, traces, and logs are not exported through this path. Use the [observability features](/docs/next/features/observability) or third-party [monitoring integrations](/docs/next/monitoring) for those signals. * **3-day rolling window.** On first connect, only the last 3 days of task history are eligible. Older local history is not backfilled. * **Append-only.** The export is one-way and append-only — the self-hosted runtime never reads from the Cloud table, and rows are not updated or deleted from the Cloud copy. * **No live-reload of the API key.** Rotating the key requires restarting the runtime. Coordinate with the chosen secret store. * **5-second export interval is not configurable.** The interval is a runtime constant. * **API key auth only.** OIDC / SSO authentication is not supported on the management path. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") | Symptom | Likely cause | Resolution | | --------------------------------------------------------------------- | ---------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Startup error: `Missing required secret: api_key. Specify a value.` | `api_key` is empty or the referenced secret is not defined. | Ensure the secret store has the key and the `${secrets:...}` reference matches. Test secret resolution with `spice run` and the [env secret store](/docs/next/components/secret-stores/env). | | Startup error: `Failed to create data connector for cloud management` | The Flight endpoint cannot be reached, or the API key is rejected. | Verify outbound TLS connectivity to `https://-prod-aws-flight.spiceai.io`. Regenerate the API key from the Spice Cloud Console; confirm `region` matches the key's region. | | No connection log line after restart | The `management:` block was not parsed — typically a YAML indentation issue. | Confirm `management:` is at the top level of `spicepod.yaml`, not nested under `runtime:`. Check `spice run` startup logs for spicepod-parse warnings. | | Records never appear in the Cloud app | Task history is disabled, or the runtime hasn't run any spans yet. | Confirm `runtime.task_history.enabled` is `true` (the default). Issue a query against the local runtime, then wait at least 5 seconds for the next export tick. | | `UNAUTHENTICATED` on Flight handshake | Wrong API key, key from a different app, or key from a different region. | Regenerate the API key from the target app's **API Keys** tab and update the secret store. Verify `params.region` matches the app's region. | | Repeated `Failed to export runtime task history records` warnings | Transient network failure between the runtime and Spice Cloud. | The runtime retries automatically on the next 5-second tick. If errors persist, check outbound network policy, DNS resolution, and the [Spice.ai status page](https://status.spice.ai). | | Records appear, but `spicepod_name` is empty for some runtimes | The runtime's spicepod has no `name:` field. | Set `name:` at the top of `spicepod.yaml` on each runtime so consolidated rows can be filtered. | For detailed export logging, set `RUST_LOG=runtime::management=trace` and watch for per-flush log lines like `Exported {n} task history records`. ## Related[​](#related "Direct link to Related") * [Task History Reference](/docs/next/reference/task_history) — schema, retention, and local querying. * [Observability & Monitoring](/docs/next/features/observability) — Prometheus metrics, OTLP export, and tracing. * [Spice.ai Data Connector](/docs/next/components/data-connectors/spiceai) — the underlying connector used for Cloud federation. * [Secret Stores](/docs/next/components/secret-stores) — managing the `api_key` securely. * [Spice Cloud Platform Deployment](/docs/next/deployment/cloud) — deploying applications on Spice Cloud. --- # Zipkin Integration Spice supports distributed tracing by integrating with [Zipkin](https://zipkin.io/) and compatible tracing systems. ![zipkin](https://github.com/user-attachments/assets/b2763cdf-1ec9-4b85-9a88-16a11a98aaf6) ![zipkin](https://github.com/user-attachments/assets/c28e042e-96d8-4aab-8da4-d1f493642e02) ## Configuration[​](#configuration "Direct link to Configuration") Enable Zipkin tracing by configuring the `runtime.tracing` section in `spicepod.yaml`: ``` runtime: tracing: zipkin_enabled: true zipkin_endpoint: "http://your_zipkin_host:9411/api/v2/spans" ``` Trace data will be available in the Zipkin UI at `http://your_zipkin_host:9411`. See the [Zipkin Quickstart](https://zipkin.io/pages/quickstart) to run a test server. ## Trace Information[​](#trace-information "Direct link to Trace Information") A trace in Spice represents a completed task (such as a SQL query, AI chat completion, or tool call). Each trace is a unique span, recording execution time, inputs, outputs, and errors. ``` +----------------------------------+------------------+---------------------+----------------------------+----------------------------+-----------------------+---------------------------------------------------------------------------------------------+ | trace_id | span_id | task | start_time | end_time | execution_duration_ms | error_message | +----------------------------------+------------------+---------------------+----------------------------+----------------------------+-----------------------+---------------------------------------------------------------------------------------------+ | 687e0970f8c49d19c5a08764ea2d4dc1 | f4f52ed29db8b151 | text_embed | 2024-11-25T05:39:37.444749 | 2024-11-25T05:39:53.577195 | 16132.446000000002 | | | 1e881188e5fd252b26adb8a8d838efb8 | 532b0019ad778094 | sql_query | 2024-11-25T05:40:38.864982 | 2024-11-25T05:40:38.871090 | 6.108 | | | ac5abd8bfec7e5aa7c19fc84772c55f1 | 316622ac359e3c00 | sql_query | 2024-11-25T05:39:39.872946 | 2024-11-25T05:39:39.872994 | 0.048 | This feature is not implemented: The context currently only supports a single SQL statement | | 701874d7282dd47791e7519b343a9694 | 5dacf75c4537ee0e | accelerated_refresh | 2024-11-25T05:39:30.452534 | 2024-11-25T05:39:30.452900 | 0.366 | | | 18d76b6389898cc5253a49294607477d | cc0d06a4e69cbcd5 | health | 2024-11-25T05:39:30.451626 | 2024-11-25T05:39:31.563876 | 1112.25 | | | 3c75d16b6b4b8da98c551d115e1c049c | 9a16dc065a95236a | sql_query | 2024-11-25T05:42:27.386754 | 2024-11-25T05:42:27.386859 | 0.10500000000000001 | SQL error: ParserError("Expected: an SQL statement, found: ELECT") | +----------------------------------+------------------+---------------------+----------------------------+----------------------------+-----------------------+---------------------------------------------------------------------------------------------+ ``` For more details, see [task\_history](/docs/next/reference/task_history). --- # Spice.ai OSS Reference Docs This section contains reference documentation for the Spice.ai runtime, CLI, API, and Spicepod configuration syntax. ## [🗃Spicepod specification](/docs/next/reference/spicepod) [11 items](/docs/next/reference/spicepod) ## [🗃SQL Reference](/docs/next/reference/sql) [12 items](/docs/next/reference/sql) ## [🗃Data Types Reference](/docs/next/reference/datatypes) [2 items](/docs/next/reference/datatypes) ## [📄️Models Grade Report](/docs/next/reference/models) [Spice AI graded Large-Language-Model (LLM) evaluation report](/docs/next/reference/models) ## [📄️Task History](/docs/next/reference/task_history) [The Spice runtime stores information about completed tasks in the spice.runtime.task\_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span.](/docs/next/reference/task_history) ## [📄️Duration](/docs/next/reference/duration) [Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted.](/docs/next/reference/duration) ## [📄️File Formats](/docs/next/reference/file_format) [File-based data connectors — including s3//, file//, sftp://, and others — support multiple structured and document file formats. This page details the format-specific parameters available for each.](/docs/next/reference/file_format) ## [📄️Cron Schedules](/docs/next/reference/cron) [The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc).](/docs/next/reference/cron) ## [📄️System Requirements](/docs/next/reference/system_requirements) [System requirements for running Spice.ai Open Source](/docs/next/reference/system_requirements) ## [📄️Distributions](/docs/next/reference/distributions) [Distribution variants of the Spice runtime for different use cases and deployment scenarios, including data-only, GPU-accelerated, NAS, and allocator variants.](/docs/next/reference/distributions) ## [📄️Memory](/docs/next/reference/memory) [Guidelines and best practices for managing memory usage and optimizing performance in Spice deployments.](/docs/next/reference/memory) ## [📄️Performance Tuning](/docs/next/reference/performance-tuning) [Comprehensive guide to optimizing query performance, acceleration, and resource utilization in Spice deployments.](/docs/next/reference/performance-tuning) --- # Cron Schedules The Runtime supports cron expressions with optional seconds, like `*/10 * * * * *` which evaluates to every 10th second (10, 20, 30, etc). Cron expressions in the Runtime evaluate according to the systems local time where the Runtime is running. The Spice Runtime uses the [`croner` Rust crate](https://github.com/hexagon/croner-rust?tab=readme-ov-file#pattern) for parsing cron expressions. ## Examples[​](#examples "Direct link to Examples") ### At 1am every Monday[​](#at-1am-every-monday "Direct link to At 1am every Monday") ``` 0 1 * * 1 ``` ### At midday every weekday (Monday-Friday)[​](#at-midday-every-weekday-monday-friday "Direct link to At midday every weekday (Monday-Friday)") ``` 0 12 * * 1-5 ``` ### Every hour, at 5 minutes past the hour[​](#every-hour-at-5-minutes-past-the-hour "Direct link to Every hour, at 5 minutes past the hour") ``` 5 * * * * ``` ### Every 10 minutes[​](#every-10-minutes "Direct link to Every 10 minutes") ``` */10 * * * * ``` ### Every 5 minutes at 30 seconds[​](#every-5-minutes-at-30-seconds "Direct link to Every 5 minutes at 30 seconds") ``` 30 */5 * * * * ``` --- # Data Types Reference Spice uses [Apache Arrow](https://arrow.apache.org/) data types internally, providing consistent type handling across different data sources and accelerators. This section documents how Arrow types map to specific accelerators and object store formats. For SQL type casting and conversion, see [Type Casting Operators](/docs/next/reference/sql/operators#type-casting-operators). ## [📄️Accelerator Data Types](/docs/next/reference/datatypes/accelerators) [Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator.](/docs/next/reference/datatypes/accelerators) ## [📄️Object Store Data Types](/docs/next/reference/datatypes/object_store) [Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice.](/docs/next/reference/datatypes/object_store) --- # Accelerator Data Types Spice adheres to Apache Arrow data [types](https://docs.rs/arrow/latest/arrow/datatypes/index.html). Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. | Arrow Type | Description | [Spice Cayenne (Vortex)](https://github.com/vortex-data/vortex) | [DuckDB](https://duckdb.org/docs/sql/data_types/overview) | [SQLite](https://sqlite.org/datatype3.html) | [Postgres](https://www.postgresql.org/docs/current/datatype.html#DATATYPE-TABLE) | | -------------------------- | ---------------------------------------------------------------------------- | --------------------------------------------------------------- | ---------------------------------------------------------- | ------------------------------------------- | -------------------------------------------------------------------------------- | | na | A NULL type having no physical storage. | `Null` | | | | | bool | Boolean as 1 bit, LSB bit-packed ordering. | `Bool` | `BOOLEAN` | `BOOL` | `BOOL` | | uint8 | Unsigned 8-bit little-endian integer. | `Primitive(U8)` | `UTINYINT` | `TINYINT` | `SMALLINT` | | int8 | Signed 8-bit little-endian integer. | `Primitive(I8)` | `TINYINT` | `TINYINT` | `SMALLINT` | | uint16 | Unsigned 16-bit little-endian integer. | `Primitive(U16)` | `USMALLINT` | `SMALLINT` | `SMALLINT` | | int16 | Signed 16-bit little-endian integer. | `Primitive(I16)` | `SMALLINT` | `SMALLINT` | `SMALLINT` | | uint32 | Unsigned 32-bit little-endian integer. | `Primitive(U32)` | `UINTEGER` | `INT` | `INTEGER` | | int32 | Signed 32-bit little-endian integer. | `Primitive(I32)` | `INTEGER` | `INT` | `INTEGER` | | uint64 | Unsigned 64-bit little-endian integer. | `Primitive(U64)` | `UBIGINT` | `BIGINT` | `BIGINT` | | int64 | Signed 64-bit little-endian integer. | `Primitive(I64)` | `BIGINT` | `BIGINT` | `BIGINT` | | half\_float | 2-byte floating point value | `Primitive(F32)` | | | | | float | 4-byte floating point value | `Primitive(F32)` | `FLOAT` | `FLOAT` | `REAL` | | double | 8-byte floating point value | `Primitive(F64)` | `DOUBLE` | `DOUBLE` | `DOUBLE PRECISION` | | string | UTF8 variable-length string as List\ | `Utf8` | `VARCHAR` | `TEXT` | `TEXT` | | binary | Variable-length bytes (no guarantee of UTF8-ness) | `Binary` | `BLOB` | | | | fixed\_size\_binary | Each value has equal bytes of binary. | | `BLOB` | | | | date32 | int32\_t days since the UNIX epoch | `Extension(Date)` | `DATE` | `DATE` | | | date64 | int64\_t milliseconds since the UNIX epoch | `Extension(Date)` | `DATE` | `TIMESTAMP` | | | timestamp | Exact timestamp encoded with int64 since UNIX epoch, seconds or milliseconds | `Extension(Timestamp)` | `TIMESTAMP_S`, `TIMESTAMP_MS`, `TIMESTAMP`, `TIMESTAMP_NS` | `TIMESTAMP` | `TIMESTAMP` | | time32 | Time as signed 32-bit integer, seconds or milliseconds since midnight. | `Extension(Time)` | `TIME` | `TIME` | | | time64 | Time as signed 64-bit integer, microseconds or nanoseconds since midnight. | `Extension(Time)` | `TIME` | `TIME` | | | interval\_months | YEAR\_MONTH interval in SQL style. | | | | | | interval\_day\_time | DAY\_TIME interval in SQL style. | | | | | | decimal128 | Precision- and scale-based decimal type with 128 bits. | `Decimal` | `DECIMAL(P, S)` | `DECIMAL(38, 10)` | `DECIMAL(38, 10)` | | decimal | Defined for backward-compatibility. | `Decimal` | | | | | decimal256 | Precision- and scale-based decimal type with 256 bits. | `Decimal` | | | | | list | A list of some logical data type. | `List` | `TYPE[]` | | `TYPE[]` | | struct | Struct of logical types. | `Struct` | `STRUCT(...)` | | | | sparse\_union | Sparse unions of logical types. | | | | | | dense\_union | Dense unions of logical types. | | | | | | dictionary | Dictionary-encoded type, | | | | | | map | Map, a repeated struct logical type. | | | | | | extension | Custom data type, implemented by user. | `Extension` | | | | | fixed\_size\_list | Fixed size list of some logical type. | `FixedSizeList` | `TYPE[N]` | | | | duration | Elapsed time in seconds, milliseconds, microseconds or nanoseconds. | | | | | | large\_string | Like STRING, but with 64-bit offsets. | `Utf8` | `VARCHAR` | | | | large\_binary | Like BINARY, but with 64-bit offsets. | `Binary` | `BLOB` | | | | large\_list | Like LIST, but with 64-bit offsets. | `List` | `TYPE[]` | | | | interval\_month\_day\_nano | Calendar interval type with three fields. | | `INTERVAL` | | | | run\_end\_encoded | Run-end encoded data. | | | | | | string\_view | UTF8 view type with 4-byte prefix & inline small string optimization. | `Utf8` | `VARCHAR` | | | | binary\_view | Bytes view type with 4-byte prefix and inline small string optimization. | `Binary` | `BLOB` | | | | list\_view | A list of some logical data type represented by offset and size. | | | | | | large\_list\_view | Like LIST\_VIEW, but with 64-bit offsets and sizes. | | | | | Note: Where `TYPE` is used (e.g. `TYPE[]`), it refers an established supported type for the specific data accelerator (e.g. `INTEGER[]`). Spice Cayenne (Vortex) provides zero-copy compatibility with Apache Arrow and supports most Arrow types through its logical type system. A `half_float` (Arrow `Float16`) column is widened to `Float32` before it is written, so it is stored — and read back — as `Primitive(F32)`; Vortex has no 2-byte float type. For DuckDB, the accelerator hands DuckDB the Arrow schema and uses the `CREATE TABLE` statement DuckDB derives from it, so unsigned Arrow integers keep their unsignedness (`uint32` becomes `UINTEGER`, not `INTEGER`) and `decimal128` keeps its precision and scale. `TYPE[N]` denotes a fixed-length list (e.g. `INTEGER[4]`), and a timezone-aware `timestamp` becomes `TIMESTAMP WITH TIME ZONE`. --- # Object Store Data Types Spice adheres to Apache Arrow data [types](https://docs.rs/arrow/latest/arrow/datatypes/index.html). The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. ## Parquet[​](#parquet "Direct link to Parquet") | Parquet Physical Type | Arrow Type | | ---------------------- | -------------------------------------------------------------------------------------------- | | `BOOLEAN` | `Boolean` | | `INT32` | `Int8`/`Int16`/`Int32`/`UInt8`/`UInt16`/`UInt32`/`Date32`/`Time32`/`Decimal128` | | `INT64` | `Int64`/`UInt64`/`Time64`/`Timestamp(Millisecond/Microsecond/Nanosecond, None)`/`Decimal128` | | `INT96` | `Timestamp(Nanosecond, None)` | | `FLOAT` | `Float32` | | `DOUBLE` | `Float64` | | `BYTE_ARRAY` | `Utf8`/`Binary`/`Decimal128`/`Decimal256` | | `FIXED_LEN_BYTE_ARRAY` | `Decimal128`/`Decimal256`/`Interval`/`Float16`/`FixedSizeBinary` | --- # Spice Runtime Distributions The Spice open source project provides multiple distribution variants to support different use cases and deployment scenarios. note The Spice runtime is **64-bit only**. 32-bit platforms are not supported. ## Image Channels[​](#image-channels "Direct link to Image Channels") | Channel | Image | Description | | -------------------------------------------------------------------------------------------------------- | --------------------------------- | --------------------------------------- | | [DockerHub](https://hub.docker.com/r/spiceai/spiceai) | `spiceai/spiceai` | Official release images | | [GitHub Container Registry](https://github.com/spiceai/spiceai/pkgs/container/spiceai) | `ghcr.io/spiceai/spiceai` | Official release images | | [GitHub Container Registry (Nightly)](https://github.com/spiceai/spiceai/pkgs/container/spiceai-nightly) | `ghcr.io/spiceai/spiceai-nightly` | Nightly builds with additional variants | | [AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-jmf6jskjvnq7i) | — | Enterprise image | | Azure Marketplace | — | Enterprise image (coming soon) | | [Spice Cloud Platform](https://spice.ai/pricing) | — | Uses Enterprise image | | [Spice.ai Enterprise](https://spice.ai/pricing) | — | Uses Enterprise image | note Some variant distributions are only available in **nightly images** (data) or exclusively through the [Spice Cloud Platform and Spice.ai Enterprise](https://spice.ai/pricing) (NAS, CUDA, allocator variants). ## Supported Platforms and Hardware Requirements[​](#supported-platforms-and-hardware-requirements "Direct link to Supported Platforms and Hardware Requirements") | Platform | Architecture | Minimum CPU Features | Build Prerequisites | | -------- | ----------------------- | ---------------------------------------- | ------------------- | | Linux | x86\_64 | AVX2, FMA, BMI1/2, LZCNT, POPCNT | — | | Linux | aarch64 (arm64) | NEON, FP16 (FEAT\_FP16), FHM (FEAT\_FHM) | `clang`, `lld` | | macOS | aarch64 (Apple Silicon) | Native (build host) | — | | Windows | x86\_64 (MSVC) | — | MSVC toolchain | note Windows support is CLI (`spice`) only. The runtime daemon (`spiced`) is not supported on Windows natively — use WSL instead. ## Distribution Availability[​](#distribution-availability "Direct link to Distribution Availability") | Distribution / Variant | Image Tag | Open Source | Spice Cloud | Enterprise | | ---------------------- | ------------------------------------- | ---------------- | ----------- | ---------- | | Default (Data + AI) | `latest` | ✅ | ✅ | ✅ | | Data-only | `latest-data` | Nightly only | ✅ | ✅ | | NFS connector | — | Local build only | ❌ | ✅ | | Metal (macOS) | — | Local build only | ✅ | ✅ | | CUDA (Linux) | `latest-cuda` | Local build only | ✅ | ✅ | | Allocator variants | `latest-{jemalloc,mimalloc,sysalloc}` | Local build only | ✅ | ✅ | | ODBC connector | — | Local build only | ✅ | ✅ | ## Default Distribution[​](#default-distribution "Direct link to Default Distribution") The default distribution includes all features including AI/ML model support. This is the recommended distribution for most users. **Included Features:** * All standard data connectors (PostgreSQL, MySQL, DuckDB, SQLite, ClickHouse, etc.) * Embedded data accelerators (Spice Cayenne, DuckDB, SQLite) * AI/ML model inference (LLMs, embeddings) * Search capabilities (Vector and BM-25 Full-Text-Search) * Default memory allocator (snmalloc) note The PostgreSQL data accelerator is only available in nightly builds. The PostgreSQL data connector is included in all distributions. **Installation:** ``` curl https://install.spiceai.org | /bin/bash ``` **Docker:** ``` docker pull ghcr.io/spiceai/spiceai:latest # or docker pull spiceai/spiceai:latest ``` ## Data Distribution[​](#data-distribution "Direct link to Data Distribution") The data distribution excludes AI/ML model support, resulting in a smaller binary size and reduced attack surface. Use this when data federation and acceleration capabilities are needed without AI features. note **Open Source:** Available in nightly builds only. **[Cloud Platform & Enterprise](https://spice.ai/pricing):** Production-ready data distribution available. **Included Features:** * All data connectors * All data accelerators * Default memory allocator (snmalloc) **Excluded Features:** * AI/ML model inference * LLM support * Embedding models **Docker (Nightly):** ``` docker pull ghcr.io/spiceai/spiceai-nightly:latest-data ``` **Local Build:** ``` make install-data-only ``` ## GPU-Accelerated Distributions[​](#gpu-accelerated-distributions "Direct link to GPU-Accelerated Distributions") ### Metal (macOS)[​](#metal-macos "Direct link to Metal (macOS)") For macOS systems with Apple Silicon, the Metal distribution enables GPU-accelerated AI/ML inference. **Included Features:** * All default features * Metal GPU acceleration for model inference **Local Build:** ``` make install-metal ``` ### CUDA (Linux)[​](#cuda-linux "Direct link to CUDA (Linux)") For Linux systems with NVIDIA GPUs, CUDA distributions enable GPU-accelerated AI/ML inference. Multiple CUDA compute capability versions are available. note CUDA distributions are available with the [Spice Cloud Platform and Spice.ai Enterprise](https://spice.ai/pricing). Open source users can build locally for development and testing. **Included Features:** * All default features * CUDA GPU acceleration for model inference **Supported Compute Capabilities:** * 80 (A100, A30) * 86 (RTX 30xx, A40, A10) * 87 (Jetson Orin) * 89 (RTX 40xx, L40, L4) * 90 (H100, H200) **Local Build:** ``` CUDA_COMPUTE_CAP=89 make install-cuda ``` ## NFS Distribution[​](#nfs-distribution "Direct link to NFS Distribution") The NFS (Network File System) distribution adds support for the NFS data connector, enabling federated queries against data stored on NFS exports. note The SMB data connector is **not** part of this variant — `smb` is in the default `spiced` feature set, so SMB shares are queryable from the standard distribution. Only the NFS connector is feature-gated. note The NFS connector is available with [Spice.ai Enterprise](https://spice.ai/pricing). Open source users can build locally for development and testing. **Included Features:** * All default features * NFS data connector **Local Build:** ``` make install-nfs ``` ## Allocator Variants[​](#allocator-variants "Direct link to Allocator Variants") Different memory allocators can significantly impact performance depending on workload characteristics. note Allocator variants are available with the [Spice Cloud Platform and Spice.ai Enterprise](https://spice.ai/pricing). Open source users can build locally for development and testing. ### snmalloc (Default)[​](#snmalloc-default "Direct link to snmalloc (Default)") The default allocator, optimized for concurrent workloads. ### jemalloc[​](#jemalloc "Direct link to jemalloc") Alternative allocator that may perform better for certain memory allocation patterns. ### mimalloc[​](#mimalloc "Direct link to mimalloc") Microsoft's mimalloc allocator, designed for performance and security. ### System Allocator[​](#system-allocator "Direct link to System Allocator") Uses the system's default allocator (glibc malloc on Linux). ## Platform Support[​](#platform-support "Direct link to Platform Support") | Platform | Default | Data | NFS | Metal | CUDA | | ----------------------------- | ------- | --------------- | --------------- | ----- | ---------------- | | Linux x86\_64 | ✅ | Nightly | Enterprise only | ❌ | Cloud/Enterprise | | Linux aarch64 | ✅ | Nightly | Enterprise only | ❌ | ❌ | | macOS aarch64 (Apple Silicon) | ✅ | Nightly | Enterprise only | ✅ | ❌ | | Windows (WSL) | ✅ | Nightly | Enterprise only | ❌ | Cloud/Enterprise | | Windows (Native) | ❌ | Enterprise only | Enterprise only | ❌ | Enterprise only | note Native Windows support for the Spice runtime is available with the [Spice Cloud Platform and Spice.ai Enterprise](https://spice.ai/pricing). Open source users on Windows should use Windows Subsystem for Linux (WSL). ## Choosing a Distribution[​](#choosing-a-distribution "Direct link to Choosing a Distribution") | Use Case | Recommended Distribution | | --------------------------------------- | ------------------------ | | General purpose with AI capabilities | Default | | Data federation only, minimal footprint | Data (nightly) | | NFS exports | NFS | | macOS with GPU acceleration | Metal | | Linux with NVIDIA GPU | CUDA | | Memory allocation benchmarking | Allocator variants | ## Additional Connectors[​](#additional-connectors "Direct link to Additional Connectors") Some connectors require additional dependencies and are available with the [Spice Cloud Platform and Spice.ai Enterprise](https://spice.ai/pricing): * **ODBC** - Connect to any ODBC-compatible data source These can be built locally for development and testing: ``` make install-odbc ``` ## Platform-Specific Notes[​](#platform-specific-notes "Direct link to Platform-Specific Notes") ### Linux arm64[​](#linux-arm64 "Direct link to Linux arm64") * **FP16 (FEAT\_FP16)** is required because the `gemm` matrix multiplication library (used by the Candle ML framework) contains half-precision ARM inline assembly that requires the `fullfp16` CPU feature. This is supported on AWS Graviton2+, Ampere Altra, Apple M-series (via Linux VM), and most ARMv8.2-A+ processors. * **lld** is required as the linker because the spiced debug binary is large enough to exceed GNU ld's ±128 MiB branch range for `R_AARCH64_CALL26` relocations. lld automatically inserts range extension thunks. * Install prerequisites on Ubuntu/Debian: `sudo apt-get install -y clang lld` ### Linux x86\_64[​](#linux-x86_64 "Direct link to Linux x86_64") * Release builds target AVX2+ for optimized SIMD performance, covering Intel Haswell (2013+) and AMD Excavator (2015+) processors, including all current AWS x86\_64 instance families (C6/C7/C8). ## Building Custom Distributions[​](#building-custom-distributions "Direct link to Building Custom Distributions") Custom distributions with specific feature combinations can be built: ``` # Build with specific features (replaces the default feature set) SPICED_CUSTOM_FEATURES="duckdb,postgres,sqlite,models" make build # Build with non-default features added to defaults SPICED_NON_DEFAULT_FEATURES="odbc" make install ``` See the project [Makefile](https://github.com/spiceai/spiceai/blob/trunk/Makefile) for all available build targets and options. --- # Duration Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. `1.5h`) are accepted. Supported time units are: | Time Unit | Identifier | Calculation | | ------------- | ---------- | ----------- | | `Nanosecond` | ns | `1` | | `Millisecond` | ms | `1000000` | | `Second` | s | `1s` | | `Minute` | m | `60s` | | `Hour` | h | `60m` | | `Day` | d | `24h` | | `Week` | w | `7d` | Microseconds are also supported, spelled `Ms` (capital `M`, lowercase `s`). Month and year units are **not** supported — their length is ambiguous, so express longer intervals in weeks or days. ## Example[​](#example "Direct link to Example") ``` # 1 second 1s # 250 milliseconds 250ms # 3 minutes 3m # 1 hour 1h ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` # 2 days 2d # 1 week 1w # 90 minutes 1.5h # 30 seconds (no suffix) 30 ``` --- # File Formats File-based [data connectors](/docs/next/components/data-connectors) — including [`s3://`](/docs/next/components/data-connectors/s3), [`abfs://`](/docs/next/components/data-connectors/abfs), [`file://`](/docs/next/components/data-connectors/file), [`ftp://`](/docs/next/components/data-connectors/ftp), [`sftp://`](/docs/next/components/data-connectors/ftp), and others — support multiple structured and document file formats. This page details the format-specific parameters available for each. ## Common Parameters[​](#common-parameters "Direct link to Common Parameters") These parameters apply across multiple file formats. | Parameter | Type | Default | Description | | --------------------------- | ------- | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `file_format` | String | Inferred | Selects the file reader. If omitted, format is inferred from the file extension. See [Supported Formats](#supported-formats). | | `file_extension` | String | Derived | Overrides the file extension filter used when listing files. Defaults to the extension matching the resolved format. Compound extensions like `.csv.gz` set both format and compression at once. | | `schema_infer_max_records` | Integer | `1000` | Maximum number of records scanned to infer the schema. | | `file_compression_type` | String | Inferred | File-level compression for CSV, TSV, and JSON files. Valid values: `GZIP`, `BZIP2`, `XZ`, `ZSTD`, `UNCOMPRESSED`. If unset, Spice infers compression from compound file extensions such as `.gz`, `.bz2`, `.xz`, and `.zst` (see [Compressed File Extensions](#compressed-file-extensions)). | | `hive_partitioning_enabled` | Boolean | `false` | Enables Hive-style partition discovery from directory structure. | ## Supported Formats[​](#supported-formats "Direct link to Supported Formats") The `file_format` parameter accepts these values: | Value | Reader | Default Extension | Notes | | --------- | ------------------- | ----------------- | ---------------------------------------------------------------- | | `parquet` | Apache Parquet | `.parquet` | | | `vortex` | Vortex | `.vortex` | Columnar format. Not available on Windows. | | `csv` | CSV | `.csv` | Uses `csv_*` parameters. | | `tsv` | TSV (tab-delimited) | `.tsv` | Uses `tsv_*` parameters. Delimiter is tab. | | `json` | JSON | `.json` | Auto-detects format. Uses `json_format` to control parsing mode. | | `jsonl` | JSON Lines | `.jsonl` | Line-delimited JSON. | When `file_format` is omitted, Spice infers the format from the dataset path extension. If the extension does not match one of the values above, a configuration error is returned. ## Compressed File Extensions[​](#compressed-file-extensions "Direct link to Compressed File Extensions") For CSV, TSV, JSON, and JSONL files, Spice auto-detects both the file format and the compression codec from compound file extensions. Files matching one of the recognized format extensions followed by a recognized compression suffix are read transparently — no `file_compression_type` parameter is required. Recognized compression suffixes: `.gz` (GZIP), `.bz2` (BZIP2), `.xz` (XZ), `.zst` (ZSTD). | Path / extension | Format inferred | Compression inferred | | ------------------ | --------------- | -------------------- | | `data.csv.gz` | `csv` | `GZIP` | | `events.jsonl.zst` | `jsonl` | `ZSTD` | | `report.tsv.xz` | `tsv` | `XZ` | | `data.json.bz2` | `json` | `BZIP2` | | `data.csv` | `csv` | `UNCOMPRESSED` | The `.ndjson` and `.ldjson` extensions are also accepted as JSON Lines aliases — `events.ndjson.gz` is read as JSONL with GZIP compression. Auto-detection works the same way when reading a single file or listing a directory: object-store listings are filtered by the compound extension, and the same compression codec is applied to every file matched. To override auto-detection — for example, to force a specific compression codec — set `file_compression_type` explicitly. Explicit values always take precedence over the inferred value. ``` datasets: - from: s3://my-bucket/exports/ name: exports params: file_format: csv # Listing only includes matching files file_extension: .csv.gz # Reads gzipped CSVs in the directory ``` Auto-detection applies to listing data connectors (S3, ABFS, GCS, HTTP/HTTPS, FTP, SFTP, SMB, NFS, file) and to the HTTPS connector's auto-format detection. Parquet handles compression internally and is not affected by this setting — see [Parquet](#parquet) for the codecs Parquet supports natively. ## Parquet[​](#parquet "Direct link to Parquet") Spice reads any Parquet file regardless of the compression codec or data encoding. Supported compression codecs: * [`UNCOMPRESSED`](https://parquet.apache.org/docs/file-format/data-pages/compression/#uncompressed) * [`SNAPPY`](https://parquet.apache.org/docs/file-format/data-pages/compression/#snappy) * [`GZIP`](https://parquet.apache.org/docs/file-format/data-pages/compression/#gzip) * [`LZO`](https://parquet.apache.org/docs/file-format/data-pages/compression/#lzo) * [`BROTLI`](https://parquet.apache.org/docs/file-format/data-pages/compression/#brotli) * [`LZ4`](https://parquet.apache.org/docs/file-format/data-pages/compression/#lz4) (deprecated in favor of `LZ4_RAW`) * [`LZ4_RAW`](https://parquet.apache.org/docs/file-format/data-pages/compression/#lz4_raw) * [`ZSTD`](https://parquet.apache.org/docs/file-format/data-pages/compression/#zstd) Supported data encodings: * [`PLAIN`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#plain-plain--0) * [`PLAIN_DICTIONARY` / `RLE_DICTIONARY`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#dictionary-encoding-plain_dictionary--2-and-rle_dictionary--8) * [`RLE`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#run-length-encoding--bit-packing-hybrid-rle--3) * [`BIT_PACKED`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#bit-packed-deprecated-bit_packed--4) (deprecated in favor of `RLE`) * [`DELTA_BINARY_PACKED`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#delta-binary-packing-delta_binary_packed--5) * [`DELTA_LENGTH_BYTE_ARRAY`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#delta-length-byte-array-delta_length_byte_array--6) * [`DELTA_BYTE_ARRAY`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#delta-strings-delta_byte_array--7) * [`BYTE_STREAM_SPLIT`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#byte-stream-split-byte_stream_split--9) ## Vortex[​](#vortex "Direct link to Vortex") [Vortex](https://github.com/vortex-data/vortex) is a columnar file format read everywhere Parquet is, including the S3, ABFS, GCS, file, and HTTPS connectors. Set `file_format: vortex` or use a `.vortex` file extension (auto-detection also resolves `.vortex` files when `file_format` is omitted or set to `auto`). Vortex has no format-specific parameters. It is not available on Windows builds. ## CSV[​](#csv "Direct link to CSV") ### Parameters[​](#csv "Direct link to Parameters") | Parameter | Type | Default | Description | | ------------------------------ | ------- | ------- | ----------------------------------------------------------------------------------------------------- | | `csv_has_header` | Boolean | `true` | Whether the first row contains column headers. | | `csv_quote` | Char | `"` | Character used to quote fields containing special characters. | | `csv_escape` | Char | *none* | Character used to escape special characters within a field. | | `csv_delimiter` | Char | `,` | Character used to separate fields. | | `csv_schema_infer_max_records` | Integer | `1000` | **Deprecated.** Use `schema_infer_max_records` instead. Maximum records scanned for schema inference. | ## TSV[​](#tsv "Direct link to TSV") TSV (tab-separated values) is a first-class format. Set `file_format: tsv` or use a `.tsv` file extension. The delimiter is always tab and cannot be changed. ### Parameters[​](#parameters "Direct link to Parameters") | Parameter | Type | Default | Description | | ------------------------------ | ------- | ------- | ----------------------------------------------------------------------------------------------------- | | `tsv_has_header` | Boolean | `true` | Whether the first row contains column headers. | | `tsv_quote` | Char | `"` | Character used to quote fields containing special characters. | | `tsv_escape` | Char | *none* | Character used to escape special characters within a field. | | `tsv_schema_infer_max_records` | Integer | `1000` | **Deprecated.** Use `schema_infer_max_records` instead. Maximum records scanned for schema inference. | ## JSON[​](#json "Direct link to JSON") Set `file_format: json` for JSON files. Use the `json_format` parameter to select the parsing mode. ### Parsing Modes[​](#json-parsing-modes "Direct link to Parsing Modes") The `json_format` parameter controls how JSON content is interpreted. | Value | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `auto` | Default. Auto-detects the format by inspecting content. Detects SODA responses, JSON arrays, single objects, and line-delimited JSON. | | `json` | Auto-detects array vs line-delimited JSON by peeking at the first byte, but does not perform SODA auto-detection. Used implicitly when `file_format: json` is set explicitly. | | `jsonl`, `ndjson`, `ldjson` | Line-delimited JSON. Each line contains one JSON value. | | `array` | The file contains a single top-level JSON array. Each element becomes a row. | | `object` | The file contains a single JSON object, producing one row. | | `soda`, `socrata` | [Socrata Open Data API (SODA)](https://dev.socrata.com/docs/endpoints.html) format. Schema is derived from `meta.view.columns` in the response. Cannot be combined with `json_pointer`. | When `file_format` is omitted and the file extension is `.json`, the default parsing mode is `auto`, which includes SODA auto-detection. When `file_format: json` is set explicitly, the default mode is `json`, which skips SODA auto-detection. ### Parameters[​](#parameters-1 "Direct link to Parameters") | Parameter | Type | Default | Description | | --------------- | ------- | ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `json_format` | String | `auto` | Parsing mode. See [Parsing Modes](#json-parsing-modes). | | `json_pointer` | String | *none* | Extracts a sub-value from the document before parsing. Alias: `json_path`. Cannot be used with `soda` format. | | `flatten_json` | Boolean | `false` | When `true`, nested JSON objects are flattened with `.` as a separator (e.g., `address.city`). | | `soda_metadata` | String | `disabled` | When `enabled`, includes Socrata metadata columns (`:sid`, `:id`, `:position`, `:created_at`, `:created_meta`, `:updated_at`, `:updated_meta`, `:meta`) in the output. Only applies when parsing SODA format data. | Setting `file_format: jsonl` uses the DataFusion JSON Lines reader directly, without `json_format`, `flatten_json`, or `json_pointer` support. ### Examples[​](#examples "Direct link to Examples") Extract a nested value from a JSON document using `json_pointer`: ``` datasets: - from: s3://my-bucket/data/ name: events params: file_format: json json_pointer: /results/events ``` Read a SODA response with metadata columns included: ``` curl -sL "https://data.ct.gov/api/views/kf98-j89e/rows.json?accessType=DOWNLOAD" -o house_price_index.json ``` ``` datasets: - from: file:house_price_index.json name: house_price_index params: json_format: soda soda_metadata: enabled ``` ``` sql> select * from house_price_index limit 5; +--------------------+--------------------------------------+-----------+-------------+---------------+-------------+---------------+-------+---------------------+---------+ | :sid | :id | :position | :created_at | :created_meta | :updated_at | :updated_meta | :meta | observation_date | ctsthpi | | varchar | varchar | int64 | int64 | varchar | int64 | varchar |varchar| timestamp[s] | float64 | +--------------------+--------------------------------------+-----------+-------------+---------------+-------------+---------------+-------+---------------------+---------+ | row-r4ag~gfrd~dqcz | 00000000-0000-0000-6A52-5730E0309BF7 | 0 | 1768216213 | | 1768216213 | | { } | 1975-01-01T00:00:00 | 62.9 | | row-65s5_stm6-jbjc | 00000000-0000-0000-3B25-D45DF23837A7 | 0 | 1768216213 | | 1768216213 | | { } | 1975-04-01T00:00:00 | 62.94 | | row-buhc_mzb7.95pa | 00000000-0000-0000-A414-989C0238E96F | 0 | 1768216213 | | 1768216213 | | { } | 1975-07-01T00:00:00 | 61.93 | | row-3fp4~bx38-mwgi | 00000000-0000-0000-58CE-D4AF76C589B0 | 0 | 1768216213 | | 1768216213 | | { } | 1975-10-01T00:00:00 | 61.85 | | row-khut~7dd3-vi9e | 00000000-0000-0000-A274-646F4876CC81 | 0 | 1768216213 | | 1768216213 | | { } | 1976-01-01T00:00:00 | 64.83 | +--------------------+--------------------------------------+-----------+-------------+---------------+-------------+---------------+-------+---------------------+---------+ Time: 0.0028895 seconds. 5 rows. ``` --- # Managing Memory Usage Effective memory management is essential for maintaining optimal performance and stability in Spice deployments. This guide outlines recommendations and best practices for managing memory usage across different [Data Accelerators](/docs/next/components/data-accelerators). ## How Spice Uses Memory[​](#how-spice-uses-memory "Direct link to How Spice Uses Memory") Spice's memory footprint has two parts, and they behave differently: ``` Total Memory = Baseline + Working Set ``` * **Baseline** — memory the runtime needs to operate at all: process overhead, per-dataset accelerator caches, results caches, task history, serialization buffers, and allocator arenas. It is driven by **how the deployment is configured** — dataset count above all — and by how much traffic it has served, rather than by how much data the datasets hold. It does not shrink when the data does. * **Working Set** — memory that scales with the work: query execution, refreshes, and concurrency. This is the part bounded by [`runtime.query.memory_limit`](#memory-limit-configuration). Two refinements matter when reading a memory graph or sizing from a measurement: * **The baseline is mostly bounded buffers that grow toward their bounds with use, not memory claimed up front.** The SQLite metastore's page cache, gRPC and HTTP serialization buffers, and the accelerator caches all have ceilings, but they reach them only as traffic drives them there. The baseline is therefore the **converged** footprint under sustained load, not what the process holds at first startup — a freshly started runtime under-reports it, sometimes by a lot. * **The working set scales with the data each query touches, not just with dataset size.** `SELECT 1` costs almost nothing, `SELECT *` materializes a result set, and a multi-way join with a sort can exceed the size of the data it reads. Whatever a single query costs is then multiplied by how many run concurrently. Almost every surprising memory result comes from treating the total as a single quantity that scales with data volume. It does not. Sizing rules expressed as a multiple of dataset size — including the ones below — describe the **working set**, and apply on top of the baseline: ``` Total Memory ≈ Baseline + (multiple × dataset size) ``` At production data volumes the working set dominates and the baseline is a rounding error. At small data volumes the reverse is true, which is why a deployment holding a few hundred megabytes of data can still need several gigabytes of RAM, and why cutting the data volume by 100x does not let you cut memory by 100x. See [Sizing a Non-Production Environment](#sizing-a-non-production-environment). Because both terms understate themselves early — buffers have not yet grown, and a light query mix has not yet hit its peak — sizing from a short observation biases low in both directions at once. Size from a sustained run, and see [Validating a Memory Configuration](#validating-a-memory-configuration). ## General Memory Recommendations[​](#general-memory-recommendations "Direct link to General Memory Recommendations") Memory requirements vary based on workload characteristics, dataset sizes, query complexity, and refresh modes. | Workload Type | Minimum RAM | Notes | | ---------------------------------------- | ----------------- | -------------------------------------------------------------- | | Typical workloads | 8 GB | Suitable for most development and small production deployments | | Large datasets (`refresh_mode: full`) | 2.5x dataset size | Requires memory for both old and new tables during refresh | | Large datasets (`refresh_mode: append`) | 1.5x dataset size | Memory for incremental data only | | Large datasets (`refresh_mode: changes`) | 1.5x dataset size | Depends on CDC event volume and frequency | The multiples are working-set figures. Add them to the baseline rather than using them as the total, and treat the 8 GB row as a floor that applies regardless of how small the data is. Memory requirements can be reduced by using file-based acceleration with [DuckDB](/docs/next/components/data-accelerators/duckdb), [SQLite](/docs/next/components/data-accelerators/sqlite), [Turso](/docs/next/components/data-accelerators/turso), or [Spice Cayenne](/docs/next/components/data-accelerators/cayenne), which store data on disk and support spilling. Datasets 10 GB or larger For any dataset of **10 GB or larger**, [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) is recommended over [DuckDB](/docs/next/components/data-accelerators/duckdb), because of DuckDB's memory requirements. Cayenne typically needs **one-third to one-half** the memory of the DuckDB accelerator for the same dataset. ### What Makes Up the Baseline[​](#what-makes-up-the-baseline "Direct link to What Makes Up the Baseline") The baseline is the sum of a fixed process cost and a per-dataset cost. The per-dataset allocations are sized as a fraction of total memory but **clamped to a floor**, so they stop shrinking once the environment is small: | Allocation | Applies to | Default size | | ------------------------------ | -------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | Runtime process overhead | Always | \~500 MB | | Cayenne segment cache | Each Cayenne-accelerated dataset | 1/128 of total memory, clamped to 256 MB–1 GB | | Cayenne PK keyset cache | Each Cayenne CDC/upsert dataset | 1/32 of total memory, clamped to 256 MiB–8 GiB, and additionally bounded by a process-wide ceiling across all datasets | | CDC coalesce buffer | Each CDC dataset | 128 MiB (`cdc_max_coalesced_bytes`) | | Results caches | Runtime | 128 MiB each for SQL, search, and embedding results | | Task history | Runtime (enabled by default) | No byte cap — bounded by time instead; scales with task rate × `retention_period` (8h) | | Metastore and protocol buffers | Runtime | Bounded, but reached only under sustained traffic | The last two rows are the ones that make a short measurement misleading. **Task history is an in-memory accelerated table**, so its footprint is a product of how many tasks the deployment completes and how long records are kept, rather than a fixed allocation — a high-throughput deployment accumulates far more than a quiet one at the same configuration. Setting `captured_plan` or `captured_output` increases the size of every record substantially. ``` runtime: task_history: retention_period: 1h # default 8h; the main lever on its footprint captured_plan: none # capturing plans materially increases per-record size ``` Prefer shortening `retention_period` to disabling task history outright — it backs `runtime.task_history`, which is the primary tool for diagnosing the very problems this page describes. Disable it only if the deployment is memory-critical and its diagnostics are served another way. Two consequences follow: * **The baseline scales with dataset count.** Ten Cayenne-accelerated datasets reserve at least 2.5 GB of segment cache between them whether each dataset holds a gigabyte or a megabyte. * **The baseline is a larger share of a smaller container.** The per-dataset floors are absolute, so halving the container's memory does not halve the baseline — it raises the baseline's share of the total. The runtime accounts for this when deriving its own defaults: the query memory limit is reduced by the per-dataset reservations so the pools plus the caches fit within the memory the process may use. Explicitly configured limits are honored as-is and are **not** reduced, so a hand-set `runtime.query.memory_limit` is the one case where the total can be over-committed. Per-dataset caps sized in isolation still add up, and the aggregate is what the kernel makes its OOM decision on. #### Accelerated catalogs[​](#accelerated-catalogs "Direct link to Accelerated catalogs") An [accelerated catalog](/docs/next/reference/spicepod/catalogs#acceleration) creates a Cayenne table for every table it discovers, so each of those tables carries the same per-table allocations as one configured by hand. Two properties follow from the catalog being configuration rather than an enumeration: * **The projected reservation counts one table's worth per catalog, not one per discovered table.** How many tables a catalog will accelerate is not knowable before the connector connects, so the projection the runtime derives its defaults from is a **floor** — it leaves the query pool larger than the discovered tables warrant. Size a catalog-backed deployment from the table count you expect it to discover rather than from the projection. * **A catalog always reaches the in-memory CDC tier.** `changes` is the only refresh mode catalog acceleration accepts, so a Spicepod whose only Cayenne acceleration is a catalog gets the reduced query pool and the compaction carve described in [How the Runtime Partitions Memory](#how-the-runtime-partitions-memory), exactly as a CDC-accelerated dataset does. ## Sizing a Non-Production Environment[​](#sizing-a-non-production-environment "Direct link to Sizing a Non-Production Environment") A common approach to staging is to take the production configuration and scale every number down by the ratio of data volume — a tenth of the data, a tenth of the memory. Because the baseline does not scale, this does not produce a smaller model of production. It produces a **different regime**, in which the baseline dominates, the working set is squeezed into whatever is left, and behavior no longer predicts what production will do. Consider a Spicepod with ten accelerated datasets, deployed at two container sizes. The per-dataset caches sit at their floor in both, so the baseline is the same number in each row: | Container memory | Baseline (overhead + 10 datasets) | Baseline share | Left for the working set | | ---------------- | --------------------------------- | -------------- | ------------------------ | | 4 GiB | \~3 GB | \~76% | \~1 GB | | 32 GiB | \~3 GB | \~9% | \~29 GB | The same Spicepod is memory-bound in the first row and comfortable in the second. Nothing about the first row's behavior — its headroom, its spill frequency, its latency profile — carries over to the second. **Guidance for lower environments:** * **Scale the working set, not the baseline.** Reduce data volume and concurrency; keep the dataset count and the accelerator configuration the same. * **Do not copy limit ratios between environments.** Whatever fraction of the container `runtime.query.memory_limit` occupies in a small environment, the correct production value is a *larger* fraction, not the same one — the baseline it has to leave room for stays roughly constant while the container grows around it. Derive each environment's limit from its own baseline rather than carrying a percentage across. * **Reduce dataset count if you must run small.** Removing datasets lowers the baseline; shrinking their contents does not. * **Treat a configuration tuned to fit an undersized environment as disposable.** Lowering per-dataset caches and buffers to fit a container that is smaller than the runtime's floors will make it run, but it tunes for a regime you will not operate in, and the values do not transfer to production. Costs of tuning below the defaults The default cache and buffer sizes are the sizes the rest of the system is tuned against. Lowering them to fit a constrained environment is supported, but expect higher and less predictable query latency, more frequent spilling, and performance characteristics that will not match production. Prefer validating at a representative size — see [Validating a Memory Configuration](#validating-a-memory-configuration). ## Accelerator-Specific Memory Management[​](#accelerator-specific-memory-management "Direct link to Accelerator-Specific Memory Management") Different acceleration engines have distinct memory characteristics and tuning options. ### Arrow (In-Memory)[​](#arrow-in-memory "Direct link to Arrow (In-Memory)") The default Arrow accelerator stores all data in memory uncompressed. Datasets must fit entirely in available RAM. * Data is stored uncompressed in Apache Arrow format * No configuration options for memory limits * Best for smaller datasets requiring maximum query speed * Consider switching to file-based accelerators for datasets exceeding available memory **Hash Index Memory (Experimental, v1.11.0-rc.2+):** When using the optional [hash index](/docs/next/features/data-acceleration/hash-index), additional memory is required: | Component | Memory per Row | | ------------ | -------------- | | Hash slot | 16 bytes | | Bloom filter | \~1.25 bytes | | **Total** | \~17.25 bytes | For a 10 million row dataset with hash index enabled, expect \~165 MB additional memory overhead. ### Spice Cayenne[​](#spice-cayenne "Direct link to Spice Cayenne") [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) stores data on disk using the [Vortex](https://github.com/vortex-data/vortex) columnar format, with configurable caches for metadata and frequently accessed data segments. The caches can be configured to reside either in memory or on disk, which impacts overall memory behavior. Spice Cayenne is DataFusion query-native, meaning all query execution adheres to the `runtime.query.memory_limit` setting. When query memory is exhausted, DataFusion spills intermediate results to disk. This architecture provides predictable memory usage while maintaining high query performance. **Memory Configuration Parameters:** | Parameter | Scope | Default | Description | | -------------------------- | --------------------- | ------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cayenne_footer_cache_mb` | `runtime.params` | `50` (unset) | Size of the engine-global in-memory Vortex footer cache in megabytes, shared by all Cayenne datasets. Larger values improve query performance for repeated scans by caching file metadata. Optional; when unset, DataFusion's default file-metadata-cache limit of 50 MB applies. | | `cayenne_segment_cache_mb` | `acceleration.params` | auto | Per-dataset size of the in-memory Vortex segment cache in megabytes. Caches decompressed data segments for improved query performance. When unset, derived as 1/128 of total memory, clamped to 256 MB–1 GB. | **Memory Usage Guidelines:** * Base memory: \~500 MB for runtime overhead * Footer cache: unset by default (DataFusion's 50 MB file-metadata-cache limit applies); increase for datasets with many files * Segment cache: auto-derived per dataset with a 256 MB floor; increase for workloads with repeated scans on the same data * Query execution memory: Depends on query complexity and concurrency The segment cache is allocated **per Cayenne-accelerated dataset**, and its floor does not scale down with the container. Count it once per dataset when sizing — see [What Makes Up the Baseline](#what-makes-up-the-baseline). **Example Configuration:** ``` runtime: params: # Engine-global footer cache (shared by all Cayenne datasets) cayenne_footer_cache_mb: 256 datasets: - from: s3://my-bucket/large-dataset/ name: large_dataset acceleration: engine: cayenne mode: file params: # Per-dataset segment cache cayenne_segment_cache_mb: 512 ``` ### DuckDB[​](#duckdb "Direct link to DuckDB") [DuckDB](/docs/next/components/data-accelerators/duckdb) manages memory through streaming execution, intermediate spilling, and buffer management. Left to itself, each DuckDB instance sizes its own limit at roughly 80% of host memory. **Memory Configuration Parameters:** | Parameter | Default | Description | | --------------------- | ---------------------------------------------- | -------------------------------------- | | `duckdb_memory_limit` | A coordinated share of the query memory budget | Maximum memory for the DuckDB instance | When `duckdb_memory_limit` is not set, Spice does not leave the instance on DuckDB's own \~80%-of-host-RAM default. At startup it computes a [coordinated memory budget](/docs/next/components/data-accelerators/duckdb#coordinated-memory-budget) across the query pool and every DuckDB instance so their combined ceilings fit within the memory the process can use, capping each un-limited instance at an equal share and reducing the query pool to match. Explicit `duckdb_memory_limit` and `runtime.query.memory_limit` values are always honored as-is. **Memory Usage Guidelines:** * For datasets of **10 GB or larger**, prefer [Spice Cayenne](#spice-cayenne), which typically needs one-third to one-half the memory of DuckDB for the same dataset * Set `duckdb_memory_limit` to control memory per DuckDB instance, rather than relying on the automatic split * DuckDB indexes do not support spilling and may consume significant memory * Allocate at least 30% additional container/machine memory for the runtime process **Example Configuration:** ``` datasets: - from: postgres:analytics.orders name: orders acceleration: engine: duckdb mode: file params: duckdb_memory_limit: 4GB ``` ### SQLite[​](#sqlite "Direct link to SQLite") [SQLite](/docs/next/components/data-accelerators/sqlite) is lightweight and efficient for smaller datasets but does not support intermediate spilling. Datasets must fit in memory or use application-level paging. ## Refresh Modes and Memory Implications[​](#refresh-modes-and-memory-implications "Direct link to Refresh Modes and Memory Implications") Refresh modes affect memory usage as follows: | Refresh Mode | Memory Behavior | | ------------ | ------------------------------------------------------------------------------------------------------------------------------------------ | | `full` | Temporarily loads data into a new table before replacing the existing table (atomic swap). Requires memory for both tables simultaneously. | | `append` | Incrementally inserts or upserts data, using memory only for the incremental batch. | | `changes` | Applies CDC events incrementally. Memory usage depends on event volume and frequency. | | `caching` | Caches query results on disk. Memory usage is limited to active queries and cache metadata. | ## DataFusion Query Memory Management[​](#datafusion-query-memory-management "Direct link to DataFusion Query Memory Management") Spice uses DataFusion as its query execution engine. By default, Spice limits query engine memory to **90% of the memory the process may use** (its cgroup memory limit when one binds — a container, a `systemd` unit's `MemoryMax=`, a capped parent slice, or a Kubernetes pod cgroup — otherwise total system memory; reduced to **70%** when Cayenne's in-memory CDC tier is reachable, to leave headroom for Cayenne's compaction memory pool and that tier). Either base is then reduced by the per-dataset cache reservations and floored at 50% — see [How the Runtime Partitions Memory](#how-the-runtime-partitions-memory). This can be tuned through the `runtime.query.memory_limit` configuration. ### Memory Limit Configuration[​](#memory-limit-configuration "Direct link to Memory Limit Configuration") The `runtime.query.memory_limit` parameter defines the maximum memory available for query execution. If not specified, it defaults to 90% of the memory the process may use — its cgroup memory limit when one binds, otherwise total system memory (70% when Cayenne's in-memory CDC tier is reachable), less the per-dataset cache reservations and floored at 50%. Once the memory limit is reached, supported query operations spill data to disk. ``` runtime: query: memory_limit: 4GiB temp_directory: /tmp/spice # Directory for spill files ``` Spice uses [Apache DataFusion](https://datafusion.apache.org/) as its query execution engine, which provides vectorized, multi-threaded query execution with automatic memory management. DataFusion's [GreedyMemoryPool](https://docs.rs/datafusion/latest/datafusion/execution/memory_pool/struct.GreedyMemoryPool.html) allows memory reservations on a first-come, first-served basis up to the configured limit, improving throughput for high-concurrency queries with many partitions. ### How the Runtime Partitions Memory[​](#how-the-runtime-partitions-memory "Direct link to How the Runtime Partitions Memory") When the limit is left unset, the runtime divides the memory the process may use into a partition that is designed to sum to 100%. Which partition applies depends on whether Cayenne's in-memory CDC tier is reachable — which the runtime decides from every Cayenne acceleration the Spicepod declares, on a dataset, a view, or a [catalog](/docs/next/reference/spicepod/catalogs#acceleration): | Slice | Standard deployment | Cayenne CDC active | | --------------------------------- | ------------------- | ------------------------------------------------- | | Query memory pool | 90% | 70% base, less the compaction carve | | Compaction memory pool | — | Carved from the query pool (20% of it by default) | | In-memory CDC tier | — | 20%, clamped to between 1/32 and 1/5 of memory | | Headroom for off-pool allocations | 10% | 10% | The headroom slice is what covers the per-dataset caches, encode buffers, and allocator overhead described below. When the projected per-dataset reservations exceed that headroom, the excess is carved out of the query pool so the total still fits — which is why the effective query pool on a dataset-heavy Spicepod is smaller than the headline percentage. The query pool is never reduced below **50%** of available memory; when that floor binds and the caches still do not fit, the runtime warns at startup that the configuration is unfittable (see [Check the Startup Budget Warnings First](#check-the-startup-budget-warnings-first)). ### Tuning the Memory Limit Safely[​](#tuning-the-memory-limit-safely "Direct link to Tuning the Memory Limit Safely") `runtime.query.memory_limit` is bounded on both sides when Cayenne CDC ingestion is active, and both failure modes are counterintuitive: * **Setting it too low does not necessarily reduce resident memory.** The in-memory CDC tier is sized from the memory left over after the pools, so lowering the query pool hands that memory to the tier, which may float up to a quarter of total memory to reclaim it. The memory moves rather than being released. On a deployment being throttled to reduce its footprint, this is the most common reason the graph does not come down. * **Setting it too high starves CDC ingestion.** A greedy explicit limit can squeeze the tier's remainder below its floor, at which point the tier degrades to a refuse-all state and ingestion falls back to its disk-based path. The runtime warns when this happens. Both are consequences of the partition being coordinated. The practical guidance is to **leave `runtime.query.memory_limit` unset and size the container instead**, which lets the runtime derive every slice together. Set it explicitly only when a co-resident accelerator manages its own pool, and treat the value as one input to a partition rather than as a ceiling on the process. ### What the Memory Limit Does Not Cover[​](#what-the-memory-limit-does-not-cover "Direct link to What the Memory Limit Does Not Cover") `runtime.query.memory_limit` bounds the query execution pool. It is not a ceiling on the process. Memory allocated outside that pool includes: * Per-dataset accelerator caches (see [What Makes Up the Baseline](#what-makes-up-the-baseline)) * Serialization and encode buffers for query results, Arrow IPC, and Flight responses * Embedded engine internals that manage their own memory, such as DuckDB's pool and SQLite's page cache * Results, search, and embedding caches * Allocator retention — pages the process has freed but not returned to the operating system The gap between the query limit and the container's memory limit is the headroom that absorbs all of the above. The defaults reserve 10% of the memory the process may use, or 30% when Cayenne acceleration is active. **That reservation is a percentage, but what it has to cover is largely fixed**, so it gets tighter as the container gets smaller: 10% of a 4 GiB container is roughly 400 MB of headroom for buffers and caches whose own floors do not shrink with it. This is the usual reason a container is OOM-killed despite having a memory limit configured. Two corollaries: * **Lowering `runtime.query.memory_limit` does not reduce off-pool memory.** It shrinks the part that was already bounded and was probably not the cause. If the process is killed while the pool gauges show plenty of unused reservation, the memory is off-pool and a lower limit will not recover it. * **Explicit limits are not reduced for you.** When the limit is left unset the runtime derives it with the per-dataset reservations already subtracted. Setting it explicitly opts out of that arithmetic, so an explicit value must leave room for the baseline itself. When an explicit limit is set alongside Cayenne acceleration, the runtime logs the projected off-pool reservation at startup (`Explicit query memory limit set; the projected per-table Cayenne cache reservation is OFF-pool and unaffected by this limit`) — check it against the container's memory limit. ### Spill-to-Disk[​](#spill-to-disk "Direct link to Spill-to-Disk") Operators such as Sort, Join, and GroupByHash spill intermediate results to disk when memory limits are exceeded, preventing out-of-memory errors. DataFusion writes spill files using the [Arrow IPC Stream format](https://arrow.apache.org/docs/format/Columnar.html#ipc-streaming-format). **Spill Compression:** The `runtime.query.spill_compression` parameter controls how spill files are compressed: | Value | Description | | ---------------- | ---------------------------------------------- | | `zstd` (default) | High compression ratio, reduces disk usage | | `lz4_frame` | Faster compression/decompression, larger files | | `uncompressed` | No compression overhead, largest files | ``` runtime: query: memory_limit: 4GiB spill_compression: lz4_frame ``` ### Spill Limitations[​](#spill-limitations "Direct link to Spill Limitations") DataFusion supports spilling for several operators, but the following operations do not currently support spilling: * HashJoin ([tracking issue](https://github.com/apache/datafusion/issues/12952)) * ExternalSorterMerge * RepartitionMerge Queries using these operators that exceed memory limits may fail. Monitor query patterns and allocate sufficient memory for workloads that rely on these operators. ### How a Memory Refusal Surfaces[​](#how-a-memory-refusal-surfaces "Direct link to How a Memory Refusal Surfaces") When the query memory pool cannot satisfy a reservation, the query is refused rather than the process being allowed to grow past `runtime.query.memory_limit`. The refusal is reported as a capacity condition, not a client error, so retry middleware and load balancers can treat it as retriable: | Surface | Reported as | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | HTTP (`/v1/sql`) | `503 Service Unavailable`, with the pool's message (including its top memory consumers) as the response body. Earlier releases answered this condition with `400 Bad Request`. | | Flight and Flight SQL | gRPC status `RESOURCE_EXHAUSTED` | | `query_failures` metric | `err_code="ResourcesExhausted"` | | Runtime log | Logged at `warn` level (`Query refused, out of memory: …`), so a runtime refusing queries is visible at the default `INFO` verbosity | A sustained rate of these means the deployment needs more memory, fewer concurrent queries, or fewer partitions — not that the client's SQL is malformed. Note that `/health` is served by a separate Tokio runtime and stays green while queries are being refused, so alert on `query_failures{err_code="ResourcesExhausted"}` rather than relying on health checks. ### Bounding Peak Memory with Concurrency[​](#bounding-peak-memory-with-concurrency "Direct link to Bounding Peak Memory with Concurrency") Concurrency is the multiplier on the working set: each executing plan holds its own reservations, so peak memory scales with how many run at once. Two settings control this, and **both derive their defaults from the CPU entitlement rather than from memory**: | Setting | Default | Effect on memory | | -------------------------------------- | ------------------------------- | -------------------------------------------------------- | | `runtime.query.max_concurrent_queries` | 4 × the CPU entitlement's cores | Caps how many plans execute at once; excess queries wait | | `runtime.query.target_partitions` | The CPU entitlement's cores | Caps per-plan fan-out and its per-partition reservations | Because the defaults follow CPU, a pod that is CPU-rich and memory-poor admits far more concurrent work than its memory can support. A pod sized for 8 cores admits 32 concurrent queries by default; one running [`runtime.cpu.cores: all`](/docs/next/reference/spicepod/runtime#runtimecpucores) on a large node admits proportionally more. Neither default consults `runtime.query.memory_limit`. ``` runtime: query: max_concurrent_queries: 8 # bound admission by memory, not by core count ``` Lowering `max_concurrent_queries` reduces peak **memory** and converts what would have been capacity refusals into queueing. The trade-off is that peak **query time** rises: excess queries now wait for admission, and that wait counts toward end-to-end duration. Measure total query duration rather than execution time alone when tuning it. It is still usually the better first move than lowering `runtime.query.memory_limit`, which shrinks the pool available to each query without reducing how many run — and which, on a CDC deployment, may not reduce resident memory at all (see [Tuning the Memory Limit Safely](#tuning-the-memory-limit-safely)). When load testing, note that **peak concurrency, not average throughput, sets the peak memory** — a test that averages the target rate but never bursts will under-report the peak. ### Client-Side Resiliency[​](#client-side-resiliency "Direct link to Client-Side Resiliency") No memory configuration produces a zero-failure system, and sizing for one is not the goal. In any distributed deployment a query can fail for reasons unrelated to how Spice is tuned — an instance is evicted or preempted, a node is reclaimed, a rolling upgrade replaces a pod, a network path drops. Clients need an error budget and a retry path regardless of memory sizing. Recommended client behavior: * **Retry a failed query against the cluster before treating it as an outage.** With more than one instance behind a load balancer, a retry usually lands on a healthy instance. Capacity refusals (`503` / `RESOURCE_EXHAUSTED`) and connection failures are both retriable. * **Bound the retry.** One or two attempts with a short backoff is normally enough; unbounded retries turn a capacity problem into an outage by adding load to a runtime that is already refusing work. * **Set an explicit client timeout** so a slow query cannot consume the caller's own request budget. * **Fall back to the underlying data source last, not first.** Where a fallback path to the source of truth exists, place it after the in-cluster retry, so a single unhealthy instance does not divert all traffic away from the accelerated path. * **Consider a circuit breaker for sustained failure.** Retries handle isolated failures; they make a sustained one worse. After a threshold of consecutive failures, stop sending traffic for a cooldown, then probe with a fraction of it and restore full load only once the probes succeed. This is usually best implemented at the load balancer or service mesh rather than in each client, so the decision is shared across callers instead of being re-learned by each one. Distinguish the two cases when alerting: an isolated failure that a retry resolves is expected operational noise, while a sustained rate of capacity refusals is a sizing signal — see the metric guidance above. ## Predicate Pushdown and Memory Reduction[​](#predicate-pushdown-and-memory-reduction "Direct link to Predicate Pushdown and Memory Reduction") Predicate pushdown reduces memory consumption by filtering data early in the query execution pipeline. Rather than reading all data and filtering afterward, Spice pushes filter predicates to the data source, reducing the volume of data materialized in memory. ### How Pushdown Reduces Memory[​](#how-pushdown-reduces-memory "Direct link to How Pushdown Reduces Memory") | Stage | Without Pushdown | With Pushdown | | ---------------- | ---------------- | ------------------ | | Read from source | All rows | Matching rows only | | Decompress | Full row groups | Pruned row groups | | Materialize | Entire dataset | Filtered subset | | Process | Full scan | Reduced scan | For a query selecting 1% of rows from a 100 GB dataset, pushdown can reduce peak memory from tens of gigabytes to hundreds of megabytes. ### Pushdown Techniques by Format[​](#pushdown-techniques-by-format "Direct link to Pushdown Techniques by Format") **Parquet and Parquet-backed sources (Iceberg, Delta Lake):** * **Row group pruning**: Skips entire row groups (typically 128 MB) based on min/max statistics * **Page Index**: Skips individual pages (typically 8 KB) within row groups * **Bloom filters**: Skips row groups for equality predicates * **Late materialization**: Filters during decoding, reducing columns materialized **Vortex (Spice Cayenne):** * **Segment pruning**: Skips segments based on per-segment min/max statistics * **Compute push-down**: Evaluates predicates on compressed data, reducing decompression overhead ### Configuration for Memory Efficiency[​](#configuration-for-memory-efficiency "Direct link to Configuration for Memory Efficiency") For memory-constrained environments, set an appropriate memory limit and use file-based acceleration: ``` runtime: query: memory_limit: 2GiB datasets: - from: s3://bucket/data/ name: filtered_data acceleration: engine: cayenne # Segment pruning + compute push-down mode: file ``` Sorting data by frequently filtered columns maximizes pushdown effectiveness. When data is sorted, entire segments or row groups have non-overlapping value ranges, enabling efficient pruning. ### Memory Impact of Data Layout[​](#memory-impact-of-data-layout "Direct link to Memory Impact of Data Layout") | Data Layout | Pushdown Effectiveness | Memory Impact | | ----------------------- | ---------------------- | ------------------ | | Sorted by filter column | Excellent | Minimal data read | | Clustered (Z-ordered) | Good | Moderate data read | | Random | Limited | Most data read | For time-series data, sort by timestamp. For multi-tenant data, consider sorting by tenant\_id or clustering by (tenant\_id, timestamp). ## Embedded Data Accelerator Comparison[​](#embedded-data-accelerator-comparison "Direct link to Embedded Data Accelerator Comparison") | Accelerator | Storage | Query Memory Control | Memory Spilling | Best For | | ------------- | -------------- | ---------------------------- | --------------- | --------------------------------------------------------------------- | | Arrow | Memory only | `runtime.query.memory_limit` | Yes | Small datasets, maximum speed | | Spice Cayenne | Disk (Vortex) | `runtime.query.memory_limit` | Yes | Datasets 10 GB and above, scalable analytics; lowest memory footprint | | DuckDB | Memory or Disk | `duckdb_memory_limit` | Yes | Datasets under 10 GB, complex queries | | SQLite | Memory or Disk | None | No | Small-medium datasets, simple queries | Spice Cayenne and Arrow both use DataFusion as the query execution engine and share the same `runtime.query.memory_limit` configuration. DuckDB manages its own memory pool separately via the `duckdb_memory_limit` parameter — though when that parameter is unset, the runtime sizes the two together rather than independently (see [Coordinated memory budget](/docs/next/components/data-accelerators/duckdb#coordinated-memory-budget)). ## Memory Allocators[​](#memory-allocators "Direct link to Memory Allocators") The Spice runtime supports multiple memory allocators that affect how the process allocates and frees memory at the system level. The choice of allocator can significantly impact performance depending on workload characteristics such as concurrency, allocation size distribution, and fragmentation behavior. | Allocator | Description | Best For | | --------------------- | ------------------------------------------------------------ | --------------------------------------------- | | snmalloc (default) | Optimized for concurrent workloads with low fragmentation | General-purpose, high-concurrency deployments | | jemalloc | Mature allocator with strong profiling support | Workloads with varied allocation patterns | | mimalloc | Microsoft's allocator, designed for performance and security | Performance-sensitive deployments | | System (glibc malloc) | Uses the OS default allocator | Compatibility testing, debugging | The default distribution uses snmalloc. Alternative allocators are available as separate [distribution variants](/docs/next/reference/distributions#allocator-variants), each published as a distinct Docker image tag. note Allocator variants are available with the [Spice Cloud Platform and Spice.ai Enterprise](https://spice.ai/pricing). Open source users can build locally for development and testing. The memory allocator operates independently from the query memory management described above. `runtime.query.memory_limit` controls DataFusion's query execution memory pool, while the allocator determines how the runtime process itself requests and releases memory from the operating system. ## Results Cache Memory[​](#results-cache-memory "Direct link to Results Cache Memory") Spice maintains in-memory caches for SQL query results, search results, and embeddings. These caches consume memory in addition to accelerator and query execution memory. ### Cache Memory Configuration[​](#cache-memory-configuration "Direct link to Cache Memory Configuration") | Cache Type | Default Max Size | Description | | ---------------- | ---------------- | ------------------------------------------ | | `sql_results` | 128 MiB | Caches SQL query results | | `search_results` | 128 MiB | Caches vector and full-text search results | | `embeddings` | 128 MiB | Caches embedding model responses | **Example Configuration:** ``` runtime: caching: sql_results: enabled: true max_size: 512MiB item_ttl: 5m eviction_policy: tiny_lfu search_results: enabled: true max_size: 256MiB item_ttl: 1m embeddings: enabled: true max_size: 256MiB item_ttl: 10m ``` ### Cache Eviction Policies[​](#cache-eviction-policies "Direct link to Cache Eviction Policies") | Policy | Description | Performance | | --------------- | ------------------------ | ------------------------------------------- | | `lru` (default) | Least Recently Used | Good general-purpose hit rates | | `tiny_lfu` | TinyLFU admission policy | Higher hit rates for skewed access patterns | TinyLFU maintains frequency information to admit only items likely to be accessed again, resulting in higher hit rates for workloads with varying query frequency patterns. ### Cache Memory Impact[​](#cache-memory-impact "Direct link to Cache Memory Impact") When sizing memory, account for cache allocations: ``` Total Memory = Runtime Overhead + (Per-Dataset Accelerator Caches × dataset count) + Query Memory Limit + Cache Memory ``` **Example calculation** — four Cayenne-accelerated datasets, none using CDC: * Runtime overhead: 500 MB * Footer cache: 50 MB (engine-global, shared by all Cayenne datasets) * Segment caches: 1 GB (256 MB × 4 datasets — per-dataset, not shared) * Results caches: 1 GB (512 MB SQL + 256 MB search + 256 MB embeddings) * Query memory limit: 4 GB * **Total: \~6.5 GB minimum** Everything above the query memory limit in that list is baseline, and it grows with the dataset count. Adding four more datasets adds roughly another 1 GB before a single query runs. Datasets using `refresh_mode: changes` add a PK keyset cache and a coalesce buffer on top of this — see [What Makes Up the Baseline](#what-makes-up-the-baseline). This is a planning floor, not a prediction A calculation like this sums the allocations that can be named and sized in advance, so it **under-estimates actual usage** and should never be used as the container's memory limit. It omits the buffers that grow with traffic, task history's throughput-dependent growth, allocator retention and fragmentation, and any transient peak above the steady state. The query memory limit in particular is a *ceiling* on the pool, not a reservation — actual usage moves within it. Treat the result as the minimum below which the deployment certainly will not fit, add headroom, and then replace the estimate with a measurement from a sustained representative run. See [Validating a Memory Configuration](#validating-a-memory-configuration). ### Cache Performance Considerations[​](#cache-performance-considerations "Direct link to Cache Performance Considerations") * Larger `max_size` values improve hit rates but consume more memory * Shorter `item_ttl` values reduce memory usage but may decrease hit rates * Use `stale_while_revalidate_ttl` to serve stale results while refreshing in the background * Monitor cache hit rates via [observability metrics](/docs/next/features/observability) to tune configuration See [Caching](/docs/next/features/caching) for complete cache configuration options. ## Kubernetes Memory Configuration[​](#kubernetes-memory-configuration "Direct link to Kubernetes Memory Configuration") Configure memory requests and limits in Kubernetes pod specifications based on expected workload: ``` apiVersion: v1 kind: Pod metadata: name: spice-pod spec: containers: - name: spice image: spiceai/spiceai:latest resources: requests: memory: '8Gi' cpu: '4' limits: memory: '16Gi' # Set higher than request for burst capacity # Do not set CPU limits - can cause throttling ``` **Recommendations:** * Set memory requests to at least 1.3x the configured `runtime.query.memory_limit` plus accelerator cache sizes, counting the per-dataset caches once **per dataset** * Set memory limits higher than requests to handle temporary spikes * Avoid setting CPU limits, as they can cause [throttling](https://home.robusta.dev/blog/stop-using-cpu-limits) even when CPU is available * Monitor actual usage with observability tools and adjust accordingly * Prefer leaving `runtime.query.memory_limit` unset so the runtime derives it from the pod's cgroup limit with the per-dataset reservations subtracted; set it explicitly only when co-locating accelerators that manage their own pools The pod's memory limit is what the kernel enforces, so it must cover the baseline and the query pool together. A pod whose `runtime.query.memory_limit` was chosen without accounting for the per-dataset baseline will be OOM-killed while the query pool still reports unused capacity. ## Monitoring and Profiling[​](#monitoring-and-profiling "Direct link to Monitoring and Profiling") Use observability tools to monitor and profile memory usage regularly. Spice exposes metrics for: * Query execution memory usage * Accelerator cache hit rates * Data refresh memory consumption When [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) acceleration is configured, three gauges are sampled every 2 seconds and can be read together to reconcile the memory budgets against actual process memory: | Metric | Description | | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | | `query_memory_pool_used_bytes` | Live bytes reserved in the query memory pool (`runtime.query.memory_limit`), excluding the in-memory CDC tier's mirror account. | | `cayenne_compaction_memory_pool_used_bytes` | Live bytes reserved in the dedicated Cayenne compaction memory pool. | | `process_resident_memory_bytes` | Resident set size of the `spiced` process — what the kernel's OOM decision is made on. | The pool gauges describe what the memory accounting believes is reserved; `process_resident_memory_bytes` describes what the process actually holds. A large and growing gap between them is off-pool memory that `runtime.query.memory_limit` does not bound — lowering the query memory limit will not shrink it. See [Observability](/docs/next/features/observability) for configuration details. ### Reading a Memory Graph[​](#reading-a-memory-graph "Direct link to Reading a Memory Graph") Resident memory that climbs and then stays high is the expected shape, not a leak. Caches fill toward their configured ceilings and are not evicted until they are full, buffers are retained rather than reallocated because reallocation is expensive, and memory allocators keep freed pages mapped instead of returning them to the operating system. **Memory returning to its startup value is not a goal, and a flat high plateau is not evidence of a problem.** The last of those is a documented property of the allocators themselves, not a Spice-specific behavior. [jemalloc](https://jemalloc.net/jemalloc.3.html) retains freed pages and returns them only on a decay schedule (`dirty_decay_ms`, `muzzy_decay_ms`). glibc's allocator returns memory to the OS only when the free block at the top of the heap exceeds [`M_TRIM_THRESHOLD`](https://man7.org/linux/man-pages/man3/mallopt.3.html), so freed memory in the middle of the heap stays resident by design. Resident set size is consequently an upper bound on what the process is actively using — and, per the [cgroup v2 documentation](https://docs.kernel.org/admin-guide/cgroup-v2.html#memory-interface-files), it is nonetheless what the kernel charges against `memory.max` when deciding whether to OOM-kill. What matters is *where* it plateaus and *whether* it plateaus: | Observed shape | Interpretation | | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | | Rises, then flat well below the container limit | Healthy. Caches are warm and the deployment has headroom. | | Rises, then flat just under the container limit | Sized too tightly. Working normally, but with no margin for a spike — the next burst is an OOM kill. | | Rises steadily under constant load and never flattens | Investigate. Compare against the pool gauges to determine whether the growth is inside or outside the pools. | | Flat, then a sharp step on a specific operation | A refresh, compaction, or large query. Size for the peak, not the steady state. | To tell a plateau from slow growth, hold the load constant and run long enough for the steady state to establish itself: caches full, at least one full refresh cycle per dataset, and any background compaction having run. Reaching that point can take hours to days depending on refresh cadence, and reading the graph before then will show a rise that has not yet finished. When growth does not flatten, the two gauges separate the causes: growth in `query_memory_pool_used_bytes` is query work inside the bounded pool, while growth in `process_resident_memory_bytes` with the pool gauges flat is off-pool memory, which `runtime.query.memory_limit` does not govern. ## Validating a Memory Configuration[​](#validating-a-memory-configuration "Direct link to Validating a Memory Configuration") Memory behavior cannot be extrapolated from an environment that is not representative. Because the baseline is fixed and the working set is not, a deployment that is healthy on small data tells you very little about the same deployment on production data — and a deployment that is unhealthy on small data may be revealing only that the environment is below the runtime's floors. Start with the startup check below — it is free and immediate — then validate under load. ### Check the Startup Budget Warnings First[​](#check-the-startup-budget-warnings-first "Direct link to Check the Startup Budget Warnings First") The runtime computes its memory partition at startup, before any data is loaded, and warns when the configuration cannot fit. These warnings are emitted at `warn` level, so they appear at the default `INFO` verbosity, and they are the cheapest possible sizing check: **an unfittable configuration is detectable in the first seconds of a run rather than after days of load.** | Startup warning (leading phrase) | Meaning | Action | | ---------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- | | `Cayenne CDC cache reservation exceeds what the query pool can yield` | The per-dataset caches do not fit beside the query pool even after the pool has been reduced to its floor. The startup commitment already exceeds available memory. | Add memory, reduce dataset count, or lower per-dataset cache parameters. Expect resident memory above the budgets until then. | | `Cayenne in-memory CDC ingestion has limited memory available` | The query pool, compaction pool, and any co-resident DuckDB reservations leave little room for CDC ingestion. | Lower `runtime.query.memory_limit` or per-dataset `duckdb_memory_limit`. Ingestion will spill to disk more often until then. | | `Detected potential memory over-commit from DuckDB accelerators` | Un-limited DuckDB instances would each default to \~80% of host RAM. The runtime capped them automatically. | Set explicit `duckdb_memory_limit` values rather than relying on the automatic split. | | `The explicit DuckDB accelerator memory limits plus the query memory limit exceed the coordinated memory budget` | Explicitly configured limits over-commit memory. These are honored as-is and **not** reduced. | Lower the explicit `duckdb_memory_limit` and/or `runtime.query.memory_limit` values. | The first warning is the important one for a constrained environment: it fires precisely in the situation described in [Sizing a Non-Production Environment](#sizing-a-non-production-environment) — an environment small enough that the per-dataset floors no longer fit beside a working query pool. Treat it as a statement that the environment is undersized for the dataset count, not as a tuning prompt. Also check the derived limit itself. At `DEBUG` verbosity the runtime logs the arithmetic it used (`No query memory limit specified; ...`), including the reservation it subtracted, which is the fastest way to see what the baseline actually costs for a given Spicepod. ### Representative Load Testing[​](#representative-load-testing "Direct link to Representative Load Testing") A load test predicts production only if it reproduces all four of these. Missing any one of them makes the result unreliable: | Property | Why it matters | | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | | **Data volume** | Determines the working set and the scan sizes. Generate synthetic data to production scale if real data cannot be copied into the environment. | | **Concurrency** | Simultaneous in-flight queries multiply per-query memory. Peak concurrency, not average throughput, sets the peak. | | **Duration** | Caches, refresh cycles, and compaction only reach steady state after a sustained run. Short runs measure the warm-up, not the plateau. | | **Query diversity** | Varied predicates and filter values exercise real scan and cache behavior. A test that replays one query measures the results cache. | Generating production-scale data into a dedicated source is usually less effort than it appears, and is the piece most often skipped — it is also the piece that makes the other three meaningful. ### Shadow or Canary Deployment[​](#shadow-or-canary-deployment "Direct link to Shadow or Canary Deployment") Where a load test is impractical, mirror production traffic to a canary instance that serves no user-facing responses. This gives genuine query distribution, concurrency, and data volume without exposing users to the result. Confirm that the shadow path cannot write to production systems or double-count downstream side effects before enabling it. ### What Not to Do[​](#what-not-to-do "Direct link to What Not to Do") * **Do not size production by scaling a small environment's numbers up proportionally.** The baseline does not scale, so the ratios do not transfer. See [Sizing a Non-Production Environment](#sizing-a-non-production-environment). * **Do not tune caches and buffers downward to make an undersized environment stop failing, and then ship that configuration.** It validates a regime you will not run in, at the cost of latency and predictability. * **Do not draw conclusions from a run that has not reached steady state**, in either direction. ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") **Spice Documentation:** * [Performance Tuning](/docs/next/reference/performance-tuning) - Comprehensive guide to optimizing Spice performance * [Data Accelerators](/docs/next/components/data-accelerators) - Accelerator configuration reference * [Runtime Configuration](/docs/next/reference/spicepod/runtime) - Runtime parameter reference **External References:** * [Apache DataFusion](https://datafusion.apache.org/) - Query execution engine used by Spice * [DataFusion Memory Usage](https://datafusion.apache.org/user-guide/configs.html#runtime-configuration-settings) - DataFusion runtime memory configuration * [DataFusion Tuning Guide](https://datafusion.apache.org/user-guide/configs.html#tuning-guide) - Memory-limited query optimization * [DuckDB Memory Management](https://duckdb.org/docs/operations_manual/limits.html) - DuckDB memory limits documentation * [jemalloc](https://jemalloc.net/jemalloc.3.html) - Page retention and the `dirty_decay_ms` / `muzzy_decay_ms` decay schedule * [`mallopt`](https://man7.org/linux/man-pages/man3/mallopt.3.html) - glibc's `M_TRIM_THRESHOLD` and when freed memory is returned to the OS * [cgroup v2 memory interface](https://docs.kernel.org/admin-guide/cgroup-v2.html#memory-interface-files) - What `memory.max` accounts for and how the OOM decision is made --- # Models Grade Report This document presents the evaluation report for various Large-Language-Models (LLMs) graded by Spice AI. The models are assessed based on their basic capabilities, quality of tool calls, and accuracy of output when integrated with Spice. For more details on how model grades are evaluated in Spice, refer to the [model grading criteria](https://github.com/spiceai/spiceai/blob/f6039123028209e20469b342791fa85d52b7771e/docs/criteria/models/grading). | Model | Spice Grade | Model Provider | Context Window
Max Output Tokens | Chat Completion | Response Format
(Structued Outputs) | Tools | Recursive
Tool Calling | Reasoning | Streaming | Model Release Date | Spice Version | | ----------------------------------------------- | ----------- | -------------- | ------------------------------------- | --------------- | ---------------------------------------- | ----- | --------------------------- | --------- | --------- | ------------------ | ------------- | | `o3-mini-2025-01-31 (Reasoning effort: high)` | **A** | `openai` | 200k tokens
100k tokens | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 2025-01-31 | v1.0.2 | | `o3-mini-2025-01-31 (Reasoning effort: medium)` | **B** | `openai` | 200k tokens
100k tokens | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 2025-01-31 | v1.0.2 | | `o3-mini-2025-01-31 (Reasoning effort: low)` | **C** | `openai` | 200k tokens
100k tokens | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 2025-01-31 | v1.0.2 | | `o1-2024-12-17 (Reasoning effort: high)` | **C** | `openai` | 200k tokens
100k tokens | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | 2024-12-17 | v1.0.2 | | `o1-2024-12-17 (Reasoning effort: medium)` | **C** | `openai` | 200k tokens
100k tokens | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | 2024-12-17 | v1.0.2 | | `o1-2024-12-17 (Reasoning effort: low)` | **C** | `openai` | 200k tokens
100k tokens | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | 2024-12-17 | v1.0.2 | | `gpt-4o-2024-08-06` | **B** | `openai` | 128k tokens
16384 tokens | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | 2024-08-06 | v1.0.2 | | `claude-3-5-sonnet-20241022` | **C** | `anthropic` | 200k tokens
8192 tokens | ✅ | ❌ | ✅ | ✅ | ❌ | ✅ | 2024-10-22 | v1.0.2 | | `grok-4.3` | Ungraded | `xai` | 1M tokens
− | ✅ | − | − | − | ✅ | − | 2026 | v2.1 | | `deepseek-ai/DeepSeek-R1-Distill-Llama-8B` | Ungraded | `huggingface` | − | ✅ | − | − | − | ✅ | − | Not Available | v1.0.2 | | `meta-llama/Llama-3.2-3B-Instruct` | Ungraded | `huggingface` | − | ✅ | − | − | − | ❌ | − | Not Available | v1.0.2 | | `deepseek-reasoner` | Ungraded | `openai` | 64k tokens
8k tokens | ✅ | − | − | − | ✅ | ✅ | 2025/01/20 | v1.0.4 | --- # Performance Tuning This guide provides recommendations for optimizing Spice performance across data acceleration, query execution, caching, and resource allocation. ## Accelerator Selection[​](#accelerator-selection "Direct link to Accelerator Selection") Choose the appropriate [Data Accelerator](/docs/next/components/data-accelerators) based on dataset characteristics and query patterns. | Scenario | Recommended Accelerator | Key Configuration | | ---------------------------------------- | -------------------------- | ------------------------------------------------------------------------------- | | Small datasets (under 1 GB), low latency | `arrow` | Default in-memory | | Small datasets (1-10 GB), complex SQL | `duckdb` with `mode: file` | Set `duckdb_memory_limit` | | Datasets 10 GB and above (up to 1+ TB) | `cayenne` | Tune cache parameters; needs 1/3 to 1/2 the memory of `duckdb` | | Write-heavy workloads | `cayenne` with `zstd` | Set `cayenne_compression_strategy: zstd` | | Point lookups, large datasets | `cayenne` | Vortex provides [100x faster random access](https://bench.vortex.dev) | | Point lookups, small-medium datasets | `arrow` with hash index | Set a `primary_key` to auto-enable the hash index (experimental, v1.11.0-rc.2+) | | Point lookups with explicit indexes | `duckdb` or `sqlite` | Configure indexes | ## Spice Cayenne Performance Optimization[​](#spice-cayenne-performance-optimization "Direct link to Spice Cayenne Performance Optimization") [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) uses the [Vortex](https://github.com/vortex-data/vortex) columnar format for high-performance analytics on large datasets. ### Point Lookups and Random Access[​](#point-lookups-and-random-access "Direct link to Point Lookups and Random Access") Vortex provides [100x faster random access](https://bench.vortex.dev) compared to Apache Parquet through: * **Segment statistics**: Per-segment min/max/null\_count for predicate pushdown (zone-map equivalent) * **Fast random access encodings**: [FSST](https://www.vldb.org/pvldb/vol13/p2649-boncz.pdf), [FastLanes](https://www.vldb.org/pvldb/vol16/p2132-afroozeh.pdf), and [ALP](https://ir.cwi.nl/pub/33334/33334.pdf) support O(1) or near-O(1) random access * **Compute push-down**: Filter execution on compressed data without full decompression * **Array statistics**: `is_sorted`, `is_constant`, `min`, `max` for query optimization For point lookups on large datasets, Spice Cayenne often matches or exceeds the performance of traditional B-tree indexes while consuming no additional memory for index structures. ### Cache Configuration[​](#cache-configuration "Direct link to Cache Configuration") Spice Cayenne maintains two in-memory caches that significantly impact query performance. The footer cache is engine-global and set under `runtime.params`; the segment cache is configured per dataset under `acceleration.params`: ``` runtime: params: cayenne_footer_cache_mb: 256 # Engine-global; increase for many files datasets: - from: s3://bucket/data/ name: analytics acceleration: engine: cayenne mode: file params: cayenne_segment_cache_mb: 512 # Per-dataset; increase for hot data patterns ``` **Footer Cache Sizing:** The footer cache stores file metadata. Size based on file count: * 1-10 KB per file * Default: unset — when omitted, DataFusion's 50 MB file-metadata-cache limit applies * Increase for datasets with more files **Segment Cache Sizing:** The segment cache stores decompressed data. Size based on working set: * Estimate the volume of frequently accessed data * Cache hits avoid decompression overhead * Monitor cache hit rates via [observability metrics](/docs/next/features/observability) ### Compression Strategy[​](#compression-strategy "Direct link to Compression Strategy") | Strategy | Read Performance | Write Performance | Compression Ratio | | --------------------- | ---------------- | ----------------- | ----------------- | | `btrblocks` (default) | Fastest | Moderate | Higher | | `zstd` | Moderate | Faster | High | Choose `btrblocks` for read-heavy analytics workloads. Use `zstd` only when size on disk is the primary concern—setting `zstd` trades query performance for reduced storage size. ## DuckDB Performance Optimization[​](#duckdb-performance-optimization "Direct link to DuckDB Performance Optimization") [DuckDB](/docs/next/components/data-accelerators/duckdb) provides mature SQL support with sophisticated query optimization. Datasets 10 GB or larger For any dataset of **10 GB or larger**, [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) is recommended over DuckDB, because of DuckDB's memory requirements. Cayenne typically needs **one-third to one-half** the memory of the DuckDB accelerator for the same dataset. See [Spice Cayenne Performance Optimization](#spice-cayenne-performance-optimization). ### Memory Configuration[​](#memory-configuration "Direct link to Memory Configuration") ``` datasets: - from: postgres:schema.table name: orders acceleration: engine: duckdb mode: file params: duckdb_memory_limit: 4GB duckdb_file: /data/orders.duckdb ``` **Guidelines:** * Set `duckdb_memory_limit` to control memory per instance * DuckDB defaults to 80% of system memory per instance * Reserve 30% of container memory for the runtime * Multiple datasets using the same `duckdb_file` share a connection pool ### Connection Pool Tuning[​](#connection-pool-tuning "Direct link to Connection Pool Tuning") ``` acceleration: engine: duckdb params: connection_pool_size: 20 # Default: 10 ``` Increase `connection_pool_size` for high-concurrency workloads. Each connection consumes memory. ### Index Configuration[​](#index-configuration "Direct link to Index Configuration") DuckDB supports [ART (Adaptive Radix Tree) indexes](https://duckdb.org/docs/stable/guides/performance/indexing) for faster point lookups: ``` datasets: - from: postgres:schema.orders name: orders acceleration: engine: duckdb mode: file indexes: order_id: enabled '(customer_id, created_at)': enabled ``` Indexes consume memory and [do not spill to disk](https://duckdb.org/docs/stable/guides/performance/indexing#indexes-and-memory). Creating an index requires the entire dataset to be loaded into memory. Monitor memory usage when adding indexes. For more details on ART index performance, see the [ART paper](https://db.in.tum.de/~leis/papers/ART.pdf). ### Zone-Maps and Sorted Data[​](#zone-maps-and-sorted-data "Direct link to Zone-Maps and Sorted Data") DuckDB automatically creates [zone-maps](https://duckdb.org/docs/stable/guides/performance/indexing#zonemaps) (min/max statistics) for each row group, enabling efficient predicate pushdown. In practice, zone-maps on sorted data often outperform ART indexes for range and equality queries while consuming no additional memory. **Why Zone-Maps Outperform Indexes:** * Zero memory overhead (statistics stored with data) * No index maintenance during writes * Automatic predicate pushdown during scans * Effective when data is sorted by query filter columns **Optimization Pattern: Sorted Views** Accelerate a view with `ORDER BY` to create sorted physical data, then set `duckdb_preserve_insertion_order: true` to maintain sort order: ``` datasets: - from: iceberg:catalog/namespace/table name: raw_data_by_arrival time_column: processed_time acceleration: enabled: true engine: duckdb mode: file refresh_mode: append primary_key: id on_conflict: id: upsert params: duckdb_memory_limit: 12GiB duckdb_preserve_insertion_order: false # Raw data doesn't need order views: - name: data_sorted sql: | SELECT id, account_id, pool_id, value, created_at FROM raw_data_by_arrival WHERE __deleted = 'false' ORDER BY account_id, pool_id acceleration: enabled: true engine: duckdb mode: file refresh_check_interval: 210s params: duckdb_file: data_sorted.duckdb duckdb_memory_limit: 6GiB duckdb_preserve_insertion_order: true # Maintains ORDER BY sort ``` **Key Configuration:** | Parameter | Value | Purpose | | --------------------------------- | ----------------------- | --------------------------------------------- | | `duckdb_preserve_insertion_order` | `true` on sorted view | Maintains physical sort order from `ORDER BY` | | `duckdb_preserve_insertion_order` | `false` on source table | Faster writes without order guarantees | | Separate `duckdb_file` | Per view | Isolates sorted data from source tables | Queries filtering on `account_id` or `(account_id, pool_id)` benefit from zone-map pruning, skipping entire row groups that don't match the filter predicates. ### Aggregate Pushdown[​](#aggregate-pushdown "Direct link to Aggregate Pushdown") Enable aggregate pushdown for improved performance on supported aggregate queries: ``` acceleration: engine: duckdb params: optimizer_duckdb_aggregate_pushdown: enabled ``` Requires `query_federation` to be disabled. Supports `count`, `sum`, `avg`, `min`, and `max` functions. ## DataFusion Query Engine[​](#datafusion-query-engine "Direct link to DataFusion Query Engine") Spice uses [Apache DataFusion](https://datafusion.apache.org/) as its query execution engine for Arrow and Spice Cayenne accelerators. DataFusion provides vectorized, multi-threaded query execution with automatic memory management and spilling. ### Query Parallelism[​](#query-parallelism "Direct link to Query Parallelism") DataFusion automatically parallelizes queries across available CPU cores. By default, the number of partitions equals the runtime's [CPU entitlement](/docs/next/reference/spicepod/runtime#runtimecpu) in whole cores, providing maximum parallelism. Override it with `runtime.query.target_partitions`, or state the entitlement itself with `runtime.cpu.cores` to size partitions and every other CPU-derived pool together. DataFusion's [GreedyMemoryPool](https://docs.rs/datafusion/latest/datafusion/execution/memory_pool/struct.GreedyMemoryPool.html) allows memory reservations on a first-come, first-served basis up to the configured `memory_limit`. This approach improves throughput for high-concurrency queries with many partitions compared to dividing memory evenly. ### Join Algorithm Selection[​](#join-algorithm-selection "Direct link to Join Algorithm Selection") DataFusion supports multiple join algorithms and automatically selects the best one based on query statistics: | Algorithm | Memory Usage | Best For | | ---------------- | ------------ | ------------------------------------------------ | | Hash Join | Higher | Fast execution with sufficient memory (default) | | Sort-Merge Join | Lower | Memory-constrained environments, pre-sorted data | | Nested Loop Join | Variable | Cross joins, non-equi joins | DataFusion prefers hash joins by default for equi-joins. Hash joins do not currently support spilling, so memory-constrained environments may benefit from sort-merge joins for large datasets. ### Dynamic Filter Pushdown[​](#dynamic-filter-pushdown "Direct link to Dynamic Filter Pushdown") DataFusion pushes filters from operators (TopK, Join, Aggregate) into file scans to prune data early. This optimization is enabled by default and can skip entire row groups or files based on statistics. For example, a `SELECT * FROM t ORDER BY timestamp DESC LIMIT 10` query pushes timestamp filters down to file scans, pruning files that cannot contain top-10 candidates. ### Parquet Read Optimizations[​](#parquet-read-optimizations "Direct link to Parquet Read Optimizations") Spice uses DataFusion's Parquet reader, which applies several optimizations automatically when reading Parquet files from S3, file, Iceberg, and Delta Lake connectors: * **Row group pruning**: Skips entire row groups (typically 128 MB) based on min/max statistics in Parquet metadata * **Page Index filtering**: Uses page-level min/max statistics (typically 8 KB chunks) for finer-grained pruning * **Bloom filter evaluation**: Checks bloom filters for equality predicates when available in Parquet files * **Projection pushdown**: Reads only the columns referenced in the query These optimizations are applied automatically and require no configuration. The effectiveness depends on data layout—sorting data by frequently filtered columns maximizes row group pruning. ## File Format Filtering and Optimization[​](#file-format-filtering-and-optimization "Direct link to File Format Filtering and Optimization") Spice connects to various file formats (Parquet, Iceberg, Delta Lake) and uses DataFusion's query execution engine to push down predicates and prune data at the file, row group, and page level. Understanding these optimizations helps when designing data layouts for optimal query performance. ### Parquet[​](#parquet "Direct link to Parquet") Apache Parquet stores data in row groups with per-column statistics. DataFusion uses these statistics to skip row groups that cannot contain matching rows. **Row Group Pruning:** Each Parquet row group contains min/max statistics for each column. When a query includes a `WHERE` clause, DataFusion evaluates whether each row group could contain matching rows based on these statistics. Row groups that cannot match are skipped entirely. For example, with a predicate `WHERE timestamp > '2024-01-01'`, DataFusion skips row groups where the maximum timestamp is before 2024-01-01. **Page Index:** Parquet's [Page Index](https://parquet.apache.org/docs/file-format/data-pages/) provides finer-grained statistics at the page level within row groups. DataFusion uses the Page Index to skip individual pages, reducing I/O for selective queries. **Bloom Filters:** Parquet files can include bloom filters for membership testing. For equality predicates like `WHERE user_id = 'abc123'`, DataFusion checks the bloom filter before reading column data. If the bloom filter indicates the value is not present, the row group is skipped. **Late Materialization (Filter Pushdown):** DataFusion supports applying filters during Parquet decoding rather than after. This optimization, called late materialization, filters rows before materializing all columns, reducing memory usage for selective queries. ### Iceberg[​](#iceberg "Direct link to Iceberg") [Apache Iceberg](https://iceberg.apache.org/) provides hidden partitioning and multi-level metadata filtering that simplifies query optimization. **Hidden Partitioning:** Iceberg automatically derives partition values from source columns using transforms like `day(timestamp)`, `month(timestamp)`, or `bucket(user_id, 16)`. Queries filter on the source column directly (e.g., `WHERE timestamp > '2024-01-01'`), and Iceberg automatically prunes partitions without requiring users to specify partition columns in predicates. **Two-Level Metadata Filtering:** Iceberg uses a hierarchical metadata structure that enables filtering at multiple levels: 1. **Manifest list filtering**: The manifest list contains partition value ranges for each manifest file. Manifests that cannot contain matching rows are skipped entirely. 2. **Manifest filtering**: Each manifest file contains per-file column statistics (min/max, null counts, value counts). Data files that cannot contain matching rows are skipped. This two-level approach can eliminate entire groups of files before reading any data, providing significant performance benefits for large tables. **Column-Level Statistics:** Iceberg manifests store column-level statistics including: * `lower_bound` and `upper_bound` for min/max filtering * `null_count` for null handling optimization * `value_count` for cardinality estimation **Partition Evolution:** Iceberg supports changing partition schemes without rewriting existing data. Historical data retains its original partitioning while new data uses the updated scheme. Queries automatically account for both partition layouts. ### Delta Lake[​](#delta-lake "Direct link to Delta Lake") [Delta Lake](https://delta.io/) provides data skipping and Z-ordering for query optimization. **Data Skipping:** Delta Lake collects column statistics (min, max, null counts) during writes. The `delta.dataSkippingNumIndexedCols` table property controls how many columns have statistics collected (counted from the first column in the schema). Queries filter using these statistics to skip files that cannot contain matching rows. **Generated Columns:** Delta Lake supports generated columns that derive values from other columns. When partitioned by a generated column (e.g., `eventDate` generated from `CAST(eventTime AS DATE)`), queries filtering on the source column automatically benefit from partition pruning. **Z-Ordering:** [Z-ordering](https://docs.delta.io/latest/optimizations-oss.html#z-ordering-multi-dimensional-clustering) colocates related data in the same files by clustering on specified columns. After running `OPTIMIZE ... ZORDER BY (column)`, queries filtering on the Z-ordered columns benefit from improved data skipping. **Compaction:** Delta Lake's `OPTIMIZE` command compacts small files into larger ones, reducing the number of files to scan and improving query performance through better I/O patterns. ### Vortex (Spice Cayenne)[​](#vortex-spice-cayenne "Direct link to Vortex (Spice Cayenne)") Spice Cayenne uses [Vortex](https://github.com/vortex-data/vortex), which provides segment-level statistics and compute push-down on compressed data. **Segment Statistics:** Vortex's ChunkedLayout maintains per-segment statistics including `min`, `max`, `null_count`, `is_sorted`, and `is_constant` for each column. These statistics function similarly to DuckDB's zone-maps, enabling segment pruning during query execution. **Compute Push-Down:** Vortex supports executing filter operations directly on compressed data. For encodings like FSST (strings), FastLanes (integers), and ALP (floats), predicates can be evaluated without full decompression, reducing CPU and memory usage. **Encoding-Aware Optimization:** Vortex tracks encoding metadata that enables additional optimizations: * `is_sorted`: Enables binary search for point lookups * `is_constant`: Returns values immediately without scanning * Encoding-specific optimizations based on data characteristics See [Spice Cayenne Performance Optimization](#spice-cayenne-performance-optimization) for cache tuning and other Cayenne-specific settings. ### Performance Implications[​](#performance-implications "Direct link to Performance Implications") | Optimization | Parquet | Iceberg | Delta Lake | Vortex | | -------------------------- | ------- | ---------------- | ---------------- | ------------------ | | Row group/file pruning | ✅ | ✅ | ✅ | ✅ | | Page-level filtering | ✅ | ✅ (via Parquet) | ✅ (via Parquet) | ✅ (segment-level) | | Bloom filters | ✅ | ✅ (via Parquet) | ❌ | ❌ | | Hidden partitioning | ❌ | ✅ | ❌ | ❌ | | Manifest-level filtering | ❌ | ✅ | ❌ | ❌ | | Compute on compressed data | ❌ | ❌ | ❌ | ✅ | | Z-ordering | ❌ | ✅ | ✅ | ❌ | **Optimization Recommendations:** * **Sort data by filter columns**: Row group and segment statistics are most effective when data is sorted by commonly filtered columns * **Use appropriate file sizes**: Larger row groups (128 MB+) provide better compression but reduce pruning granularity * **Collect statistics on filter columns**: Ensure filter columns are within the statistics collection limit (e.g., `delta.dataSkippingNumIndexedCols`) * **Consider Z-ordering for multi-column filters**: When queries filter on multiple columns, Z-ordering colocates related data ## Query Memory Management[​](#query-memory-management "Direct link to Query Memory Management") Configure DataFusion query memory limits to prevent out-of-memory errors: ``` runtime: query: memory_limit: 8GiB temp_directory: /tmp/spice spill_compression: zstd ``` ### Memory Limit[​](#memory-limit "Direct link to Memory Limit") If not specified, `memory_limit` defaults to 90% of the memory the process may use — its cgroup memory limit when one binds (a container, a `systemd` unit's `MemoryMax=`, a capped parent slice, or a Kubernetes pod cgroup), otherwise total system memory; when a Cayenne acceleration can reach the in-memory CDC tier (`refresh_mode: changes` or `caching`, or `append` with `refresh_check_interval` ≤ 5m), the default is 70% instead, reserving headroom for Cayenne's compaction memory pool and in-memory CDC tier. A pod whose Cayenne accelerations only bulk-write — `refresh_mode: full`, or `append` on a slower cadence — cannot fill that tier and keeps the 90% base, reduced by the off-pool per-table scan caches Cayenne holds for each accelerated table. Datasets, views and [catalogs](/docs/next/reference/spicepod/catalogs#acceleration) are all classified this way; because `changes` is the only refresh mode catalog acceleration accepts, an accelerated catalog always reaches the tier, and it contributes one table's worth of off-pool reservation however many tables it goes on to discover (see [Accelerated catalogs](/docs/next/reference/memory#accelerated-catalogs)). Either way the derived limit is floored at 50% of the same memory. For deployments with co-located accelerators, set an explicit limit based on available memory: ``` runtime memory_limit = Total Memory - Accelerator Memory - OS/Runtime Overhead (30%) ``` Accelerator memory in that expression is **per dataset**, not per deployment: each accelerated dataset holds its own caches, sized as a fraction of total memory but clamped to a floor that does not shrink with the container. The limit therefore has to be derived for the dataset count, and the ratio that works in a small environment does not carry to a large one. See [How Spice Uses Memory](/docs/next/reference/memory#how-spice-uses-memory) and [Sizing a Non-Production Environment](/docs/next/reference/memory#sizing-a-non-production-environment). Note also that this limit bounds the query execution pool, not the process. Serialization buffers, accelerator caches, and allocator retention sit outside it, which is why lowering it does not always reduce resident memory — see [What the Memory Limit Does Not Cover](/docs/next/reference/memory#what-the-memory-limit-does-not-cover). When the goal is to reduce peak memory, prefer bounding concurrency over lowering this limit: | Intent | Setting | | ------------------------------------------------- | ---------------------------------------------------------------------------------------------------- | | Reduce peak memory under burst | `runtime.query.max_concurrent_queries` (defaults to 4× the CPU entitlement's cores, not from memory) | | Reduce per-plan fan-out and its reservations | `runtime.query.target_partitions` (defaults to the CPU entitlement's cores) | | Change how much memory a single query may reserve | `runtime.query.memory_limit` | Lowering `max_concurrent_queries` reduces peak **memory** usage directly, by bounding how many plans hold reservations at once. The trade-off is latency: excess queries wait for admission rather than being refused, so peak end-to-end query time rises with the time spent queueing. Tune it against the workload's tolerance for wait, and watch total query duration — not just execution time — when you change it. Lowering `memory_limit` shrinks the pool each query draws from without reducing how many run concurrently, and on a Cayenne CDC deployment it can leave resident memory unchanged because the in-memory CDC tier expands into the freed budget — see [Tuning the Memory Limit Safely](/docs/next/reference/memory#tuning-the-memory-limit-safely). ### Spill Compression[​](#spill-compression "Direct link to Spill Compression") | Compression | Disk Usage | CPU Overhead | | ---------------- | ---------- | ------------ | | `zstd` (default) | Lowest | Moderate | | `lz4_frame` | Medium | Lowest | | `uncompressed` | Highest | None | Choose `lz4_frame` when CPU is limited and disk is abundant. Choose `uncompressed` for debugging or when spill files are rare. DataFusion uses [Arrow IPC Stream format](https://arrow.apache.org/docs/format/Columnar.html#ipc-streaming-format) for spill files. ### Batch Processing[​](#batch-processing "Direct link to Batch Processing") DataFusion processes data in batches of 8192 rows by default. This batch size balances memory usage with vectorized execution efficiency. Larger batches improve CPU cache utilization and SIMD operations but consume more memory per partition. ### Temporary Directory[​](#temporary-directory "Direct link to Temporary Directory") Place spill files on fast storage (SSD/NVMe) separate from data files: ``` runtime: query: temp_directory: /fast-ssd/spice-temp ``` The `max_temp_directory_size` setting limits the total size of temporary files (default: 100 GB). ## Caching Configuration[​](#caching-configuration "Direct link to Caching Configuration") Spice supports multiple caching layers for query acceleration. ### SQL Results Cache[​](#sql-results-cache "Direct link to SQL Results Cache") Cache query results for repeated queries: ``` runtime: caching: sql_results: enabled: true max_size: 512MiB item_ttl: 5m eviction_policy: tiny_lfu # Higher hit rate than lru cache_key_type: plan # Matches semantically equivalent queries ``` **Cache Key Types:** | Type | Behavior | Use Case | | ---------------- | ----------------------- | -------------------------- | | `plan` (default) | Uses query logical plan | Varied query formatting | | `sql` | Uses exact SQL string | Identical repeated queries | Use `sql` for lowest latency with identical queries. Use `plan` for semantic query matching. ### Stale-While-Revalidate[​](#stale-while-revalidate "Direct link to Stale-While-Revalidate") Serve stale cached results while refreshing in the background: ``` runtime: caching: sql_results: enabled: true item_ttl: 1m stale_while_revalidate_ttl: 5m # Serve stale for 5m while refreshing ``` This pattern reduces query latency spikes during cache refresh. ## Data Refresh Optimization[​](#data-refresh-optimization "Direct link to Data Refresh Optimization") ### Refresh Mode Selection[​](#refresh-mode-selection "Direct link to Refresh Mode Selection") | Mode | Memory Impact | Use Case | | --------- | ------------- | --------------------------------------- | | `full` | 2.5x dataset | Small-medium datasets, complete updates | | `append` | Minimal | Time-series, logs, immutable data | | `changes` | Minimal | CDC-enabled sources | | `caching` | Minimal | Dynamic content, API data | ### Append Mode Optimization[​](#append-mode-optimization "Direct link to Append Mode Optimization") Use `time_column` for efficient incremental updates: ``` datasets: - from: s3://bucket/events/ name: events time_column: event_time acceleration: engine: cayenne mode: file refresh_mode: append refresh_check_interval: 5m ``` ### Partitioned Data[​](#partitioned-data "Direct link to Partitioned Data") For data that has a granular time column and a separate column for partitioning (i.e. day buckets), set `time_partition_column` to the partitioning column based on time: ``` datasets: - from: s3://bucket/events/ name: events time_column: event_time time_partition_column: event_date # Physical partition column acceleration: refresh_mode: append ``` In this scenario, `event_date` is the day bucket used for physical partitioning (e.g., s3://bucket/events/event\_date=2025-10-01/), while event\_time provides the granular timestamp for precise filtering. | event\_id | event\_time (time\_column) | event\_date (time\_partition\_column) | event\_type | user\_id | | --------- | -------------------------- | ------------------------------------- | ----------- | -------- | | 8f2a-1 | 2025-10-01 08:14:22.123 | 2025-10-01 | page\_view | u\_442 | | 8f2a-2 | 2025-10-01 22:01:05.884 | 2025-10-01 | click | u\_901 | | 9c11-a | 2025-10-02 01:12:44.001 | 2025-10-02 | purchase | u\_442 | | 9c11-b | 2025-10-02 14:30:12.550 | 2025-10-02 | page\_view | u\_118 | ### Last-Modified Optimization[​](#last-modified-optimization "Direct link to Last-Modified Optimization") For object storage with append-only files, use `last_modified` to skip unchanged files: ``` datasets: - from: s3://bucket/logs/ name: logs time_column: last_modified # Special value using file metadata acceleration: refresh_mode: append ``` ## Resource Allocation[​](#resource-allocation "Direct link to Resource Allocation") ### Kubernetes[​](#kubernetes "Direct link to Kubernetes") Configure resource requests and limits based on workload: ``` apiVersion: v1 kind: Pod metadata: name: spice spec: containers: - name: spice image: spiceai/spiceai:latest resources: requests: memory: '8Gi' cpu: '4' limits: memory: '12Gi' # Do not set CPU limits - can cause throttling volumeMounts: - name: data mountPath: /data - name: temp mountPath: /tmp/spice volumes: - name: data persistentVolumeClaim: claimName: spice-data - name: temp emptyDir: medium: Memory # Use RAM for temp files if available ``` CPU Limits Avoid setting CPU limits. CPU limits can cause [throttling](https://home.robusta.dev/blog/stop-using-cpu-limits) even when CPU is available, degrading query performance. Set CPU requests to guarantee scheduling. To bound how wide Spice builds its thread pools without imposing a CFS quota, use [`runtime.cpu.cores`](/docs/next/reference/spicepod/runtime#runtimecpucores) rather than `resources.limits.cpu`. It caps how much machine the runtime organizes itself around; it does not cap how much CPU the process may consume, so there is no throttling. #### Sizing the Runtime for its CPU Entitlement[​](#sizing-the-runtime-for-its-cpu-entitlement "Direct link to Sizing the Runtime for its CPU Entitlement") The runtime sizes its thread pools, query fan-out, and accelerator concurrency from one CPU entitlement. Every CPU-derived pool scales with that number, and so does memory: worker threads, partitions, and per-plan operator reservations all grow with it, roughly linearly. A value far above the cores actually available buys nothing and costs both scheduling overhead and memory. **A pod that follows the advice above — requests, no limits — is sized for a bounded multiple of its CPU request**, not for the whole node. A `requests.cpu: 4` pod on a 64-core node sizes for 8 cores. This is the default and needs no configuration: the Spice Helm chart and the Spice Kubernetes Operator both pass the pod's CPU request through automatically whenever one is set. That default suits the common case — a pod scheduled against a request, sharing a node. Two deployments want something else: | Intent | Configuration | | -------------------------------------------------------------------------------------- | ------------------------------------------- | | Pack many mostly-idle instances on a node, each free to burst across the whole machine | `runtime.cpu.cores: all`, small request | | Size for a specific number regardless of what the pod requests | `runtime.cpu.cores: 6` | | Hard-cap CPU consumption because cluster policy requires it | `resources.limits.cpu` (accepts throttling) | ``` runtime: cpu: cores: all # every available core, regardless of the CPU request # (a CPU limit, if one is set, is still respected) ``` A hand-written pod spec that sets a CPU request but does not pass it through gets neither behavior — it falls back to sizing for the machine, and the runtime warns at startup naming the variable to set. See [Sizing from a CPU request](/docs/next/reference/spicepod/runtime#sizing-from-a-cpu-request). Compare `spiced_cpu_budget_cores` against `spiced_cpu_request_millicores` and `spiced_cpu_limit_millicores` to see what a pod sized for and what it was chosen against; the `source` label says which rung produced it. See [`runtime.cpu`](/docs/next/reference/spicepod/runtime#runtimecpu). ### Storage Recommendations[​](#storage-recommendations "Direct link to Storage Recommendations") | Storage Type | Use Case | | ------------ | ---------------------------------------------- | | SSD/NVMe | Acceleration data files, query spill files | | HDD | Cold data archival, infrequent access | | RAM (tmpfs) | Temporary files for high-performance workloads | ## Monitoring and Profiling[​](#monitoring-and-profiling "Direct link to Monitoring and Profiling") ### Task History[​](#task-history "Direct link to Task History") Enable query plan capture for slow query analysis: ``` runtime: task_history: enabled: true captured_plan: explain analyze min_sql_duration: 1s # Only capture plans for queries >1s ``` ### Metrics[​](#metrics "Direct link to Metrics") Monitor key performance metrics: * `query_duration_ms` - Query execution time * `query_executions` - Query throughput * `dataset_load_state` - Acceleration status * Cache hit rates See [Observability](/docs/next/features/observability) for metric configuration. ## Performance Checklist[​](#performance-checklist "Direct link to Performance Checklist") Use this checklist when optimizing Spice deployments: * Select appropriate accelerator based on dataset size and query patterns * Configure memory limits for DuckDB and/or Spice Cayenne caches * Set `runtime.query.memory_limit` with spill directory on fast storage * Enable caching for repeated queries * Use `refresh_mode: append` for time-series data * Configure indexes for point lookup queries (DuckDB/SQLite) * Set resource limits in Kubernetes * Enable observability for monitoring ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") **Spice Documentation:** * [Managing Memory Usage](/docs/next/reference/memory) - Memory configuration reference * [Data Accelerators](/docs/next/components/data-accelerators) - Accelerator documentation * [Spice Cayenne Data Accelerator](/docs/next/components/data-accelerators/cayenne) - Spice Cayenne-specific tuning * [DuckDB Data Accelerator](/docs/next/components/data-accelerators/duckdb) - DuckDB-specific tuning * [Caching](/docs/next/features/caching) - Cache configuration * [Observability](/docs/next/features/observability) - Metrics and monitoring **External References:** * [Apache DataFusion](https://datafusion.apache.org/) - Query execution engine for Arrow and Spice Cayenne * [DataFusion Configuration](https://datafusion.apache.org/user-guide/configs.html) - DataFusion configuration settings * [DataFusion Tuning Guide](https://datafusion.apache.org/user-guide/configs.html#tuning-guide) - Performance tuning for DataFusion * [DuckDB Indexing](https://duckdb.org/docs/stable/guides/performance/indexing) - Zone-maps and ART index documentation * [Vortex](https://github.com/vortex-data/vortex) - Columnar format used by Spice Cayenne * [Vortex Benchmarks](https://bench.vortex.dev) - Performance benchmarks for Vortex * [Apache Parquet File Format](https://parquet.apache.org/docs/file-format/) - Row groups, statistics, and Page Index * [Iceberg Partitioning](https://iceberg.apache.org/docs/latest/partitioning/) - Hidden partitioning and partition evolution * [Iceberg Performance](https://iceberg.apache.org/docs/latest/performance/) - Metadata filtering and column statistics * [Delta Lake Optimizations](https://docs.delta.io/latest/optimizations-oss.html) - Data skipping, Z-ordering, and compaction --- # YAML syntax for Spicepod manifests Spicepod manifests use YAML syntax. By default they are stored in the root directory of the application and named `spicepod.yaml` or `spicepod.yml` — when the runtime is given a directory path, it looks for one of those two filenames. The runtime also accepts any YAML file path directly (e.g. `spiced ./configs/my-app.yaml`); files passed by path must declare `kind: Spicepod` and a recognized `version` and are otherwise rejected with an explicit error. Tip Readers who are new to YAML can find a primer in "[Learn YAML in Y minutes](https://learnxinyminutes.com/docs/yaml/)." ## `version`[​](#version "Direct link to version") The version of the Spicepod manifest. The current version is `v2`. ## `kind`[​](#kind "Direct link to kind") The kind of Spicepod manifest. The kind is `Spicepod`. ## `name`[​](#name "Direct link to name") The name of the Spicepod. ## `secrets`[​](#secrets "Direct link to secrets") The secrets section in the Spicepod manifest is optional and is used to configure how secrets are stored and accessed by the Spicepod. For more information, see [Secret Stores](/docs/next/components/secret-stores). ### `secrets.from`[​](#secretsfrom "Direct link to secretsfrom") The `from` field is a string that represents the Uniform Resource Identifier (URI) for the secret store. This URI is composed of two parts: a prefix indicating the Secret Store to use, and an optional selector that specifies the secret to retrieve. The syntax for the `from` field is as follows: ``` from: : ``` Where: * ``: The Secret Store to use Currently supported secret stores: * [`env`](/docs/next/components/secret-stores/env) * [`kubernetes`](/docs/next/components/secret-stores/kubernetes) * [`keyring`](/docs/next/components/secret-stores/keyring) * [`aws-secrets-manager`](/docs/next/components/secret-stores/aws-secrets-manager) If no secret stores are explicitly specified, it defaults to `env`. * ``: The secret within the secret store to load. The type of secret store for reading secrets. Example ``` secrets: - from: env name: env ``` ### `secrets.name`[​](#secretsname "Direct link to secretsname") The name of the secret store. This is used to reference the store in the secret replacement syntax, `${:}`. ## `runtime`[​](#runtime "Direct link to runtime") The `runtime` section specifies configuration settings for the Spice runtime. For detailed documentation, see the [Runtime YAML reference](/docs/next/reference/spicepod/runtime). ## `metadata`[​](#metadata "Direct link to metadata") An optional `map` of metadata. **Example** ``` metadata: epoch_time: 1605312000 period: 72h interval: 1m granularity: 10s episodes: 10 ``` ## `datasets`[​](#datasets "Direct link to datasets") A Spicepod can contain one or more [datasets](/docs/next/reference/spicepod/datasets) referenced by relative path. **Example** A datasets referenced by relative path. ``` datasets: - ref: datasets/uniswap_v2_eth_usdc ``` A dataset defined inline. ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true refresh_mode: full refresh_check_interval: 1h ``` ## `snapshots`[​](#snapshots "Direct link to snapshots") Optional. Configure managed acceleration snapshots that Spice can use to bootstrap file-based accelerations. When enabled, datasets that opt in with [`acceleration.snapshots`](/docs/next/reference/spicepod/datasets#accelerationsnapshots) will download database files from the snapshot location if the local file is missing, and will optionally write new snapshots after each refresh. DuckDB, SQLite, Cayenne, and Turso accelerations running in `mode: file` are supported, and each dataset must write to its own file path. ``` snapshots: enabled: true location: s3://my_bucket/snapshots/ bootstrap_on_failure_behavior: warn # warn | retry | fallback params: s3_auth: iam_role ``` ### `snapshots.enabled`[​](#snapshotsenabled "Direct link to snapshotsenabled") Enable or disable snapshot management globally. Defaults to `true`. ### `snapshots.location`[​](#snapshotslocation "Direct link to snapshotslocation") The folder where snapshots are stored. Supports S3 bucket URIs (`s3://bucket/prefix/`), Azure ADLS Gen2 URIs (`abfss://container@account.dfs.core.windows.net/path/`), Google Cloud Storage URIs (`gs://bucket/prefix/`), and absolute or relative filesystem paths. The path must resolve to a single folder; Spice creates per-dataset folders underneath using Hive-style partitions (`month=YYYY-MM/day=YYYY-MM-DD/dataset=`). ### `snapshots.bootstrap_on_failure_behavior`[​](#snapshotsbootstrap_on_failure_behavior "Direct link to snapshotsbootstrap_on_failure_behavior") Controls what happens when Spice cannot load the most recent snapshot on startup. Valid values: * `warn` (default) – Log a warning and continue with an empty acceleration. * `retry` – Retry the newest snapshot until it loads successfully. * `fallback` – Attempt older snapshots in the same dataset folder until one works. ### `snapshots.params`[​](#snapshotsparams "Direct link to snapshotsparams") Optional key-value map passed to the snapshot storage layer. When `location` points to S3, the configuration accepts any of the [S3 dataset parameters](/docs/next/components/data-connectors/s3). Snapshots default to `s3_auth: iam_role`, which differs from the S3 dataset default of `public`. Azure ADLS and GCS locations also accept their respective connector parameters for explicit credential overrides; when no overrides are supplied, Spice reads standard environment variables for each cloud provider. ## `models`[​](#models "Direct link to models") A Spicepod can contain one or more [models](/docs/next/reference/spicepod/models) referenced by relative path. **Example** A model referenced by path. ``` models: - from: models/drive_stats ``` A model defined inline. ``` models: - from: spice.ai:openai/gpt-4o name: cloud_llm params: spiceai_api_key: ${secrets:SPICEAI_API_KEY} ``` ## `embeddings`[​](#embeddings "Direct link to embeddings") A Spicepod can contain one or more [embeddings](/docs/next/reference/spicepod/embeddings) referenced by relative path. **Example** An embeddings model referenced by path. ``` embeddings: - from: embeddings/openai_text_embedding_3 ``` An embedding defined inline. ``` embeddings: - name: hf_baai_bge from: huggingface:huggingface.co/BAAI/bge-small-en-v1.5 ``` ## `dependencies`[​](#dependencies "Direct link to dependencies") A list of dependent Spicepods. ``` dependencies: - lukekim/demo - spicehq/nfts ``` ## `views`[​](#views "Direct link to views") A Spicepod can contain one or more views which are virtual tables defined by SQL queries. **Example** ``` views: - name: rankings sql: | WITH a AS ( SELECT products.id, SUM(count) AS count FROM orders INNER JOIN products ON orders.product_id = products.id GROUP BY products.id ) SELECT name, count FROM products LEFT JOIN a ON products.id = a.id ORDER BY count DESC LIMIT 5 ``` ## `workers`[​](#workers "Direct link to workers") A Spicepod can contain one or more [workers](/docs/next/reference/spicepod/workers) defining configurable units of compute. **Example** ``` workers: - name: round-robin description: | Distributes requests between 'llama3_2' and 'gpt4_1' models in a round-robin fashion. load_balance: routing: - from: llama3_2 - from: gpt4_1 - name: fallback description: | Attempts 'gpt4_1' first, then 'llama3_2', then 'anth_haiku' if previous models fail. load_balance: routing: - from: llama3_2 order: 2 - from: gpt4_1 order: 1 - from: anth_haiku order: 3 - name: weighted description: | Routes 80% of traffic to 'llama3_2'. load_balance: routing: - from: llama3_2 weight: 4 - from: gpt4_1 weight: 1 ``` For a complete specification of worker configuration, see the [Workers Reference](/docs/next/reference/spicepod/workers). --- The `catalogs` section of a Spicepod defines connections to external data catalogs, such as Databricks Unity Catalog or Spice.ai Cloud. Catalogs expose multiple schemas and tables through a single configuration, making it easier to work with large numbers of datasets. # `catalogs` Example: `spicepod.yaml` ``` catalogs: - from: spice.ai name: spiceai include: - 'tpch.*' # Include only the "tpch" tables. ``` ## `from`[​](#from "Direct link to from") The `from` field is a string that represents the Uniform Resource Identifier (URI) for the catalog provider. This URI is composed of two parts: a prefix indicating the Catalog Connector to use, and the catalog path within the source. The syntax for the `from` field is as follows: ``` from: : ``` Where: * ``: The Catalog Connector to use to connect to the dataset Currently supported catalog connectors: * [`spice.ai`](/docs/next/components/catalogs/spiceai) * [`databricks`](/docs/next/components/catalogs/databricks) * [`unity_catalog`](/docs/next/components/catalogs/unity-catalog) If the Data Connector is not explicitly specified, it defaults to `spiceai`. * ``: The path to the catalog within the provider. ## `ref`[​](#ref "Direct link to ref") An alternative to adding the catalog definition inline in the `spicepod.yaml` file. `ref` can be use to point to a directory with a catalog defined in a `catalog.yaml` file. For example, a catalog configured in a catalog.yaml in the "catalogs/sample" directory can be referenced with the following: **catalogs/sample/catalog.yaml** ``` from: spice.ai name: spiceai include: - 'tpch.*' # Include only the "tpch" tables. ``` **ref used in spicepod.yaml** ``` version: v1 kind: Spicepod name: duckdb catalogs: - ref: catalogs/sample ``` ## `name`[​](#name "Direct link to name") The name of the catalog to register in Spice. The schema hierarchy of the external catalog is preserved in Spice. It doesn't need to match the name of the catalog in the external provider. ## `include`[​](#include "Direct link to include") Optional. The `include` field is used to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. ## `exclude`[​](#exclude "Direct link to exclude") Optional. The `exclude` field specifies tables to omit from the catalog, matched against the same table name as `include` and using the same glob syntax. Multiple `exclude` patterns are OR'ed together, and `exclude` takes precedence over `include` — a table matched by both is omitted. Every Catalog Connector applies it. A common use is to keep tables that cannot be [CDC-accelerated](/docs/next/components/catalogs/postgres#catalog-level-cdc-acceleration) out of an accelerated catalog's scope. ## `access`[​](#access "Direct link to access") Optional. Specifies the access level for the catalog. Supported values are: * `read` (default): Read-only access. * `read_write`: Enables both read and write operations. Only supported for [write-capable catalogs](/docs/next/tags/write). ## `params`[​](#params "Direct link to params") Optional. Parameters to pass to the catalog connector for retrieving the metadata on the schemas and tables to be included. The parameters are specific to the connector used. ## `dataset_params`[​](#dataset_params "Direct link to dataset_params") Optional. Parameters used when constructing the individual datasets that are registered in Spice from the catalog. The parameters are specific to the connector used. ## `acceleration`[​](#acceleration "Direct link to acceleration") Optional. Bootstraps and accelerates every table discovered by the catalog (subject to `include`/`exclude`), with no per-table configuration. Currently supported for the [PostgreSQL catalog connector](/docs/next/components/catalogs/postgres#catalog-level-cdc-acceleration) only. ``` catalogs: - from: pg name: my_pg acceleration: engine: cayenne # optional; cayenne is the only supported engine refresh_mode: changes # required ``` * `engine`: Optional. The accelerator engine used for every table. Defaults to `cayenne`, currently the only supported value. * `refresh_mode`: Required. The only supported value is `changes` (CDC); there is no catalog-level `full` mode. Per-table-only acceleration settings (`primary_key`, `on_conflict`, `indexes`, and other per-dataset overrides) are not configurable at the catalog level — they remain on an individual [dataset's `acceleration` block](/docs/next/components/data-accelerators). An accelerated catalog is included in the runtime's memory budget: the engine, refresh mode, storage mode and params above are the configuration every table it accelerates is built from, and the runtime classifies that same configuration when it derives `runtime.query.memory_limit`. A Spicepod whose only Cayenne acceleration is a catalog therefore gets the reduced query pool and the compaction reservation, and the projected off-pool reservation counts one table per catalog rather than one per discovered table — see [Accelerated catalogs](/docs/next/reference/memory#accelerated-catalogs) for what that means when sizing. --- # Datasets A Spicepod can contain one or more `datasets` referenced by relative path or defined inline. Inline example: `spicepod.yaml` ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true mode: memory # / file engine: arrow # / cayenne / duckdb / sqlite / postgres / turso refresh_check_interval: 1h refresh_mode: full / append # / changes / caching / snapshot ``` `spicepod.yaml` ``` datasets: - from: databricks:spiceai.datasets.specific_table name: uniswap_eth_usd params: environment: prod acceleration: enabled: true mode: memory # / file engine: arrow # / cayenne / duckdb / sqlite / postgres / turso refresh_check_interval: 1h refresh_mode: full / append # / changes / caching / snapshot ``` Relative path example: `spicepod.yaml` ``` datasets: - ref: datasets/taxi_trips ``` `datasets/taxi_trips/dataset.yaml` ``` from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true refresh_check_interval: 1h ``` ## `from`[​](#from "Direct link to from") The `from` field is a string that represents the Uniform Resource Identifier (URI) for the dataset. This URI is composed of two parts: a prefix indicating the Data Connector to use to connect to the dataset, a delimiter, and the path to the dataset within the source. The syntax for the `from` field is as follows: ``` from: : # OR from: / # OR from: :// ``` Where: * ``: The Data Connector to use to connect to the dataset Currently supported data connectors: * [`spiceai`](/docs/next/components/data-connectors/spiceai) * [`dremio`](/docs/next/components/data-connectors/dremio) * [`spark`](/docs/next/components/data-connectors/spark) * [`databricks`](/docs/next/components/data-connectors/databricks) * [`s3`](/docs/next/components/data-connectors/s3) * [`postgres`](/docs/next/components/data-connectors/postgres) * [`mysql`](/docs/next/components/data-connectors/mysql) * [`flightsql`](/docs/next/components/data-connectors/flightsql) * [`snowflake`](/docs/next/components/data-connectors/snowflake) * [`ftp`, `sftp`](/docs/next/components/data-connectors/ftp) * [`http`, `https`](/docs/next/components/data-connectors/https) * [`clickhouse`](/docs/next/components/data-connectors/clickhouse) * [`graphql`](/docs/next/components/data-connectors/graphql) * [`cosmosdb`](/docs/next/components/data-connectors/cosmosdb) If the Data Connector is not explicitly specified, it defaults to `spiceai`. * ``: The delimiter between the Data Connector and the path. Currently supported delimiters are `:`, `/`, and `://`. Some connectors place additional restrictions on the allowed delimiters to better conform to the expected syntax of the underlying data source, i.e. `s3://` is the only supported delimiter for the `s3` connector. * ``: The path to the dataset within the source. Identifier Case Sensitivity Unquoted identifiers in the `` are normalized to lowercase. To reference a table or schema with mixed-case or uppercase characters, wrap each case-sensitive part in double quotes: ``` datasets: # Case is preserved for "ActionExecutions" - from: postgres:my_schema."ActionExecutions" name: action_executions # Quote each part individually as needed - from: databricks:my_catalog."MySchema"."MyTable" name: my_table ``` This applies to all federated database connectors where `` references a table identifier. Connectors that interpret `` as a file path (e.g. `s3`, `delta_lake`, `ftp`) do not apply identifier normalization. See [Identifier Case Sensitivity and Quoting](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting) for details. ## `ref`[​](#ref "Direct link to ref") An alternative to adding the dataset definition inline in the `spicepod.yaml` file. `ref` can be use to point to a directory with a dataset defined in a `dataset.yaml` file. For example, a dataset configured in a dataset.yaml in the "datasets/sample" directory can be referenced with the following: **dataset.yaml** ``` from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true refresh_check_interval: 1h ``` **ref used in spicepod.yaml** ``` version: v1 kind: Spicepod name: duckdb datasets: - ref: datasets/sample ``` ## `name`[​](#name "Direct link to name") The name of the dataset. Used to reference the dataset in the pod manifest, as well as in external data sources. The name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). Spice follows PostgreSQL SQL syntax conventions, which normalize unquoted identifiers to lowercase. A dataset named `LINEITEM` is accessible in queries as `lineitem`. To preserve uppercase or mixed-case names, wrap the name in double quotes. In YAML, this requires an extra layer of quoting: ``` datasets: - from: snowflake:SNOWFLAKE_SAMPLE_DATA.TPCH_SF100.LINEITEM name: '"LINEITEM"' params: snowflake_account: JYFGIWYEFBW snowflake_warehouse: snowflake_wh snowflake_password: ${secrets:SNOWFLAKE_PASSWORD} snowflake_username: ${secrets:SNOWFLAKE_USERNAME} ``` ``` -- Query using the preserved uppercase name SELECT * FROM "LINEITEM"; ``` Without the double quotes, the same dataset would be queryable only as `lineitem`. See [Identifier Case Sensitivity and Quoting](/docs/next/components/data-connectors#identifier-case-sensitivity-and-quoting) for full details on quoting in both the `from` and `name` fields. ## `description`[​](#description "Direct link to description") The description of the dataset. Used as part of the [Semantic Data Model](/docs/next/features/semantic-model). ## `access`[​](#access "Direct link to access") Optional. Specifies the access level for the dataset. Supported values are: * `read` (default): Read-only access. * `read_write`: Enables both read and write operations. Only supported for [write-capable connectors](/docs/next/tags/write). To enable write operations, configure your dataset with `read_write` access: ``` datasets: - from: glue:my_catalog.my_schema.my_table name: my_table access: read_write params: # ... connector-specific parameters ``` ## `time_column`[​](#time_column "Direct link to time_column") Optional. The name of the column that represents the temporal (time) ordering of the dataset. Required to enable a retention policy on the dataset. ## `time_format`[​](#time_format "Direct link to time_format") Optional. The format of the `time_column`. The following values are supported: * `timestamp` - Default. Timestamp without a timezone. E.g. `2016-06-22 19:10:25` with data type `timestamp`. * `timestamptz` - Timestamp with a timezone. E.g. `2016-06-22 19:10:25-07` with data type `timestamptz`. * `unix_seconds` - Unix timestamp in seconds. E.g. `1718756687`. * `unix_millis` - Unix timestamp in milliseconds. E.g. `1718756687000`. * `unix_nanos` - Unix timestamp in nanoseconds. E.g. `1718756687000000000`. Use this for OpenTelemetry's `time_unix_nano` columns — see [OpenTelemetry Data Ingestion](/docs/next/features/data-ingestion#opentelemetry-data-ingestion). * `ISO8601` - [ISO 8601](https://en.wikipedia.org/wiki/ISO_8601) format. * `date` - Date in YYYY-MM-DD format. E.g. `2024-01-01`. Spice emits a warning if the `time_column` from the data source is incompatible with the `time_format` config. Limitations * String-based columns are assumed to be ISO8601 format. ## `time_partition_column`[​](#time_partition_column "Direct link to time_partition_column") (Optional) Specify the column that represents the physical partitioning of the dataset when using append-based acceleration. When the defined `time_column` is a fine-grained timestamp and the dataset is physically partitioned by a coarser granularity (for example, by date), setting `time_partition_column` to the partition column (e.g. date\_col) improves partition pruning, excludes irrelevant partitions during refreshes, and optimizes scan efficiency. ## `time_partition_format`[​](#time_partition_format "Direct link to time_partition_format") (Optional) Define the format of the `time_partition_column`. For instance, if the physical partitions follow a date format (YYYY-MM-DD), set this value to `date`. The same format options as `time_format` are supported for `time_partition_column`. ## Schema Inference and Evolution[​](#schema-inference-and-evolution "Direct link to Schema Inference and Evolution") Spice infers the dataset schema from the data source at startup. The inferred schema defines the column names, data types, and nullability used for the lifetime of that runtime process. By default, schema changes at the source are not applied at runtime — data refreshes will fail if the source schema drifts, and you must restart the runtime to re-infer the schema. Accelerated datasets can opt into automatic, in-place schema evolution with the [`on_schema_change`](#on_schema_change) policy, which adopts lossless, widening-compatible source changes without a restart — and can additionally drop and recreate the accelerated table on incompatible changes (`drop_and_recreate`) when `refresh_mode: full` is set. For connector-specific inference parameters, runtime schema change behavior, and recommendations, see [Schema Inference](/docs/next/components/data-connectors#schema-inference). ## `on_schema_change`[​](#on_schema_change "Direct link to on_schema_change") Optional. Controls how the runtime reacts when the source schema changes after the dataset is registered. Applies to **accelerated datasets only** — federated (non-accelerated) queries always reflect the live source schema, so the policy is inert for them (a non-default value logs a warning and otherwise has no effect). The following values are supported: * `block` - Default. Schema changes are not applied automatically. The dataset stays healthy and continues serving queries using the registered schema; this preserves the historical behavior. * `fail` - Set the dataset to an error status with an actionable message when the projected source schema diverges from the registered schema. Self-heals if the source reverts. * `append_new_columns` - Adopt new nullable source columns in place; type changes and relaxed/tightened nullability are treated as `block` (the dataset keeps serving on the old schema) and a warning is logged. * `sync_all_columns` - Adopt the full set of lossless, widening changes in place: new nullable columns, widened column types (for example `Int32`→`Int64`, an increase in decimal precision, or `Utf8`→`LargeUtf8`), and relaxed nullability. Non-widening changes (column removals, narrowing type changes) remain `block`-equivalent and are warned. * `drop_and_recreate` - Everything `sync_all_columns` does (adopt lossless widening changes in place), and additionally **recreate** the accelerated table for changes that cannot be applied in place — column removals, narrowing or otherwise incompatible type changes, and new non-nullable columns. Recreation is **destructive**: the accelerated data is dropped and rebuilt from the source, so it is only performed with `refresh_mode: full` (a full refresh re-fetches every row). With `refresh_mode: append` or `changes`, an incompatible change is rejected and the existing table is preserved (recreating would drop rows that cannot be re-fetched). Supported on the `duckdb`, `sqlite`, `turso`, and `cayenne` engines. ``` datasets: - from: postgres:public.events name: events on_schema_change: drop_and_recreate acceleration: engine: duckdb mode: file refresh_mode: full ``` note In-place evolution (no restart) is supported for the `duckdb`, `sqlite`, `turso`, and Spice Cayenne (`cayenne`) acceleration engines, including for PostgreSQL CDC (`refresh_mode: changes`). Other engines (for example `arrow` and the PostgreSQL accelerator) log a clear unsupported message and degrade safely, applying additive changes on restart. Constraint and primary-key columns cannot be widened in place. For destructive schema changes (column removals or narrowing), set `on_schema_change: drop_and_recreate` with `refresh_mode: full` to drop and recreate the accelerated table from the source, or use [`mode: file_update`](#accelerationmode), which recreates the acceleration file on any change. ## `unsupported_type_action`[​](#unsupported_type_action "Direct link to unsupported_type_action") Optional. Specifies the action to take when a data type that is not supported by the data connector is encountered. The following values are supported: * `error` - Default. Return an error when an unsupported data type is encountered. * `warn` - Log a warning and ignore the column containing the unsupported data type. * `ignore` - Log nothing and ignore the column containing the unsupported data type. * `string` - Attempt to convert the unsupported data type to a string. Currently only supports converting the PostgreSQL JSONB type. Limitations Not all connectors support specifying an `unsupported_type_action`. When specified on a connector that does not support the option, the connector will fail to register. The following connectors support `unsupported_type_action`: * [DuckDB](/docs/next/components/data-connectors/duckdb) * [PostgreSQL](/docs/next/components/data-connectors/postgres) ## `schema_inference`[​](#schema_inference "Direct link to schema_inference") Removed The `schema_inference` dataset field was removed. Schema inference is now **always on** — Spice attempts the deepest inference each source permits and no longer requires an opt-in. A Spicepod that still specifies `schema_inference` fails to load with an unknown-field error; remove the field. For connectors that expose catalog metadata — [PostgreSQL](/docs/next/components/data-connectors/postgres), [MySQL](/docs/next/components/data-connectors/mysql), and [MongoDB](/docs/next/components/data-connectors/mongodb) — inference additionally detects the source's primary key (and, where the source exposes them, secondary indexes and sort/clustering columns) and applies them to any acceleration settings left unset. It degrades gracefully — with info-level logs — when the connection role lacks catalog read access. Other connectors infer the base column schema (names, types, nullability) only. For the previous opt-in `standard` / `extended` behavior, see the [v2.1.x documentation](https://docs.spiceai.org/docs/2.1.x/reference/spicepod/datasets#schema_inference). ## `ready_state`[​](#ready_state "Direct link to ready_state") Supports one of three values (defaults to `on_load`): * `on_registration`: Mark the dataset as ready immediately, and queries on this table will fall back to the underlying source directly until the initial acceleration is complete. When combined with fully declared [`columns[].type`](#columnstype) entries, enables [deferred dataset initialization](#deferred-dataset-initialization) — the source connector is not created until the first query. * `on_load`: (default) Mark the dataset as ready only after the initial acceleration. Queries against the dataset will return an error before the load has been completed. * `on_schema_resolved`: Mark the dataset as ready once the federated source's schema has been resolved (which also verifies access to the source), without waiting for the initial data refresh. Queries fall back to the federated source until the initial load completes; subsequent refresh failures are still reported via dataset status and metrics. ``` datasets: - from: s3://my_bucket/my_dataset/ name: my_dataset ready_state: on_registration # or on_load params: ... acceleration: enabled: true ``` ## `check_availability`[​](#check_availability "Direct link to check_availability") Spice can monitor whether the source backing a non-accelerated dataset is still reachable, marking the dataset `Error` while it is not. Availability monitoring is **opt-in**: it runs only for datasets that set [`check_availability_interval`](#check_availability_interval). Note that probing may trigger the startup of compute resources (for example, Databricks or Snowflake), potentially incurring additional costs. * `auto`: Monitor the dataset when it is not accelerated and `check_availability_interval` is set. This is the default value. Accelerated datasets are never monitored. * `disabled`: Never monitor the dataset, even with `check_availability_interval` set. ``` datasets: - from: databricks:catalog.schema.table name: my_dataset check_availability: disabled params: ... ``` ## `check_availability_interval`[​](#check_availability_interval "Direct link to check_availability_interval") Optional. How often the runtime probes the source backing this dataset to confirm it is still reachable. Accepts a duration string (for example `60s`, `5m`). There is no default — leave it unset and the dataset is not monitored at all. ``` datasets: - from: postgres:public.orders name: orders check_availability_interval: 60s params: ... ``` Only non-accelerated datasets with `check_availability: auto` are monitored. Setting the interval on an accelerated dataset logs a warning and has no effect, because an accelerated dataset keeps serving from its accelerator even when the source is unavailable. An unparseable duration fails dataset load rather than silently disabling the check. A recent successful query counts as proof that the source is reachable: when a dataset was queried successfully more recently than its own interval, that query stands in for the probe, so actively-queried datasets are not probed redundantly. This uses [task history](/docs/next/reference/task_history), and looks back at most one hour regardless of the configured interval; datasets with a longer interval are probed directly. Availability monitoring is not supported for Iceberg datasets and is skipped for them (see [spiceai/spiceai#6994](https://github.com/spiceai/spiceai/issues/6994)). The monitoring works by executing a query that selects one row and all columns from the dataset. i.e.: ``` SELECT "p_partkey", "p_name", "p_mfgr", "p_brand", "p_type", "p_size", "p_container", "p_retailprice", "p_comment" FROM spiceai_sandbox.tpch.part LIMIT 1; ``` If the monitoring query fails, a warning is emitted in the logs, an error is propagated to the `task_history` table, the `dataset_unavailable_time_ms` metric is recorded for the failing dataset, and the dataset's status becomes `Error` — visible via `GET /v1/datasets?status=true`. A later successful probe restores the dataset to `Ready`. The status transition applies only to a dataset that is otherwise `Ready` (or already in `Error` from a previous failed probe); transient lifecycle states such as `Initializing` and `Refreshing` are left to the load and refresh paths that own them, and a recovery only restores `Ready` for a dataset the availability monitor itself put into `Error`. ## Deferred dataset initialization[​](#deferred-dataset-initialization "Direct link to Deferred dataset initialization") Datasets can defer connector creation and schema inference until the first query by combining `ready_state: on_registration` with fully declared `columns[].type` entries. When every column has an explicit type, the runtime registers a placeholder table with the declared Arrow schema at startup — SQL planning and federation analysis work against this schema **without contacting the source**. On first query, the placeholder is swapped for the real provider. A dataset is eligible for deferred initialization when: * It is read-only. * `ready_state: on_registration` is set. * It has no embedding or full-text-search columns. * Every column has an explicit [`columns[].type`](#columnstype). ``` datasets: - from: https://api.example.com/data.json name: my_data ready_state: on_registration columns: - name: id type: bigint - name: name type: text - name: created_at type: timestamptz ``` When deferred initialization is active, the runtime: * Registers the dataset immediately with the declared schema — queries can reference the table in planning before the source is contacted. * On the first query that references the dataset, initializes the connector and loads the real data transparently. * Coordinates concurrent triggers so the dataset is only initialized once. * Supports acceleration — after deferred initialization, the acceleration table, refresh loop, and health monitor are set up normally. Breaking change `load: on_demand` has been removed. Replace it with `ready_state: on_registration` combined with explicit `columns[].type` declarations. ## `acceleration`[​](#acceleration "Direct link to acceleration") Optional. Accelerate queries to the dataset by caching data locally. ## `acceleration.enabled`[​](#accelerationenabled "Direct link to accelerationenabled") Enable or disable acceleration, defaults to `true`. ## `acceleration.engine`[​](#accelerationengine "Direct link to accelerationengine") The acceleration engine to use, defaults to `arrow`. The following engines are supported: * `arrow` - Accelerated in-memory backed by Apache Arrow DataTables. * [`cayenne`](/docs/next/components/data-accelerators/cayenne) - Accelerated by Spice Cayenne (Vortex) engine (Stable, v1.9.0-rc.1+). * [`duckdb`](/docs/next/components/data-accelerators/duckdb) - Accelerated by an embedded DuckDB database. * [`postgres`](/docs/next/components/data-accelerators/postgres) - Accelerated by a Postgres database. * [`sqlite`](/docs/next/components/data-accelerators/sqlite) - Accelerated by an embedded SQLite database. * [`turso`](/docs/next/components/data-accelerators/turso) - Accelerated by an embedded Turso (libSQL) database (Beta). ## `acceleration.mode`[​](#accelerationmode "Direct link to accelerationmode") Optional. The mode of acceleration. The following values are supported: * `memory` - Store acceleration data in-memory. Supported for Spice Cayenne (`cayenne`), where the acceleration is ephemeral and reloads from its source on restart. * `file` - Store acceleration data in a file. Reuses any existing file on startup. Supported for Spice Cayenne (`cayenne`), `duckdb`, `sqlite`, and `turso` acceleration engines. * `file_create` - Always create a new acceleration file on startup, removing any existing file. When [snapshots](/docs/next/features/data-acceleration/snapshots) are enabled, the existing file is snapshotted before deletion. Supported for Spice Cayenne (`cayenne`), `duckdb`, `sqlite`, and `turso` acceleration engines. * `file_update` - Open an existing acceleration file if it exists, then check schema compatibility on refresh. If the source schema change is additive (new columns only), the existing file is kept. If the schema change is incompatible (columns removed, renamed, or type changed), the file is snapshotted (if [snapshots](/docs/next/features/data-acceleration/snapshots) are enabled) and recreated from scratch. Supported for Spice Cayenne (`cayenne`), `duckdb`, `sqlite`, and `turso` acceleration engines. ## `acceleration.storage_profile`[​](#accelerationstorage_profile "Direct link to accelerationstorage_profile") Optional. The storage profile for file-backed acceleration. The runtime uses this hint to tune connection-pool sizing, checkpoint thresholds, and file-size defaults for the underlying medium. Only applies to file-mode accelerators (`duckdb`, `sqlite`, `turso`, and Spice Cayenne); memory-mode accelerators ignore this setting. Supported values: * `auto` (default) – Detect the storage profile from the resolved acceleration file path. On Linux, detection reads `/proc/self/mountinfo` and inspects block-device metadata to recognize Amazon EBS, Azure Managed Disks, Amazon EC2 NVMe instance storage, `tmpfs`/`ramfs`, and generic NVMe/SSD. On other platforms, detection returns unknown and the engine defaults apply. * `local_ssd` (aliases: `ssd`, `nvme`) – Treat the acceleration file location as local SSD/NVMe (for example, EC2 instance store or Azure temporary/NVMe local storage). Uses the engine defaults for connection pool size and checkpoint thresholds. * `ebs` (aliases: `azure_disk`, `managed_disk`, `network_disk`) – Treat the acceleration file location as network-attached block storage (for example, Amazon EBS or Azure Managed Disks). Reduces connection-pool size and raises DuckDB's checkpoint threshold so per-IO latency is amortized across larger flushes. Spice Cayenne uses smaller per-file targets to reduce write amplification. * `tmpfs` (aliases: `ram`, `ramdisk`, `ramfs`, `memory`) – Treat the acceleration file location as RAM-backed storage. Raises DuckDB's checkpoint threshold so steady-state workloads don't pay checkpoint cost on small amounts of dirty data; Spice Cayenne uses larger per-file targets to improve scan throughput. Example: ``` datasets: - from: s3://bucket/data/ name: analytics acceleration: engine: duckdb mode: file storage_profile: ebs params: duckdb_file: /mnt/ebs/analytics.db ``` ## `acceleration.snapshots`[​](#accelerationsnapshots "Direct link to accelerationsnapshots") Optional. Controls how this dataset participates in managed acceleration snapshots. Requires the Spicepod to configure the top-level [`snapshots` block](/docs/next/reference/spicepod/#snapshots), the acceleration engine to be `duckdb`, `sqlite`, `cayenne`, or `turso`, and `mode: file` with a dataset-specific file path (for example `acceleration.params.duckdb_file: /nvme/my_dataset.db`). Supported values: * `enabled` – Download the newest snapshot on startup when the acceleration file is missing and write a fresh snapshot after each refresh. * `bootstrap_only` – Download snapshots on startup but never write new ones. * `create_only` – Write snapshots after refreshes but never download them on startup. * `disabled` (default) – Do not use snapshots for this dataset. Snapshots are written beneath the configured snapshot location using Hive-style partitioning (`month=YYYY-MM/day=YYYY-MM-DD/dataset=`). For more background, see [Acceleration snapshots](/docs/next/features/data-acceleration/snapshots). ## `acceleration.snapshots_trigger`[​](#accelerationsnapshots_trigger "Direct link to accelerationsnapshots_trigger") Optional. Controls when Spice creates new snapshots. The available triggers depend on the dataset's refresh mode. **For batch-based datasets** (`refresh_mode: full`, or `refresh_mode: append` with `time_column`): * `refresh_complete` (default) – Create a snapshot after each data refresh completes. * `time_interval` – Create snapshots at a fixed time interval specified by `snapshots_trigger_threshold`. **For caching datasets** (`refresh_mode: caching`): * `time_interval` (default) – The only supported trigger. Create snapshots at a fixed time interval, defaulting to `10m` if `snapshots_trigger_threshold` is not specified. `refresh_complete` and `stream_batches` are not supported in caching mode and cause a configuration error. **For stream-based datasets** (`refresh_mode: changes`, or `refresh_mode: append` without `time_column`): * `time_interval` (default) – Create snapshots at a fixed time interval. Defaults to `10m` if `snapshots_trigger_threshold` is not specified. * `stream_batches` – Create a snapshot after a specified number of batches are processed. See [Acceleration snapshots](/docs/next/features/data-acceleration/snapshots) for more details. ## `acceleration.snapshots_trigger_threshold`[​](#accelerationsnapshots_trigger_threshold "Direct link to accelerationsnapshots_trigger_threshold") Optional. The threshold value for snapshot creation, interpreted based on the configured `snapshots_trigger`: * When `snapshots_trigger: time_interval` – A duration specifying how often to create snapshots (e.g., `10m`, `1h`). Defaults to `10m` for stream-based datasets. * When `snapshots_trigger: stream_batches` – An integer specifying the number of batch updates after which to create a snapshot. Not applicable when `snapshots_trigger: refresh_complete`. ## `acceleration.snapshots_compaction`[​](#accelerationsnapshots_compaction "Direct link to accelerationsnapshots_compaction") Optional. Enable database compaction before uploading snapshots. Only supported for the `duckdb` acceleration engine. Defaults to `disabled`. When enabled, Spice uses DuckDB's internal compaction mechanism (`COPY DATABASE`) to optimize the database file before uploading, reducing snapshot size and improving bootstrap performance. Supported values: * `enabled` – Compact the database before creating each snapshot. * `disabled` (default) – Upload snapshots without compaction. ## `acceleration.snapshots_creation_policy`[​](#accelerationsnapshots_creation_policy "Direct link to accelerationsnapshots_creation_policy") Optional. Controls when new snapshots are created after data refreshes. Defaults to `on_change`. Supported values: * `on_change` (default) – Only create a new snapshot when the data has changed since the last snapshot. * `always` – Create a new snapshot after every refresh, regardless of whether the data has changed. ## `acceleration.snapshots_reset_expiry_on_load`[​](#accelerationsnapshots_reset_expiry_on_load "Direct link to accelerationsnapshots_reset_expiry_on_load") Optional. Controls whether the snapshot expiry timer is reset when loading from a snapshot during bootstrap. Defaults to `disabled`. Supported values: * `disabled` (default) – Snapshot expiry is not reset on load; the original expiry time is preserved. * `enabled` – Reset the snapshot expiry when it is loaded during bootstrap. ## `acceleration.refresh_mode`[​](#accelerationrefresh_mode "Direct link to accelerationrefresh_mode") Optional. How to refresh the dataset. The following values are supported: * `full` - Refresh the entire dataset. * `append` - Append new data to the dataset. When `time_column` is specified, new records are fetched from the latest timestamp in the accelerated data at the `acceleration.refresh_check_interval`. * `changes` - Apply change data capture (CDC) events to incrementally update the dataset. * `caching` - Cache data based on request metadata (HTTP requests). Uses row-level replacement based on cache keys. See [Caching Mode](/docs/next/features/data-acceleration/refresh-modes/caching) for details. * `snapshot` - Reload exclusively from the [snapshot store](/docs/next/features/data-acceleration/snapshots). The federated source is never queried; the runtime polls for newer snapshots at `refresh_check_interval` (default: 60s). Requires `acceleration.snapshots: enabled` or `bootstrap_only` and a snapshot-capable file-based engine (DuckDB, SQLite, Cayenne, or Turso). Writes (`INSERT INTO`) are rejected. See [Snapshot Refresh Mode](/docs/next/features/data-acceleration/data-refresh#snapshot). ## `acceleration.write_mode`[​](#accelerationwrite_mode "Direct link to accelerationwrite_mode") Optional. Controls how writes to a `read_write` accelerated dataset propagate between the local accelerator and the federated source. Only applies when the dataset has `access: read_write` and the source connector supports writes. Supported values: * `write_through` (default) – Writes are sent to the federated source synchronously. The client receives an ACK only after the source commits the change, providing ACID guarantees. The local accelerator is updated through the configured refresh path (for example, the WAL stream when `refresh_mode: changes`). * `write_back` – Writes are applied to the local accelerator first (fast ACK), then forwarded asynchronously to the federated source. Choose this for write throughput when eventual consistency at the source is acceptable. ## `acceleration.refresh_check_interval`[​](#accelerationrefresh_check_interval "Direct link to accelerationrefresh_check_interval") Optional. How often data should be refreshed. For `append` datasets without a specific `time_column`, this config is not used. If not defined, the accelerator will not refresh after it initially loads data. Cannot be specified in conjunction with a `refresh_cron`. See [Duration](/docs/next/reference/duration) ## `acceleration.refresh_cron`[​](#accelerationrefresh_cron "Direct link to accelerationrefresh_cron") Optional. Specifies a cron schedule which controls how often data is refreshed. For `append` datasets without a specific `time_column`, this config is not used. If not defined, the accelerator will not refresh after it initially loads data. See the [cron schedule reference](/docs/next/reference/cron). ## `acceleration.params.caching_ttl`[​](#accelerationparamscaching_ttl "Direct link to accelerationparamscaching_ttl") Optional. The time-to-live (TTL) for cached data before it is considered stale. Only applicable when `refresh_mode: caching`. Defaults to `30s`. When cached data exceeds this age (measured from the `fetched_at` timestamp), it becomes stale. If `caching_stale_while_revalidate_ttl` is also configured, stale data is immediately served to queries (no delay) while a background refresh is triggered to update the cache, implementing the Stale-While-Revalidate (SWR) pattern. If `caching_stale_while_revalidate_ttl` is not set, queries wait for fresh data once the TTL expires. **Example**: ``` datasets: - from: https://api.tvmaze.com name: tv_shows acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file # Persist cache to disk params: caching_ttl: 15s # Cache data is fresh for 15 seconds refresh_check_interval: 30s # Periodic background refresh ``` See [Caching Mode](/docs/next/features/data-acceleration/refresh-modes/caching#cache-ttl-time-to-live) for detailed TTL configuration and behavior. See [Duration](/docs/next/reference/duration) ## `acceleration.params.caching_stale_while_revalidate_ttl`[​](#accelerationparamscaching_stale_while_revalidate_ttl "Direct link to accelerationparamscaching_stale_while_revalidate_ttl") Optional. The duration after `caching_ttl` expires during which stale data is served while refreshing in the background. Only applicable when `refresh_mode: caching`. Defaults to none (stale data is not served). When `caching_ttl` expires and data becomes stale, this parameter controls how long stale data continues to be served immediately while a background refresh occurs. After the combined `caching_ttl + caching_stale_while_revalidate_ttl` period, queries wait for fresh data instead of returning stale results. If omitted, cached data becomes "rotten" immediately after `caching_ttl` expires, and queries will wait for fresh data rather than returning stale results. **Example**: ``` datasets: - from: https://api.tvmaze.com name: tv_shows acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 15s # Cache data is fresh for 15 seconds caching_stale_while_revalidate_ttl: 30s # Serve stale data for 30 seconds while refreshing refresh_check_interval: 60s ``` See [Caching Mode](/docs/next/features/data-acceleration/refresh-modes/caching#cache-ttl-time-to-live) for detailed TTL configuration and behavior. See [Duration](/docs/next/reference/duration) ## `acceleration.params.caching_stale_if_error`[​](#accelerationparamscaching_stale_if_error "Direct link to accelerationparamscaching_stale_if_error") Optional. Controls whether expired cached data is served when the upstream data source returns an error. Only applicable when `refresh_mode: caching`. Defaults to `disabled`. When set to `enabled`, queries return expired cached data instead of failing if the upstream source returns an error during a refresh attempt. This provides fault tolerance for APIs with intermittent availability or rate limits. Valid values: * `enabled` - Serve expired cached data when upstream errors occur * `disabled` (default) - Propagate upstream errors to queries **Example**: ``` datasets: - from: https://api.tvmaze.com name: tv_shows acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 15s caching_stale_while_revalidate_ttl: 30s caching_stale_if_error: enabled # Serve stale data on upstream errors refresh_check_interval: 60s ``` See [Caching Mode](/docs/next/features/data-acceleration/refresh-modes/caching#stale-if-error-behavior) for detailed behavior. ## `acceleration.refresh_sql`[​](#accelerationrefresh_sql "Direct link to accelerationrefresh_sql") Optional. Filters the data fetched from the source to be stored in the accelerator engine. Supported for `full` and `append` refresh mode datasets. Must be of the form `SELECT * FROM {name} WHERE {refresh_filter}`. `{name}` is the dataset name declared above, `{refresh_filter}` is any SQL expression that can be used to filter the data, i.e. `WHERE city = 'Seattle'` to reduce the working set of data that is accelerated within Spice from the data source. Limitations * The refresh SQL only supports filtering data from the current dataset - joining across other datasets is not supported. * Queries for data that have been filtered out will not fall back to querying against the federated table. ## `acceleration.refresh_data_window`[​](#accelerationrefresh_data_window "Direct link to accelerationrefresh_data_window") Optional. A duration to filter dataset refresh source queries to recent data (duration into past from now). Requires `time_column` and `time_format` to also be configured. Supported for `full` and `append` refresh mode datasets. For example, `refresh_data_window: 24h` will include only records with a timestamp within the last 24 hours. See [Duration](/docs/next/reference/duration) ## `acceleration.refresh_append_overlap`[​](#accelerationrefresh_append_overlap "Direct link to accelerationrefresh_append_overlap") Optional. A duration to specify how far back to include records based on the most recent timestamp found in the accelerated data. Requires `time_column` to also be configured. Only supported for `append` refresh mode datasets. This setting can help mitigate missing data issues caused by late arriving data. Example: If the latest timestamp in the accelerated data table is `2020-01-01T02:00:00Z`, setting `refresh_append_overlap: 1h` will include records starting from `2020-01-01T01:00:00Z`. See [Duration](/docs/next/reference/duration) ## `acceleration.refresh_retry_enabled`[​](#accelerationrefresh_retry_enabled "Direct link to accelerationrefresh_retry_enabled") Optional. Specifies whether an accelerated dataset should retry data refresh in the event of transient errors. The default setting is true. Retries follow a [Fibonacci backoff strategy](https://en.wikipedia.org/wiki/Fibonacci_sequence). To disable refresh retries, set `refresh_retry_enabled: false`. ## `acceleration.refresh_retry_max_attempts`[​](#accelerationrefresh_retry_max_attempts "Direct link to accelerationrefresh_retry_max_attempts") Optional. Defines the maximum number of retry attempts when refresh retries are enabled. The default is undefined, with no upper limit on attempts. ## `acceleration.refresh_on_startup`[​](#accelerationrefresh_on_startup "Direct link to accelerationrefresh_on_startup") Optional. Controls the refresh behavior of an accelerated dataset across restarts. Defaults to `auto`. ### Supported Values[​](#supported-values "Direct link to Supported Values") * **`auto` (Default)** – Maintains refresh state across restarts: * With `refresh_check_interval`: Schedules next refresh based on last successful refresh time, triggering immediately if interval has already elapsed * Without `refresh_check_interval`: No refresh (on-demand only) * **`always`** – Forces a dataset refresh on every startup, regardless of the existing acceleration state. Setting `refresh_on_startup: always` ensures that accelerated data is always refreshed to match the source when the service restarts. This is useful in **development environments** or when **data consistency is critical** after deployment. ## `acceleration.params`[​](#accelerationparams "Direct link to accelerationparams") Optional. Parameters to pass to the acceleration engine. The parameters are specific to the acceleration engine used. ## `acceleration.retention_check_enabled`[​](#accelerationretention_check_enabled "Direct link to accelerationretention_check_enabled") Optional. Enable or disable retention policy check, defaults to `false`. ## `acceleration.retention_period`[​](#accelerationretention_period "Direct link to accelerationretention_period") Optional. The retention period for the dataset. Combine with `time_column` and `time_format` to determine if the data should be retained or not. `retention_period` or `retention_sql` must be specified when `acceleration.retention_check_enabled` is `true`. When both `retention_period` and `retention_sql` are configured, both retention policies will be applied during each retention check. See [Duration](/docs/next/reference/duration) ## `acceleration.retention_sql`[​](#accelerationretention_sql "Direct link to accelerationretention_sql") Optional. Custom SQL statement to define data retention logic. Takes the form of a `DELETE FROM
WHERE ` statement. This parameter is useful for scenarios like soft-deleting rows in append-only datasets or removing data based on complex business logic that goes beyond simple time-based retention. `retention_period` or `retention_sql` must be specified when `acceleration.retention_check_enabled` is `true`. When both `retention_period` and `retention_sql` are configured, both retention policies will be applied during each retention check. ## `acceleration.retention_check_interval`[​](#accelerationretention_check_interval "Direct link to accelerationretention_check_interval") Optional. How often the retention policy should be checked. Required when `acceleration.retention_check_enabled` is `true`. See [Duration](/docs/next/reference/duration) ## `acceleration.refresh_jitter_enabled`[​](#accelerationrefresh_jitter_enabled "Direct link to accelerationrefresh_jitter_enabled") Optional. Enable or disable refresh jitter, defaults to `false`. The refresh jitter adds/substracts a randomized time period from the `refresh_check_interval`. ## `acceleration.refresh_jitter_max`[​](#accelerationrefresh_jitter_max "Direct link to accelerationrefresh_jitter_max") Optional. The maximum amount of jitter to add to the refresh interval. The jitter is a random value between 0 and `refresh_jitter_max`. Defaults to 10% of `refresh_check_interval`. ## `metrics`[​](#metrics "Direct link to metrics") Optional. Enable component-specific metrics for the dataset. Each component can expose its own set of metrics that can be enabled selectively to monitor specific aspects of its operation. Most component metrics are disabled by default and can be enabled by adding a `metrics` section to the dataset configuration. Each metric can be enabled individually by specifying its name in the metrics list. Some metrics are **auto-registered**: they export without any `metrics` configuration, and an entry with `enabled: false` is what turns one off. See [Component Metrics](/docs/next/features/observability/component_metrics) and each component's own documentation for which metrics are auto-registered. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - from: mysql:my_table name: my_dataset metrics: - name: connection_count enabled: true - name: connections_in_pool enabled: true - name: active_wait_requests enabled: true params: mysql_host: localhost mysql_tcp_port: 3306 mysql_user: root mysql_pass: ${secrets:MYSQL_PASS} ``` For detailed information about metrics available for specific components, see the [component metrics documentation](/docs/next/features/observability/component_metrics). ## `acceleration.indexes`[​](#accelerationindexes "Direct link to accelerationindexes") Optional. Specify which indexes should be applied to the locally accelerated table. Not supported for in-memory Arrow acceleration engine. The `indexes` field is a map where the key is the column reference and the value is the index type. A column reference can be a single column name or a multicolumn key. The column reference must be enclosed in parentheses if it is a multicolumn key. See [Indexes](/docs/next/features/data-acceleration/indexes) ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite indexes: number: enabled # Index the `number` column '(hash, timestamp)': unique # Add a unique index with a multicolumn key comprised of the `hash` and `timestamp` columns ``` ## `acceleration.primary_key`[​](#accelerationprimary_key "Direct link to accelerationprimary_key") Optional. Specify the primary key constraint on the locally accelerated table. Not supported for in-memory Arrow acceleration engine. The `primary_key` field is a string that represents the column reference that should be used as the primary key. The column reference can be a single column name or a multicolumn key. The column reference must be enclosed in parentheses if it is a multicolumn key. See [Constraints](/docs/next/features/data-acceleration/constraints) ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash # Define a primary key on the `hash` column ``` ## `acceleration.on_conflict`[​](#accelerationon_conflict "Direct link to accelerationon_conflict") Optional. Specify what should happen when a constraint is violated. Not supported for in-memory Arrow acceleration engine. The `on_conflict` field is a map where the key is the column reference and the value is the conflict resolution strategy. A column reference can be a single column name or a multicolumn key. The column reference must be enclosed in parentheses if it is a multicolumn key. Only a single `on_conflict` target can be specified, unless all `on_conflict` targets are specified with `drop`. The possible conflict resolution strategies are: * `upsert` - Upsert the incoming data when the primary key constraint is violated. * `upsert_dedup` - Same as `upsert`, but also deduplicates the data if there are duplicate rows that trigger a violation constraint within a single update. See [Advanced upsert behavior](/docs/next/features/data-acceleration/constraints#advanced-upsert-options). * `upsert_dedup_by_row_id` - Same as `upsert`, but resolves any violations by arbitrarily choosing the row with the highest row id. See [Advanced upsert behavior](/docs/next/features/data-acceleration/constraints#advanced-upsert-options). * `drop` - Drop the data when the primary key constraint is violated. See [Constraints](/docs/next/features/data-acceleration/constraints) ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash indexes: '(number, timestamp)': unique on_conflict: # Upsert the incoming data when the primary key constraint on "hash" is violated, # alternatively "drop" can be used instead of "upsert" to drop the data update. hash: upsert ``` ## `acceleration.on_zero_results`[​](#accelerationon_zero_results "Direct link to accelerationon_zero_results") Optional. Controls the behavior when an accelerated query returns zero results. Defaults to `return_empty`. The following values are supported: * `return_empty` - Default. Return an empty result set when the accelerated query returns no rows. * `use_source` - Fall back to querying the original data source when the accelerated query returns no rows. ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true on_zero_results: use_source ``` ## `acceleration.partition_by`[​](#accelerationpartition_by "Direct link to accelerationpartition_by") Optional. Specifies columns to partition the accelerated data by, enabling partition-level operations and optimized storage. Defaults to no partitioning (empty). ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true partition_by: block_date ``` ## `columns`[​](#columns "Direct link to columns") Optional. Define metadata, semantic details and features (e.g. embeddings, or table indexes) for specific columns in the dataset. ``` datasets: - from: file:sales_data.parquet name: sales columns: - name: address_line1 description: The first line of the address. embeddings: - from: hf_minilm row_id: order_number chunking: enabled: true target_chunk_size: 256 overlap_size: 32 full_text_search: enabled: true ``` ## `columns[*].name`[​](#columnsname "Direct link to columnsname") The name of the column in the table schema. ## `columns[*].type`[​](#columnstype "Direct link to columnstype") Optional. Declares the expected data type for the column, overriding the type inferred from the data source. Accepts three families of type expressions: * **Arrow display forms** — `Int64`, `Utf8`, `Float64`, `Bool`, `Date32`, `Timestamp(Microsecond, UTC)`, `List`, `Decimal128(p, s)`, `Map`, etc. * **SQL / Postgres forms** — `BIGINT`, `INTEGER`, `TEXT`, `VARCHAR(n)`, `BOOLEAN`, `DOUBLE PRECISION`, `NUMERIC(p,s)`, `DATE`, `TIMESTAMP`, `TIMESTAMP WITH TIME ZONE`, `BYTEA`, etc. * **Postgres aliases** — `int2`, `int4`, `int8`, `float4`, `float8`, `serial`, `bigserial`, `timestamptz`, `uuid`, and the `T[]` array suffix (e.g. `int4[]`). ``` datasets: - from: postgres:public.events name: events columns: - name: event_id type: bigint nullable: false - name: payload type: Map - name: tags type: text[] - name: amount type: numeric(18,4) ``` ## `columns[*].nullable`[​](#columnsnullable "Direct link to columnsnullable") Optional. Declares whether the column allows null values. When omitted, the column defaults to nullable. Set to `false` to enforce a non-null constraint. ## `columns[*].description`[​](#columnsdescription "Direct link to columnsdescription") Optional. A description of the column's contents and purpose. Used as part of the [Semantic Data Model](/docs/next/features/semantic-model). ## `columns[*].embeddings`[​](#columnsembeddings "Direct link to columnsembeddings") Optional. Create vector embeddings for this column. ## `columns[*].embeddings[*].from`[​](#columnsembeddingsfrom "Direct link to columnsembeddingsfrom") The embedding model to use, specify the component name. ## `columns[*].embeddings[*].row_id`[​](#columnsembeddingsrow_id "Direct link to columnsembeddingsrow_id") Optional. For datasets without a primary key, used to explicitly specify column(s) that uniquely identify a row. Specifying a `row_id` enables unique identifier lookups for datasets from external systems that may not have a primary key. ## `columns[*].embeddings[*].chunking`[​](#columns-embeddings-chunking "Direct link to columns-embeddings-chunking") Optional. The configuration to enable and define the chunking strategy for the embedding column. ``` columns: - name: description embeddings: - from: hf_minilm chunking: enabled: true target_chunk_size: 512 overlap_size: 128 trim_whitespace: false ``` See [`embeddings[*].chunking`](#embeddingschunking) for details. ## `columns[*].embeddings[*].vector_size`[​](#columnsembeddingsvector_size "Direct link to columnsembeddingsvector_size") Optional. Specifies the size (number of dimensions) of the embedding vector for use in federated queries to databases that do not support arrays with fixed lengths. ``` columns: - name: review_body embeddings: - from: embed-static-retrieval vector_size: 1024 ``` ## `columns[*].embeddings[*].aggregation`[​](#columnsembeddingsaggregation "Direct link to columnsembeddingsaggregation") Optional. For multi-vector columns (`List` source), the strategy used to combine per-element similarities into a single per-row score during vector search. Only meaningful when the underlying column is list-typed. | Value | Description | | ------ | ---------------------------------------------------------------------------------- | | `max` | ColBERT-style `MaxSim`. Row scores as high as its best-matching element (default). | | `mean` | Average similarity across elements. | | `sum` | Sum of similarities across elements. | See [Multi-Vector Search](/docs/next/features/search/multi-vector) for details. ``` columns: - name: tags embeddings: - from: local_embedding_model aggregation: max ``` ## `columns[*].embeddings[*].max_elements_per_row`[​](#columnsembeddingsmax_elements_per_row "Direct link to columnsembeddingsmax_elements_per_row") Optional. For multi-vector columns, the maximum number of list elements embedded per row. Defaults to `32`; hard-capped at `1024`. Excess elements are dropped with a warning log. ``` columns: - name: tags embeddings: - from: local_embedding_model max_elements_per_row: 128 ``` ## `columns[*].full_text_search`[​](#columns-search-full-text "Direct link to columns-search-full-text") ## `columns[*].full_text_search.enabled`[​](#columnsfull_text_searchenabled "Direct link to columnsfull_text_searchenabled") Optional. Enable or disable full text search support for specific column in the dataset. Default `false`. ## `columns[*].full_text_search.row_id`[​](#columnsfull_text_searchrow_id "Direct link to columnsfull_text_searchrow_id") Optional. For datasets without a primary key, used to explicitly specify column(s) that uniquely identify a row. Specifying a `row_id` enables unique identifier lookups for datasets from external systems that may not have a primary key. ## `columns[*].metadata`[​](#columnsmetadata "Direct link to columnsmetadata") Optional. Specific metadata associated to the column. ## `columns[*].metadata.vectors`[​](#columnsmetadatavectors "Direct link to columnsmetadatavectors") Optional. If provided, a vector engine (see [below](#vectors)) should store this column for a particular use, determined by the value, which is one of: * `non-filterable`: Store the column in the vector engine. * `filterable`: Store the column in the vector engine, and ensure the engine can filter on the column (if possible in the engine). For the vector engine, only applicable if `vectors.enabled` is both defined and `true`. The distinction between the two values is the vector engine's. When the dataset also has a [full-text search index](/docs/next/features/search/full-text), either value carries the column into that index, where it is returned by searches and [filterable](/docs/next/reference/sql/search#full-text-filter-pushdown) — whether or not a vector engine is configured. A column whose type the full-text index cannot represent (a date or timestamp) is skipped there, with a warning logged at startup. ## `embeddings`[​](#embeddings "Direct link to embeddings") Optional. Create vector embeddings for specific columns of the dataset. ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks embeddings: - column: extra_data use: hf_minilm ``` ## `embeddings[*].column`[​](#embeddingscolumn "Direct link to embeddingscolumn") The column name to create an embedding for. ## `embeddings[*].use`[​](#embeddingsuse "Direct link to embeddingsuse") The embedding model to use, specific the component name `embeddings[*].name`. ## `embeddings[*].column_pk`[​](#embeddingscolumn_pk "Direct link to embeddingscolumn_pk") Optional. For datasets without a primary key, explicitly specify column(s) that uniquely identify a row. ## `embeddings[*].chunking`[​](#embeddingschunking "Direct link to embeddingschunking") Optional. The configuration to enable and define the chunking strategy for the embedding column. ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks embeddings: - column: extra_data use: hf_minilm chunking: enabled: true target_chunk_size: 512 overlap_size: 128 trim_whitespace: false ``` ## `embeddings[*].chunking.enabled`[​](#embeddingschunkingenabled "Direct link to embeddingschunkingenabled") Optional. Enable or disable chunking for the embedding column. Defaults to `false`. ## `embeddings[*].chunking.target_chunk_size`[​](#embeddingschunkingtarget_chunk_size "Direct link to embeddingschunkingtarget_chunk_size") The desired size of each chunk, in tokens. If the desired chunk size is larger than the maximum size of the embedding model, the maximum size will be used. ## `embeddings[*].chunking.overlap_size`[​](#embeddingschunkingoverlap_size "Direct link to embeddingschunkingoverlap_size") Optional. The number of tokens to overlap between chunks. Defaults to `0`. ## `embeddings[*].chunking.trim_whitespace`[​](#embeddingschunkingtrim_whitespace "Direct link to embeddingschunkingtrim_whitespace") Optional. If enabled, the content of each chunk will be trimmed to remove leading and trailing whitespace. Defaults to `false`. ## `metadata`[​](#metadata "Direct link to metadata") Optional. Additional key-value metadata for the dataset. The `metadata` field serves two purposes: 1. **Semantic metadata** — Arbitrary key-value pairs used as part of the [Semantic Data Model](/docs/next/features/semantic-model). ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks metadata: instructions: The last 128 blocks. ``` 2. **File metadata columns** — For [file-based connectors](/docs/next/components/data-connectors#metadata-columns) (S3, ABFS, File, FTP, SFTP, SMB, NFS, HTTP/HTTPS), the following reserved keys enable virtual columns that expose per-file object store metadata in query results: | Key | Value | Column Type | Description | | ---------------- | --------- | ---------------------- | ------------------------------- | | `_location` | `enabled` | `Utf8` | Full URI of the source file | | `_last_modified` | `enabled` | `Timestamp(µs, "UTC")` | When the file was last modified | | `_size` | `enabled` | `UInt64` | File size in bytes | ``` datasets: - from: s3://bucket/data/ name: my_data params: file_format: parquet metadata: _location: enabled _last_modified: enabled _size: enabled ``` If a data file already contains a column with the same name as a metadata column, the metadata column is not added. ## `full_text_search`[​](#dataset-full-text-search "Direct link to dataset-full-text-search") Optional. Dataset-level full-text search engine configuration. When absent, the built-in Tantivy in-process engine is used (controlled by column-level [`columns[*].full_text_search`](#columns-search-full-text) settings). ## `full_text_search.enabled`[​](#full_text_searchenabled "Direct link to full_text_searchenabled") Enable or disable the dataset-level FTS engine, defaults to `true`. ## `full_text_search.engine`[​](#full_text_searchengine "Direct link to full_text_searchengine") The full-text search engine to use. Currently only `elasticsearch` is supported. When absent, the built-in Tantivy engine is used. ## `full_text_search.params`[​](#full_text_searchparams "Direct link to full_text_searchparams") Optional. Engine-specific connection and tuning parameters. See [Full-Text Search — Elasticsearch](/docs/next/features/search/full-text#using-elasticsearch-as-the-fts-engine) for available parameters. ``` datasets: - from: file:./articles.parquet name: articles acceleration: enabled: true full_text_search: engine: elasticsearch params: elasticsearch_endpoint: http://localhost:9200 elasticsearch_index: articles-fts columns: - name: body full_text_search: enabled: true row_id: - id ``` ## `vectors`[​](#vectors "Direct link to vectors") ## `vectors.enabled`[​](#vectorsenabled "Direct link to vectorsenabled") Enable or disable vector storage, defaults to `true`. ## `vectors.engine`[​](#vectorsengine "Direct link to vectorsengine") The vector engine to use. The following engines are supported for datasets: * [`s3_vectors`](/docs/next/components/vectors/s3_vectors) - Vectors are created and indexed into [Amazon S3 Vectors](https://aws.amazon.com/s3/features/vectors/). * [`elasticsearch`](/docs/next/components/vectors/elasticsearch) - Vectors are created and indexed into an [Elasticsearch](https://www.elastic.co/) cluster. Dataset and view support may differ. If a view uses `vectors`, refer to the views reference for the engines supported there. ## `vectors.params`[​](#vectorsparams "Direct link to vectorsparams") Optional. Parameters to pass to the vector engine. The parameters are specific to the vector engine used. --- # Embeddings Embeddings convert text or other data into vector representations for machine learning and natural language processing tasks. ## `embeddings`[​](#embeddings "Direct link to embeddings") The `embeddings` section in your configuration specifies one or more embedding models for your datasets. Example: ``` embeddings: - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2:latest name: text_embedder params: max_length: '128' datasets: - my_text_dataset ``` ### `from`[​](#from "Direct link to from") The `from` field specifies the source of the embedding model. It supports the following prefixes: * `huggingface:huggingface.co` - Models from Hugging Face * `file:` - Local file paths * `openai` - OpenAI (or OpenAI-compatible) models * `azure` - Azure OpenAI models * `databricks` - Databricks-hosted models * `bedrock` - Amazon Bedrock models * `google` - Google AI models * `model2vec` - Model2Vec static embedding models Follows the same convention as [`models.from`](/docs/next/reference/spicepod/models#from). ### `name`[​](#name "Direct link to name") A unique identifier for this embedding component. ### `files`[​](#files "Direct link to files") Optional. A list of files associated with this model. Each file has: * `path`: The path to the file * `name`: Optional. A name for the file * `type`: Optional. The type of the file (automatically determined if not specified) Follows the same convention as [`models.files`](/docs/next/reference/spicepod/models#files). ### `params`[​](#params "Direct link to params") Optional. A map of key-value pairs for additional parameters specific to the embedding model. ### `dependsOn`[​](#dependson "Direct link to dependson") Optional. A list of dependencies that must be loaded and available before this embedding model. --- # Evals (Deprecated) Deprecated The `evals` Spicepod component is no longer supported and was removed in Spice v2.0 ([spiceai/spiceai#9420](https://github.com/spiceai/spiceai/pull/9420)). For the evals YAML reference in the last version that supported it, see the [v1.11.x Evals reference](https://docs.spiceai.org/docs/1.11.x/reference/spicepod/evals). --- # Functions (User-Defined Functions) Functions extend Spice's SQL engine with custom scalar and table logic. Each entry in the top-level `functions:` block is registered as a callable SQL function and (for scalar functions, by default) as an LLM tool. For an overview, examples, and execution-tier details, see [Functions](/docs/next/features/functions). Registration is off by default The `functions:` section is only honored when [`runtime.functions.enabled`](/docs/next/reference/spicepod/runtime#runtimefunctions) is set to `true`. ## `functions`[​](#functions "Direct link to functions") The `functions:` section in your configuration declares one or more scalar or table functions. Example: ``` runtime: functions: enabled: true functions: - name: double_it from: sql description: Double a 64-bit integer. volatility: immutable signature: args: - { name: x, type: int64 } returns: int64 body: 'x * 2' ``` ### `name`[​](#name "Direct link to name") A unique identifier for the function. The name is used to register the UDF in the SQL session and as the tool name when surfaced to LLMs. ### `from`[​](#from "Direct link to from") Source URI selecting the execution tier: * `sql` — Inline SQL body executed in-process. * `http://...` / `https://...` — Remote endpoint invoked over HTTP + JSON. Other schemes are rejected at startup. See [Execution Tiers](/docs/next/features/functions#execution-tiers). ### `enabled`[​](#enabled "Direct link to enabled") Optional. Defaults to `true`. Set to `false` to keep the declaration in the spicepod without registering the function — useful for staged rollouts. Disabled functions are also hidden from `list_udfs()` and `GET /v1/functions`. ### `description`[​](#description "Direct link to description") Optional. Free-form description surfaced in `list_udfs()`, `GET /v1/functions`, and the LLM tool registry. ### `kind`[​](#kind "Direct link to kind") Optional. Defaults to `scalar`. | Value | Description | | -------- | ---------------------------------------------------------------------------------------------------------------- | | `scalar` | (default) Returns a single value per row. Called as `SELECT my_fn(x) FROM ...`. | | `table` | Returns multiple rows and columns. Called as `SELECT ... FROM my_fn(x)`. Always SQL-only (`as_tool` is ignored). | ### `volatility`[​](#volatility "Direct link to volatility") Optional. Defaults to `volatile`. Controls how the optimizer treats the function across calls. | Value | Description | | ----------- | ----------------------------------------------------------------------------- | | `immutable` | Same inputs always yield the same output. May be constant-folded and cached. | | `stable` | Stable within a single query but may change across queries. Cached per query. | | `volatile` | (default) Unpredictable on every call. Never cached. | Choose the strongest level that's actually true. See [Volatility](/docs/next/features/functions#volatility). ### `signature`[​](#signature "Direct link to signature") Required. Typed argument list and return type. ``` signature: args: - { name: lat1, type: float64 } - { name: lon1, type: float64 } returns: float64 ``` #### `signature.args`[​](#signatureargs "Direct link to signatureargs") Optional. Positional argument list. Empty for niladic functions. Each entry has: * `name` — Argument name. Used inside SQL `body` expressions and as the JSON key for remote tier requests. * `type` — Arrow type. Accepts Spicepod aliases (`int64`, `utf8`, `list`, `decimal(38,10)`, `timestamp(us, utc)`) and Arrow display forms (`Int64`, `List(Int64)`, etc.). See [Types](/docs/next/features/functions#types). #### `signature.returns`[​](#signaturereturns "Direct link to signaturereturns") Required. For **scalar functions**, a single Arrow type string (e.g., `int64`, `utf8`). For **table functions**, a list of output column definitions: ``` # Scalar function returns: int64 # Table function returns: - { name: value, type: int64 } - { name: label, type: utf8 } ``` ### `body`[​](#body "Direct link to body") Inline SQL body. Required when `from: sql` (unless `body_ref` is set instead). For **scalar functions**, must be a single SQL expression referencing the function's arguments by name. For **table functions**, must be a single `SELECT` query; scalar arguments are available via a virtual `args` table. ``` body: | 6371 * acos( cos(radians(lat1)) * cos(radians(lat2)) * cos(radians(lon2) - radians(lon1)) + sin(radians(lat1)) * sin(radians(lat2)) ) ``` The body can call any DataFusion built-in scalar function (math, string, datetime, JSON, regex, etc.). A table function body may also call the [search functions](/docs/next/reference/sql/search) — `vector_search()`, `text_search()`, or either nested in `rrf()` — with one of the function's own `utf8` arguments as the search query. Those arguments are substituted as literals before the body is planned rather than being read from the `args` table; see [Wrapping search functions](/docs/next/features/functions#wrapping-search-functions). Mutually exclusive with `body_ref`. Must not be set for non-SQL `from:` schemes. ### `body_ref`[​](#body_ref "Direct link to body_ref") Path to a local filesystem file whose contents are the SQL body. Resolved relative to the runtime's current working directory at registration time. ``` body_ref: ./functions/shipping_class.sql ``` `body_ref` is not read from object stores when a spicepod is loaded remotely — use inline `body:` for portable remote spicepods. Mutually exclusive with `body`. ### `params`[​](#params "Direct link to params") Optional. Tier-specific parameters. Supports `${secrets:KEY}` and `${env:KEY}` interpolation at registration time. For the **Remote tier** (`from: http://...` / `https://...`): | Parameter | Default | Description | | ------------------- | ------- | ------------------------------------------------------------------------------------- | | `timeout` | `30s` | Per-call timeout. Plain integer seconds or `Ns` / `Nms` strings. | | `batch_size` | `1024` | Maximum rows per HTTP request. Capped at `100 000`. | | `batch_concurrency` | `4` | Maximum in-flight HTTP batches per invocation. Capped at `64`. | | `auth_bearer` | unset | When set, adds `Authorization: Bearer ` to each request. Use `${secrets:...}`. | The SQL tier does not currently use `params`. ### `metadata`[​](#metadata "Direct link to metadata") Optional. Free-form key/value pairs surfaced alongside the function definition. Useful for tagging, ownership, or custom metadata consumers. ``` metadata: owner: data-platform jira: DATA-1234 ``` ### `as_tool`[​](#as_tool "Direct link to as_tool") Optional. Defaults to `true` for scalar functions. When `true`, the function is registered as an LLM tool with the same name and description and becomes callable from chat completions, `POST /v1/tools/`, and the `/v1/tools` listing. Table functions (`kind: table`) are always SQL-only regardless of this setting. Set to `false` to keep the function SQL-only: ``` functions: - name: internal_hash from: sql as_tool: false signature: args: [{ name: x, type: int64 }] returns: int64 body: 'x * 2654435761' ``` ### `dependsOn`[​](#dependson "Direct link to dependson") Optional. Names of other Spicepod components that must be loaded before this function (e.g. a dataset the function queries internally). ### `metrics`[​](#metrics "Direct link to metrics") Optional. Per-function metrics configuration. See [Component Metrics](/docs/next/features/observability/component_metrics). ## Discovering registered functions[​](#discovering-registered-functions "Direct link to Discovering registered functions") After startup, registered functions can be inspected: * **From SQL** — `SELECT * FROM list_udfs() WHERE source = 'user';` * **From the HTTP API** — `GET /v1/functions` returns a JSON array of user functions. See [Discovering registered functions](/docs/next/features/functions#discovering-registered-functions). --- # Reserved Keywords The following keywords cannot be used as names for datasets. ## General Protected Keywords[​](#general-protected-keywords "Direct link to General Protected Keywords") These keywords apply to all data connectors: * `COUNT` * `FALSE` * `NULL` * `TRUE` * `END-EXEC` * `LATERAL` * `TABLE` * `UNNEST` ## Connector-Specific Protected Keywords[​](#connector-specific-protected-keywords "Direct link to Connector-Specific Protected Keywords") These data connectors have reserved keywords beyond the ones mentioned above. * [ClickHouse Keywords](/docs/next/components/data-connectors/clickhouse#name) * [Snowflake Keywords](/docs/next/components/data-connectors/snowflake#name) * [Microsoft SQL Server Keywords](/docs/next/components/data-connectors/mssql#name) * [MySQL Keywords](/docs/next/components/data-connectors/mysql#name) --- # Models The `models` section of a Spicepod defines large language models (LLMs) for use with Spice. Models can be loaded from Hugging Face, OpenAI, local files, or other supported providers. | Field | Description | | ------------- | ------------------------------------------------------------------------ | | `name` | Unique, readable name for the model within the Spicepod. | | `from` | Source-specific address to uniquely identify a model. | | `description` | Additional details about the model, useful for displaying to users. | | `datasets` | Datasets the model's tools may access, forming a table allowlist. | | `files` | Specify additional files, or override default files needed by the model. | | `params` | Additional parameters to be passed to the model. | ## `models`[​](#models "Direct link to models") The `models` section in your configuration specifies one or more models to be used with your datasets. Example: ``` models: - from: huggingface:huggingface.co/gpt4:latest name: text_generator files: - path: model.safetensors type: weights - path: config.json type: config - path: tokenizer.json type: tokenizer params: max_length: '128' datasets: - my_text_dataset ``` ### `from`[​](#from "Direct link to from") The `from` field specifies both the source of the model (e.g Huggingface, or a local file), and the unique identifier of the model (relative to the source). The `from` value expects the following format ``` - from: / ``` #### Model Source[​](#model-source "Direct link to Model Source") The `` prefix of the `from` field indicates where the model is sourced from: * `huggingface:huggingface.co` - Models from Hugging Face * `file:` - Local file paths * `openai` - OpenAI (or compatible) models * `spiceai` - Spice AI models #### Model ID[​](#model-id "Direct link to Model ID") The `` suffix of the `from` field is a unique (per source) identifier for the model: * For Spice AI: The identifier of a model served by the [Spice.ai Cloud Platform](/docs/next/components/models/spiceai) (or by another Spice runtime), in the form `/`. * Example: `spice.ai:openai/gpt-4o` * For Hugging Face: A repo\_id and, optionally, revision hash or tag. * `Qwen/Qwen1.5-0.5B` (no revision) * `meta-llama/Meta-Llama-3-8B:cd892e8f4da1043d4b01d5ea182a2e8412bf658f` (with revision hash) * For local files: Represents the absolute or relative path to the model weights file on the local file system. See [below](#files) for the accepted model weight types and formats. * For OpenAI: Only supports LMs. For OpenAI models, valid IDs can be found in their model [documentation](https://platform.openai.com/docs/models/continuous-model-upgrades). For OpenAI compatible providers, specify the value required in their `v1/chat/completion` [payload](https://platform.openai.com/docs/api-reference/chat/create#chat-create-model). ### `name`[​](#name "Direct link to name") A unique identifier for this model component. ### `description`[​](#description "Direct link to description") Additional details about the model, useful for displaying to users ### `files`[​](#files "Direct link to files") Optional. A list of files associated with this model. Each file has: * `path`: The path to the file * `name`: Optional. A name for the file * `type`: Optional. The type of the file (automatically determined if not specified) File types include: * `weights`: Model weights * For LLMs: `.gguf`, `.ggml`, or `.safetensors` files * Pickle-based checkpoints (`.bin`, `.pt`, `.pth`, `.ckpt`) execute arbitrary code when loaded and are **rejected by default**. Set [`trust_pickle: true`](/docs/next/components/models/filesystem#params-optional) to load them from a fully trusted source. * These files contain the trained parameters of the model * `config`: Model configuration * Usually a `config.json` file * Contains model architecture and hyperparameters * `tokenizer`: Tokenizer file * Usually a `tokenizer.json` file * Defines how input text is converted into tokens for the model * `tokenizer_config`: Tokenizer configuration * Usually a `tokenizer_config.json` file * Contains additional configuration for the tokenizer The system attempts to automatically determine the file type based on the file name and extension. If the type cannot be determined automatically, you can explicitly specify it in the configuration. ### `params`[​](#params "Direct link to params") Optional. A map of key-value pairs for additional parameters specific to the model. Example uses include: * Setting default OpenAI request parameters for language models, see [parameter overrides](/docs/next/features/large-language-models/parameter_overrides). * Enabling language models to perform actions against Spice (e.g. making SQL queries), via language model tool use, see [runtime tools](/docs/next/features/large-language-models/tools). * Invoking language models directly from SQL queries using the [`ai()` function](/docs/next/reference/sql/scalar_functions#ai-and-embed). #### `params.tools`[​](#paramstools "Direct link to paramstools") Which tools should be made available to the model. Supported values: `auto`, `all`, `search_registry`, or a comma-separated list of specific tool names. See [Tool Modes](/docs/next/features/large-language-models/tools#tool-modes). #### `params.tool_embedding_model`[​](#paramstool_embedding_model "Direct link to paramstool_embedding_model") The name of an embedding model (defined in the `embeddings` section) to use for searchable tool discovery. Required when `tools: search_registry` is set. When `tools: auto` is used, this model enables registry-based discovery if the tool count exceeds the auto-search threshold (20 tools). If only one embedding model is configured, it is used automatically. #### `params.prompt_cache_key`[​](#paramsprompt_cache_key "Direct link to paramsprompt_cache_key") Optional. A stable key forwarded to the LLM provider to enable prompt/prefix caching. When set, Spice maps this key into the provider-native caching mechanism: | Provider | Behavior | | -------------------------- | --------------------------------------------------------------------------------------- | | OpenAI / Azure OpenAI | Passed through on Chat and Responses API requests | | Anthropic | Adds `cache_control: { type: "ephemeral" }` to the request | | Google Gemini | Maps to `cached_content.name` (must be a valid cached-content resource name) | | xAI (Grok) | Sent as the `x-grok-conv-id` HTTP header | | AWS Bedrock (Converse) | Appends a native `CachePoint` block | | Databricks (hosted Claude) | Adds Claude-style `cache_control` to the last content part | | Local (mistral-rs) | Paged-attention scheduling is enabled automatically on supported backends (CUDA + Unix) | ``` models: - from: openai:gpt-4o name: my_model params: prompt_cache_key: "schema-context" ``` #### `params.prompt_cache_retention`[​](#paramsprompt_cache_retention "Direct link to paramsprompt_cache_retention") Optional. Retention hint for prompt caching, applicable to the OpenAI Responses API only. For example, `"24h"` requests that the cached content be retained for 24 hours. ``` models: - from: openai:gpt-4o name: my_model params: prompt_cache_key: "schema-context" prompt_cache_retention: "24h" ``` ### `datasets`[​](#datasets "Direct link to datasets") Optional. A list of [dataset names](/docs/next/reference/spicepod/datasets#name) that scope the model's tool access, forming a table allowlist for SQL and NSQL tool use. When omitted, the model's tools are not restricted to a specific set of datasets. ### `dependsOn`[​](#dependson "Direct link to dependson") Optional. A list of dependencies that must be loaded and available before this model. --- # Runtime The `runtime` section specifies configuration settings for the Spice runtime. ## Reload Behavior[​](#reload-behavior "Direct link to Reload Behavior") Editing a Spicepod on disk reloads it into the running process, but **most of `runtime.*` is consumed once when `spiced` starts and cannot be rebuilt in place**. A reload installs the new value in the app while the process keeps running the old one; `spiced` logs a warning naming each start-time-only section that changed and telling you to restart. Restart `spiced` to apply them. **Applied when `spiced` starts — a reload logs a warning and the previous value stays in effect:** `runtime.auth` · `runtime.caching` · `runtime.cors` · `runtime.cpu` · `runtime.dataset_load_parallelism` · `runtime.mcp` · `runtime.metrics` · `runtime.output_level` · `runtime.query` · `runtime.ready_state` · `runtime.scheduler` · `runtime.task_history` · `runtime.telemetry` · `runtime.tls` · `runtime.tracing` **Applied when `spiced` starts, except for components the same reload recreates** — a connector rebuilt by the reload reads the new value, while the process-wide use of it does not change until a restart: `runtime.flight` · `runtime.params` · `runtime.source_rate_control` **Applied on reload — no restart needed:** | Setting | Why it applies | | -------------------------------------------- | ------------------------------------------------------ | | `runtime.shutdown_timeout` | Read from the current app when the runtime shuts down. | | `runtime.functions` | Reconciled by the function diff on each reload. | | `runtime.caching.sql_results.cache_key_type` | Resolved per request from the live app. | | `runtime.query.timeout` | Resolved per request from the live app. | | `runtime.telemetry.user_agent_collection` | Resolved per request from the live app. | The three per-request settings sit inside otherwise start-time-only sections. Changing one of them alone takes effect on the next request and is not reported as requiring a restart; changing any *other* key in the same section is. note `runtime.tls` certificate and CA **files** are separately hot-reloadable without a restart — see [Certificate Hot-Reload](#certificate-hot-reload). Inline PEM material is loaded once at startup. ## `runtime.auth`[​](#runtimeauth "Direct link to runtimeauth") ### `runtime.auth.api-key`[​](#runtimeauthapi-key "Direct link to runtimeauthapi-key") Spice supports adding optional authentication to its API endpoints via configurable API keys. [Learn more](/docs/next/api/auth). ``` runtime: auth: api-key: enabled: true keys: - ${ secrets:api_key } # Use the secret replacement syntax to load the API key from a secret store - 1234567890 # Or specify the API key directly ``` API key authentication supports the following configuration parameters: | Parameter name | Optional | Default | Description | | -------------- | -------- | ------- | ------------------------------------------------------------- | | `enabled` | Yes | `true` | Defaults to `true`. Whether API key authentication is enabled | | `keys` | Yes | `[]` | A list of API keys used to authenticate requests. | ## `runtime.dataset_load_parallelism`[​](#runtimedataset_load_parallelism "Direct link to runtimedataset_load_parallelism") This setting specifies the maximum number of datasets that can be loaded in parallel during startup. By default, the number of parallel datasets is unlimited. ## `runtime.caching`[​](#runtimecaching "Direct link to runtimecaching") This setting specifies cache settings for supported Runtime components: * `sql_results`: Specifies cache settings for results from SQL queries. * `search_results`: Specifies cache settings for results from searches. * `embeddings`: Specifies cache settings for embeddings requests. Runtime caches support common configuration parameters: | Parameter name | Optional | Default | Description | | ------------------- | -------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `enabled` | Yes | `true` | Defaults to `true`. | | `max_size` | Yes | `128MiB` | Maximum cache size. Defaults to `128MiB`. | | `eviction_policy` | Yes | `lru` | Cache replacement policy when the cache reaches `max_size`. Defaults to `lru`. Supports `lru` (Least Recently Used) and `tiny_lfu` (Tiny Least Frequently Used, higher hit rate for skewed access patterns). | | `item_ttl` | Yes | `1s` | Cache entry expiration duration (Time to Live). Defaults to 1 second. | | `hashing_algorithm` | Yes | `xxh3` | Selects which hashing algorithm is used to hash the cache keys when storing the results. Defaults to `xxh3`. Supports `xxh3`, `ahash`, `siphash`, `blake3`, `xxh32`, `xxh64`, or `xxh128`. | ### `runtime.caching.search_results`[​](#runtimecachingsearch_results "Direct link to runtimecachingsearch_results") The search results cache section specifies runtime search cache configuration. [Learn more](/docs/next/features/caching). ``` runtime: caching: search_results: enabled: true max_size: 128MiB item_ttl: 1s ``` The search results cache supports the common cache configuration parameters. ### `runtime.caching.embeddings`[​](#runtimecachingembeddings "Direct link to runtimecachingembeddings") The embeddings cache section specifies runtime embeddings requests cache configuration. [Learn more](/docs/next/features/caching). ``` runtime: caching: embeddings: enabled: true max_size: 128MiB item_ttl: 1s ``` The embeddings cache supports the common cache configuration parameters. ### `runtime.caching.sql_results`[​](#runtimecachingsql_results "Direct link to runtimecachingsql_results") The SQL results cache section specifies runtime SQL query cache configuration. [Learn more](/docs/next/features/caching). ``` runtime: caching: sql_results: enabled: true max_size: 128MiB item_ttl: 1s ``` In addition to the common cache configuration parameters, `sql_results` also supports the following parameters: | Parameter name | Optional | Default | Description | | ---------------------------- | -------- | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cache_key_type` | Yes | `plan` | Determines how cache keys are generated. Defaults to `plan`. `plan` uses the query's logical plan, while `sql` uses the raw SQL query string. | | `encoding` | Yes | `none` | Compression algorithm for cached results. Defaults to `none`. Supports `none` or `zstd`. | | `stale_while_revalidate_ttl` | Yes | `0s` | Duration to serve stale cache entries while revalidating in the background. When set to a non-zero value, expired cache entries continue to be served while a background refresh occurs. Defaults to `0s` (disabled). | info `runtime.results_cache` has been deprecated and will be removed in a future release. If `runtime.results_cache` is specifed in the spicepod it will override the `runtime.caching.sql_results` settings if it is not defined. #### Choosing a `cache_key_type`[​](#choosing-a-cache_key_type "Direct link to choosing-a-cache_key_type") * **`plan` (Default):** Uses the query's logical plan as the cache key. Matches semantically equivalent queries but requires query parsing. * **`sql`:** Uses the raw SQL string as the cache key. Provides faster lookups but requires exact string matches. Queries with dynamic functions, such as `NOW()`, may produce unexpected results. Use `sql` only when results are predictable. Use `sql` for the lowest latency with identical queries that do not include dynamic functions. Use `plan` for greater flexibility. ### Choosing a `hashing_algorithm`[​](#choosing-a-hashing_algorithm "Direct link to choosing-a-hashing_algorithm") * **`xxh3` (Default):** Uses the [XXH3](https://cyan4973.github.io/xxHash/) algorithm for hashing the cache keys. XXH3 is a fast, non-cryptographic hash algorithm that provides high performance and good distribution. It is suitable for scenarios where speed is critical and cryptographic security is not required. * **`siphash`:** Uses the SipHash1-3 algorithm for hashing the cache keys, the [default hashing algorithm of Rust](https://github.com/rust-lang/rust/commit/db1b1919baba8be48d997d9f70a6a5df7e31612a). This hashing algorithm is a secure algorithm that implements verified protections against ["hash flooding"](https://v8.dev/blog/hash-flooding) denial of service (DoS) attacks. Reasonably performant, and provides a high level of security. * **`ahash`:** Uses the [AHash](https://github.com/tkaitchuck/ahash) algorithm for hashing the cache keys. The AHash algorithm is a [high quality](https://github.com/tkaitchuck/aHash/blob/master/compare/readme#Quality) hashing algorithm, and has claimed resistance against hashing DoS attacks. AHash has higher performance than SipHash1-3, especially when used with `cache_key_type: plan`. * **`blake3`:** Uses the [BLAKE3](https://github.com/BLAKE3-team/BLAKE3) cryptographic hash function. BLAKE3 is a fast, parallelizable hash function that provides cryptographic security while maintaining high performance. It is suitable for scenarios requiring both speed and cryptographic guarantees. * **`xxh32`, `xxh64`, `xxh128`:** Variants of the XXH hashing algorithm with different output sizes. These algorithms offer a balance between speed and collision resistance, with larger hash sizes providing better collision resistance at the cost of performance. Use `xxh3` (the default) for its superior speed in most scenarios. Use `ahash`, `xxh64` or `xxh128` for reduced collision probability when caching a large number of queries. Use `blake3` when cryptographic security is required. Use `siphash` when protection against hash flooding attacks is a priority. ## `runtime.params`[​](#runtimeparams "Direct link to runtimeparams") Optional. Global key-value parameters for the runtime. ### HTTP Rate Control[​](#http-rate-control "Direct link to HTTP Rate Control") HTTP-based connectors (HTTP/HTTPS, GraphQL, GitHub) support the following rate control defaults: | Parameter Name | Description | | -------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `http_max_concurrent_requests` | Default maximum concurrent HTTP requests per upstream origin. Can be overridden per-dataset with `max_concurrent_requests`. | | `http_requests_per_second_limit` | Default maximum HTTP requests per second per upstream origin. Can be overridden per-dataset with `requests_per_second_limit`. | | `http_requests_per_minute_limit` | Default maximum HTTP requests per minute per upstream origin. Can be overridden per-dataset with `requests_per_minute_limit`. | | `http_rate_control_jitter_min` | Default minimum random delay before HTTP requests when rate control is active. Defaults to `5ms` when a rate limit is configured. Can be overridden per-dataset. | | `http_rate_control_jitter_max` | Default maximum random delay before HTTP requests when rate control is active. Defaults to `10ms` when a rate limit is configured. Can be overridden per-dataset. | ``` runtime: params: http_max_concurrent_requests: 10 http_requests_per_second_limit: 5 http_requests_per_minute_limit: 200 ``` ### Spatial SQL Functions (opt-in)[​](#spatial-sql-functions-opt-in "Direct link to Spatial SQL Functions (opt-in)") PostGIS-style spatial `ST_*` SQL functions (via [`geodatafusion`](https://github.com/datafusion-contrib/geodatafusion)) can be optionally registered with the SQL engine. | Parameter Name | Description | | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `geo` | Set to `enabled` to register `ST_*` spatial functions. Requires a `spiced` binary built with the `geo` Cargo feature (`cargo build -p spiced --features geo`). Unset by default. | Both gates must be satisfied: the binary must be built with `--features geo` **and** `runtime.params.geo: enabled` must be set in the Spicepod. Standard distributions of `spiced` do not include the `geo` feature, so spatial functions remain unregistered unless you produce a custom build. ``` runtime: params: geo: enabled ``` ``` SELECT ST_AsText(ST_Point(0.0, 0.0)) AS geom; -- POINT(0 0) ``` ### Spice Cayenne (engine-global)[​](#spice-cayenne-engine-global "Direct link to Spice Cayenne (engine-global)") Engine-global tuning for the [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) data accelerator. These apply to every Cayenne-accelerated dataset in the instance and are **not** valid under a dataset's `acceleration.params` (per-dataset Cayenne parameters are documented on the [Cayenne accelerator page](/docs/next/components/data-accelerators/cayenne#acceleration-parameters-accelerationparams)). | Parameter Name | Description | | ----------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cayenne_footer_cache_mb` | Size of the engine-wide in-memory Vortex footer cache in megabytes, shared across all Cayenne datasets. Optional; when unset, DataFusion's default file-metadata-cache limit of 50 MB applies (there is no fixed 128 MB default). | | `cayenne_filter_propagation` | Enables Cayenne's filter-propagation optimizer rules. Accepts `enabled` or `disabled`; defaults to `disabled`. | | `cayenne_optimizer_rules` | Selects which Cayenne optimizer rules run. Accepts `auto` (default), `all`, `none` / `disabled`, or a comma-separated list of rule names. | | `cayenne_compaction_memory_fraction` | Fraction of the query memory pool reserved for the dedicated Cayenne compaction pool. Defaults to `0.2` (clamped to a supported range). Applied only when an enabled Cayenne dataset can accumulate files to compact — a file acceleration `mode` on a refresh mode that is not a whole-table replace — and dedicated thread pools are not disabled. A pod whose Cayenne datasets are all `refresh_mode: full`, or all `mode: memory`, carves no compaction pool. | | `cayenne_sort_merge_min_rows` | Advanced anti-join tuning: row-count threshold above which filter propagation switches to a sort-merge strategy. Internally tuned default. | | `cayenne_sort_merge_memory_pool_fraction` | Advanced anti-join tuning: fraction of the memory pool the sort-merge anti-join strategy may use. Internally tuned default. | ``` runtime: params: cayenne_footer_cache_mb: 512 cayenne_filter_propagation: enabled ``` ## `runtime.source_rate_control`[​](#runtimesource_rate_control "Direct link to runtimesource_rate_control") Optional. Configures how Spice limits outbound requests to upstream data sources, and optionally enables cluster-wide coordination through persisted state in object storage. Without `state_location`, rate limits are local to each Spice instance. When `state_location` is set, Spice instances coordinate through object storage so that a configured limit is shared across the cluster. For example, `requests_per_second_limit: 20` means approximately 20 RPS total across all replicas, not 20 RPS per replica. ``` runtime: source_rate_control: state_location: s3://my-bucket/spice/rate-control/ refresh_interval: 30s params: s3_region: us-west-2 s3_key: ${ secrets:AWS_ACCESS_KEY_ID } s3_secret: ${ secrets:AWS_SECRET_ACCESS_KEY } github_concurrent_connections_limit: 10 ``` | Parameter Name | Optional | Default | Description | | ------------------------------------- | -------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `state_location` | Yes | - | Root URI for globally persisted rate-control state (e.g. `s3://bucket/path/`). Enables cluster-wide rate control when set. Without this, limits are local to each Spice instance. | | `params` | Yes | - | Object-store authentication parameters for `state_location`. Supports the same keys as other object-store configurations (e.g. `s3_region`, `s3_key`, `s3_secret` for S3; `account`, `access_key` for Azure). Supports `${ secrets:NAME }` references. | | `refresh_interval` | Yes | `30s` | How often each instance refreshes and persists per-source rate-control state. Longer intervals reduce object-store writes but adapt more slowly to demand changes. | | `github_concurrent_connections_limit` | Yes | `10` | Maximum number of concurrent GitHub HTTP requests per authentication context. Replaces the deprecated `runtime.params.github_max_concurrent_connections`. | HTTP/API rate limits are configured through [`runtime.params`](#runtimeparams) (cluster defaults) and per-dataset overrides. Precedence is: ``` dataset param > runtime.params.http_* default > unset ``` When `state_location` is set, the configured RPS/RPM quota is converted into a token budget per lease window and distributed across replicas using a demand-weighted leased token-bucket model. ## `runtime.functions`[​](#runtimefunctions "Direct link to runtimefunctions") Controls whether [functions](/docs/next/features/functions) declared in the top-level `functions:` section (and `tools:` entries with `as_sql: true`) are registered with the SQL engine. Defaults to disabled. ``` runtime: functions: enabled: true ``` | Parameter | Optional | Default | Description | | --------- | -------- | ------- | ----------------------------------------------------------------------------------------------------- | | `enabled` | Yes | `false` | When `true`, the runtime registers `functions:` entries and exposes them via SQL and `/v1/functions`. | When disabled, the `functions:` block is parsed but not registered, `list_udfs()` returns no `user`-source rows, and `GET /v1/functions` returns an empty array. See the [Functions Spicepod reference](/docs/next/reference/spicepod/functions) for the function declaration schema. ## `runtime.shutdown_timeout`[​](#runtimeshutdown_timeout "Direct link to runtimeshutdown_timeout") Controls how long Spice waits for connections to be gracefully drained and for components to shut down cleanly during runtime termination. Defaults to 30 seconds. ``` runtime: shutdown_timeout: 1m ``` ## `runtime.tls`[​](#runtimetls "Direct link to runtimetls") The TLS section specifies the configuration for enabling Transport Layer Security (TLS) for all endpoints exposed by the runtime. [Learn more about enabling TLS](/docs/next/api/tls). In addition to configuring TLS via the manifest, TLS can also be configured via `spiced` command line arguments using the `--tls-enabled true` flag along with `--tls-certificate`/`--tls-certificate-file` and `--tls-key`/`--tls-key-file`. ### Certificate Hot-Reload[​](#certificate-hot-reload "Direct link to Certificate Hot-Reload") Spice can hot-reload TLS certificates and client CA files for runtime endpoints. Update the certificate, key, or CA file on disk, then send `SIGHUP` to the Spice process to reload without restart. Only file-based certificates/keys/CA are hot-reloaded (not inline PEM). Existing connections are not interrupted; only new connections use the updated files. If reload fails, the previous certificate remains active and a warning is logged. **Steps:** 1. Replace the certificate/key/CA file on disk. 2. Send `SIGHUP` to the Spice process (e.g., `kill -SIGHUP `). 3. Check logs for reload confirmation or errors. ### `runtime.tls.enabled`[​](#runtimetlsenabled "Direct link to runtimetlsenabled") Enables or disables TLS for the runtime endpoints. ``` runtime: tls: ... enabled: true # or false ``` ### `runtime.tls.certificate`[​](#runtimetlscertificate "Direct link to runtimetlscertificate") The TLS certificate to use for securing the runtime endpoints. The certificate can also come from [secrets](/docs/next/components/secret-stores). ``` runtime: tls: certificate: | -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ``` runtime: tls: ... certificate: ${secrets:tls_cert} ``` ### `runtime.tls.certificate_file`[​](#runtimetlscertificate_file "Direct link to runtimetlscertificate_file") The path to the TLS PEM-encoded certificate file. Only one of `certificate` or `certificate_file` must be used. ``` runtime: tls: certificate_file: /path/to/cert.pem ``` ### `runtime.tls.key`[​](#runtimetlskey "Direct link to runtimetlskey") The TLS key to use for securing the runtime endpoints. The key can also come from [secrets](/docs/next/components/secret-stores). ``` runtime: tls: key: | -----BEGIN PRIVATE KEY----- (private key contents) -----END PRIVATE KEY----- ``` ``` runtime: tls: ... key: ${secrets:tls_key} ``` ### `runtime.tls.key_file`[​](#runtimetlskey_file "Direct link to runtimetlskey_file") The path to the TLS PEM-encoded key file. Only one of `key` or `key_file` must be used. ``` runtime: tls: key_file: /path/to/key.pem ``` ### `runtime.tls.client_auth_mode`[​](#runtimetlsclient_auth_mode "Direct link to runtimetlsclient_auth_mode") Enterprise Feature mTLS (client certificate authentication) is included in the Enterprise distribution of Spice.ai. [Learn more](https://docs.spice.ai/docs/enterprise). Controls whether the runtime requires, requests, or ignores client certificates on its public endpoints (HTTP, Flight, Metrics). Defaults to `none`. | Mode | Behavior | | ------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `none` *(default)* | Standard one-way TLS. No client certificate is requested. | | `request` | The server sends a `CertificateRequest` but accepts connections without a certificate. Presented certificates are verified against the configured CA. Useful for migration or audit-only deployments. | | `required` | A valid client certificate is required. The Flight (gRPC) listener rejects connections without a certificate at the TLS handshake. The HTTP listener admits no-cert connections so `/health` and `/v1/ready` remain accessible for Kubernetes probes, but all other HTTP endpoints return 401 without a verified client certificate. The metrics listener has no client-auth gate. | Requires `client_auth_ca_file` or `client_auth_ca` to be set when mode is `request` or `required`. ``` runtime: tls: enabled: true certificate_file: /path/to/cert.pem key_file: /path/to/key.pem client_auth_mode: required client_auth_ca_file: /path/to/client-ca.pem ``` ### `runtime.tls.client_auth_ca_file`[​](#runtimetlsclient_auth_ca_file "Direct link to runtimetlsclient_auth_ca_file") Path to a PEM-encoded CA bundle used to verify client certificates. The file is watched for changes and reloaded atomically alongside the server certificate and key. ``` runtime: tls: client_auth_ca_file: /path/to/client-ca.pem ``` ### `runtime.tls.client_auth_ca`[​](#runtimetlsclient_auth_ca "Direct link to runtimetlsclient_auth_ca") Inline PEM (or `${ secrets:... }`) form of the client CA bundle. Mutually exclusive with `client_auth_ca_file`. Inline material is loaded once at startup and is not hot-reloaded. ``` runtime: tls: client_auth_ca: | -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ## `runtime.task_history`[​](#runtimetask_history "Direct link to runtimetask_history") The task history section specifies runtime task history configuration. For more details, see the [Task History documentation](/docs/next/reference/task_history). ``` runtime: task_history: enabled: true captured_output: none retention_period: 8h retention_check_interval: 15m min_sql_duration: 5s ``` | Parameter name | Optional | Description | | -------------------------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `enabled` | Yes | Defaults to `true`. | | `captured_output` | Yes | Specifies the level of output captured by the task history table. Defaults to `none`. | | `captured_context` | Yes | Controls how much of the `input` and `captured_output` payload is stored for AI and search tasks. Options: `truncated` (default), `redacted`, or `full`. Values are case-sensitive. | | `captured_plan` | Yes | Controls SQL query plan capture. Options: `none` (default), `explain`, or `explain analyze`. Query plans are captured asynchronously after query completion. | | `min_sql_duration` | Yes | Minimum query execution duration before a plan is captured. Only queries exceeding this threshold are captured. Example: `5s`. | | `min_plan_duration` | Yes | Minimum plan execution duration before a plan is captured. This threshold applies to the execution time of the `EXPLAIN` operation itself. Example: `10s`. | | `retention_period` | Yes | Specifies how long records in the task history table are retained. Defaults to `8h` (8 hours). | | `retention_check_interval` | Yes | Specifies how often old records are checked for removal. Defaults to `15m` (15 minutes). | ## `runtime.cors`[​](#runtimecors "Direct link to runtimecors") The CORS section specifies the configuration for enabling Cross-Origin Resource Sharing (CORS) for the HTTP endpoint. By default, CORS is disabled. Default configuration: ``` runtime: cors: enabled: false ``` ### `runtime.cors.enabled`[​](#runtimecorsenabled "Direct link to runtimecorsenabled") Enables or disables CORS for the HTTP endpoint. Defaults to `false`. ### `runtime.cors.allowed_origins`[​](#runtimecorsallowed_origins "Direct link to runtimecorsallowed_origins") A list of allowed origins for CORS requests. Defaults to `["*"]`, which permits all origins. Example: ``` runtime: cors: enabled: true allowed_origins: ['https://example.com'] ``` This configuration permits requests only from the `https://example.com` origin. ## `runtime.cpu`[​](#runtimecpu "Direct link to runtimecpu") The CPU section states how many CPUs the runtime should behave as though it has. That single entitlement sizes every CPU-derived pool coherently — the tokio runtimes' worker threads, DataFusion's query fan-out (`runtime.query.target_partitions`) and query admission bound (`runtime.query.max_concurrent_queries`), the Cayenne encode, compaction, upload and file-scan concurrency defaults, the Cayenne SQLite metastore pool, the embedding inference pool, DuckDB's per-instance `threads`, and a cluster executor's concurrent-task advertisement. ### `runtime.cpu.cores`[​](#runtimecpucores "Direct link to runtimecpucores") ``` runtime: cpu: cores: 4 # `auto` (the default) detects it ``` Accepts a Kubernetes CPU quantity — a whole number of cores (`4`), a fraction (`3.5`), or millicores (`3500m`) — or one of two named values: | Value | Meaning | | ---------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `auto` | The default. Detect the entitlement. On a pod that declares a CPU request and no CPU limit, that is **twice the request** — `min(max(2 cores, request × 2), available CPUs)`. The multiple is deliberate: it exceeds the request so the pod can still burst above its scheduling floor. | | `all` | Every CPU this process may use, regardless of any CPU request. A CPU **limit**, if set, still applies. | | A quantity | Use exactly this, whatever the pod declares. | Millicores must be integral, so `3500m` is valid and `3.5m` is not. A value of `0`, a negative value, or an unparseable one fails startup with an actionable error rather than being clamped silently — `0` is not a spelling of `all`, so a typo cannot resolve to full-machine sizing. Three configuration surfaces set the same value. Precedence, highest first: | Surface | Form | | -------------------- | ---------------------- | | Command-line flag | `--cpu-cores 4` | | Environment variable | `SPICE_CPU_CORES=4` | | Spicepod | `runtime.cpu.cores: 4` | A surface set to `auto` still takes precedence over the surfaces below it; it simply resolves to detection. `all` is the exception: it states that a surface imposes no ceiling of its own, so it **defers to a quantity named on a lower-precedence surface**. A platform that sets `SPICE_CPU_CORES=all` on every deployment therefore does not silence an operator who wrote `runtime.cpu.cores: 4` in their spicepod. It does not defer to `auto`, which is an instruction ("detect it") rather than the absence of one. Applied at startup only. The thread pools it sizes cannot be resized afterwards, so changing `runtime.cpu` and reloading the spicepod logs a warning that a restart is required rather than taking effect. It is one of several start-time-only sections — see [Reload Behavior](#reload-behavior). #### Detection[​](#detection "Direct link to Detection") With nothing configured (`auto`), the entitlement is detected. First match wins: 1. A cgroup CPU **quota** — cgroup v2 `cpu.max`, cgroup v1 `cpu.cfs_quota_us`; Kubernetes `resources.limits.cpu`. Read along the whole cgroup path, taking the smallest quota found at any level, and capped by the CPU affinity mask. 2. The pod's **declared CPU request**, as `min(max(2 cores, request × 2), available CPUs)` — see [Sizing from a CPU request](#sizing-from-a-cpu-request). 3. The process's CPU affinity mask (`sched_getaffinity` on Linux, the logical CPU count elsewhere) — the CPUs the process may run on, which a `cpuset` or `taskset` may narrow below the host's core count. 4. One core, when nothing can be determined. A CPU limit outranks a request: bursting past a quota does not produce CPU, it produces CFS throttling. `runtime.cpu.cores: all` suppresses rung 2 only, so it resolves exactly as it would on a pod that declared no request — a CPU limit if one is set, otherwise every available CPU. A process that declares **no** CPU request skips rung 2 entirely and is sized for every CPU it can see. That covers every bare-metal deployment, `docker run` without CPU flags, and every benchmark. #### Sizing from a CPU request[​](#sizing-from-a-cpu-request "Direct link to Sizing from a CPU request") A pod that sets `resources.requests.cpu` without `resources.limits.cpu` has no cgroup quota. Sizing for the whole node would build thread pools and query fan-out for a machine the pod does not own, so the entitlement is derived from the request instead — as a **bounded multiple** of it, currently 2×. The multiple is what makes this safe to automate. A request is a scheduling floor rather than a ceiling, so sizing *at* the request would remove the bursting that is the reason the limit was omitted. The result floors at **2 cores**, so a `requests.cpu: 100m` pod still has enough parallelism to overlap a scan with something, and yields to a genuinely smaller host. The request must be **declared** — a cgroup CPU *share* is never an input. Every cgroup carries a share whether or not a request was expressed (a plain `docker run` reports `cpu.weight: 100`), and the conversion back to a request varies by container runtime, so a share is read for reporting only and never interpreted as a number. Declaring it is the deployment surface's job, through `SPICE_CPU_REQUEST_MILLICORES`: ``` env: - name: SPICE_CPU_REQUEST_MILLICORES valueFrom: resourceFieldRef: containerName: spiceai resource: requests.cpu divisor: 1m # required: makes the value millicores — a `requests.cpu` of 4 arrives as "4000" ``` The [Spice Helm chart](https://github.com/spiceai/spiceai/tree/trunk/deploy/chart) and the Spice Kubernetes Operator both emit this automatically whenever the pod sets a CPU request, so neither needs configuring. A hand-written pod spec must include it, or the pod falls through to rung 3 and sizes for the machine — the runtime warns at startup when it detects that case. Two details in that block are load-bearing. The `divisor: 1m` is what makes the value millicores, which is what the variable's name states; without it a `requests.cpu` of 4 arrives as `4` and reads as four millicores. And the block must be emitted **only when a CPU request is actually set**: with no request declared, `resourceFieldRef` reports the node's *allocatable* CPU, which is exactly the over-sizing this exists to prevent. See [Resource Allocation](/docs/next/reference/performance-tuning#resource-allocation) for the Kubernetes guidance. #### Observability[​](#observability "Direct link to Observability") At startup the runtime logs the effective entitlement, the rung of the ladder or the setting it came from, the readings it sits between, and the defaults derived from it: ``` CPU budget: 8 cores (source: the declared CPU request (x2); host reports 64, declared CPU request 4 cores, cgroup share unset, cgroup limit unset) → 8 main worker threads, 7 per dedicated runtime pool, 8 target partitions CPU budget derived sizing: main_runtime_worker_threads=8, dedicated_runtime_worker_threads=7, target_partitions=8, max_concurrent_queries=32, ... ``` These two lines are the primary diagnostic: between them they name what the runtime sized for, which rung produced it, and every quantity derived from it. A pod sized differently than expected is answered here rather than by inference. The derived line reports **defaults**, not necessarily the values in force: several are overridable by their own setting (`runtime.query.target_partitions`, `runtime.query.max_concurrent_queries`, DuckDB's `threads`, a model's parallelism), and the line is logged before that configuration is resolved. Each overridable consumer separately logs the value it used and where that value came from. Three warnings cover the cases the summary cannot state on its own. Each names a cause and an action; none of them fires for a deployment that is merely sized small on purpose, which the summary already records. | Warning | Fires when | | ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | CPU request present but not passed through | Running under Kubernetes with a cgroup share but no `SPICE_CPU_REQUEST_MILLICORES` — the deployment surface is not emitting the block above, so sizing fell through to the machine. | | Declared request implausibly small | A declared request below 10 millicores, which is what a `resourceFieldRef` missing its `divisor: 1m` produces for a request of one to nine cores. | | CPU share changed after startup | The cgroup share moved from its value at startup — the pod was resized in place. The entitlement cannot change without a restart, so this reports the drift rather than acting on it. | The `spiced_cpu_budget_cores`, `spiced_cpu_budget_millicores`, `spiced_cpu_limit_millicores`, and `spiced_cpu_request_millicores` gauges report the same figures — see [Observability](/docs/next/features/observability). The `source` label on `spiced_cpu_budget_cores` is the authority on which rung won, which is what makes a fleet greppable for pods that resolved somewhere unexpected. `tokio_runtime_workers` is the cross-check on the thread pools the entitlement sized. ## `runtime.query.memory_limit`[​](#runtimequerymemory_limit "Direct link to runtimequerymemory_limit") The `memory_limit` parameter sets a memory usage cap for the Spice runtime query engine. This limit applies **only** to the query engine and should be used in addition to other memory configuration options, such as `duckdb_memory_limit`. When the limit is reached, DataFusion spills intermediate data to disk using the directory configured in `runtime.query.temp_directory`. If not specified, defaults to **90% of the memory the process may use** — its cgroup memory limit when one binds, otherwise total system memory. The limit is read from the process's own cgroup path (cgroup v2 `memory.max`, cgroup v1 `memory.limit_in_bytes`), taking the smallest limit found at any level of that path, so a container limit, a `systemd` unit's `MemoryMax=`, a capped parent slice, and a Kubernetes pod cgroup all size the default. When a cgroup limit binds, the runtime logs the figure it sized from at startup. When Cayenne acceleration is active, the default is reduced to **70%** to reserve headroom for Cayenne's dedicated compaction memory pool and its in-memory CDC tier. When DuckDB accelerators are configured, that default is reduced further so the query pool and the DuckDB instance ceilings together fit within available memory — see [Coordinated memory budget](/docs/next/components/data-accelerators/duckdb#coordinated-memory-budget). An explicitly configured `memory_limit` is always honored verbatim and is never reduced. ``` runtime: query: memory_limit: 4GiB ``` Specify the value as a size, for example `4GiB` or `1024MiB`. For detailed memory information, see [Memory](/docs/next/reference/memory). ## `runtime.query.max_concurrent_queries`[​](#runtimequerymax_concurrent_queries "Direct link to runtimequerymax_concurrent_queries") The `max_concurrent_queries` parameter bounds how many query-executing plans may run concurrently. Excess queries wait (admission control) rather than oversubscribing the shared query runtime and memory pool, which can otherwise cause queries to starve each other under load — for example, analytical queries running alongside CDC ingestion and compaction. ``` runtime: query: max_concurrent_queries: 8 ``` Behavior: * Applies to ordinary queries, DDL/DML, and `EXECUTE`. Lightweight session-state statements (`PREPARE`, `DEALLOCATE`, `SET`) are not gated. * A permit is held for the plan's full execution and result-streaming lifetime. A results-cache hit is never gated. * If not set, the bound is sized from the CPU entitlement at **four concurrent plans per core** — `16` on a 4-core budget. Above one per core so a query blocked on I/O does not idle its core, and low enough that the query memory pool is shared between a countable number of plans. See [`runtime.cpu`](#runtimecpu). * `max_concurrent_queries: 0` opts out and leaves concurrency **unbounded**. Every other configured value is a limit, applied verbatim. Behavior change Prior to this release the default was **unbounded**, and an explicit `0` was clamped to a minimum of `1` (a single concurrent query). Both have changed: an unset value is now bounded by the CPU entitlement, and `0` now means unbounded rather than maximally throttled. A deployment that set `0` to disable admission control gets the behavior it intended; one that set `0` expecting a throttle gets the opposite. ## `runtime.query.timeout`[​](#runtimequerytimeout "Direct link to runtimequerytimeout") The `timeout` parameter sets a maximum wall-clock duration a query may run before it is automatically cancelled, expressed as a human-readable duration (for example `30s` or `5m`). The clock covers the query's full lifetime: planning, admission-control waits (`max_concurrent_queries`), execution, and streaming results to the client. ``` runtime: query: timeout: 30s ``` Behavior: * Applies to queries issued through the runtime's query APIs (HTTP, Flight, and Flight SQL). Internal runtime queries — acceleration refreshes and health checks — are exempt. * Enforcement is cooperative (best-effort): the query is cancelled at its next cancellation checkpoint, so actual runtime can slightly exceed the configured value. * On expiry, the query fails with a timeout error. If the timeout is observed before the response starts, the client receives an HTTP `504` / gRPC `DEADLINE_EXCEEDED`. If results are already streaming, the status can no longer change, so the in-progress stream is terminated with the error — data streamed before expiry will have been delivered, but the stream never ends silently as if complete. * If not set, queries run with **no timeout** (the default behavior). The value must be a positive duration greater than `0`. ## `runtime.query.spill_compression`[​](#runtimequeryspill_compression "Direct link to runtimequeryspill_compression") The `spill_compression` parameter configures compression for spill files generated during large query execution in the Spice runtime. **Supported values:** * `zstd` (default): Enables high compression ratios for spill files, reducing disk usage but with moderate (de)compression speed. * `lz4_frame`: Provides faster (de)compression, resulting in larger spill files and potentially higher disk usage. * `uncompressed`: Disables compression. Spill files will be the largest, but with no (de)compression overhead. ``` runtime: query: spill_compression: lz4_frame ``` This setting controls the trade-off between disk space usage and query performance for large-scale analytics workloads. ## `runtime.query.temp_directory`[​](#runtimequerytemp_directory "Direct link to runtimequerytemp_directory") []() The path to a temporary directory that Spice uses for query and acceleration operations that spill to disk. For more details, see the [Managing Memory Usage documentation](/docs/next/reference/memory) and the [DuckDB Data Accelerator documentation](/docs/next/components/data-accelerators/duckdb). ``` runtime: query: temp_directory: /tmp/spice ``` ## `runtime.output_level`[​](#runtimeoutput_level "Direct link to runtimeoutput_level") Controls verbosity in addition to the existing [CLI and environment variable support.](https://spiceai.org/docs/cli/tracing). Supported values are `info`, `verbose`, and `very_verbose`. The value is applied in the following priority: CLI, environment variables, then YAML configuration. ``` runtime: output_level: info # or verbose, very_verbose ``` ## `runtime.telemetry`[​](#runtimetelemetry "Direct link to runtimetelemetry") The telemetry section configures runtime telemetry collection and export. [Learn more](/docs/next/features/observability). ``` runtime: telemetry: enabled: true otel_exporter: enabled: true endpoint: 'localhost:4317' push_interval: '5m' ``` ### `runtime.telemetry.enabled`[​](#runtimetelemetryenabled "Direct link to runtimetelemetryenabled") Enables or disables runtime telemetry collection. Defaults to `true`. ### `runtime.telemetry.metric_prefix`[​](#runtimetelemetrymetric_prefix "Direct link to runtimetelemetrymetric_prefix") Optional string prepended to every exported metric name. Useful for namespacing Spice metrics in shared backends (e.g. Datadog, Grafana Cloud, New Relic) so they do not collide with metrics from other services. Defaults to no prefix. The prefix applies to **all** metric readers — the Prometheus scrape endpoint (`--metrics`), the cluster on-demand OTLP reader, and the `otel_exporter` push exporter — because OpenTelemetry views are configured at the meter-provider level rather than per reader. ``` runtime: telemetry: metric_prefix: 'spiceai.' ``` With this configuration, the runtime metric `query_duration_ms` is exported as `spiceai.query_duration_ms`. The prefix is validated at spicepod load against the [OpenTelemetry instrument name syntax](https://opentelemetry.io/docs/specs/otel/metrics/api/#instrument-name-syntax) so `{prefix}{instrument}` stays a valid OTLP/Prometheus metric name. A non-empty prefix must: * start with an ASCII letter (`A`–`Z` or `a`–`z`); * contain only ASCII letters, digits, `_`, `.`, `-`, or `/`; and * be at most 128 characters (leaving ≥127 characters of headroom under the OpenTelemetry 255-character instrument-name limit for the base metric name). An invalid prefix fails fast with an actionable error rather than producing malformed metric names. An empty or unset `metric_prefix` applies no prefix. ### `runtime.telemetry.properties`[​](#runtimetelemetryproperties "Direct link to runtimetelemetryproperties") Map of custom key/value attributes attached to telemetry metrics emitted by `spiced`. Applied as OpenTelemetry resource attributes on the runtime's `MeterProvider`, so they appear as dimensions/tags on every metric exported via the Prometheus scrape endpoint, the cluster on-demand OTLP reader, and the `otel_exporter` push exporter. Defaults to empty. ``` runtime: telemetry: properties: environment: prod region: us-west-2 team: data-platform ``` The standard OpenTelemetry environment variables (`OTEL_SERVICE_NAME`, `OTEL_RESOURCE_ATTRIBUTES`) are still honored and act as defaults; explicit `properties` entries take precedence on key conflicts. For backends that map OTLP resource attributes to tags through additional configuration (e.g. Datadog), see the [Datadog OTLP guide](/docs/next/monitoring/datadog#opentelemetry-otlp-export). ### `runtime.telemetry.otel_exporter`[​](#runtimetelemetryotel_exporter "Direct link to runtimetelemetryotel_exporter") Configures an [OpenTelemetry](https://opentelemetry.io/) metrics exporter to push metrics to an OpenTelemetry collector. The exporter automatically infers the protocol (gRPC or HTTP) based on the endpoint configuration. | Parameter name | Optional | Default | Description | | --------------- | -------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `enabled` | Yes | `true` | Whether the OpenTelemetry exporter is enabled. | | `endpoint` | No | - | The OpenTelemetry collector endpoint. Protocol is inferred from the format (see examples below). | | `push_interval` | Yes | `60s` | How frequently metrics are pushed to the collector. Specify as a [duration](/docs/next/reference/duration). | | `metrics` | Yes | `[]` | List of metric names to export. When empty (default), all metrics are exported. | | `headers` | Yes | `{}` | Map of headers to send with each export request. For HTTP these are sent as HTTP headers; for gRPC they are sent as metadata entries (keys must be lowercase ASCII). Values support the `${secrets:...}` [replacement syntax](/docs/next/components/secret-stores#using-secrets) for loading credentials from a [secret store](/docs/next/components/secret-stores). | **Protocol inference:** * **gRPC (default):** Use a bare host :port endpoint without a scheme (e.g., `localhost:4317`). gRPC uses port 4317 by default. * **HTTP:** Include the `http://` or `https://` scheme and the `/v1/metrics` path (e.g., `http://localhost:4318/v1/metrics`). HTTP uses port 4318 by default. **Examples:** gRPC configuration: ``` runtime: telemetry: enabled: true otel_exporter: # gRPC - no scheme or path needed endpoint: 'localhost:4317' push_interval: '30s' ``` HTTP configuration: ``` runtime: telemetry: enabled: true otel_exporter: enabled: true # HTTP - include scheme and /v1/metrics path endpoint: 'http://localhost:4318/v1/metrics' push_interval: '30s' ``` With metric filtering (export only specific metrics): ``` runtime: telemetry: enabled: true otel_exporter: endpoint: 'localhost:4317' push_interval: '30s' metrics: - query_duration_ms - query_executions - dataset_load_state ``` Filtering happens after `metric_prefix` is applied The whitelist is matched against the **final** metric name, after [`runtime.telemetry.metric_prefix`](#runtimetelemetrymetric_prefix) has been prepended. If you set `metric_prefix: 'spiceai.'`, the entries under `metrics:` must include the prefix (e.g. `spiceai.query_duration_ms`), otherwise nothing will match and no metrics will be exported. **Authenticated exporters:** For collectors that require authentication, set the `headers` map. Load credentials from a [secret store](/docs/next/components/secret-stores) via `${secrets:...}` rather than committing them to source. Datadog (OTLP/HTTP) — replace `us3` with your Datadog site: ``` runtime: telemetry: otel_exporter: endpoint: 'https://otlp.us3.datadoghq.com/v1/metrics' headers: DD-API-KEY: ${secrets:dd_api_key} ``` Grafana Cloud (OTLP/HTTP) — use the base64 `instanceID:accessPolicyToken` from the Grafana Cloud OpenTelemetry connection page: ``` runtime: telemetry: otel_exporter: endpoint: 'https://otlp-gateway-us-central2.grafana.net/otlp/v1/metrics' headers: Authorization: 'Basic ${secrets:grafana_cloud_auth}' ``` gRPC collector with auth metadata (keys must be lowercase ASCII): ``` runtime: telemetry: otel_exporter: endpoint: 'otel-collector.internal:4317' headers: api-key: ${secrets:collector_api_key} ``` ## `runtime.metrics`[​](#runtimemetrics "Direct link to runtimemetrics") Specifies metrics that are disabled by default. Following metrics are disabled by default: * `dataset_acceleration_max_timestamp_before_refresh_ms` * `dataset_acceleration_max_timestamp_after_refresh_ms` * `dataset_acceleration_refresh_lag_ms` * `dataset_acceleration_ingestion_lag_ms` For details about these metrics, see [Observability](/docs/next/features/observability). ``` runtime: metrics: - name: dataset_acceleration_max_timestamp_before_refresh_ms - name: dataset_acceleration_max_timestamp_after_refresh_ms enabled: true - name: dataset_acceleration_refresh_lag_ms enabled: false - name: dataset_acceleration_ingestion_lag_ms ``` ## `runtime.flight`[​](#runtimeflight "Direct link to runtimeflight") Configures Arrow Flight protocol settings for the runtime. ``` runtime: flight: max_message_size: 16MiB do_put_rate_limit_enabled: true ``` | Parameter name | Optional | Default | Description | | --------------------------- | -------- | ------- | -------------------------------------------------------------------- | | `max_message_size` | Yes | - | Maximum size of a single Arrow Flight message. | | `do_put_rate_limit_enabled` | Yes | `true` | Whether rate limiting is applied to `DoPut` Arrow Flight operations. | ## `runtime.mcp`[​](#runtimemcp "Direct link to runtimemcp") Configures settings for the Spice MCP server endpoint (`/v1/mcp`). ### `runtime.mcp.allowed_hosts`[​](#runtimemcpallowed_hosts "Direct link to runtimemcpallowed_hosts") Controls which `Host` header values are accepted on the `/v1/mcp` endpoint. This prevents [DNS rebinding](https://en.wikipedia.org/wiki/DNS_rebinding) attacks against the MCP server. | Behavior | Configuration | | ---------------------- | --------------------------------------------------------------------------------------------------------------------- | | **Default** (not set) | Only `localhost`, `127.0.0.1`, and `::1` are permitted. Requests with any other `Host` value receive `403 Forbidden`. | | **Explicit list** | Replaces the defaults entirely. Only the listed hosts are accepted. | | **Wildcard** (`["*"]`) | Disables host checking — all `Host` header values are accepted. | ``` runtime: mcp: allowed_hosts: - localhost - my-host.internal:8090 ``` To disable host checking entirely: ``` runtime: mcp: allowed_hosts: - "*" ``` Each entry can be a bare hostname (`example.com`), a host-port pair (`example.com:8090`), or a full origin URL (`https://example.com`). ## `runtime.ready_state`[​](#runtimeready_state "Direct link to runtimeready_state") Controls when the runtime readiness probe (`/v1/ready`) reports the runtime as ready. This is particularly useful for Kubernetes readiness probes. ``` runtime: ready_state: on_load ``` | Value | Description | | ------------------- | ----------------------------------------------------------------------------------------------------- | | `on_load` (default) | The runtime reports ready after all components (datasets, models, etc.) have loaded successfully. | | `on_registration` | The runtime reports ready as soon as all components have been registered, before they finish loading. | ## `runtime.scheduler`[​](#runtimescheduler "Direct link to runtimescheduler") Configures the cluster scheduler when running Spice in [cluster mode](/docs/next/deployment/architectures/cluster). This section is relevant only when using `--role scheduler`. ``` runtime: scheduler: state_location: s3://my-bucket/spice-cluster-state/ params: s3_region: us-east-1 partition_assignment_interval: 30s max_partition_assignments_per_interval: 100 max_partitions_per_executor: 1000 partition_discovery_timeout: 60s ``` | Parameter name | Optional | Default | Description | | ---------------------------------------- | -------- | ------- | ---------------------------------------------------------------------- | | `state_location` | No | - | Root URI for shared cluster state storage (e.g. `s3://bucket/path/`). | | `params` | Yes | - | Object store parameters (e.g. `aws_region`). | | `partition_assignment_interval` | Yes | `30s` | How often the scheduler runs partition assignment cycles. | | `max_partition_assignments_per_interval` | Yes | `100` | Maximum number of partition assignments per interval. | | `max_partitions_per_executor` | Yes | `1000` | Maximum number of partitions assigned to a single executor. | | `partition_discovery_timeout` | Yes | `60s` | How long the scheduler waits for executor discovery before timing out. | --- # Tools (Function Calling) Tools define functions that can be invoked within the Spice runtime, either directly or by a [language model](/docs/next/features/large-language-models) (LLMs). These tools provide access to different functionalities and can be customized in the `tools` section of `spicepod.yaml`. ## `tools`[​](#tools "Direct link to tools") The `tools` section in your configuration specifies one or more tools available for use in the runtime. Example: ``` tools: - name: my_mcp_tool from: mcp:http://localhost:8090/v1/mcp description: 'An MCP tool.' ``` ### `name`[​](#name "Direct link to name") A unique identifier for this tool. ### `from`[​](#from "Direct link to from") Defines the source of the tool, or the specific built-in tool to customise. See [Available Tools](/docs/next/components/tools#available-tools) for a list of available tools. For [MCP tools](/docs/next/components/tools/mcp), the target is part of the `from` value itself, in the form `mcp:`: * `mcp:` connects to an MCP server over Streamable HTTP, e.g. `mcp:http://localhost:8090/v1/mcp`. * `mcp:` runs a stdio-based MCP server as a subprocess, e.g. `mcp:npx`. A bare `from: mcp` (without a `:` suffix) is rejected when the spicepod is loaded. ### `description`[​](#description "Direct link to description") Optional. A textual description of the tool's function. ### `params`[​](#params "Direct link to params") Optional. A map of key-value pairs for additional parameters specific to the tool. For MCP tools these are `mcp_args` (stdio), and `mcp_auth_token` / `mcp_headers` (Streamable HTTP) — see [MCP Tools](/docs/next/components/tools/mcp#params). The MCP server address is set via `from`, not via `params`. ### `env`[​](#env "Direct link to env") Optional. A map of key-value pairs of arbitrary environment variables to set when running the tool. Only useable if the tool requires a subprocess to run (e.g. MCP over stdio) . ### `dependsOn`[​](#dependson "Direct link to dependson") Optional. A list of dependencies that must be available before this tool can be used. --- # Views A Spicepod can contain one or more `views` referenced by relative path or defined inline. Inline example: `spicepod.yaml` ``` views: - name: votes sql_ref: ./a/file/path.sql - name: rankings sql: | WITH a AS ( SELECT products.id, SUM(count) AS count FROM orders INNER JOIN products ON orders.product_id = products.id GROUP BY products.id ) SELECT name, count FROM products LEFT JOIN a ON products.id = a.id ORDER BY count DESC LIMIT 5 acceleration: enabled: true ``` ## `name`[​](#name "Direct link to name") The name of the view. Used to reference the view in the pod manifest, as well as in external data sources. The name cannot be a [reserved keyword](/docs/next/reference/spicepod/keywords). ## `description`[​](#description "Direct link to description") The description of the view. Used as part of the [Semantic Data Model](/docs/next/features/semantic-model). ## `params`[​](#params "Direct link to params") Optional. A map of view-level parameters, mirroring the `params` map available on [datasets](/docs/next/reference/spicepod/datasets). Currently the only supported key is `file_format`, which selects the [file format](/docs/next/reference/file_format) used when chunking the view's columns for [embeddings](/docs/next/components/embeddings). It accepts the same values as the dataset `file_format` parameter (for example `parquet`, `csv`, `json`, `md`). ``` views: - name: my_view sql_ref: ./my_view.sql params: file_format: md columns: - name: body embeddings: - from: my_embedding_model chunking: enabled: true ``` ## `ready_state`[​](#ready_state "Direct link to ready_state") Supports one of two values: * `on_registration`: Mark the view as ready immediately, and queries on this table will fall back to the underlying source directly until the initial acceleration is complete * `on_load`: Mark the view as ready only after the initial acceleration. Queries against the view will return an error before the load has been completed. ``` views: - name: my_view sql_ref: ./my_view.sql ready_state: on_registration # or on_load acceleration: enabled: true ``` ## `acceleration`[​](#acceleration "Direct link to acceleration") Optional. Accelerate queries to the view by caching data locally. ## `acceleration.enabled`[​](#accelerationenabled "Direct link to accelerationenabled") Enable or disable acceleration, defaults to `true`. ## `acceleration.engine`[​](#accelerationengine "Direct link to accelerationengine") The acceleration engine to use, defaults to `arrow`. The following engines are supported: * `arrow` - Accelerated in-memory backed by Apache Arrow DataTables. * [`duckdb`](/docs/next/components/data-accelerators/duckdb) - Accelerated by an embedded DuckDB database. * [`postgres`](/docs/next/components/data-accelerators/postgres) - Accelerated by a Postgres database. * [`sqlite`](/docs/next/components/data-accelerators/sqlite) - Accelerated by an embedded SQLite database. * [`turso`](/docs/next/components/data-accelerators/turso) - Accelerated by an embedded Turso (libSQL) database (Beta). ## `acceleration.mode`[​](#accelerationmode "Direct link to accelerationmode") Optional. The mode of acceleration. The following values are supported: * `memory` - Store acceleration data in-memory. Supported for Spice Cayenne (`cayenne`), where the acceleration is ephemeral and reloads from its source on restart. * `file` - Store acceleration data in a file. Supported for Spice Cayenne (`cayenne`), `duckdb`, `sqlite`, and `turso` acceleration engines. ## `acceleration.snapshots`[​](#accelerationsnapshots "Direct link to accelerationsnapshots") Optional. Controls how this view participates in managed acceleration snapshots. Requires the Spicepod to configure the top-level [`snapshots` block](/docs/next/reference/spicepod/#snapshots), the acceleration engine to be `duckdb` or `sqlite`, and `mode: file` with a view-specific file path (for example `acceleration.params.duckdb_file: /nvme/my_view.db`). Supported values: * `enabled` – Download the newest snapshot on startup when the acceleration file is missing and write a fresh snapshot after each refresh. * `bootstrap_only` – Download snapshots on startup but never write new ones. * `create_only` – Write snapshots after refreshes but never download them on startup. * `disabled` (default) – Do not use snapshots for this view. Snapshots are written beneath the configured snapshot location using Hive-style partitioning (`month=YYYY-MM/day=YYYY-MM-DD/view=`). For more background, see [Acceleration snapshots](/docs/next/features/data-acceleration/snapshots). ## `acceleration.refresh_mode`[​](#accelerationrefresh_mode "Direct link to accelerationrefresh_mode") Optional. How to refresh the view. The following values are supported: * `full` - Refresh the entire view. * `append` - Append new data to the view. When `time_column` is specified, new records are fetched from the latest timestamp in the accelerated data at the `acceleration.refresh_check_interval`. ## `acceleration.refresh_check_interval`[​](#accelerationrefresh_check_interval "Direct link to accelerationrefresh_check_interval") Optional. How often data should be refreshed. For `append` views without a specific `time_column`, this config is not used. If not defined, the accelerator will not refresh after it initially loads data. Cannot be specified in conjunction with a `refresh_cron`. See [Duration](/docs/next/reference/duration) ## `acceleration.refresh_cron`[​](#accelerationrefresh_cron "Direct link to accelerationrefresh_cron") Optional. Specifies a cron schedule which controls how often data is refreshed. For `append` views without a specific `time_column`, this config is not used. If not defined, the accelerator will not refresh after it initially loads data. See the [cron schedule reference](/docs/next/reference/cron). ## `acceleration.refresh_sql`[​](#accelerationrefresh_sql "Direct link to accelerationrefresh_sql") Optional. Filters the data fetched from the source to be stored in the accelerator engine. Only supported for `full` refresh\_mode views. Must be of the form `SELECT * FROM {name} WHERE {refresh_filter}`. `{name}` is the view name declared above, `{refresh_filter}` is any SQL expression that can be used to filter the data, i.e. `WHERE city = 'Seattle'` to reduce the working set of data that is accelerated within Spice from the data source. Limitations * The refresh SQL only supports filtering data from the current view - joining across other views is not supported. * Queries for data that have been filtered out will not fall back to querying against the federated table. ## `acceleration.refresh_data_window`[​](#accelerationrefresh_data_window "Direct link to accelerationrefresh_data_window") Optional. A duration to filter view refresh source queries to recent data (duration into past from now). Requires `time_column` and `time_format` to also be configured. Only supported for `full` refresh mode views. For example, `refresh_data_window: 24h` will include only records with a timestamp within the last 24 hours. See [Duration](/docs/next/reference/duration) ## `acceleration.refresh_append_overlap`[​](#accelerationrefresh_append_overlap "Direct link to accelerationrefresh_append_overlap") Optional. A duration to specify how far back to include records based on the most recent timestamp found in the accelerated data. Requires `time_column` to also be configured. Only supported for `append` refresh mode views. This setting can help mitigate missing data issues caused by late arriving data. Example: If the latest timestamp in the accelerated data table is `2020-01-01T02:00:00Z`, setting `refresh_append_overlap: 1h` will include records starting from `2020-01-01T01:00:00Z`. See [Duration](/docs/next/reference/duration) ## `acceleration.refresh_retry_enabled`[​](#accelerationrefresh_retry_enabled "Direct link to accelerationrefresh_retry_enabled") Optional. Specifies whether an accelerated view should retry data refresh in the event of transient errors. The default setting is true. Retries follow a [Fibonacci backoff strategy](https://en.wikipedia.org/wiki/Fibonacci_sequence). To disable refresh retries, set `refresh_retry_enabled: false`. ## `acceleration.refresh_retry_max_attempts`[​](#accelerationrefresh_retry_max_attempts "Direct link to accelerationrefresh_retry_max_attempts") Optional. Defines the maximum number of retry attempts when refresh retries are enabled. The default is undefined, with no upper limit on attempts. ## `acceleration.refresh_on_startup`[​](#accelerationrefresh_on_startup "Direct link to accelerationrefresh_on_startup") Optional. Controls the refresh behavior of an accelerated view across restarts. Defaults to `auto`. ### Supported Values[​](#supported-values "Direct link to Supported Values") * **`auto` (Default)** – Maintains refresh state across restarts: * With `refresh_check_interval`: Schedules next refresh based on last successful refresh time, triggering immediately if interval has already elapsed * Without `refresh_check_interval`: No refresh (on-demand only) * **`always`** – Forces a view refresh on every startup, regardless of the existing acceleration state. Setting `refresh_on_startup: always` ensures that accelerated data is always refreshed to match the source when the service restarts. This is useful in **development environments** or when **data consistency is critical** after deployment. ## `acceleration.params`[​](#accelerationparams "Direct link to accelerationparams") Optional. Parameters to pass to the acceleration engine. The parameters are specific to the acceleration engine used. ## `acceleration.retention_check_enabled`[​](#accelerationretention_check_enabled "Direct link to accelerationretention_check_enabled") Optional. Enable or disable retention policy check, defaults to `false`. ## `acceleration.retention_period`[​](#accelerationretention_period "Direct link to accelerationretention_period") Optional. The retention period for the view. Combine with `time_column` and `time_format` to determine if the data should be retained or not. `retention_period` or `retention_sql` must be specified when `acceleration.retention_check_enabled` is `true`. When both `retention_period` and `retention_sql` are configured, both retention policies will be applied during each retention check. See [Duration](/docs/next/reference/duration) ## `acceleration.retention_sql`[​](#accelerationretention_sql "Direct link to accelerationretention_sql") Optional. Custom SQL statement to define data retention logic. Takes the form of a `DELETE FROM
WHERE ` statement. This parameter is useful for scenarios like soft-deleting rows in append-only views or removing data based on complex business logic that goes beyond simple time-based retention. `retention_period` or `retention_sql` must be specified when `acceleration.retention_check_enabled` is `true`. When both `retention_period` and `retention_sql` are configured, both retention policies will be applied during each retention check. ## `acceleration.retention_check_interval`[​](#accelerationretention_check_interval "Direct link to accelerationretention_check_interval") Optional. How often the retention policy should be checked. Required when `acceleration.retention_check_enabled` is `true`. See [Duration](/docs/next/reference/duration) ## `acceleration.refresh_jitter_enabled`[​](#accelerationrefresh_jitter_enabled "Direct link to accelerationrefresh_jitter_enabled") Optional. Enable or disable refresh jitter, defaults to `false`. The refresh jitter adds/subtracts a randomized time period from the `refresh_check_interval`. ## `acceleration.refresh_jitter_max`[​](#accelerationrefresh_jitter_max "Direct link to accelerationrefresh_jitter_max") Optional. The maximum amount of jitter to add to the refresh interval. The jitter is a random value between 0 and `refresh_jitter_max`. Defaults to 10% of `refresh_check_interval`. ## `acceleration.indexes`[​](#accelerationindexes "Direct link to accelerationindexes") Optional. Specify which indexes should be applied to the locally accelerated table. Not supported for in-memory Arrow acceleration engine. The `indexes` field is a map where the key is the column reference and the value is the index type. A column reference can be a single column name or a multicolumn key. The column reference must be enclosed in parentheses if it is a multicolumn key. See [Indexes](/docs/next/features/data-acceleration/indexes) ``` views: - name: my_view sql_ref: ./my_view.sql acceleration: enabled: true engine: sqlite indexes: number: enabled # Index the `number` column '(hash, timestamp)': unique # Add a unique index with a multicolumn key comprised of the `hash` and `timestamp` columns ``` ## `acceleration.primary_key`[​](#accelerationprimary_key "Direct link to accelerationprimary_key") Optional. Specify the primary key constraint on the locally accelerated table. Not supported for in-memory Arrow acceleration engine. The `primary_key` field is a string that represents the column reference that should be used as the primary key. The column reference can be a single column name or a multicolumn key. The column reference must be enclosed in parentheses if it is a multicolumn key. See [Constraints](/docs/next/features/data-acceleration/constraints) ``` views: - name: my_view sql_ref: ./my_view.sql acceleration: enabled: true engine: sqlite primary_key: hash # Define a primary key on the `hash` column ``` ## `acceleration.on_conflict`[​](#accelerationon_conflict "Direct link to accelerationon_conflict") Optional. Specify what should happen when a constraint is violated. Not supported for in-memory Arrow acceleration engine. The `on_conflict` field is a map where the key is the column reference and the value is the conflict resolution strategy. A column reference can be a single column name or a multicolumn key. The column reference must be enclosed in parentheses if it is a multicolumn key. Only a single `on_conflict` target can be specified, unless all `on_conflict` targets are specified with `drop`. The possible conflict resolution strategies are: * `upsert` - Upsert the incoming data when the primary key constraint is violated. * `upsert_dedup` - Same as `upsert`, but also deduplicates the data if there are duplicate rows that trigger a constraint violation within a single update. See [Advanced upsert behavior](/docs/next/features/data-acceleration/constraints#advanced-upsert-options). * `upsert_dedup_by_row_id` - Same as `upsert`, but resolves any violations by arbitrarily choosing the row with the highest row id. See [Advanced upsert behavior](/docs/next/features/data-acceleration/constraints#advanced-upsert-options). * `drop` - Drop the data when the primary key constraint is violated. See [Constraints](/docs/next/features/data-acceleration/constraints) ``` views: - name: my_view sql_ref: ./my_view.sql acceleration: enabled: true engine: sqlite primary_key: hash indexes: '(number, timestamp)': unique on_conflict: # Upsert the incoming data when the primary key constraint on "hash" is violated, # alternatively "drop" can be used instead of "upsert" to drop the data update. hash: upsert ``` ## `columns`[​](#columns "Direct link to columns") Optional. Define metadata, semantic details and features (e.g. embeddings, or table indexes) for specific columns in the view. ``` views: - name: my_view sql_ref: ./my_view.sql columns: - name: address_line1 description: The first line of the address. embeddings: - from: hf_minilm row_id: order_number chunking: enabled: true target_chunk_size: 256 overlap_size: 32 full_text_search: enabled: true ``` ## `columns[*].name`[​](#columnsname "Direct link to columnsname") The name of the column in the table schema. ## `columns[*].description`[​](#columnsdescription "Direct link to columnsdescription") Optional. A description of the column's contents and purpose. Used as part of the [Semantic Data Model](/docs/next/features/semantic-model). ## `columns[*].embeddings`[​](#columnsembeddings "Direct link to columnsembeddings") Optional. Create vector embeddings for this column. ## `columns[*].embeddings[*].from`[​](#columnsembeddingsfrom "Direct link to columnsembeddingsfrom") The embedding model to use, specify the component name. ## `columns[*].embeddings[*].row_id`[​](#columnsembeddingsrow_id "Direct link to columnsembeddingsrow_id") Optional. For views without a primary key, used to explicitly specify column(s) that uniquely identify a row. Specifying a `row_id` enables unique identifier lookups for views from external systems that may not have a primary key. ## `columns[*].embeddings[*].chunking`[​](#columns-embeddings-chunking "Direct link to columns-embeddings-chunking") Optional. The configuration to enable and define the chunking strategy for the embedding column. ``` columns: - name: description embeddings: - from: hf_minilm chunking: enabled: true target_chunk_size: 512 overlap_size: 128 trim_whitespace: false ``` See [`embeddings[*].chunking`](#columns-embeddings-chunking) for details. ## `columns[*].embeddings[*].vector_size`[​](#columnsembeddingsvector_size "Direct link to columnsembeddingsvector_size") Optional. Specifies the size (number of dimensions) of the embedding vector for use in federated queries to databases that do not support arrays with fixed lengths. ``` columns: - name: review_body embeddings: - from: embed-static-retrieval vector_size: 1024 ``` ## `columns[*].full_text_search`[​](#columns-search-full-text "Direct link to columns-search-full-text") ## `columns[*].full_text_search.enabled`[​](#columnsfull_text_searchenabled "Direct link to columnsfull_text_searchenabled") Optional. Enable or disable full text search support for specific column in the view. Default `false`. ## `columns[*].full_text_search.row_id`[​](#columnsfull_text_searchrow_id "Direct link to columnsfull_text_searchrow_id") Optional. For views without a primary key, used to explicitly specify column(s) that uniquely identify a row. Specifying a `row_id` enables unique identifier lookups for views from external systems that may not have a primary key. ## `columns[*].metadata`[​](#columnsmetadata "Direct link to columnsmetadata") Optional. Specific metadata associated to the column. ## `columns[*].metadata.vectors`[​](#columnsmetadatavectors "Direct link to columnsmetadatavectors") Optional. If provided, a vector engine (see [below](#vectors)) should store this column for a particular use, determined by the value, which is one of: * `non-filterable`: Store the column in the vector engine. * `filterable`: Store the column in the vector engine, and ensure the engine can filter on the column (if possible in the engine). Only applicable if `vectors.enabled` is both defined and `true`. ## `metadata`[​](#metadata "Direct link to metadata") Optional. Additional key-value metadata for the view. Used as part of the [Semantic Data Model](/docs/next/features/semantic-model). ``` views: - name: my_view sql_ref: ./my_view.sql metadata: instructions: The last 128 blocks. ``` ## `vectors`[​](#vectors "Direct link to vectors") ## `vectors.enabled`[​](#vectorsenabled "Direct link to vectorsenabled") Enable or disable vector storage, defaults to `true`. ## `vectors.engine`[​](#vectorsengine "Direct link to vectorsengine") The vector engine to use. The following engines are supported: * [`s3_vectors`](/docs/next/components/vectors/s3_vectors) - Vectors are created and indexed into [Amazon S3 Vectors](https://aws.amazon.com/s3/features/vectors/). ## `vectors.params`[​](#vectorsparams "Direct link to vectorsparams") Optional. Parameters to pass to the vector engine. The parameters are specific to the vector engine used. --- # Workers Workers in the Spice runtime represent configurable units of compute that help coordinate and manage interactions between models and tools. Currently, workers define how one or more [llms](/docs/next/reference/models) can be combined into a logically single model. ## `workers`[​](#workers "Direct link to workers") The `workers` section in your configuration specifies one or more workers. Example: ``` workers: - name: round-robin description: | Distributes requests between 'llama3_2' and 'gpt4_1' models in a round-robin fashion. load_balance: routing: - from: llama3_2 - from: gpt4_1 - name: fallback description: | Attempts 'gpt4_1' first, then 'llama3_2', then 'anth_haiku' if previous models fail. load_balance: routing: - from: llama3_2 order: 2 - from: gpt4_1 order: 1 - from: anth_haiku order: 3 - name: weighted description: | Routes 80% of traffic to 'llama3_2'. load_balance: routing: - from: llama3_2 weight: 4 - from: gpt4_1 weight: 1 ``` ### `name`[​](#name "Direct link to name") A unique identifier for this worker component. ### `description`[​](#description "Direct link to description") Additional details about the worker, useful for displaying to users and providing to LLM context. ### `cron`[​](#cron "Direct link to cron") Specifies a cron schedule to automatically run the worker at the specified times. The worker action controls the behavior of the schedule. See the [cron schedule reference](/docs/next/reference/cron) for more information on cron schedules. #### `cron` with a `load_balance` action[​](#cron-with-a-load_balance-action "Direct link to cron-with-a-load_balance-action") When a `load_balance` action is specified with a cron schedule, the `params.prompt` parameter is used to automatically request a chat completion. When no `params.prompt` parameter is specified, the cron schedule is not activated. ##### Worker with a round-robin balancer, that is automatically prompted on a schedule[​](#worker-with-a-round-robin-balancer-that-is-automatically-prompted-on-a-schedule "Direct link to Worker with a round-robin balancer, that is automatically prompted on a schedule") ``` workers: - name: round-robin description: | Call models 'llama3_2' & 'gpt4_1' in round robin. load_balance: routing: - from: llama3_2 - from: gpt4_1 cron: '* * * * *' # every minute params: prompt: "What's the date today?" ``` #### `cron` with a `sql` action[​](#cron-with-a-sql-action "Direct link to cron-with-a-sql-action") When a `sql` action is specified with a cron schedule, the worker runs the SQL at the specified scheduled times. ##### Worker with a SQL action, that automatically executes on a schedule[​](#worker-with-a-sql-action-that-automatically-executes-on-a-schedule "Direct link to Worker with a SQL action, that automatically executes on a schedule") ``` workers: - name: sql-worker cron: '* * * * *' # every minute sql: 'SELECT COUNT(*) FROM orders' ``` ### `load_balance`[​](#load_balance "Direct link to load_balance") Specifies the configuration for a `load_balance` worker. When a `load_balance` section is present, other worker actions cannot be specified (e.g. `sql`). ### `load_balance.routing`[​](#load_balancerouting "Direct link to load_balancerouting") A list of model configurations that define how the load balancing behaves. The elements' structure uniquely determine the model worker algorithm. List elements should be of consistent type. | Key name | Key type | Description | | -------- | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | | from | String | The `model.name` of a defined `model` spicepod component. | | order | Integer, positive | The priority of the model in order. The lowest value is used first, followed by increasing order. The ordering of models with equal `order` is undefined. | | weight | Integer | The weight for weighted routing. Traffic is distributed proportionally based on each model's weight relative to the total. | #### Worker with round-robin routing across models[​](#worker-with-round-robin-routing-across-models "Direct link to Worker with round-robin routing across models") Example ``` workers: - name: round-robin description: | Call models 'llama3_2' & 'gpt4_1' in round robin. load_balance: routing: - from: llama3_2 - from: gpt4_1 ``` The worker selects each model in turn for subsequent requests. #### Worker with fallback model routing[​](#worker-with-fallback-model-routing "Direct link to Worker with fallback model routing") Example ``` workers: - name: fallback description: | Call 'gpt4_1'. On error, call 'llama3_2'. Failing that 'anth_haiku'. load_balance: routing: - from: llama3_2 order: 2 - from: gpt4_1 order: 1 - from: anth_haiku order: 3 ``` The worker uses the models in increasing order, returning the first result that is not an error. #### Worker with weighted model routing[​](#worker-with-weighted-model-routing "Direct link to Worker with weighted model routing") Example ``` workers: - name: weighted description: | Routes 80% of traffic to 'llama3_2' (20% to 'gpt4_1'). load_balance: routing: - from: llama3_2 weight: 4 - from: gpt4_1 weight: 1 ``` The worker routes traffic to the models in accordance to the weighting (i.e. 80% to `llama3_2`, 20% to `gpt4_1`). ### `sql`[​](#sql "Direct link to sql") Specifies an SQL query action to run for this worker. When specified without a `cron` parameter, the worker does nothing. When this parameter is present, other worker actions cannot be specified (e.g. `load_balance`) ### `params`[​](#params "Direct link to params") Optional, additional parameters for the specified worker action. #### `params.prompt`[​](#paramsprompt "Direct link to paramsprompt") Valid only when the `load_balance` worker action is specified with a `cron` schedule, otherwise ignored. The value specified by this parameter is used as the input to a new chat completion request on the specified `cron` schedule for the worker. --- # SQL Reference This section provides a comprehensive reference for SQL support in Spice.ai, including syntax, data types, operators, functions, and system features. The reference is organized by topic for ease of navigation. ## Table of Contents[​](#table-of-contents "Direct link to Table of Contents") ### [SELECT Syntax](/docs/next/reference/sql/select)[​](#select-syntax "Direct link to select-syntax") * [WITH Clause](/docs/next/reference/sql/select#with-clause) * [SELECT Clause](/docs/next/reference/sql/select#select-clause) * [FROM Clause](/docs/next/reference/sql/select#from-clause) * [WHERE Clause](/docs/next/reference/sql/select#where-clause) * [JOIN Clause](/docs/next/reference/sql/select#join-clause) * [GROUP BY Clause](/docs/next/reference/sql/select#group-by-clause) * [HAVING Clause](/docs/next/reference/sql/select#having-clause) * [QUALIFY Clause](/docs/next/reference/sql/select#qualify-clause) * [UNION Clause](/docs/next/reference/sql/select#union-clause) * [ORDER BY Clause](/docs/next/reference/sql/select#order-by-clause) * [LIMIT Clause](/docs/next/reference/sql/select#limit-clause) * [EXCLUDE and EXCEPT Clause](/docs/next/reference/sql/select#exclude-except-replace-and-ilike-clauses) ### [Subqueries](/docs/next/reference/sql/subqueries)[​](#subqueries "Direct link to subqueries") * [Subquery Operators](/docs/next/reference/sql/subqueries#subquery-operators) * [SELECT Clause Subqueries](/docs/next/reference/sql/subqueries#select-clause-subqueries) * [FROM Clause Subqueries](/docs/next/reference/sql/subqueries#from-clause-subqueries) * [WHERE Clause Subqueries](/docs/next/reference/sql/subqueries#where-clause-subqueries) * [HAVING Clause Subqueries](/docs/next/reference/sql/subqueries#having-clause-subqueries) * [Subquery Categories](/docs/next/reference/sql/subqueries#subquery-categories) ### [EXPLAIN](/docs/next/reference/sql/explain)[​](#explain "Direct link to explain") * [EXPLAIN ANALYZE](/docs/next/reference/sql/explain#explain-analyze) ### [Information Schema](/docs/next/reference/sql/information_schema)[​](#information-schema "Direct link to information-schema") * [SHOW TABLES](/docs/next/reference/sql/information_schema#show-tables) * [SHOW COLUMNS](/docs/next/reference/sql/information_schema#show-columns) * [SHOW ALL (configuration options)](/docs/next/reference/sql/information_schema#show-all-configuration-options) ### [AI Functions](/docs/next/reference/sql/ai)[​](#ai-functions "Direct link to ai-functions") * [ai (LLM Text Generation)](/docs/next/reference/sql/ai#ai) * [embed (Vector Embeddings)](/docs/next/reference/sql/ai#embed) ### [Operators and Literals](/docs/next/reference/sql/operators)[​](#operators-and-literals "Direct link to operators-and-literals") * [Numerical Operators](/docs/next/reference/sql/operators#numerical-operators) * [Comparison Operators](/docs/next/reference/sql/operators#comparison-operators) * [Logical Operators](/docs/next/reference/sql/operators#logical-operators) * [Bitwise Operators](/docs/next/reference/sql/operators#bitwise-operators) * [Type Casting Operators](/docs/next/reference/sql/operators#type-casting-operators) * [Other Operators](/docs/next/reference/sql/operators#other-operators) * [Literals](/docs/next/reference/sql/operators#literals) ### [Scalar Functions](/docs/next/reference/sql/scalar_functions)[​](#scalar-functions "Direct link to scalar-functions") * [Math Functions](/docs/next/reference/sql/scalar_functions#math-functions) * [Conditional Functions](/docs/next/reference/sql/scalar_functions#conditional-functions) * [String Functions](/docs/next/reference/sql/scalar_functions#string-functions) * [Binary String Functions](/docs/next/reference/sql/scalar_functions#binary-string-functions) * [Regular Expression Functions](/docs/next/reference/sql/scalar_functions#regular-expression-functions) * [Time and Date Functions](/docs/next/reference/sql/scalar_functions#time-and-date-functions) * [Array Functions](/docs/next/reference/sql/scalar_functions#array-functions) * [Struct Functions](/docs/next/reference/sql/scalar_functions#struct-functions) * [Map Functions](/docs/next/reference/sql/scalar_functions#map-functions) * [Hashing Functions](/docs/next/reference/sql/scalar_functions#hashing-functions) * [Encoding Functions](/docs/next/reference/sql/scalar_functions#encoding-functions) * [Union Functions](/docs/next/reference/sql/scalar_functions#union-functions) * [Other Functions](/docs/next/reference/sql/scalar_functions#other-functions) Spark-compatible scalar functions such as `array`, `bit_get`, `date_add`, `like`, and `parse_url` follow the semantics documented in the [Spark SQL built-in function reference](https://spark.apache.org/docs/latest/api/sql/index.html). ### [Aggregate Functions](/docs/next/reference/sql/aggregate_functions)[​](#aggregate-functions "Direct link to aggregate-functions") * [Filter Clause](/docs/next/reference/sql/aggregate_functions#filter-clause) * [WITHIN GROUP / Ordered-set Aggregates](/docs/next/reference/sql/aggregate_functions#within-group--ordered-set-aggregates) * [General Aggregate Functions](/docs/next/reference/sql/aggregate_functions#general-functions) * [Statistical Aggregate Functions](/docs/next/reference/sql/aggregate_functions#statistical-functions) * [Approximate Aggregate Functions](/docs/next/reference/sql/aggregate_functions#approximate-functions) ### Window Functions[​](#window-functions "Direct link to Window Functions") Window functions perform calculations across sets of rows related to the current row. Spice supports window functions using the `OVER` clause with aggregate and ranking functions. See [Aggregate Functions](/docs/next/reference/sql/aggregate_functions) for functions that support the `OVER` clause, including `ROW_NUMBER`, `RANK`, `DENSE_RANK`, `PERCENT_RANK`, `CUME_DIST`, `NTILE`, `LAG`, `LEAD`, `FIRST_VALUE`, `LAST_VALUE`, and `NTH_VALUE`. **Example:** ``` SELECT dept_id, employee_name, salary, ROW_NUMBER() OVER (PARTITION BY dept_id ORDER BY salary DESC) AS rank FROM employees; ``` ### [JSON Functions and Operators](/docs/next/reference/sql/json)[​](#json-functions-and-operators "Direct link to json-functions-and-operators") * [JSON Functions](/docs/next/reference/sql/json#json-functions) * [JSON Operators](/docs/next/reference/sql/json#json-operators) * [Usage Examples](/docs/next/reference/sql/json#usage-examples) ### [Search](/docs/next/reference/sql/search)[​](#search "Direct link to search") * [Vector Search (`vector_search`)](/docs/next/reference/sql/search#vector-search-vector_search) * [Full-Text Search (`text_search`)](/docs/next/reference/sql/search#full-text-search-text_search) * [Reciprocal Rank Fusion (`rrf`)](/docs/next/reference/sql/search#reciprocal-rank-fusion-rrf) * [Reranking (`rerank`)](/docs/next/reference/sql/search#reranking-rerank) * [Lexical Search: LIKE, =, and Regex](/docs/next/reference/sql/search#lexical-search-like--and-regex) ### [Prepared Statements](/docs/next/reference/sql/prepared_statements)[​](#prepared-statements "Direct link to prepared-statements") * [Positional Arguments](/docs/next/reference/sql/prepared_statements#positional-arguments) ### [DML (Data Manipulation Language)](/docs/next/reference/sql/dml)[​](#dml-data-manipulation-language "Direct link to dml-data-manipulation-language") * [INSERT Statement](/docs/next/reference/sql/dml#insert) ### Data Types[​](#data-types "Direct link to Data Types") Spice uses Apache Arrow data types internally. For data type compatibility with accelerators, see [Data Type Reference](/docs/next/reference/datatypes). Common SQL types include: | SQL Type | Description | | ----------------- | ----------------------------------- | | `INT`, `BIGINT` | Integer types | | `FLOAT`, `DOUBLE` | Floating-point types | | `VARCHAR`, `TEXT` | String types | | `BOOLEAN` | Boolean type | | `TIMESTAMP` | Timestamp with nanosecond precision | | `DATE` | Date type | | `DECIMAL` | Arbitrary precision numeric | Use `CAST(expression AS type)` or `expression::type` to convert between types. Refer to each section for detailed syntax, supported features, and examples. --- # Aggregate Functions info Spice is built on [Apache DataFusion](https://datafusion.apache.org/) and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. Aggregate functions operate on a set of values to compute a single result. Window Functions Many aggregate functions can also be used as **window functions** by adding an `OVER` clause. Window functions perform calculations across a set of rows related to the current row without collapsing them into a single output row. ``` SELECT employee_id, salary, AVG(salary) OVER (PARTITION BY dept_id) AS dept_avg_salary FROM employees; ``` Supported window functions include: `ROW_NUMBER`, `RANK`, `DENSE_RANK`, `PERCENT_RANK`, `CUME_DIST`, `NTILE`, `LAG`, `LEAD`, `FIRST_VALUE`, `LAST_VALUE`, `NTH_VALUE`, and all general aggregate functions. ## Filter clause[​](#filter-clause "Direct link to Filter clause") Aggregate functions support the SQL `FILTER (WHERE ...)` clause to restrict which input rows contribute to the aggregate result. ``` function([exprs]) FILTER (WHERE condition) ``` Example: ``` SELECT sum(salary) FILTER (WHERE salary > 0) AS sum_positive_salaries, count(*) FILTER (WHERE active) AS active_count FROM employees; ``` Note: When no rows pass the filter, `COUNT` returns `0` while `SUM`/`AVG`/`MIN`/`MAX` return `NULL`. ## WITHIN GROUP / Ordered-set Aggregates[​](#within-group--ordered-set-aggregates "Direct link to WITHIN GROUP / Ordered-set Aggregates") Some aggregate functions accept the SQL `WITHIN GROUP (ORDER BY ...)` clause to specify the ordering the aggregate relies on. These are called ordered-set aggregate functions. ``` percentile_cont(0.5) WITHIN GROUP (ORDER BY value) ``` The built-in aggregate functions that support `WITHIN GROUP` are: * `percentile_cont` — exact percentile aggregate * `approx_percentile_cont` — approximate percentile using the t-digest algorithm * `approx_percentile_cont_with_weight` — approximate weighted percentile Using `WITHIN GROUP` with other aggregates (such as `SUM` or `COUNT`) results in an error. ## General Functions[​](#general-functions "Direct link to General Functions") * [array\_agg](#array_agg) * [avg](#avg) * [bit\_and](#bit_and) * [bit\_or](#bit_or) * [bit\_xor](#bit_xor) * [bool\_and](#bool_and) * [bool\_or](#bool_or) * [count](#count) * [first\_value](#first_value) * [grouping](#grouping) * [last\_value](#last_value) * [max](#max) * [mean](#mean) * [median](#median) * [min](#min) * [percentile\_cont](#percentile_cont) * [string\_agg](#string_agg) * [sum](#sum) * [var](#var) * [var\_pop](#var_pop) * [var\_population](#var_population) * [var\_samp](#var_samp) * [var\_sample](#var_sample) ### `array_agg`[​](#array_agg "Direct link to array_agg") Returns an array created from the expression elements. If ordering is required, elements are inserted in the specified order. This aggregation function can only mix DISTINCT and ORDER BY if the ordering expression is exactly the same as the argument expression. ``` array_agg(expression [ORDER BY expression]) ``` #### Arguments[​](#arguments "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example "Direct link to Example") ``` > SELECT array_agg(column_name ORDER BY other_column) FROM table_name; +-----------------------------------------------+ | array_agg(column_name ORDER BY other_column) | +-----------------------------------------------+ | [element1, element2, element3] | +-----------------------------------------------+ > SELECT array_agg(DISTINCT column_name ORDER BY column_name) FROM table_name; +--------------------------------------------------------+ | array_agg(DISTINCT column_name ORDER BY column_name) | +--------------------------------------------------------+ | [element1, element2, element3] | +--------------------------------------------------------+ ``` ### `avg`[​](#avg "Direct link to avg") Returns the average of numeric values in the specified column. ``` avg(expression) ``` #### Arguments[​](#arguments-1 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-1 "Direct link to Example") ``` > SELECT avg(column_name) FROM table_name; +---------------------------+ | avg(column_name) | +---------------------------+ | 42.75 | +---------------------------+ ``` #### Aliases[​](#aliases "Direct link to Aliases") * mean ### `bit_and`[​](#bit_and "Direct link to bit_and") Computes the bitwise AND of all non-null input values. ``` bit_and(expression) ``` #### Arguments[​](#arguments-2 "Direct link to Arguments") * **expression**: Integer expression to operate on. Can be a constant, column, or function, and any combination of operators. ### `bit_or`[​](#bit_or "Direct link to bit_or") Computes the bitwise OR of all non-null input values. ``` bit_or(expression) ``` #### Arguments[​](#arguments-3 "Direct link to Arguments") * **expression**: Integer expression to operate on. Can be a constant, column, or function, and any combination of operators. ### `bit_xor`[​](#bit_xor "Direct link to bit_xor") Computes the bitwise exclusive OR of all non-null input values. ``` bit_xor(expression) ``` #### Arguments[​](#arguments-4 "Direct link to Arguments") * **expression**: Integer expression to operate on. Can be a constant, column, or function, and any combination of operators. ### `bool_and`[​](#bool_and "Direct link to bool_and") Returns true if all non-null input values are true, otherwise false. ``` bool_and(expression) ``` #### Arguments[​](#arguments-5 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-2 "Direct link to Example") ``` > SELECT bool_and(column_name) FROM table_name; +----------------------------+ | bool_and(column_name) | +----------------------------+ | true | +----------------------------+ ``` ### `bool_or`[​](#bool_or "Direct link to bool_or") Returns true if any non-null input value is true, otherwise false. ``` bool_or(expression) ``` #### Arguments[​](#arguments-6 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-3 "Direct link to Example") ``` > SELECT bool_or(column_name) FROM table_name; +----------------------------+ | bool_or(column_name) | +----------------------------+ | true | +----------------------------+ ``` ### `count`[​](#count "Direct link to count") Returns the number of non-null values in the specified column. To include null values in the total count, use `count(*)`. ``` count(expression) ``` #### Arguments[​](#arguments-7 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-4 "Direct link to Example") ``` > SELECT count(column_name) FROM table_name; +-----------------------+ | count(column_name) | +-----------------------+ | 100 | +-----------------------+ > SELECT count(*) FROM table_name; +------------------+ | count(*) | +------------------+ | 120 | +------------------+ ``` ### `first_value`[​](#first_value "Direct link to first_value") Returns the first element in an aggregation group according to the requested ordering. If no ordering is given, returns an arbitrary element from the group. ``` first_value(expression [ORDER BY expression]) ``` #### Arguments[​](#arguments-8 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-5 "Direct link to Example") ``` > SELECT first_value(column_name ORDER BY other_column) FROM table_name; +-----------------------------------------------+ | first_value(column_name ORDER BY other_column)| +-----------------------------------------------+ | first_element | +-----------------------------------------------+ ``` ### `grouping`[​](#grouping "Direct link to grouping") Returns 1 if the data is aggregated across the specified column, or 0 if it is not aggregated in the result set. ``` grouping(expression) ``` #### Arguments[​](#arguments-9 "Direct link to Arguments") * **expression**: Expression to evaluate whether data is aggregated across the specified column. Can be a constant, column, or function. #### Example[​](#example-6 "Direct link to Example") ``` > SELECT column_name, GROUPING(column_name) AS group_column FROM table_name GROUP BY GROUPING SETS ((column_name), ()); +-------------+-------------+ | column_name | group_column | +-------------+-------------+ | value1 | 0 | | value2 | 0 | | NULL | 1 | +-------------+-------------+ ``` ### `last_value`[​](#last_value "Direct link to last_value") Returns the last element in an aggregation group according to the requested ordering. If no ordering is given, returns an arbitrary element from the group. ``` last_value(expression [ORDER BY expression]) ``` #### Arguments[​](#arguments-10 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-7 "Direct link to Example") ``` > SELECT last_value(column_name ORDER BY other_column) FROM table_name; +-----------------------------------------------+ | last_value(column_name ORDER BY other_column) | +-----------------------------------------------+ | last_element | +-----------------------------------------------+ ``` ### `max`[​](#max "Direct link to max") Returns the maximum value in the specified column. ``` max(expression) ``` #### Arguments[​](#arguments-11 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-8 "Direct link to Example") ``` > SELECT max(column_name) FROM table_name; +----------------------+ | max(column_name) | +----------------------+ | 150 | +----------------------+ ``` ### `mean`[​](#mean "Direct link to mean") *Alias of [avg](#avg).* ### `median`[​](#median "Direct link to median") Returns the median value in the specified column. ``` median(expression) ``` #### Arguments[​](#arguments-12 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-9 "Direct link to Example") ``` > SELECT median(column_name) FROM table_name; +----------------------+ | median(column_name) | +----------------------+ | 45.5 | +----------------------+ ``` ### `min`[​](#min "Direct link to min") Returns the minimum value in the specified column. ``` min(expression) ``` #### Arguments[​](#arguments-13 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-10 "Direct link to Example") ``` > SELECT min(column_name) FROM table_name; +----------------------+ | min(column_name) | +----------------------+ | 12 | +----------------------+ ``` ### `percentile_cont`[​](#percentile_cont "Direct link to percentile_cont") Returns the exact percentile of input values, interpolating between values if needed. ``` percentile_cont(percentile) WITHIN GROUP (ORDER BY expression) ``` #### Arguments[​](#arguments-14 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. * **percentile**: Percentile to compute. Must be a float value between 0 and 1 (inclusive). #### Example[​](#example-11 "Direct link to Example") ``` > SELECT percentile_cont(0.75) WITHIN GROUP (ORDER BY column_name) FROM table_name; +----------------------------------------------------------+ | percentile_cont(0.75) WITHIN GROUP (ORDER BY column_name) | +----------------------------------------------------------+ | 45.5 | +----------------------------------------------------------+ ``` An alternate syntax is also supported: ``` > SELECT percentile_cont(column_name, 0.75) FROM table_name; +---------------------------------------+ | percentile_cont(column_name, 0.75) | +---------------------------------------+ | 45.5 | +---------------------------------------+ ``` #### Aliases[​](#aliases-1 "Direct link to Aliases") * quantile\_cont ### `string_agg`[​](#string_agg "Direct link to string_agg") Concatenates the values of string expressions and places separator values between them. If ordering is required, strings are concatenated in the specified order. This aggregation function can only mix DISTINCT and ORDER BY if the ordering expression is exactly the same as the first argument expression. ``` string_agg([DISTINCT] expression, delimiter [ORDER BY expression]) ``` #### Arguments[​](#arguments-15 "Direct link to Arguments") * **expression**: The string expression to concatenate. Can be a column or any valid string expression. * **delimiter**: A literal string used as a separator between the concatenated values. #### Example[​](#example-12 "Direct link to Example") ``` > SELECT string_agg(name, ', ') AS names_list FROM employee; +--------------------------+ | names_list | +--------------------------+ | Alice, Bob, Bob, Charlie | +--------------------------+ > SELECT string_agg(name, ', ' ORDER BY name DESC) AS names_list FROM employee; +--------------------------+ | names_list | +--------------------------+ | Charlie, Bob, Bob, Alice | +--------------------------+ > SELECT string_agg(DISTINCT name, ', ' ORDER BY name DESC) AS names_list FROM employee; +--------------------------+ | names_list | +--------------------------+ | Charlie, Bob, Alice | +--------------------------+ ``` ### `sum`[​](#sum "Direct link to sum") Returns the sum of all values in the specified column. ``` sum(expression) ``` #### Arguments[​](#arguments-16 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-13 "Direct link to Example") ``` > SELECT sum(column_name) FROM table_name; +-----------------------+ | sum(column_name) | +-----------------------+ | 12345 | +-----------------------+ ``` ### `var`[​](#var "Direct link to var") Returns the statistical sample variance of a set of numbers. ``` var(expression) ``` #### Arguments[​](#arguments-17 "Direct link to Arguments") * **expression**: Numeric expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Aliases[​](#aliases-2 "Direct link to Aliases") * var\_sample * var\_samp ### `var_pop`[​](#var_pop "Direct link to var_pop") Returns the statistical population variance of a set of numbers. ``` var_pop(expression) ``` #### Arguments[​](#arguments-18 "Direct link to Arguments") * **expression**: Numeric expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Aliases[​](#aliases-3 "Direct link to Aliases") * var\_population ### `var_population`[​](#var_population "Direct link to var_population") *Alias of [var\_pop](#var_pop).* ### `var_samp`[​](#var_samp "Direct link to var_samp") *Alias of [var](#var).* ### `var_sample`[​](#var_sample "Direct link to var_sample") *Alias of [var](#var).* ## Statistical Functions[​](#statistical-functions "Direct link to Statistical Functions") * [corr](#corr) * [covar](#covar) * [covar\_pop](#covar_pop) * [covar\_samp](#covar_samp) * [nth\_value](#nth_value) * [regr\_avgx](#regr_avgx) * [regr\_avgy](#regr_avgy) * [regr\_count](#regr_count) * [regr\_intercept](#regr_intercept) * [regr\_r2](#regr_r2) * [regr\_slope](#regr_slope) * [regr\_sxx](#regr_sxx) * [regr\_sxy](#regr_sxy) * [regr\_syy](#regr_syy) * [stddev](#stddev) * [stddev\_pop](#stddev_pop) * [stddev\_samp](#stddev_samp) ### `corr`[​](#corr "Direct link to corr") Returns the coefficient of correlation between two numeric values. ``` corr(expression1, expression2) ``` #### Arguments[​](#arguments-19 "Direct link to Arguments") * **expression1**: First expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression2**: Second expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-14 "Direct link to Example") ``` > SELECT corr(column1, column2) FROM table_name; +--------------------------------+ | corr(column1, column2) | +--------------------------------+ | 0.85 | +--------------------------------+ ``` ### `covar`[​](#covar "Direct link to covar") *Alias of [covar\_samp](#covar_samp).* ### `covar_pop`[​](#covar_pop "Direct link to covar_pop") Returns the sample covariance of a set of number pairs. ``` covar_samp(expression1, expression2) ``` #### Arguments[​](#arguments-20 "Direct link to Arguments") * **expression1**: First expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression2**: Second expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-15 "Direct link to Example") ``` > SELECT covar_samp(column1, column2) FROM table_name; +-----------------------------------+ | covar_samp(column1, column2) | +-----------------------------------+ | 8.25 | +-----------------------------------+ ``` ### `covar_samp`[​](#covar_samp "Direct link to covar_samp") Returns the sample covariance of a set of number pairs. ``` covar_samp(expression1, expression2) ``` #### Arguments[​](#arguments-21 "Direct link to Arguments") * **expression1**: First expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression2**: Second expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-16 "Direct link to Example") ``` > SELECT covar_samp(column1, column2) FROM table_name; +-----------------------------------+ | covar_samp(column1, column2) | +-----------------------------------+ | 8.25 | +-----------------------------------+ ``` #### Aliases[​](#aliases-4 "Direct link to Aliases") * covar ### `nth_value`[​](#nth_value "Direct link to nth_value") Returns the nth value in a group of values. ``` nth_value(expression, n ORDER BY expression) ``` #### Arguments[​](#arguments-22 "Direct link to Arguments") * **expression**: The column or expression to retrieve the nth value from. * **n**: The position (nth) of the value to retrieve, based on the ordering. #### Example[​](#example-17 "Direct link to Example") ``` > SELECT dept_id, salary, NTH_VALUE(salary, 2) OVER (PARTITION BY dept_id ORDER BY salary ASC) AS second_salary_by_dept FROM employee; +---------+--------+-------------------------+ | dept_id | salary | second_salary_by_dept | +---------+--------+-------------------------+ | 1 | 30000 | NULL | | 1 | 40000 | 40000 | | 1 | 50000 | 40000 | | 2 | 35000 | NULL | | 2 | 45000 | 45000 | +---------+--------+-------------------------+ ``` ### `regr_avgx`[​](#regr_avgx "Direct link to regr_avgx") Computes the average of the independent variable (input) expression\_x for the non-null paired data points. ``` regr_avgx(expression_y, expression_x) ``` #### Arguments[​](#arguments-23 "Direct link to Arguments") * **expression\_y**: Dependent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression\_x**: Independent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-18 "Direct link to Example") ``` create table daily_sales(day int, total_sales int) as values (1,100), (2,150), (3,200), (4,NULL), (5,250); select * from daily_sales; +-----+-------------+ | day | total_sales | | --- | ----------- | | 1 | 100 | | 2 | 150 | | 3 | 200 | | 4 | NULL | | 5 | 250 | +-----+-------------+ SELECT regr_avgx(total_sales, day) AS avg_day FROM daily_sales; +----------+ | avg_day | +----------+ | 2.75 | +----------+ ``` ### `regr_avgy`[​](#regr_avgy "Direct link to regr_avgy") Computes the average of the dependent variable (output) expression\_y for the non-null paired data points. ``` regr_avgy(expression_y, expression_x) ``` #### Arguments[​](#arguments-24 "Direct link to Arguments") * **expression\_y**: Dependent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression\_x**: Independent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-19 "Direct link to Example") ``` create table daily_temperature(day int, temperature int) as values (1,30), (2,32), (3, NULL), (4,35), (5,36); select * from daily_temperature; +-----+-------------+ | day | temperature | | --- | ----------- | | 1 | 30 | | 2 | 32 | | 3 | NULL | | 4 | 35 | | 5 | 36 | +-----+-------------+ -- temperature as Dependent Variable(Y), day as Independent Variable(X) SELECT regr_avgy(temperature, day) AS avg_temperature FROM daily_temperature; +-----------------+ | avg_temperature | +-----------------+ | 33.25 | +-----------------+ ``` ### `regr_count`[​](#regr_count "Direct link to regr_count") Counts the number of non-null paired data points. ``` regr_count(expression_y, expression_x) ``` #### Arguments[​](#arguments-25 "Direct link to Arguments") * **expression\_y**: Dependent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression\_x**: Independent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-20 "Direct link to Example") ``` create table daily_metrics(day int, user_signups int) as values (1,100), (2,120), (3, NULL), (4,110), (5,NULL); select * from daily_metrics; +-----+---------------+ | day | user_signups | | --- | ------------- | | 1 | 100 | | 2 | 120 | | 3 | NULL | | 4 | 110 | | 5 | NULL | +-----+---------------+ SELECT regr_count(user_signups, day) AS valid_pairs FROM daily_metrics; +-------------+ | valid_pairs | +-------------+ | 3 | +-------------+ ``` ### `regr_intercept`[​](#regr_intercept "Direct link to regr_intercept") Computes the y-intercept of the linear regression line. For the equation (y = kx + b), this function returns b. ``` regr_intercept(expression_y, expression_x) ``` #### Arguments[​](#arguments-26 "Direct link to Arguments") * **expression\_y**: Dependent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression\_x**: Independent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-21 "Direct link to Example") ``` create table weekly_performance(week int, productivity_score int) as values (1,60), (2,65), (3, 70), (4,75), (5,80); select * from weekly_performance; +------+---------------------+ | week | productivity_score | | ---- | ------------------- | | 1 | 60 | | 2 | 65 | | 3 | 70 | | 4 | 75 | | 5 | 80 | +------+---------------------+ SELECT regr_intercept(productivity_score, week) AS intercept FROM weekly_performance; +----------+ |intercept| |intercept | +----------+ | 55 | +----------+ ``` ### `regr_r2`[​](#regr_r2 "Direct link to regr_r2") Computes the square of the correlation coefficient between the independent and dependent variables. ``` regr_r2(expression_y, expression_x) ``` #### Arguments[​](#arguments-27 "Direct link to Arguments") * **expression\_y**: Dependent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression\_x**: Independent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-22 "Direct link to Example") ``` create table weekly_performance(day int ,user_signups int) as values (1,60), (2,65), (3, 70), (4,75), (5,80); select * from weekly_performance; +-----+--------------+ | day | user_signups | +-----+--------------+ | 1 | 60 | | 2 | 65 | | 3 | 70 | | 4 | 75 | | 5 | 80 | +-----+--------------+ SELECT regr_r2(user_signups, day) AS r_squared FROM weekly_performance; +---------+ |r_squared| +---------+ | 1.0 | +---------+ ``` ### `regr_slope`[​](#regr_slope "Direct link to regr_slope") Returns the slope of the linear regression line for non-null pairs in aggregate columns. Given input column Y and X: regr\_slope(Y, X) returns the slope (k in Y = k\*X + b) using minimal RSS fitting. ``` regr_slope(expression_y, expression_x) ``` #### Arguments[​](#arguments-28 "Direct link to Arguments") * **expression\_y**: Dependent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression\_x**: Independent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-23 "Direct link to Example") ``` create table weekly_performance(day int, user_signups int) as values (1,60), (2,65), (3, 70), (4,75), (5,80); select * from weekly_performance; +-----+--------------+ | day | user_signups | +-----+--------------+ | 1 | 60 | | 2 | 65 | | 3 | 70 | | 4 | 75 | | 5 | 80 | +-----+--------------+ SELECT regr_slope(user_signups, day) AS slope FROM weekly_performance; +--------+ | slope | +--------+ | 5.0 | +--------+ ``` ### `regr_sxx`[​](#regr_sxx "Direct link to regr_sxx") Computes the sum of squares of the independent variable. ``` regr_sxx(expression_y, expression_x) ``` #### Arguments[​](#arguments-29 "Direct link to Arguments") * **expression\_y**: Dependent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression\_x**: Independent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-24 "Direct link to Example") ``` create table study_hours(student_id int, hours int, test_score int) as values (1,2,55), (2,4,65), (3,6,75), (4,8,85), (5,10,95); select * from study_hours; +------------+-------+------------+ | student_id | hours | test_score | +------------+-------+------------+ | 1 | 2 | 55 | | 2 | 4 | 65 | | 3 | 6 | 75 | | 4 | 8 | 85 | | 5 | 10 | 95 | +------------+-------+------------+ SELECT regr_sxx(test_score, hours) AS sxx FROM study_hours; +------+ | sxx | +------+ | 40.0 | +------+ ``` ### `regr_sxy`[​](#regr_sxy "Direct link to regr_sxy") Computes the sum of products of paired data points. ``` regr_sxy(expression_y, expression_x) ``` #### Arguments[​](#arguments-30 "Direct link to Arguments") * **expression\_y**: Dependent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression\_x**: Independent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-25 "Direct link to Example") ``` create table employee_productivity(week int, productivity_score int) as values(1,60), (2,65), (3,70); select * from employee_productivity; +------+--------------------+ | week | productivity_score | +------+--------------------+ | 1 | 60 | | 2 | 65 | | 3 | 70 | +------+--------------------+ SELECT regr_sxy(productivity_score, week) AS sum_product_deviations FROM employee_productivity; +------------------------+ | sum_product_deviations | +------------------------+ | 10.0 | +------------------------+ ``` ### `regr_syy`[​](#regr_syy "Direct link to regr_syy") Computes the sum of squares of the dependent variable. ``` regr_syy(expression_y, expression_x) ``` #### Arguments[​](#arguments-31 "Direct link to Arguments") * **expression\_y**: Dependent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. * **expression\_x**: Independent variable expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-26 "Direct link to Example") ``` create table employee_productivity(week int, productivity_score int) as values (1,60), (2,65), (3,70); select * from employee_productivity; +------+--------------------+ | week | productivity_score | +------+--------------------+ | 1 | 60 | | 2 | 65 | | 3 | 70 | +------+--------------------+ SELECT regr_syy(productivity_score, week) AS sum_squares_y FROM employee_productivity; +---------------+ | sum_squares_y | +---------------+ | 50.0 | +---------------+ ``` ### `stddev`[​](#stddev "Direct link to stddev") Returns the standard deviation of a set of numbers. ``` stddev(expression) ``` #### Arguments[​](#arguments-32 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-27 "Direct link to Example") ``` > SELECT stddev(column_name) FROM table_name; +----------------------+ | stddev(column_name) | +----------------------+ | 12.34 | +----------------------+ ``` #### Aliases[​](#aliases-5 "Direct link to Aliases") * stddev\_samp ### `stddev_pop`[​](#stddev_pop "Direct link to stddev_pop") Returns the population standard deviation of a set of numbers. ``` stddev_pop(expression) ``` #### Arguments[​](#arguments-33 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-28 "Direct link to Example") ``` > SELECT stddev_pop(column_name) FROM table_name; +--------------------------+ | stddev_pop(column_name) | +--------------------------+ | 10.56 | +--------------------------+ ``` ### `stddev_samp`[​](#stddev_samp "Direct link to stddev_samp") *Alias of [stddev](#stddev).* ## Approximate Functions[​](#approximate-functions "Direct link to Approximate Functions") * [approx\_distinct](#approx_distinct) * [approx\_median](#approx_median) * [approx\_percentile\_cont](#approx_percentile_cont) * [approx\_percentile\_cont\_with\_weight](#approx_percentile_cont_with_weight) ### `approx_distinct`[​](#approx_distinct "Direct link to approx_distinct") Returns the approximate number of distinct input values calculated using the HyperLogLog algorithm. ``` approx_distinct(expression) ``` #### Arguments[​](#arguments-34 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-29 "Direct link to Example") ``` > SELECT approx_distinct(column_name) FROM table_name; +-----------------------------------+ | approx_distinct(column_name) | +-----------------------------------+ | 42 | +-----------------------------------+ ``` ### `approx_median`[​](#approx_median "Direct link to approx_median") Returns the approximate median (50th percentile) of input values. It is an alias of `approx_percentile_cont(0.5) WITHIN GROUP (ORDER BY x)`. ``` approx_median(expression) ``` #### Arguments[​](#arguments-35 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. #### Example[​](#example-30 "Direct link to Example") ``` > SELECT approx_median(column_name) FROM table_name; +-----------------------------------+ | approx_median(column_name) | +-----------------------------------+ | 23.5 | +-----------------------------------+ ``` ### `approx_percentile_cont`[​](#approx_percentile_cont "Direct link to approx_percentile_cont") Returns the approximate percentile of input values using the t-digest algorithm. ``` approx_percentile_cont(percentile [, centroids]) WITHIN GROUP (ORDER BY expression) ``` #### Arguments[​](#arguments-36 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. * **percentile**: Percentile to compute. Must be a float value between 0 and 1 (inclusive). * **centroids**: Number of centroids to use in the t-digest algorithm. *Default is 100*. A higher number results in more accurate approximation but requires more memory. #### Example[​](#example-31 "Direct link to Example") ``` > SELECT approx_percentile_cont(0.75) WITHIN GROUP (ORDER BY column_name) FROM table_name; +------------------------------------------------------------------+ | approx_percentile_cont(0.75) WITHIN GROUP (ORDER BY column_name) | +------------------------------------------------------------------+ | 65.0 | +------------------------------------------------------------------+ > SELECT approx_percentile_cont(0.75, 100) WITHIN GROUP (ORDER BY column_name) FROM table_name; +-----------------------------------------------------------------------+ | approx_percentile_cont(0.75, 100) WITHIN GROUP (ORDER BY column_name) | +-----------------------------------------------------------------------+ | 65.0 | +-----------------------------------------------------------------------+ ``` An alternate syntax is also supported: ``` > SELECT approx_percentile_cont(column_name, 0.75) FROM table_name; +-----------------------------------------------+ | approx_percentile_cont(column_name, 0.75) | +-----------------------------------------------+ | 65.0 | +-----------------------------------------------+ > SELECT approx_percentile_cont(column_name, 0.75, 100) FROM table_name; +----------------------------------------------------------+ | approx_percentile_cont(column_name, 0.75, 100) | +----------------------------------------------------------+ | 65.0 | +----------------------------------------------------------+ ``` ### `approx_percentile_cont_with_weight`[​](#approx_percentile_cont_with_weight "Direct link to approx_percentile_cont_with_weight") Returns the weighted approximate percentile of input values using the t-digest algorithm. ``` approx_percentile_cont_with_weight(weight, percentile [, centroids]) WITHIN GROUP (ORDER BY expression) ``` #### Arguments[​](#arguments-37 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. * **weight**: Expression to use as weight. Can be a constant, column, or function, and any combination of arithmetic operators. * **percentile**: Percentile to compute. Must be a float value between 0 and 1 (inclusive). * **centroids**: Number of centroids to use in the t-digest algorithm. *Default is 100*. A higher number results in more accurate approximation but requires more memory. #### Example[​](#example-32 "Direct link to Example") ``` > SELECT approx_percentile_cont_with_weight(weight_column, 0.90) WITHIN GROUP (ORDER BY column_name) FROM table_name; +---------------------------------------------------------------------------------------------+ | approx_percentile_cont_with_weight(weight_column, 0.90) WITHIN GROUP (ORDER BY column_name) | +---------------------------------------------------------------------------------------------+ | 78.5 | +---------------------------------------------------------------------------------------------+ > SELECT approx_percentile_cont_with_weight(weight_column, 0.90, 100) WITHIN GROUP (ORDER BY column_name) FROM table_name; +--------------------------------------------------------------------------------------------------+ | approx_percentile_cont_with_weight(weight_column, 0.90, 100) WITHIN GROUP (ORDER BY column_name) | +--------------------------------------------------------------------------------------------------+ | 78.5 | +--------------------------------------------------------------------------------------------------+ ``` An alternative syntax is also supported: ``` > SELECT approx_percentile_cont_with_weight(column_name, weight_column, 0.90) FROM table_name; +--------------------------------------------------+ | approx_percentile_cont_with_weight(column_name, weight_column, 0.90) | +--------------------------------------------------+ | 78.5 | +--------------------------------------------------+ ``` --- # AI Functions AI functions in Spice provide direct integration with large language models (LLMs) and embedding models within SQL queries. These functions process text through configured model providers and return generated responses or vector embeddings. ## `ai`[​](#ai "Direct link to ai") Invokes large language models (LLMs) directly within SQL queries for text generation tasks. This asynchronous function processes prompts through configured model providers and returns generated text responses. ``` ai(message) ai(message, model_name) ``` ### Arguments[​](#arguments "Direct link to Arguments") * **message**: String prompt to send to the language model. * **model\_name** (optional): Name of the model to use as configured in your Spicepod. If omitted, the default model is used (only valid when exactly one model is configured). ### Return Type[​](#return-type "Direct link to Return Type") Returns a string containing the generated text response. Returns NULL for NULL input messages or when the model returns an empty response. If the underlying model call fails, the query errors out (errors are logged). ### Behavior[​](#behavior "Direct link to Behavior") Queries execute asynchronously, processing LLM calls in parallel across rows for improved performance. Each invocation issues an asynchronous call to the specified model provider. **Concurrency**: When a per-model rate controller is configured in the Spicepod, it manages concurrency and backpressure. Otherwise, concurrency falls back to DataFusion's `execution.target_partitions` setting. When multiple models with different providers are configured (e.g., OpenAI and Anthropic), each provider's requests are controlled independently. **Limits**: * Maximum batch size: 1000 rows per invocation * Maximum input message size: 1,000,000 bytes (1 MB) per message * Maximum model name length: 256 characters ### Configuration[​](#configuration "Direct link to Configuration") Models must be configured in `spicepod.yaml` under the `models` section. See [Large Language Models](/docs/next/features/large-language-models) for detailed configuration. ``` models: - name: gpt-4o from: openai:gpt-4o params: openai_api_key: ${secrets:openai_key} ``` ### Task History[​](#task-history "Direct link to Task History") Each `ai()` function invocation creates an `ai` span in the [task\_history](/docs/next/reference/task_history) table, which tracks execution time, input prompts, model name, and row count. Child `ai_completion` spans capture per-row model call details. ### Examples[​](#examples "Direct link to Examples") #### Using default model[​](#using-default-model "Direct link to Using default model") When only one model is configured, the model name can be omitted: ``` SELECT zone, ai(concat_ws(' ', 'Categorize the zone', zone, 'in a single word. Only return the word.')) AS category FROM taxi_zones LIMIT 10; ``` Example output: ``` +-----------------------+-------------+ | zone | category | +-----------------------+-------------+ | Newark Airport | Transport | | Jamaica Bay | Nature | | Allerton/Pelham... | Residential | | Alphabet City | Urban | | Arden Heights | Suburban | +-----------------------+-------------+ ``` #### Specifying a model explicitly[​](#specifying-a-model-explicitly "Direct link to Specifying a model explicitly") When multiple models are configured, specify which model to use: ``` SELECT product_name, ai('Summarize this product in 5 words: ' || product_description, 'gpt-4o') AS summary FROM products LIMIT 5; ``` #### Generating search terms[​](#generating-search-terms "Direct link to Generating search terms") Use the `ai()` function to generate search terms for enhanced search functionality: ``` SELECT user_query, ai('Generate 3 alternative search terms for: ' || user_query, 'gpt-4o-mini') AS alt_terms FROM user_searches WHERE created_at > NOW() - INTERVAL '1 hour'; ``` #### Batch text processing[​](#batch-text-processing "Direct link to Batch text processing") Process multiple rows efficiently with parallel LLM calls. Each `ai()` invocation is capped at 1000 rows, so use `LIMIT` to stay within this bound: ``` SELECT customer_id, feedback, ai('Classify sentiment as positive, negative, or neutral: ' || feedback) AS sentiment FROM customer_feedback WHERE processed = false LIMIT 1000; ``` ## `embed`[​](#embed "Direct link to embed") Generates vector embeddings for text using specified embedding models. Supports both single text strings and arrays of text for batch processing. ``` embed(text, model_name) ``` ### Arguments[​](#arguments-1 "Direct link to Arguments") * **text**: String or array of strings to generate embeddings for. `Utf8` / `LargeUtf8` scalars and lists/arrays of strings are supported. * **model\_name**: Name of the embedding model to use (e.g., 'potion\_2m', 'xl\_embed') as configured in your Spicepod. Required — unlike `ai()`, `embed()` does not auto-select a default when only one model is configured. ### Return Type[​](#return-type-1 "Direct link to Return Type") * For a scalar string input: `List` — a single embedding vector. * For an array of strings: `List>` — one embedding vector per element, preserving the input length. NULL input elements produce NULL output elements. ### Configuration[​](#configuration-1 "Direct link to Configuration") Embedding models must be configured in `spicepod.yaml`. See [Embeddings](/docs/next/features/embeddings) for configuration details. ``` embeddings: - from: openai:text-embedding-3-small name: openai_embed params: openai_api_key: ${secrets:openai_key} ``` ### Examples[​](#examples-1 "Direct link to Examples") #### Single text embedding[​](#single-text-embedding "Direct link to Single text embedding") Generate an embedding for a single piece of text: ``` SELECT embed('hello world', 'potion_2m') AS embedding; ``` #### Multiple text embeddings[​](#multiple-text-embeddings "Direct link to Multiple text embeddings") Generate embeddings for multiple strings in one call: ``` SELECT embed(['hey', 'there', 'sunshine'], 'potion_2m') AS embeddings; ``` #### Embedding for vector search[​](#embedding-for-vector-search "Direct link to Embedding for vector search") Generate embeddings to use with vector search: ``` SELECT title, embed(content, 'xl_embed') AS content_embedding FROM documents WHERE embedding IS NULL LIMIT 1000; ``` *** For more information on configuring and using AI models, see: * [Large Language Models](/docs/next/features/large-language-models) * [Embeddings](/docs/next/features/embeddings) * [Vector Search](/docs/next/features/search/vector-search) * [Task History](/docs/next/reference/task_history) --- # DML (Data Manipulation Language) Data Manipulation Language (DML) statements are used to insert, update, and delete data in tables. Spice supports DML operations on [write-capable data connectors](/docs/next/tags/write) configured with `access: read_write`. Supported Operations Spice supports `INSERT` for write-capable connectors and `MERGE INTO` for [Spice Cayenne](/docs/next/components/data-accelerators/cayenne) catalog tables. `UPDATE` and `DELETE` statements are not yet supported as standalone operations. For data modifications, use `MERGE INTO` or the source database directly. info Spice is built on [Apache DataFusion](https://datafusion.apache.org/) and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. ## INSERT[​](#insert "Direct link to INSERT") Insert new rows into a table. ### Syntax[​](#syntax "Direct link to Syntax") ``` INSERT INTO table_name [ ( column_name [, ...] ) ] { VALUES ( expression [, ...] ) [, ...] | query } ``` ### Parameters[​](#parameters "Direct link to Parameters") * **`table_name`**: The name of the target table * **`column_name`**: Optional list of column names to insert into. If omitted, values must be provided for all columns in table order * **`expression`**: Values to insert into the corresponding columns * **`query`**: A SELECT statement to insert results from another table or query ### Examples[​](#examples "Direct link to Examples") #### Insert Single or Multiple Rows[​](#insert-single-or-multiple-rows "Direct link to Insert Single or Multiple Rows") ``` INSERT INTO customers (id, name, email) VALUES (1, 'Alice Smith', 'alice@example.com'); ``` ``` +-------+ | count | +-------+ | 1 | +-------+ ``` ``` INSERT INTO customers (id, name, email) VALUES (2, 'Bob Johnson', 'bob@example.com'), (3, 'Carol Wilson', 'carol@example.com'), (4, 'David Brown', 'david@example.com'); ``` ``` +-------+ | count | +-------+ | 3 | +-------+ ``` #### Insert All Columns (Optional Column List)[​](#insert-all-columns-optional-column-list "Direct link to Insert All Columns (Optional Column List)") ``` INSERT INTO products VALUES (101, 'Laptop', 999.99, 'Electronics'); ``` #### Insert from Query[​](#insert-from-query "Direct link to Insert from Query") ``` INSERT INTO archive_orders (order_id, customer_id, total, order_date) SELECT order_id, customer_id, total, order_date FROM orders WHERE order_date < '2024-01-01'; ``` --- # Explain info Spice is built on [Apache DataFusion](https://datafusion.apache.org/) and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. The `EXPLAIN` command shows the logical and physical execution plan of a SQL statement. `EXPLAIN [ANALYZE] [VERBOSE] statement` Shows the execution plan of a statement. Use `EXPLAIN VERBOSE` if more detailed output is needed. ``` EXPLAIN SELECT SUM(x) FROM table GROUP BY b; +---------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------+ | plan_type | plan | +---------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------+ | logical_plan | Projection: #SUM(table.x) | | | Aggregate: groupBy=[[#table.b]], aggr=[[SUM(#table.x)]] | | | TableScan: table projection=[x, b] | | physical_plan | ProjectionExec: expr=[SUM(table.x)@1 as SUM(table.x)] | | | AggregateExec: mode=FinalPartitioned, gby=[b@0 as b], aggr=[SUM(table.x)] | | | CoalesceBatchesExec: target_batch_size=4096 | | | RepartitionExec: partitioning=Hash([Column { name: "b", index: 0 }], 16) | | | AggregateExec: mode=Partial, gby=[b@1 as b], aggr=[SUM(table.x)] | | | RepartitionExec: partitioning=RoundRobinBatch(16) | | | DataSourceExec: file_groups={1 group: [[/tmp/table.csv]]}, projection=[x, b], has_header=false | | | | +---------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------+ ``` ## EXPLAIN ANALYZE[​](#explain-analyze "Direct link to EXPLAIN ANALYZE") Shows the execution plan of a statement. Use `EXPLAIN ANALYZE VERBOSE` if more detailed output is needed. ``` EXPLAIN ANALYZE SELECT SUM(x) FROM table GROUP BY b; +-------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------+ | plan_type | plan | +-------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------+ | Plan with Metrics | CoalescePartitionsExec, metrics=[] | | | ProjectionExec: expr=[SUM(table.x)@1 as SUM(x)], metrics=[] | | | HashAggregateExec: mode=FinalPartitioned, gby=[b@0 as b], aggr=[SUM(x)], metrics=[outputRows=2] | | | CoalesceBatchesExec: target_batch_size=4096, metrics=[] | | | RepartitionExec: partitioning=Hash([Column { name: "b", index: 0 }], 16), metrics=[sendTime=839560, fetchTime=122528525, repartitionTime=5327877] | | | HashAggregateExec: mode=Partial, gby=[b@1 as b], aggr=[SUM(x)], metrics=[outputRows=2] | | | RepartitionExec: partitioning=RoundRobinBatch(16), metrics=[fetchTime=5660489, repartitionTime=0, sendTime=8012] | | | DataSourceExec: file_groups={1 group: [[/tmp/table.csv]]}, has_header=false, metrics=[] | +-------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------+ ``` --- # Information Schema info Spice is built on [Apache DataFusion](https://datafusion.apache.org/) and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. Spice supports display metadata about available tables and views. This information is accessible through the ISO SQL `information_schema` schema or the `SHOW TABLES` and `SHOW COLUMNS` commands. ## `SHOW TABLES`[​](#show-tables "Direct link to show-tables") Use `SHOW TABLES` or query `information_schema.tables` to list the tables in the Spice catalog: ``` > show tables; or > select * from information_schema.tables; +---------------+--------------+--------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------+--------------+------------+ | spice | runtime | task_history | BASE TABLE | | spice | runtime | metrics | BASE TABLE | +---------------+--------------+--------------+------------+ ``` ## `SHOW COLUMNS`[​](#show-columns "Direct link to show-columns") Use `SHOW COLUMNS` or query `information_schema.columns` to see a table’s column definitions: ``` > show columns from t; or > select table_catalog, table_schema, table_name, column_name, data_type, is_nullable from information_schema.columns; +---------------+--------------+--------------+-----------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------+ | table_catalog | table_schema | table_name | column_name | data_type | is_nullable | +---------------+--------------+--------------+-----------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------+ | spice | runtime | task_history | trace_id | Utf8 | NO | | spice | runtime | task_history | span_id | Utf8 | NO | | spice | runtime | task_history | parent_span_id | Utf8 | YES | | spice | runtime | task_history | task | Utf8 | NO | | spice | runtime | task_history | input | Utf8 | NO | | spice | runtime | task_history | captured_output | Utf8 | YES | | spice | runtime | task_history | start_time | Timestamp(Nanosecond, Some("UTC")) | NO | | spice | runtime | task_history | end_time | Timestamp(Nanosecond, Some("UTC")) | NO | | spice | runtime | task_history | execution_duration_ms | Float64 | NO | | spice | runtime | task_history | error_message | Utf8 | YES | | spice | runtime | task_history | labels | Map(Field { name: "entries", data_type: Struct([Field { name: "keys", data_type: Utf8, nullable: false, dict_id: 0, dict_is_ordered: false, metadata: {} }, Field { name: "values", data_type: Utf8, nullable: false, dict_id: 0, dict_is_ordered: false, metadata: {} }]), nullable: false, dict_id: 0, dict_is_ordered: false, metadata: {} }, false) | NO | +---------------+--------------+--------------+-----------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------+ ``` ## `SHOW ALL` (configuration options)[​](#show-all-configuration-options "Direct link to show-all-configuration-options") Use `SHOW ALL` or query `information_schema.df_settings` to view current session configuration parameters: ``` select * from information_schema.df_settings; +-------------------------------------------------------------------------+---------------------------+ | name | value | +-------------------------------------------------------------------------+---------------------------+ | datafusion.catalog.create_default_catalog_and_schema | false | | datafusion.catalog.default_catalog | spice | | datafusion.catalog.default_schema | public | | datafusion.catalog.format | | | datafusion.catalog.has_header | true | | datafusion.catalog.information_schema | true | | datafusion.catalog.location | | | datafusion.catalog.newlines_in_values | false | | datafusion.execution.batch_size | 8192 | ... | datafusion.sql_parser.parse_float_as_decimal | false | | datafusion.sql_parser.support_varchar_with_length | true | +-------------------------------------------------------------------------+---------------------------+ ``` --- # JSON Functions and Operators JSON support in Spice is based on [datafusion-functions-json](https://github.com/datafusion-contrib/datafusion-functions-json), which provides functions and operators to extract, query, and manipulate JSON data stored as strings. Advanced features for JSON creation, modification, or complex path expressions are not supported. Limitations * JSON functions and operators are supported only **during DataFusion (Arrow)** execution. * **Federated or accelerated sources** (non-Arrow) may **not support all JSON functions**. - [JSON Functions](#json-functions) * [`json_contains`](#json_contains) * [Arguments](#arguments) * [Example](#example) * [`json_get`](#json_get) * [Arguments](#arguments-1) * [Example](#example-1) * [`json_get_str`](#json_get_str) * [Arguments](#arguments-2) * [Example](#example-2) * [`json_get_int`](#json_get_int) * [Arguments](#arguments-3) * [Example](#example-3) * [`json_get_float`](#json_get_float) * [Arguments](#arguments-4) * [Example](#example-4) * [`json_get_bool`](#json_get_bool) * [Arguments](#arguments-5) * [Example](#example-5) * [`json_get_json`](#json_get_json) * [Arguments](#arguments-6) * [Example](#example-6) * [`json_get_array`](#json_get_array) * [Arguments](#arguments-7) * [Example](#example-7) * [`json_as_text`](#json_as_text) * [Arguments](#arguments-8) * [Example](#example-8) * [`json_length`](#json_length) * [Arguments](#arguments-9) * [Example](#example-9) * [`json_object_keys`](#json_object_keys) * [Arguments](#arguments-13) * [Example](#example-13) - [JSON Operators](#json-operators) * [`->`](#op_json_get) * [Arguments](#arguments-10) * [Example](#example-10) * [`->>`](#op_json_as_text) * [Arguments](#arguments-11) * [Example](#example-11) * [`?`](#op_json_contains) * [Arguments](#arguments-12) * [Example](#example-12) - [Usage Examples](#usage-examples) * [Nested Object Access](#nested-object-access) * [Array Access](#array-access) * [Conditional JSON Queries](#conditional-json-queries) * [Using JSON Functions in Views](#using-json-functions-in-views) - [Further Reading](#further-reading) *** ## JSON Functions[​](#json-functions "Direct link to JSON Functions") Enables extracting and manipulating data from JSON strings. Each function takes a JSON string as the first argument, followed by one or more keys or indices to specify the path. ### `json_contains`[​](#json_contains "Direct link to json_contains") Returns `true` if a JSON string contains a specific key at the specified path. ``` json_contains(json_string, key1[, key2, ...]) ``` #### Arguments[​](#arguments "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Path to the key to check. Can be string keys for objects or integer indices for arrays. #### Example[​](#example "Direct link to Example") ``` > SELECT json_contains('{"a": 1, "b": 2}', 'a'); +-------------------------------------------------+ | json_contains(Utf8("{\"a\": 1, \"b\": 2}"),Utf8("a")) | +-------------------------------------------------+ | true | +-------------------------------------------------+ ``` ### `json_get`[​](#json_get "Direct link to json_get") Retrieves a value from a JSON string based on its path. ``` json_get(json_string, key1[, key2, ...]) ``` #### Arguments[​](#arguments-1 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Path to the value. Can be string keys for objects or integer indices for arrays. #### Example[​](#example-1 "Direct link to Example") ``` > SELECT json_get('{"a": 1, "b": 2}', 'a'); +----------------------------------------------+ | json_get(Utf8("{"a": 1, "b": 2}"),Utf8("a")) | +----------------------------------------------+ | {int=1} | +----------------------------------------------+ ``` ### `json_get_str`[​](#json_get_str "Direct link to json_get_str") Retrieves a string value from a JSON string based on its path. ``` json_get_str(json_string, key1[, key2, ...]) ``` #### Arguments[​](#arguments-2 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Path to the string value. #### Example[​](#example-2 "Direct link to Example") ``` > SELECT json_get_str('{"name": "John", "age": 30}', 'name'); +----------------------------------------------------------------+ | json_get_str(Utf8("{"name": "John", "age": 30}"),Utf8("name")) | +----------------------------------------------------------------+ | John | +----------------------------------------------------------------+ ``` ### `json_get_int`[​](#json_get_int "Direct link to json_get_int") Retrieves an integer value from a JSON string based on its path. ``` json_get_int(json_string, key1[, key2, ...]) ``` #### Arguments[​](#arguments-3 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Path to the integer value. #### Example[​](#example-3 "Direct link to Example") ``` > SELECT json_get_int('{"name": "John", "age": 30}', 'age'); +---------------------------------------------------------------+ | json_get_int(Utf8("{"name": "John", "age": 30}"),Utf8("age")) | +---------------------------------------------------------------+ | 30 | +---------------------------------------------------------------+ ``` ### `json_get_float`[​](#json_get_float "Direct link to json_get_float") Retrieves a float value from a JSON string based on its path. ``` json_get_float(json_string, key1[, key2, ...]) ``` #### Arguments[​](#arguments-4 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Path to the float value. #### Example[​](#example-4 "Direct link to Example") ``` > SELECT json_get_float('{"price": 19.99, "quantity": 2}', 'price'); +-----------------------------------------------------------------------+ | json_get_float(Utf8("{"price": 19.99, "quantity": 2}"),Utf8("price")) | +-----------------------------------------------------------------------+ | 19.99 | +-----------------------------------------------------------------------+ ``` ### `json_get_bool`[​](#json_get_bool "Direct link to json_get_bool") Retrieves a boolean value from a JSON string based on its path. ``` json_get_bool(json_string, key1[, key2, ...]) ``` #### Arguments[​](#arguments-5 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Path to the boolean value. #### Example[​](#example-5 "Direct link to Example") ``` > SELECT json_get_bool('{"active": true, "visible": false}', 'active'); +--------------------------------------------------------------------------+ | json_get_bool(Utf8("{"active": true, "visible": false}"),Utf8("active")) | +--------------------------------------------------------------------------+ | true | +--------------------------------------------------------------------------+ ``` ### `json_get_json`[​](#json_get_json "Direct link to json_get_json") Retrieves a nested JSON object or array as a raw JSON string from a JSON string based on its path. ``` json_get_json(json_string, key1[, key2, ...]) ``` #### Arguments[​](#arguments-6 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Path to the nested JSON value. #### Example[​](#example-6 "Direct link to Example") ``` > SELECT json_get_json('{"user": {"name": "John", "age": 30}}', 'user'); +---------------------------------------------------------------------------+ | json_get_json(Utf8("{"user": {"name": "John", "age": 30}}"),Utf8("user")) | +---------------------------------------------------------------------------+ | {"name": "John", "age": 30} | +---------------------------------------------------------------------------+ ``` ### `json_get_array`[​](#json_get_array "Direct link to json_get_array") Retrieves an arrow array from a JSON string based on its path. ``` json_get_array(json_string, key1[, key2, ...]) ``` #### Arguments[​](#arguments-7 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Path to the array value. #### Example[​](#example-7 "Direct link to Example") ``` > SELECT json_get_array('{"numbers": [1, 2, 3, 4]}', 'numbers'); +-------------------------------------------------------------------+ | json_get_array(Utf8("{"numbers": [1, 2, 3, 4]}"),Utf8("numbers")) | +-------------------------------------------------------------------+ | [1, 2, 3, 4] | +-------------------------------------------------------------------+ ``` ### `json_as_text`[​](#json_as_text "Direct link to json_as_text") Retrieves any value from a JSON string based on its path and represents it as a string. This is useful for converting JSON values to text format. ``` json_as_text(json_string, key1[, key2, ...]) ``` #### Arguments[​](#arguments-8 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Path to the value to convert to text. #### Example[​](#example-8 "Direct link to Example") ``` > SELECT json_as_text('{"age": 30, "active": true}', 'age'); +---------------------------------------------------------------+ | json_as_text(Utf8("{"age": 30, "active": true}"),Utf8("age")) | +---------------------------------------------------------------+ | 30 | +---------------------------------------------------------------+ ``` ### `json_length`[​](#json_length "Direct link to json_length") Returns the length of a JSON string, array, or object. For objects, returns the number of key-value pairs. For arrays, returns the number of elements. For strings, returns the character count. ``` json_length(json_string[, key1, key2, ...]) ``` #### Arguments[​](#arguments-9 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Optional path to a nested value. If omitted, returns the length of the root JSON value. #### Example[​](#example-9 "Direct link to Example") ``` > SELECT json_length('{"a": 1, "b": 2, "c": 3}'); +-----------------------------------------------+ | json_length(Utf8("{"a": 1, "b": 2, "c": 3}")) | +-----------------------------------------------+ | 3 | +-----------------------------------------------+ > SELECT json_length('[1, 2, 3, 4, 5]'); +--------------------------------------+ | json_length(Utf8("[1, 2, 3, 4, 5]")) | +--------------------------------------+ | 5 | +--------------------------------------+ ``` ### `json_object_keys`[​](#json_object_keys "Direct link to json_object_keys") Returns the top-level keys of a JSON object as an array of strings. If a path is provided, returns the keys of the object at that path. Returns `NULL` if the value at the path is not an object. ``` json_object_keys(json_string[, key1, key2, ...]) ``` Alias: `json_keys`. #### Arguments[​](#arguments-10 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key1, key2, ...**: Optional path to a nested object. If omitted, returns the keys of the root object. #### Example[​](#example-10 "Direct link to Example") ``` > SELECT json_object_keys('{"a": 1, "b": 2, "c": 3}'); +-----------------------------------------------------+ | json_object_keys(Utf8("{"a": 1, "b": 2, "c": 3}")) | +-----------------------------------------------------+ | [a, b, c] | +-----------------------------------------------------+ > SELECT json_object_keys('{"user": {"name": "John", "age": 30}}', 'user'); +-----------------------------------------------------------------------------+ | json_object_keys(Utf8("{"user": {"name": "John", "age": 30}}"),Utf8("user")) | +-----------------------------------------------------------------------------+ | [name, age] | +-----------------------------------------------------------------------------+ ``` *** ## JSON Operators[​](#json-operators "Direct link to JSON Operators") ### `->`[​](#op_json_get "Direct link to op_json_get") JSON access operator. Retrieves a value from a JSON string based on its path. This operator is an alias for [`json_get`](#json_get). ``` json_string -> key json_string -> key1 -> key2 ``` #### Arguments[​](#arguments-11 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key**: Object key (string) or array index (integer). #### Example[​](#example-11 "Direct link to Example") ``` > SELECT '{"user": {"name": "John", "age": 30}}' -> 'user' -> 'name'; +-------------------------------------------------------------+ | '{"user": {"name": "John", "age": 30}}' -> 'user' -> 'name' | +-------------------------------------------------------------+ | {str=John} | +-------------------------------------------------------------+ ``` ### `->>`[​](#op_json_as_text "Direct link to op_json_as_text") JSON access operator for text extraction. Retrieves any value from a JSON string and converts it to text. This operator is an alias for [`json_as_text`](#json_as_text). ``` json_string ->> key json_string -> key1 ->> key2 ``` #### Arguments[​](#arguments-12 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key**: Object key (string) or array index (integer). #### Example[​](#example-12 "Direct link to Example") ``` > SELECT '{"user": {"name": "John", "age": 30}}' -> 'user' ->> 'age'; +-------------------------------------------------------------+ | '{"user": {"name": "John", "age": 30}}' -> 'user' ->> 'age' | +-------------------------------------------------------------+ | 30 | +-------------------------------------------------------------+ ``` ### `?`[​](#op_json_contains "Direct link to op_json_contains") JSON containment operator. Returns `true` if a JSON string contains the specified key. This operator is an alias for [`json_contains`](#json_contains). ``` json_string ? 'key' ``` #### Arguments[​](#arguments-13 "Direct link to Arguments") * **json\_string**: String containing valid JSON data. * **key**: Key to check for existence. #### Example[​](#example-13 "Direct link to Example") ``` > SELECT '{"user": {"name": "John", "age": 30}}' ? 'user'; +--------------------------------------------------+ | '{"user": {"name": "John", "age": 30}}' ? 'user' | +--------------------------------------------------+ | true | +--------------------------------------------------+ ``` *** ## Usage Examples[​](#usage-examples "Direct link to Usage Examples") ### Nested Object Access[​](#nested-object-access "Direct link to Nested Object Access") ``` > SELECT '{"inventory": {"stock": {"S": 12, "M": 20}}}' -> 'inventory' -> 'stock' ->> 'S' as size_s_stock; +---------------+ | size_s_stock | +---------------+ | 12 | +---------------+ ``` ### Array Access[​](#array-access "Direct link to Array Access") ``` > SELECT '{"sizes": ["S", "M", "L", "XL"]}' -> 'sizes' ->> 0 as first_size; +------------+ | first_size | +------------+ | S | +------------+ ``` ### Conditional JSON Queries[​](#conditional-json-queries "Direct link to Conditional JSON Queries") ``` > SELECT name, properties ->> 'color' as color FROM products WHERE properties ? 'color' AND properties ->> 'color' IN ('black', 'white'); ``` ### Using JSON Functions in Views[​](#using-json-functions-in-views "Direct link to Using JSON Functions in Views") JSON functions can be used in views to simplify access to nested JSON data: ``` CREATE VIEW products_with_color AS SELECT id, name, properties ->> 'color' as color, json_get_int(properties, 'stock') as stock_count FROM products; ``` ## JSON Table Functions (UDTFs)[​](#json-table-functions-udtfs "Direct link to JSON Table Functions (UDTFs)") Spice includes table-valued functions for decomposing JSON structures into relational rows. Each function is available as both a UDTF (in the `FROM` clause with literal input) and a scalar UDF returning a list of structs (for per-row use with `UNNEST`). ### `flatten_json`[​](#flatten_json "Direct link to flatten_json") Walks an arbitrary JSON value and emits one row per reachable leaf. ``` flatten_json(input Utf8 [, options...]) -> TABLE( path Utf8, parent_path Utf8, key Utf8, value Utf8, type Utf8 -- "object"|"array"|"string"|"number"|"integer"|"boolean"|"null" ) ``` **Options (named arguments):** | Option | Type | Default | Description | | ------------------ | ---- | --------- | ------------------------------------------------------------- | | `max_depth` | UInt | `64` | Maximum recursion depth. | | `max_rows` | UInt | `1000000` | Per-document row cap. | | `max_bytes` | UInt | `8388608` | Input size limit (bytes). | | `path_style` | Utf8 | `"dot"` | `"dot"` or `"json-pointer"`. | | `include_internal` | Bool | `false` | Also emit interior object/array rows. | | `array_wildcard` | Bool | `false` | Collapse array indices to `[*]` instead of `[0]`, `[1]`, etc. | **UDTF example:** ``` SELECT path, value, type FROM flatten_json('{"user": {"name": "Alice", "scores": [95, 87]}}'); ``` | path | value | type | | ---------------- | ------- | --------- | | `user.name` | `Alice` | `string` | | `user.scores[0]` | `95` | `integer` | | `user.scores[1]` | `87` | `integer` | **Scalar UDF example (per-row with `UNNEST`):** ``` SELECT rows.path, rows.value, rows.type FROM (SELECT UNNEST(flatten_json(body)) AS rows FROM documents); ``` ### `flatten_json_properties`[​](#flatten_json_properties "Direct link to flatten_json_properties") Decomposes a JSON Schema document into one row per field, extracting metadata such as types, descriptions, required status, enums, and format. ``` flatten_json_properties(input Utf8 [, options...]) -> TABLE( path Utf8, parent_path Utf8, name Utf8, description Utf8, type Utf8, required Boolean, format Utf8, enum_values List, metadata Utf8 ) ``` Handles `properties` recursion, `items.properties` (arrays of objects), `additionalProperties` maps, `allOf`/`oneOf`/`anyOf` merging, and local `$ref` pointers with cycle detection. **Options (named arguments):** | Option | Type | Default | Description | | ------------------ | ---- | --------------- | --------------------------------------------------------------------------------------------------------- | | `max_depth` | UInt | `32` | Maximum recursion depth. | | `max_rows` | UInt | `100000` | Per-document row cap. | | `max_bytes` | UInt | `8388608` | Input size limit (bytes). | | `path_style` | Utf8 | `"dot"` | `"dot"` or `"json-pointer"`. | | `dialect` | Utf8 | `"json-schema"` | `"json-schema"` or `"openapi"` (metrics tagging). | | `include_internal` | Bool | `false` | Also emit container rows (objects, arrays). | | `expand_maps` | Bool | `false` | Walk into `additionalProperties` and emit child paths with a wildcard segment (e.g., `parent.[*].child`). | | `map_wildcard` | Utf8 | `"[*]"` | Wildcard segment for map values when `expand_maps` is `true`. | **Example:** ``` SELECT path, type, required, description FROM flatten_json_properties('{ "type": "object", "properties": { "name": {"type": "string", "description": "User name"}, "age": {"type": "integer"} }, "required": ["name"] }'); ``` | path | type | required | description | | ------ | --------- | -------- | ----------- | | `name` | `string` | `true` | `User name` | | `age` | `integer` | `false` | | **Expanding maps:** When a JSON Schema uses `additionalProperties` to describe map values, enable `expand_maps` to produce JSONPath-style paths: ``` SELECT path, type FROM flatten_json_properties( '{"type": "object", "additionalProperties": {"type": "object", "properties": {"id": {"type": "string"}, "primary": {"type": "boolean"}}}}', expand_maps => true ); ``` | path | type | | ------------- | --------- | | `[*].id` | `string` | | `[*].primary` | `boolean` | ### `json_tree`[​](#json_tree "Direct link to json_tree") Recursive depth-first walk of an arbitrary JSON document. Schema-agnostic sibling of `flatten_json_properties` that mirrors the `json_tree` table function in DuckDB and SQLite: one row per node (interior and leaf) in depth-first order, with JSON-Path addresses and a parent pointer for reconstructing the tree. ``` json_tree(input Utf8 [, options...]) -> TABLE( key Utf8, -- key under the parent (object field name) or array index; NULL for the root value Utf8, -- JSON-encoded value of the node type Utf8, -- "object"|"array"|"string"|"number"|"integer"|"boolean"|"null" atom Utf8, -- scalar value for primitive nodes; NULL for objects and arrays id Int64, -- depth-first row id (root = 0) parent Int64, -- id of the parent node; NULL at the root fullkey Utf8, -- absolute JSON-Path of this node, e.g. $.user.scores[0] path Utf8 -- JSON-Path of the parent node; NULL at the root ) ``` **Options (named arguments, UDTF form only):** | Option | Type | Default | Description | | ----------- | ---- | --------- | ------------------------- | | `max_depth` | UInt | `64` | Maximum recursion depth. | | `max_rows` | UInt | `1000000` | Per-document row cap. | | `max_bytes` | UInt | `8388608` | Input size limit (bytes). | **UDTF example:** ``` SELECT id, parent, fullkey, type, atom FROM json_tree('{"user": {"name": "Alice", "scores": [95, 87]}}'); ``` | id | parent | fullkey | type | atom | | --- | ------ | ------------------ | --------- | ------- | | `0` | | `$` | `object` | | | `1` | `0` | `$.user` | `object` | | | `2` | `1` | `$.user.name` | `string` | `Alice` | | `3` | `1` | `$.user.scores` | `array` | | | `4` | `3` | `$.user.scores[0]` | `integer` | `95` | | `5` | `3` | `$.user.scores[1]` | `integer` | `87` | **Scalar UDF example (per-row with `UNNEST`):** ``` SELECT rows.fullkey, rows.atom FROM (SELECT UNNEST(json_tree(body)) AS rows FROM documents); ``` The scalar form takes only the JSON argument and always runs with default caps; the named options above are only accepted in the UDTF (`FROM` clause) form. ## Further Reading[​](#further-reading "Direct link to Further Reading") * [datafusion-functions-json](https://github.com/datafusion-contrib/datafusion-functions-json) - The underlying JSON manipulation library * [Spice JSON Cookbook](https://github.com/spiceai/cookbook/tree/trunk/json_strings) - Spice Cookbook demonstrating how to work with JSON strings in Spice --- # Operators info Spice is built on [Apache DataFusion](https://datafusion.apache.org/) and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. ## Numerical Operators[​](#numerical-operators "Direct link to Numerical Operators") * [+ (plus)](#op_plus) * [- (minus)](#op_minus) * [\* (multiply)](#op_multiply) * [/ (divide)](#op_divide) * [% (modulo)](#op_modulo) ### `+`[​](#op_plus "Direct link to op_plus") Addition ``` > SELECT 1 + 2; +---------------------+ | Int64(1) + Int64(2) | +---------------------+ | 3 | +---------------------+ ``` ### `-`[​](#op_minus "Direct link to op_minus") Subtraction ``` > SELECT 4 - 3; +---------------------+ | Int64(4) - Int64(3) | +---------------------+ | 1 | +---------------------+ ``` ### `*`[​](#op_multiply "Direct link to op_multiply") Multiplication ``` > SELECT 2 * 3; +---------------------+ | Int64(2) * Int64(3) | +---------------------+ | 6 | +---------------------+ ``` ### `/`[​](#op_divide "Direct link to op_divide") Division (integer division truncates toward zero) ``` > SELECT 8 / 4; +---------------------+ | Int64(8) / Int64(4) | +---------------------+ | 2 | +---------------------+ ``` ### `%`[​](#op_modulo "Direct link to op_modulo") Modulo (remainder) ``` > SELECT 7 % 3; +---------------------+ | Int64(7) % Int64(3) | +---------------------+ | 1 | +---------------------+ ``` ## Comparison Operators[​](#comparison-operators "Direct link to Comparison Operators") * [= (equal)](#op_eq) * [!= (not equal)](#op_neq) * [< (less than)](#op_lt) * [<= (less than or equal to)](#op_le) * [> (greater than)](#op_gt) * [>= (greater than or equal to)](#op_ge) * [<=> (three-way comparison, alias for IS NOT DISTINCT FROM)](#op_spaceship) * [IS DISTINCT FROM](#is-distinct-from) * [IS NOT DISTINCT FROM](#is-not-distinct-from) * [\~ (regex match)](#op_re_match) * [\~\* (regex case-insensitive match)](#op_re_match_i) * [!\~ (not regex match)](#op_re_not_match) * [!\~\* (not regex case-insensitive match)](#op_re_not_match_i) ### `=`[​](#op_eq "Direct link to op_eq") Equal ``` > SELECT 1 = 1; +---------------------+ | Int64(1) = Int64(1) | +---------------------+ | true | +---------------------+ ``` ### `!=`[​](#op_neq "Direct link to op_neq") Not Equal ``` > SELECT 1 != 2; +----------------------+ | Int64(1) != Int64(2) | +----------------------+ | true | +----------------------+ ``` ### `<`[​](#op_lt "Direct link to op_lt") Less Than ``` > SELECT 3 < 4; +---------------------+ | Int64(3) < Int64(4) | +---------------------+ | true | +---------------------+ ``` ### `<=`[​](#op_le "Direct link to op_le") Less Than or Equal To ``` > SELECT 3 <= 3; +----------------------+ | Int64(3) <= Int64(3) | +----------------------+ | true | +----------------------+ ``` ### `>`[​](#op_gt "Direct link to op_gt") Greater Than ``` > SELECT 6 > 5; +---------------------+ | Int64(6) > Int64(5) | +---------------------+ | true | +---------------------+ ``` ### `>=`[​](#op_ge "Direct link to op_ge") Greater Than or Equal To ``` > SELECT 5 >= 5; +----------------------+ | Int64(5) >= Int64(5) | +----------------------+ | true | +----------------------+ ``` ### `<=>`[​](#op_spaceship "Direct link to op_spaceship") Three-way comparison operator. A NULL-safe operator that returns true if both operands are equal or both are NULL, false otherwise. ``` > SELECT NULL <=> NULL; +--------------------------------+ | NULL IS NOT DISTINCT FROM NULL | +--------------------------------+ | true | +--------------------------------+ ``` ``` > SELECT 1 <=> NULL; +------------------------------------+ | Int64(1) IS NOT DISTINCT FROM NULL | +------------------------------------+ | false | +------------------------------------+ ``` ``` > SELECT 1 <=> 2; +----------------------------------------+ | Int64(1) IS NOT DISTINCT FROM Int64(2) | +----------------------------------------+ | false | +----------------------------------------+ ``` ``` > SELECT 1 <=> 1; +----------------------------------------+ | Int64(1) IS NOT DISTINCT FROM Int64(1) | +----------------------------------------+ | true | +----------------------------------------+ ``` ### `IS DISTINCT FROM`[​](#is-distinct-from "Direct link to is-distinct-from") Guarantees the result of a comparison is `true` or `false` and not an empty set ``` > SELECT 0 IS DISTINCT FROM NULL; +--------------------------------+ | Int64(0) IS DISTINCT FROM NULL | +--------------------------------+ | true | +--------------------------------+ ``` ### `IS NOT DISTINCT FROM`[​](#is-not-distinct-from "Direct link to is-not-distinct-from") The negation of `IS DISTINCT FROM` ``` > SELECT NULL IS NOT DISTINCT FROM NULL; +--------------------------------+ | NULL IS NOT DISTINCT FROM NULL | +--------------------------------+ | true | +--------------------------------+ ``` ### `~`[​](#op_re_match "Direct link to op_re_match") Regex Match ``` > SELECT 'foo' ~ '^foo(-cli)*'; +-----------------------------------+ | Utf8("foo") ~ Utf8("^foo(-cli)*") | +-----------------------------------+ | true | +-----------------------------------+ ``` ### `~*`[​](#op_re_match_i "Direct link to op_re_match_i") Regex Case-Insensitive Match ``` > SELECT 'foo' ~* '^foo(-cli)*'; +------------------------------------+ | Utf8("foo") ~* Utf8("^foo(-cli)*") | +------------------------------------+ | true | +------------------------------------+ ``` ### `!~`[​](#op_re_not_match "Direct link to op_re_not_match") Not Regex Match ``` > SELECT 'foo' !~ '^foo(-cli)*'; +------------------------------------+ | Utf8("foo") !~ Utf8("^foo(-cli)*") | +------------------------------------+ | false | +------------------------------------+ ``` ### `!~*`[​](#op_re_not_match_i "Direct link to op_re_not_match_i") Not Regex Case-Insensitive Match ``` > SELECT 'foo' !~* '^FOO(-cli)+'; +-------------------------------------+ | Utf8("foo") !~* Utf8("^FOO(-cli)+") | +-------------------------------------+ | true | +-------------------------------------+ ``` ### `~~` Like Match ``` SELECT 'foobar' ~~ 'f_o%r'; +---------------------------------+ | Utf8("foobar") ~~ Utf8("f_o%r") | +---------------------------------+ | true | +---------------------------------+ ``` ### `~~*`[​](#-1 "Direct link to -1") Case-Insensitive Like Match ``` SELECT 'foobar' ~~* 'F_o%r'; +----------------------------------+ | Utf8("foobar") ~~* Utf8("F_o%r") | +----------------------------------+ | true | +----------------------------------+ ``` ### `!~~`[​](#-2 "Direct link to -2") Not Like Match ``` SELECT 'foobar' !~~ 'F_o%r'; +----------------------------------+ | Utf8("foobar") !~~ Utf8("F_o%r") | +----------------------------------+ | true | +----------------------------------+ ``` ### `!~~*`[​](#-3 "Direct link to -3") Not Case-Insensitive Like Match ``` SELECT 'foobar' !~~* 'F_o%Br'; +------------------------------------+ | Utf8("foobar") !~~* Utf8("F_o%Br") | +------------------------------------+ | true | +------------------------------------+ ``` ## Logical Operators[​](#logical-operators "Direct link to Logical Operators") * [AND](#and) * [OR](#or) ### `AND`[​](#and "Direct link to and") Logical And ``` > SELECT true AND true; +---------------------------------+ | Boolean(true) AND Boolean(true) | +---------------------------------+ | true | +---------------------------------+ ``` ### `OR`[​](#or "Direct link to or") Logical Or ``` > SELECT false OR true; +---------------------------------+ | Boolean(false) OR Boolean(true) | +---------------------------------+ | true | +---------------------------------+ ``` ## Bitwise Operators[​](#bitwise-operators "Direct link to Bitwise Operators") * [& (bitwise and)](#op_bit_and) * [| (bitwise or)](#op_bit_or) * [# (bitwise xor)](#op_bit_xor) * [>> (bitwise shift right)](#op_shift_r) * [<< (bitwise shift left)](#op_shift_l) ### `&`[​](#op_bit_and "Direct link to op_bit_and") Bitwise And ``` > SELECT 5 & 3; +---------------------+ | Int64(5) & Int64(3) | +---------------------+ | 1 | +---------------------+ ``` ### `|`[​](#op_bit_or "Direct link to op_bit_or") Bitwise Or ``` > SELECT 5 | 3; +---------------------+ | Int64(5) | Int64(3) | +---------------------+ | 7 | +---------------------+ ``` ### `#`[​](#op_bit_xor "Direct link to op_bit_xor") Bitwise Xor (interchangeable with `^`) ``` > SELECT 5 # 3; +---------------------+ | Int64(5) # Int64(3) | +---------------------+ | 6 | +---------------------+ ``` ### `>>`[​](#op_shift_r "Direct link to op_shift_r") Bitwise Shift Right ``` > SELECT 5 >> 3; +----------------------+ | Int64(5) >> Int64(3) | +----------------------+ | 0 | +----------------------+ ``` ### `<<`[​](#op_shift_l "Direct link to op_shift_l") Bitwise Shift Left ``` > SELECT 5 << 3; +----------------------+ | Int64(5) << Int64(3) | +----------------------+ | 40 | +----------------------+ ``` ## Type Casting Operators[​](#type-casting-operators "Direct link to Type Casting Operators") * [`CAST(expr AS type)`](#op_cast) – Explicit type conversion * [`::` (PostgreSQL-style cast)](#op_double_colon) – Shorthand type conversion ### `CAST(expr AS type)`[​](#op_cast "Direct link to op_cast") Converts an expression to the specified data type. ``` > SELECT CAST('123' AS INT); +------------------------+ | CAST(Utf8("123") AS Int32) | +------------------------+ | 123 | +------------------------+ > SELECT CAST(3.14159 AS INT); +--------------------------+ | CAST(Float64(3.14159) AS Int32) | +--------------------------+ | 3 | +--------------------------+ > SELECT CAST('2024-01-15' AS DATE); +----------------------------+ | CAST(Utf8("2024-01-15") AS Date32) | +----------------------------+ | 2024-01-15 | +----------------------------+ ``` ### `::` (PostgreSQL-style cast)[​](#op_double_colon "Direct link to op_double_colon") Shorthand syntax for type conversion, equivalent to `CAST`. ``` > SELECT '123'::INT; +-----------------------+ | Utf8("123") AS Int32 | +-----------------------+ | 123 | +-----------------------+ > SELECT '2024-01-15'::DATE; +---------------------------+ | Utf8("2024-01-15") AS Date32 | +---------------------------+ | 2024-01-15 | +---------------------------+ > SELECT 100::TEXT; +-------------------+ | Int64(100) AS Utf8 | +-------------------+ | 100 | +-------------------+ ``` **Supported Types:** | Type | Description | | ----------------------------- | ---------------------- | | `INT` / `INTEGER` / `INT4` | 32-bit signed integer | | `BIGINT` / `INT8` | 64-bit signed integer | | `SMALLINT` / `INT2` | 16-bit signed integer | | `FLOAT` / `REAL` / `FLOAT4` | 32-bit floating point | | `DOUBLE` / `FLOAT8` | 64-bit floating point | | `TEXT` / `VARCHAR` / `STRING` | Variable-length string | | `BOOLEAN` / `BOOL` | True/false value | | `DATE` | Calendar date | | `TIMESTAMP` | Date and time | | `INTERVAL` | Time duration | ## Other Operators[​](#other-operators "Direct link to Other Operators") * [|| (string concatenation)](#op_str_cat) * [\`@>\` (array contains)](#op_arr_contains) * [`<@` (array is contained by)](#op_arr_contained_by) ### `||`[​](#op_str_cat "Direct link to op_str_cat") String or array concatenation. When both operands are strings, concatenates them. When both are arrays (lists), invokes [`array_concat`](/docs/next/reference/sql/scalar_functions#array_concat); when one side is a list and the other is a scalar of the element type, invokes [`array_append`](/docs/next/reference/sql/scalar_functions#array_append) or [`array_prepend`](/docs/next/reference/sql/scalar_functions#array_prepend). ``` > SELECT 'Hello, ' || 'Spice!'; +-----------------------------------+ | Utf8("Hello, ") || Utf8("Spice!") | +-----------------------------------+ | Hello, Spice! | +-----------------------------------+ > SELECT [1, 2] || [3, 4]; +-------------------------+ | List([1,2]) || List([3,4]) | +-------------------------+ | [1, 2, 3, 4] | +-------------------------+ > SELECT [1, 2] || 3; +----------------------+ | List([1,2]) || Int64(3) | +----------------------+ | [1, 2, 3] | +----------------------+ ``` ### `@>`[​](#op_arr_contains "Direct link to op_arr_contains") Array contains. Returns `true` if every element of the right array is present in the left. Equivalent to [`array_has_all`](/docs/next/reference/sql/scalar_functions#array_has_all). Only supported with list/array arguments. ``` > SELECT make_array(1,2,3) @> make_array(1,3); +-------------------------------------------------------------------------+ | make_array(Int64(1),Int64(2),Int64(3)) @> make_array(Int64(1),Int64(3)) | +-------------------------------------------------------------------------+ | true | +-------------------------------------------------------------------------+ ``` ### `<@`[​](#op_arr_contained_by "Direct link to op_arr_contained_by") Array is contained by. Returns `true` if every element of the left array is present in the right. Equivalent to `array_has_all(right, left)`. Only supported with list/array arguments. ``` > SELECT make_array(1,3) <@ make_array(1,2,3); +-------------------------------------------------------------------------+ | make_array(Int64(1),Int64(3)) <@ make_array(Int64(1),Int64(2),Int64(3)) | +-------------------------------------------------------------------------+ | true | +-------------------------------------------------------------------------+ ``` ## Literals[​](#literals "Direct link to Literals") Use single quotes for literal string values. ``` SELECT 'foo'; ``` ### Escaping[​](#escaping "Direct link to Escaping") SQL literals do not support C-style escape sequences such as `\n` for newline by default. All characters in a `'` string are treated literally. To escape `'` in SQL literals, use `''`: ``` > SELECT 'it''s escaped'; +----------------------+ | Utf8("it's escaped") | +----------------------+ | it's escaped | +----------------------+ ``` Strings such as `'foo\nbar'` contain a literal backslash followed by `n`, not a newline: ``` > SELECT 'foo\nbar'; +------------------+ | Utf8("foo\nbar") | +------------------+ | foo\nbar | +------------------+ ``` ### E-String Escape Sequences[​](#e-string-escape-sequences "Direct link to E-String Escape Sequences") To include escaped characters such as newline or tab, use `E`-prefixed strings: ``` > SELECT E'foo\nbar'; +-----------------+ | Utf8("foo bar") | +-----------------+ | foo bar | +-----------------+ ``` Supported escape sequences: | Escape | Character | | ------ | --------------- | | `\n` | Newline | | `\t` | Tab | | `\r` | Carriage return | | `\\` | Backslash | | `\'` | Single quote | --- # Prepared Statements info Spice is built on [Apache DataFusion](https://datafusion.apache.org/) and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. ## Positional Arguments[​](#positional-arguments "Direct link to Positional Arguments") Prepared statements can use positional arguments to support multiple parameters. Each parameter is referenced by its position in the statement. **SQL Example** To create a prepared statement named `greater_than` with two parameters: ``` PREPARE greater_than(INT, DOUBLE) AS SELECT * FROM example WHERE a > $1 AND b > $2; ``` To execute the prepared statement with integer and double arguments: ``` EXECUTE greater_than(20, 23.3); ``` **Python Example** ``` import adbc_driver_flightsql.dbapi with adbc_driver_flightsql.dbapi.connect("grpc://localhost:50051") as conn: with conn.cursor() as cur: cur.execute("PREPARE greater_than(INT, DOUBLE) AS SELECT * FROM example WHERE a > $1 AND b > $2;") cur.execute("EXECUTE greater_than(?, ?)", (20, 23.3)) result = cur.fetchall() print(result) ``` Limitations * Positional arguments are not supported with the `date` keyword to construct a date value, like `date $1`. Specify the date value in the query instead: `l_shipdate > date '1995-01-01'`. --- # Scalar Functions info Spice is built on [Apache DataFusion](https://datafusion.apache.org/) and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. Scalar functions help transform, compute, and manipulate data at the row level. These functions are evaluated for each row in a query result and return a single value per invocation. Spice.ai supports a broad set of scalar functions, including math, string, conditional, date/time, array, struct, map, regular expression, and hashing functions. The function set closely follows the PostgreSQL dialect. ## Function Categories[​](#function-categories "Direct link to Function Categories") * [Math Functions](#math-functions) * [Conditional Functions](#conditional-functions) * [String Functions](#string-functions) * [Binary String Functions](#binary-string-functions) * [Regular Expression Functions](#regular-expression-functions) * [Time and Date Functions](#time-and-date-functions) * [Array Functions](#array-functions) * [Struct Functions](#struct-functions) * [Map Functions](#map-functions) * [Hashing Functions](#hashing-functions) * [Encoding Functions](#encoding-functions) * [Union Functions](#union-functions) * [Metadata Functions](#metadata-functions) * [Other Functions](#other-functions) Spark-compatible scalar functions registered by Spice are documented here only when their Spark-specific behavior differs from the PostgreSQL equivalent, or when they have no PostgreSQL analogue. For functions not listed below, refer to the [Spark SQL built-in function reference](https://spark.apache.org/docs/latest/api/sql/index.html) for semantics. *** ## Math Functions[​](#math-functions "Direct link to Math Functions") Math functions in Spice.ai SQL help perform numeric calculations, transformations, and analysis. These functions operate on numeric expressions, which can be constants, columns, or results of other functions and operators. The following math functions are supported: * [abs](#abs) * [acos](#acos) * [acosh](#acosh) * [asin](#asin) * [asinh](#asinh) * [atan](#atan) * [atan2](#atan2) * [atanh](#atanh) * [cbrt](#cbrt) * [ceil](#ceil) * [cos](#cos) * [cosh](#cosh) * [cot](#cot) * [degrees](#degrees) * [exp](#exp) * [factorial](#factorial) * [floor](#floor) * [gcd](#gcd) * [isnan](#isnan) * [iszero](#iszero) * [lcm](#lcm) * [mod](#mod) * [pmod](#pmod) * [ln](#ln) * [log](#log) * [log10](#log10) * [log2](#log2) * [nanvl](#nanvl) * [pi](#pi) * [pow](#pow-and-power) * [power](#pow-and-power) * [radians](#radians) * [random](#random) * [round](#round) * [rint](#rint) * [signum](#signum) * [sin](#sin) * [sinh](#sinh) * [sqrt](#sqrt) * [tan](#tan) * [tanh](#tanh) * [trunc](#trunc) * [width\_bucket](#width_bucket) ### `abs`[​](#abs "Direct link to abs") Returns the absolute value of a numeric expression. If the input is negative, the result is its positive equivalent; if the input is positive or zero, the result is unchanged. ``` abs(numeric_expression) ``` #### Arguments[​](#arguments "Direct link to Arguments") * **numeric\_expression**: Numeric value to evaluate. Accepts constants, columns, or expressions. #### Example[​](#example "Direct link to Example") ``` > select abs(-5); +-------------+ | abs(Int64(-5)) | +-------------+ | 5 | +-------------+ ``` ### `acos`[​](#acos "Direct link to acos") Returns the arc cosine (inverse cosine) of a numeric expression. The input must be in the range \[-1, 1]. The result is in radians. ``` acos(numeric_expression) ``` #### Arguments[​](#arguments-1 "Direct link to Arguments") * **numeric\_expression**: Value between -1 and 1. ### `acosh`[​](#acosh "Direct link to acosh") Returns the inverse hyperbolic cosine of a numeric expression. The input must be greater than or equal to 1. ``` acosh(numeric_expression) ``` #### Arguments[​](#arguments-2 "Direct link to Arguments") * **numeric\_expression**: Value greater than or equal to 1. ### `asin`[​](#asin "Direct link to asin") Returns the arc sine (inverse sine) of a numeric expression. The input must be in the range \[-1, 1]. The result is in radians. ``` asin(numeric_expression) ``` #### Arguments[​](#arguments-3 "Direct link to Arguments") * **numeric\_expression**: Value between -1 and 1. ### `asinh`[​](#asinh "Direct link to asinh") Returns the inverse hyperbolic sine of a numeric expression. ``` asinh(numeric_expression) ``` #### Arguments[​](#arguments-4 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `atan`[​](#atan "Direct link to atan") Returns the arc tangent (inverse tangent) of a numeric expression. The result is in radians. ``` atan(numeric_expression) ``` #### Arguments[​](#arguments-5 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `atan2`[​](#atan2 "Direct link to atan2") Returns the arc tangent of the quotient of its arguments, that is, atan(expression\_y / expression\_x). The result is in radians and takes into account the signs of both arguments to determine the correct quadrant. ``` atan2(expression_y, expression_x) ``` #### Arguments[​](#arguments-6 "Direct link to Arguments") * **expression\_y**: Numerator value. * **expression\_x**: Denominator value. ### `atanh`[​](#atanh "Direct link to atanh") Returns the inverse hyperbolic tangent of a numeric expression. The input must be in the range (-1, 1). ``` atanh(numeric_expression) ``` #### Arguments[​](#arguments-7 "Direct link to Arguments") * **numeric\_expression**: Value between -1 and 1 (exclusive). ### `cbrt`[​](#cbrt "Direct link to cbrt") Returns the cube root of a numeric expression. ``` cbrt(numeric_expression) ``` #### Arguments[​](#arguments-8 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `ceil`[​](#ceil "Direct link to ceil") Returns the smallest integer greater than or equal to the input value. ``` ceil(numeric_expression) ``` #### Arguments[​](#arguments-9 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `cos`[​](#cos "Direct link to cos") Returns the cosine of a numeric expression, where the input is in radians. ``` cos(numeric_expression) ``` #### Arguments[​](#arguments-10 "Direct link to Arguments") * **numeric\_expression**: Value in radians. ### `cosh`[​](#cosh "Direct link to cosh") Returns the hyperbolic cosine of a numeric expression. ``` cosh(numeric_expression) ``` #### Arguments[​](#arguments-11 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `cot`[​](#cot "Direct link to cot") Returns the cotangent of a numeric expression, where the input is in radians. ``` cot(numeric_expression) ``` #### Arguments[​](#arguments-12 "Direct link to Arguments") * **numeric\_expression**: Value in radians. ### `degrees`[​](#degrees "Direct link to degrees") Converts radians to degrees. ``` degrees(numeric_expression) ``` #### Arguments[​](#arguments-13 "Direct link to Arguments") * **numeric\_expression**: Value in radians. ### `exp`[​](#exp "Direct link to exp") Returns the value of e (Euler's number) raised to the power of the input value. ``` exp(numeric_expression) ``` #### Arguments[​](#arguments-14 "Direct link to Arguments") * **numeric\_expression**: Exponent value. ### `factorial`[​](#factorial "Direct link to factorial") Returns the factorial of a non-negative integer. For values less than 2, returns 1. ``` factorial(numeric_expression) ``` #### Arguments[​](#arguments-15 "Direct link to Arguments") * **numeric\_expression**: Non-negative integer value. ### `floor`[​](#floor "Direct link to floor") Returns the largest integer less than or equal to the input value. ``` floor(numeric_expression) ``` #### Arguments[​](#arguments-16 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `gcd`[​](#gcd "Direct link to gcd") Returns the greatest common divisor of two integer expressions. If both inputs are zero, returns 0. ``` gcd(expression_x, expression_y) ``` #### Arguments[​](#arguments-17 "Direct link to Arguments") * **expression\_x**: First integer value. * **expression\_y**: Second integer value. ### `isnan`[​](#isnan "Direct link to isnan") Returns true if the input is NaN (not a number), otherwise returns false. ``` isnan(numeric_expression) ``` #### Arguments[​](#arguments-18 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `iszero`[​](#iszero "Direct link to iszero") Returns true if the input is +0.0 or -0.0, otherwise returns false. ``` iszero(numeric_expression) ``` #### Arguments[​](#arguments-19 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `lcm`[​](#lcm "Direct link to lcm") Returns the least common multiple of two integer expressions. If either input is zero, returns 0. ``` lcm(expression_x, expression_y) ``` #### Arguments[​](#arguments-20 "Direct link to Arguments") * **expression\_x**: First integer value. * **expression\_y**: Second integer value. ### `mod`[​](#mod "Direct link to mod") Returns the remainder after dividing the first argument by the second, matching the Spark SQL `%`/`mod` semantics for signed values. ``` mod(dividend, divisor) ``` #### Arguments[​](#arguments-21 "Direct link to Arguments") * **dividend**: Numeric expression to divide. * **divisor**: Numeric expression that cannot be zero. #### Example[​](#example-1 "Direct link to Example") ``` > select mod(-10, 3); +----------------------------+ | mod(Int64(-10),Int64(3)) | +----------------------------+ | -1 | +----------------------------+ ``` Reference: [Spark SQL `mod`](https://spark.apache.org/docs/latest/api/sql/index.html#mod). ### `pmod`[​](#pmod "Direct link to pmod") Returns a positive remainder for integer or floating-point division. When the standard remainder is negative, the divisor is added to produce a non-negative result, mirroring Spark SQL behavior. ``` pmod(dividend, divisor) ``` #### Arguments[​](#arguments-22 "Direct link to Arguments") * **dividend**: Numeric expression to divide. * **divisor**: Numeric expression that cannot be zero. #### Example[​](#example-2 "Direct link to Example") ``` > select pmod(-10, 3); +-----------------------------+ | pmod(Int64(-10),Int64(3)) | +-----------------------------+ | 2 | +-----------------------------+ ``` Reference: [Spark SQL `pmod`](https://spark.apache.org/docs/latest/api/sql/index.html#pmod). ### `ln`[​](#ln "Direct link to ln") Returns the natural logarithm (base e) of a numeric expression. ``` ln(numeric_expression) ``` #### Arguments[​](#arguments-23 "Direct link to Arguments") * **numeric\_expression**: Positive numeric value. ### `log`[​](#log "Direct link to log") Returns the logarithm of a numeric expression. If a base is provided, returns the logarithm to that base; otherwise, returns the base-10 logarithm. ``` log(base, numeric_expression) log(numeric_expression) ``` #### Arguments[​](#arguments-24 "Direct link to Arguments") * **base**: Base of the logarithm (optional). * **numeric\_expression**: Positive numeric value. ### `log10`[​](#log10 "Direct link to log10") Returns the base-10 logarithm of a numeric expression. ``` log10(numeric_expression) ``` #### Arguments[​](#arguments-25 "Direct link to Arguments") * **numeric\_expression**: Positive numeric value. ### `log2`[​](#log2 "Direct link to log2") Returns the base-2 logarithm of a numeric expression. ``` log2(numeric_expression) ``` #### Arguments[​](#arguments-26 "Direct link to Arguments") * **numeric\_expression**: Positive numeric value. ### `nanvl`[​](#nanvl "Direct link to nanvl") Returns the first argument if it is not NaN; otherwise, returns the second argument. ``` nanvl(expression_x, expression_y) ``` #### Arguments[​](#arguments-27 "Direct link to Arguments") * **expression\_x**: Value to return if not NaN. * **expression\_y**: Value to return if the first argument is NaN. ### `pi`[​](#pi "Direct link to pi") Returns an approximate value of π (pi). ``` pi() ``` ### `pow` and `power`[​](#pow-and-power "Direct link to pow-and-power") Returns the value of the first argument raised to the power of the second argument. `pow` is an alias for `power`. ``` power(base, exponent) pow(base, exponent) ``` #### Arguments[​](#arguments-28 "Direct link to Arguments") * **base**: Numeric value to raise. * **exponent**: Power to raise the base to. ### `radians`[​](#radians "Direct link to radians") Converts degrees to radians. ``` radians(numeric_expression) ``` #### Arguments[​](#arguments-29 "Direct link to Arguments") * **numeric\_expression**: Value in degrees. ### `random`[​](#random "Direct link to random") Returns a random floating-point value in the range \[0, 1). The random seed is unique for each row. ``` random() ``` ### `round`[​](#round "Direct link to round") Rounds a numeric expression to the nearest integer or to a specified number of decimal places. ``` round(numeric_expression[, decimal_places]) ``` #### Arguments[​](#arguments-30 "Direct link to Arguments") * **numeric\_expression**: Value to round. * **decimal\_places**: Optional. Number of decimal places to round to. Defaults to 0. ### `rint`[​](#rint "Direct link to rint") Rounds a double-precision value to the nearest integer using IEEE-754 "round to nearest, ties to even" rules and returns the rounded value as a floating-point number, matching Spark SQL semantics. ``` rint(numeric_expression) ``` #### Arguments[​](#arguments-31 "Direct link to Arguments") * **numeric\_expression**: Floating-point expression to round. Integers are implicitly cast to double. #### Example[​](#example-3 "Direct link to Example") ``` > select rint(12.5); +-------------------------+ | rint(Float64(12.5)) | +-------------------------+ | 12.0 | +-------------------------+ ``` Reference: [Spark SQL `rint`](https://spark.apache.org/docs/latest/api/sql/index.html#rint). ### `signum`[​](#signum "Direct link to signum") Returns the sign of a numeric expression. Returns -1 for negative numbers, 1 for zero and positive numbers. ``` signum(numeric_expression) ``` #### Arguments[​](#arguments-32 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `sin`[​](#sin "Direct link to sin") Returns the sine of a numeric expression, where the input is in radians. ``` sin(numeric_expression) ``` #### Arguments[​](#arguments-33 "Direct link to Arguments") * **numeric\_expression**: Value in radians. ### `sinh`[​](#sinh "Direct link to sinh") Returns the hyperbolic sine of a numeric expression. ``` sinh(numeric_expression) ``` #### Arguments[​](#arguments-34 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `sqrt`[​](#sqrt "Direct link to sqrt") Returns the square root of a numeric expression. ``` sqrt(numeric_expression) ``` #### Arguments[​](#arguments-35 "Direct link to Arguments") * **numeric\_expression**: Non-negative numeric value. ### `tan`[​](#tan "Direct link to tan") Returns the tangent of a numeric expression, where the input is in radians. ``` tan(numeric_expression) ``` #### Arguments[​](#arguments-36 "Direct link to Arguments") * **numeric\_expression**: Value in radians. ### `tanh`[​](#tanh "Direct link to tanh") Returns the hyperbolic tangent of a numeric expression. ``` tanh(numeric_expression) ``` #### Arguments[​](#arguments-37 "Direct link to Arguments") * **numeric\_expression**: Numeric value. ### `trunc`[​](#trunc "Direct link to trunc") Truncates a numeric expression to a whole number or to a specified number of decimal places. If `decimal_places` is positive, truncates digits to the right of the decimal point; if negative, truncates digits to the left. ``` trunc(numeric_expression[, decimal_places]) ``` #### Arguments[​](#arguments-38 "Direct link to Arguments") * **numeric\_expression**: Value to truncate. * **decimal\_places**: Optional. Number of decimal places to truncate to. Defaults to 0. ### `width_bucket`[​](#width_bucket "Direct link to width_bucket") Assigns a value to an equiwidth histogram bucket. Returns 0 when the value is below `min_value`, `num_bucket + 1` when it is above `max_value`, and otherwise the 1-based bucket index, mirroring Spark SQL behavior. ``` width_bucket(value, min_value, max_value, num_bucket) ``` #### Arguments[​](#arguments-39 "Direct link to Arguments") * **value**: Numeric or interval expression to bin. * **min\_value**: Lower bound of the histogram range. * **max\_value**: Upper bound of the histogram range. * **num\_bucket**: Positive integer specifying the number of buckets. #### Example[​](#example-4 "Direct link to Example") ``` > select width_bucket(5.3, 0.2, 10.6, 5); +--------------------------------------------------+ | width_bucket(Float64(5.3),Float64(0.2),Float64(10.6),Int64(5)) | +--------------------------------------------------+ | 3 | +--------------------------------------------------+ ``` Reference: [Spark SQL `width_bucket`](https://spark.apache.org/docs/latest/api/sql/index.html#width_bucket). *** ## Conditional Functions[​](#conditional-functions "Direct link to Conditional Functions") Conditional functions help handle null values, select among alternatives, and compare multiple expressions. These are useful for data cleaning and conditional logic in queries. * [CASE](#case) * [coalesce](#coalesce) * [greatest](#greatest) * [if](#if) * [least](#least) * [nullif](#nullif) * [nvl](#nvl) * [nvl2](#nvl2) ### `CASE`[​](#case "Direct link to case") Standard SQL `CASE` expression, supported in both simple and searched forms. ``` CASE expression WHEN value1 THEN result1 [WHEN value2 THEN result2 ...] [ELSE default_result] END CASE WHEN condition1 THEN result1 [WHEN condition2 THEN result2 ...] [ELSE default_result] END ``` #### Example[​](#example-5 "Direct link to Example") ``` > SELECT CASE WHEN score >= 90 THEN 'A' WHEN score >= 80 THEN 'B' ELSE 'C' END AS grade FROM (VALUES (95), (82), (70)) AS t(score); +-------+ | grade | +-------+ | A | | B | | C | +-------+ ``` ### `coalesce`[​](#coalesce "Direct link to coalesce") Returns the first non-null value from its arguments. Returns `NULL` only if every argument is `NULL`. ``` coalesce(expression1[, ..., expression_n]) ``` #### Arguments[​](#arguments-40 "Direct link to Arguments") * **expression1, ..., expression\_n**: Expressions to evaluate in order. All arguments must share a common type. #### Example[​](#example-6 "Direct link to Example") ``` > SELECT coalesce(NULL, NULL, 'spice', 'datafusion'); +----------------------------------------------------+ | coalesce(NULL,NULL,Utf8("spice"),Utf8("datafusion")) | +----------------------------------------------------+ | spice | +----------------------------------------------------+ ``` ### `greatest`[​](#greatest "Direct link to greatest") Returns the largest value among the arguments, ignoring `NULL`s. Returns `NULL` only if every argument is `NULL`. ``` greatest(expression1[, ..., expression_n]) ``` #### Arguments[​](#arguments-41 "Direct link to Arguments") * **expression1, ..., expression\_n**: Expressions to compare. Must share a common type. #### Example[​](#example-7 "Direct link to Example") ``` > SELECT greatest(3, 7, NULL, 4); +------------------------------------------------+ | greatest(Int64(3),Int64(7),NULL,Int64(4)) | +------------------------------------------------+ | 7 | +------------------------------------------------+ ``` ### `if`[​](#if "Direct link to if") Evaluates a boolean condition and returns one of two expressions, matching the Spark SQL `if` function semantics. ``` if(condition, true_value, false_value) ``` #### Arguments[​](#arguments-42 "Direct link to Arguments") * **condition**: Boolean expression determining which branch to take. * **true\_value**: Expression to return when the condition evaluates to true. * **false\_value**: Expression to return when the condition evaluates to false or NULL. #### Example[​](#example-8 "Direct link to Example") ``` > select if(temperature > 70, 'warm', 'cool') as label from (values (65), (72)) as t(temperature); +------------+ | label | +------------+ | cool | | warm | +------------+ ``` Reference: [Spark SQL `if`](https://spark.apache.org/docs/latest/api/sql/index.html#if). ### `least`[​](#least "Direct link to least") Returns the smallest value among the arguments, ignoring `NULL`s. Returns `NULL` only if every argument is `NULL`. ``` least(expression1[, ..., expression_n]) ``` #### Arguments[​](#arguments-43 "Direct link to Arguments") * **expression1, ..., expression\_n**: Expressions to compare. Must share a common type. #### Example[​](#example-9 "Direct link to Example") ``` > SELECT least(3, 7, NULL, 4); +--------------------------------------------+ | least(Int64(3),Int64(7),NULL,Int64(4)) | +--------------------------------------------+ | 3 | +--------------------------------------------+ ``` ### `nullif`[​](#nullif "Direct link to nullif") Returns `NULL` if `expression1` equals `expression2`; otherwise returns `expression1`. Useful for converting sentinel values to `NULL`. ``` nullif(expression1, expression2) ``` #### Arguments[​](#arguments-44 "Direct link to Arguments") * **expression1**: Value to return if it differs from `expression2`. * **expression2**: Value to compare against. #### Example[​](#example-10 "Direct link to Example") ``` > SELECT nullif('unknown', 'unknown'); +-------------------------------------------+ | nullif(Utf8("unknown"),Utf8("unknown")) | +-------------------------------------------+ | | +-------------------------------------------+ > SELECT nullif('spice', 'unknown'); +-------------------------------------------+ | nullif(Utf8("spice"),Utf8("unknown")) | +-------------------------------------------+ | spice | +-------------------------------------------+ ``` ### `nvl`[​](#nvl "Direct link to nvl") Returns `expression2` when `expression1` is `NULL`; otherwise returns `expression1`. Equivalent to `coalesce(expression1, expression2)`. ``` nvl(expression1, expression2) ``` Alias: `ifnull`. ### `nvl2`[​](#nvl2 "Direct link to nvl2") Returns `expression2` if `expression1` is not `NULL`; otherwise returns `expression3`. ``` nvl2(expression1, expression2, expression3) ``` ## String Functions[​](#string-functions "Direct link to String Functions") String functions in Spice.ai SQL help manipulate, analyze, and transform text data. These functions operate on string expressions, which can be constants, columns, or results of other functions. The implementation closely follows the PostgreSQL dialect. The following string functions are supported: * [ascii](#ascii) * [bit\_length](#bit_length) * [btrim](#btrim) * [char\_length](#char_length) * [character\_length](#character_length) * [chr](#chr) * [concat](#concat) * [concat\_ws](#concat_ws) * [contains](#contains) * [like](#like) * [ilike](#ilike) * [ends\_with](#ends_with) * [find\_in\_set](#find_in_set) * [initcap](#initcap) * [instr](#instr) * [left](#left) * [length](#length) * [levenshtein](#levenshtein) * [lower](#lower) * [luhn\_check](#luhn_check) * [lpad](#lpad) * [ltrim](#ltrim) * [octet\_length](#octet_length) * [overlay](#overlay) * [parse\_url](#parse_url) * [position](#position) * [repeat](#repeat) * [replace](#replace) * [reverse](#reverse) * [right](#right) * [rpad](#rpad) * [rtrim](#rtrim) * [split\_part](#split_part) * [starts\_with](#starts_with) * [strpos](#strpos) * [substr](#substr) * [substr\_index](#substr_index) * [substring](#substring) * [substring\_index](#substring_index) * [to\_hex](#to_hex) * [translate](#translate) * [trim](#trim) * [upper](#upper) * [uuid](#uuid) ### `ascii`[​](#ascii "Direct link to ascii") Returns the Unicode code point of the first character in a string. If the string is empty, returns 0. ``` ascii(str) ``` #### Arguments[​](#arguments-45 "Direct link to Arguments") * **str**: String expression. Accepts constants, columns, or expressions. #### Example[​](#example-11 "Direct link to Example") ``` > select ascii('abc'); +--------------------+ | ascii(Utf8("abc")) | +--------------------+ | 97 | +--------------------+ > select ascii('🚀'); +-------------------+ | ascii(Utf8("🚀")) | +-------------------+ | 128640 | +-------------------+ ``` Related function: [`chr`](#chr) ### `bit_length`[​](#bit_length "Direct link to bit_length") Returns the number of bits in the string. Each character is counted according to its byte representation (8 bits per byte). ``` bit_length(str) ``` #### Arguments[​](#arguments-46 "Direct link to Arguments") * **str**: String expression. #### Example[​](#example-12 "Direct link to Example") ``` > select bit_length('datafusion'); +--------------------------------+ | bit_length(Utf8("datafusion")) | +--------------------------------+ | 80 | +--------------------------------+ ``` Related functions: [`length`](#length), [`octet_length`](#octet_length) ### `btrim`[​](#btrim "Direct link to btrim") Removes the longest string containing only characters in `trim_str` from the start and end of `str`. If `trim_str` is omitted, whitespace is removed. ``` btrim(str[, trim_str]) ``` #### Arguments[​](#arguments-47 "Direct link to Arguments") * **str**: String expression. * **trim\_str**: Optional string of characters to trim. Defaults to whitespace. #### Example[​](#example-13 "Direct link to Example") ``` > select btrim('__datafusion____', '_'); +-------------------------------------------+ | btrim(Utf8("__datafusion____"),Utf8("_")) | +-------------------------------------------+ | datafusion | +-------------------------------------------+ ``` Alternative syntax: `trim(BOTH trim_str FROM str)` or `trim(trim_str FROM str)` Aliases: `trim` Related functions: [`ltrim`](#ltrim), [`rtrim`](#rtrim) ### `char_length`[​](#char_length "Direct link to char_length") Alias of [`character_length`](#character_length). ### `character_length`[​](#character_length "Direct link to character_length") Returns the number of characters in a string, not bytes. Handles Unicode correctly. ``` character_length(str) ``` #### Arguments[​](#arguments-48 "Direct link to Arguments") * **str**: String expression. #### Example[​](#example-14 "Direct link to Example") ``` > select character_length('Ångström'); +------------------------------------+ | character_length(Utf8("Ångström")) | +------------------------------------+ | 8 | +------------------------------------+ ``` Aliases: `length`, `char_length` Related functions: [`bit_length`](#bit_length), [`octet_length`](#octet_length) ### `chr`[​](#chr "Direct link to chr") Returns the character with the specified Unicode code point. ``` chr(expression) ``` #### Arguments[​](#arguments-49 "Direct link to Arguments") * **expression**: Integer code point. #### Example[​](#example-15 "Direct link to Example") ``` > select chr(128640); +--------------------+ | chr(Int64(128640)) | +--------------------+ | 🚀 | +--------------------+ ``` Related function: [`ascii`](#ascii) ### `concat`[​](#concat "Direct link to concat") Concatenates two or more strings into a single string. ``` concat(str[, ..., str_n]) ``` #### Arguments[​](#arguments-50 "Direct link to Arguments") * **str**: String expression. * **str\_n**: Additional string expressions. #### Example[​](#example-16 "Direct link to Example") ``` > select concat('data', 'f', 'us', 'ion'); +-------------------------------------------------------+ | concat(Utf8("data"),Utf8("f"),Utf8("us"),Utf8("ion")) | +-------------------------------------------------------+ | datafusion | +-------------------------------------------------------+ ``` Related function: [`concat_ws`](#concat_ws) ### `concat_ws`[​](#concat_ws "Direct link to concat_ws") Concatenates strings using a separator between each value. ``` concat_ws(separator, str[, ..., str_n]) ``` #### Arguments[​](#arguments-51 "Direct link to Arguments") * **separator**: String separator. * **str**: String expression. * **str\_n**: Additional string expressions. #### Example[​](#example-17 "Direct link to Example") ``` > select concat_ws('_', 'data', 'fusion'); +--------------------------------------------------+ | concat_ws(Utf8("_"),Utf8("data"),Utf8("fusion")) | +--------------------------------------------------+ | data_fusion | +--------------------------------------------------+ ``` Related function: [`concat`](#concat) ### `contains`[​](#contains "Direct link to contains") Returns true if `search_str` is found within `str`. The search is case-sensitive. ``` contains(str, search_str) ``` #### Arguments[​](#arguments-52 "Direct link to Arguments") * **str**: String expression. * **search\_str**: Substring to search for. #### Example[​](#example-18 "Direct link to Example") ``` > select contains('the quick brown fox', 'row'); +---------------------------------------------------+ | contains(Utf8("the quick brown fox"),Utf8("row")) | +---------------------------------------------------+ | true | +---------------------------------------------------+ ``` ### `like`[​](#like "Direct link to like") Performs SQL pattern matching using `%` to match zero or more characters and `_` to match a single character. The optional SQL `ESCAPE` clause can be used to treat a wildcard literally, matching Spark SQL behavior. ``` like(str, pattern) ``` #### Arguments[​](#arguments-53 "Direct link to Arguments") * **str**: String expression to compare. * **pattern**: Pattern containing literal text plus `%` and `_` wildcards. #### Example[​](#example-19 "Direct link to Example") ``` > select like('spice.ai', 'spice%'); +----------------------------------+ | like(Utf8("spice.ai"),Utf8("spice%")) | +----------------------------------+ | true | +----------------------------------+ ``` Reference: [Spark SQL `like`](https://spark.apache.org/docs/latest/api/sql/index.html#like). ### `ilike`[​](#ilike "Direct link to ilike") Case-insensitive variant of [`like`](#like) that treats ASCII characters in `str` and `pattern` without regard to case. The optional SQL `ESCAPE` clause may be used to treat `%` or `_` literally. ``` ilike(str, pattern) ``` #### Arguments[​](#arguments-54 "Direct link to Arguments") * **str**: String expression to compare. * **pattern**: Case-insensitive pattern containing literal text plus `%` and `_` wildcards. #### Example[​](#example-20 "Direct link to Example") ``` > select ilike('Spice.AI', 'spice%'); +-----------------------------------+ | ilike(Utf8("Spice.AI"),Utf8("spice%")) | +-----------------------------------+ | true | +-----------------------------------+ ``` Reference: [Spark SQL `ilike`](https://spark.apache.org/docs/latest/api/sql/index.html#ilike). ### `ends_with`[​](#ends_with "Direct link to ends_with") Returns true if `str` ends with the substring `substr`. ``` ends_with(str, substr) ``` #### Arguments[​](#arguments-55 "Direct link to Arguments") * **str**: String expression. * **substr**: Substring to test for. #### Example[​](#example-21 "Direct link to Example") ``` > select ends_with('datafusion', 'soin'); +--------------------------------------------+ | ends_with(Utf8("datafusion"),Utf8("soin")) | +--------------------------------------------+ | false | +--------------------------------------------+ > select ends_with('datafusion', 'sion'); +--------------------------------------------+ | ends_with(Utf8("datafusion"),Utf8("sion")) | +--------------------------------------------+ | true | +--------------------------------------------+ ``` ### `find_in_set`[​](#find_in_set "Direct link to find_in_set") Returns the position (1-based) of `str` in the comma-separated list `strlist`. Returns 0 if not found. ``` find_in_set(str, strlist) ``` #### Arguments[​](#arguments-56 "Direct link to Arguments") * **str**: String to find. * **strlist**: Comma-separated list of substrings. #### Example[​](#example-22 "Direct link to Example") ``` > select find_in_set('b', 'a,b,c,d'); +----------------------------------------+ | find_in_set(Utf8("b"),Utf8("a,b,c,d")) | +----------------------------------------+ | 2 | +----------------------------------------+ ``` ### `initcap`[​](#initcap "Direct link to initcap") Capitalizes the first character of each word in the string. Words are delimited by non-alphanumeric characters. ``` initcap(str) ``` #### Arguments[​](#arguments-57 "Direct link to Arguments") * **str**: String expression. #### Example[​](#example-23 "Direct link to Example") ``` > select initcap('apache datafusion'); +------------------------------------+ | initcap(Utf8("apache datafusion")) | +------------------------------------+ | Apache Datafusion | +------------------------------------+ ``` Related functions: [`lower`](#lower), [`upper`](#upper) ### `instr`[​](#instr "Direct link to instr") Alias of [`strpos`](#strpos). ### `left`[​](#left "Direct link to left") Returns the first `n` characters from the left side of the string. ``` left(str, n) ``` #### Arguments[​](#arguments-58 "Direct link to Arguments") * **str**: String expression. * **n**: Number of characters to return. #### Example[​](#example-24 "Direct link to Example") ``` > select left('datafusion', 4); +-----------------------------------+ | left(Utf8("datafusion"),Int64(4)) | +-----------------------------------+ | data | +-----------------------------------+ ``` Related function: [`right`](#right) ### `length`[​](#length "Direct link to length") Alias of [`character_length`](#character_length). ### `levenshtein`[​](#levenshtein "Direct link to levenshtein") Returns the [Levenshtein distance](https://en.wikipedia.org/wiki/Levenshtein_distance) between two strings. ``` levenshtein(str1, str2) ``` #### Arguments[​](#arguments-59 "Direct link to Arguments") * **str1**: First string. * **str2**: Second string. #### Example[​](#example-25 "Direct link to Example") ``` > select levenshtein('kitten', 'sitting'); +---------------------------------------------+ | levenshtein(Utf8("kitten"),Utf8("sitting")) | +---------------------------------------------+ | 3 | +---------------------------------------------+ ``` ### `lower`[​](#lower "Direct link to lower") Converts all characters in the string to lower case. ``` lower(str) ``` #### Arguments[​](#arguments-60 "Direct link to Arguments") * **str**: String expression. #### Example[​](#example-26 "Direct link to Example") ``` > select lower('Ångström'); +-------------------------+ | lower(Utf8("Ångström")) | +-------------------------+ | ångström | +-------------------------+ ``` Related functions: [`initcap`](#initcap), [`upper`](#upper) ### `luhn_check`[​](#luhn_check "Direct link to luhn_check") Validates that a string of digits satisfies the [Luhn checksum](https://en.wikipedia.org/wiki/Luhn_algorithm), returning true for valid numbers and false otherwise. This matches the Spark SQL implementation and is useful for validating identifiers such as credit card numbers. ``` luhn_check(str) ``` #### Arguments[​](#arguments-61 "Direct link to Arguments") * **str**: String expression containing digits. #### Example[​](#example-27 "Direct link to Example") ``` > select luhn_check('79927398713'); +--------------------------------------+ | luhn_check(Utf8("79927398713")) | +--------------------------------------+ | true | +--------------------------------------+ ``` Reference: [Spark SQL `luhn_check`](https://spark.apache.org/docs/latest/api/sql/index.html#luhn_check). ### `lpad`[​](#lpad "Direct link to lpad") Pads the left side of the string with another string until the result reaches the specified length. If the padding string is omitted, a space is used. ``` lpad(str, n[, padding_str]) ``` #### Arguments[​](#arguments-62 "Direct link to Arguments") * **str**: String expression. * **n**: Target length. * **padding\_str**: Optional string to pad with. #### Example[​](#example-28 "Direct link to Example") ``` > select lpad('Dolly', 10, 'hello'); +---------------------------------------------+ | lpad(Utf8("Dolly"),Int64(10),Utf8("hello")) | +---------------------------------------------+ | helloDolly | +---------------------------------------------+ ``` Related function: [`rpad`](#rpad) ### `ltrim`[​](#ltrim "Direct link to ltrim") Removes the longest string containing only characters in `trim_str` from the start of `str`. If `trim_str` is omitted, whitespace is removed. ``` ltrim(str[, trim_str]) ``` #### Arguments[​](#arguments-63 "Direct link to Arguments") * **str**: String expression. * **trim\_str**: Optional string of characters to trim. Defaults to whitespace. #### Example[​](#example-29 "Direct link to Example") ``` > select ltrim(' datafusion '); +-------------------------------+ | ltrim(Utf8(" datafusion ")) | +-------------------------------+ | datafusion | +-------------------------------+ > select ltrim('___datafusion___', '_'); +-------------------------------------------+ | ltrim(Utf8("___datafusion___"),Utf8("_")) | +-------------------------------------------+ | datafusion___ | +-------------------------------------------+ ``` Alternative syntax: `trim(LEADING trim_str FROM str)` Related functions: [`btrim`](#btrim), [`rtrim`](#rtrim) ### `octet_length`[​](#octet_length "Direct link to octet_length") Returns the number of bytes in the string. ``` octet_length(str) ``` #### Arguments[​](#arguments-64 "Direct link to Arguments") * **str**: String expression. #### Example[​](#example-30 "Direct link to Example") ``` > select octet_length('Ångström'); +--------------------------------+ | octet_length(Utf8("Ångström")) | +--------------------------------+ | 10 | +--------------------------------+ ``` Related functions: [`bit_length`](#bit_length), [`length`](#length) ### `overlay`[​](#overlay "Direct link to overlay") Replaces a substring of `str` with `substr`, starting at position `pos` for `count` characters. If `count` is omitted, uses the length of `substr`. ``` overlay(str PLACING substr FROM pos [FOR count]) ``` #### Arguments[​](#arguments-65 "Direct link to Arguments") * **str**: String expression. * **substr**: Replacement string. * **pos**: Start position (1-based). * **count**: Optional number of characters to replace. #### Example[​](#example-31 "Direct link to Example") ``` > select overlay('Txxxxas' placing 'hom' from 2 for 4); +--------------------------------------------------------+ | overlay(Utf8("Txxxxas"),Utf8("hom"),Int64(2),Int64(4)) | +--------------------------------------------------------+ | Thomas | +--------------------------------------------------------+ ``` ### `parse_url`[​](#parse_url "Direct link to parse_url") Extracts a component from a URL, or retrieves an individual query parameter when provided a key, following Spark SQL semantics. Supported parts include `HOST`, `PATH`, `QUERY`, `REF`, `PROTOCOL`, `FILE`, and `AUTHORITY`. ``` parse_url(url, part_to_extract[, key]) ``` #### Arguments[​](#arguments-66 "Direct link to Arguments") * **url**: URL string expression. * **part\_to\_extract**: Case-insensitive token identifying which component to extract. * **key**: Optional query parameter key to extract from the `QUERY` part. #### Example[​](#example-32 "Direct link to Example") ``` > select parse_url('https://spice.ai/blog?id=42', 'HOST'); +-------------------------------------------------------+ | parse_url(Utf8("https://spice.ai/blog?id=42"),Utf8("HOST")) | +-------------------------------------------------------+ | spice.ai | +-------------------------------------------------------+ > select parse_url('https://spice.ai/blog?id=42', 'QUERY', 'id'); +------------------------------------------------------------------+ | parse_url(Utf8("https://spice.ai/blog?id=42"),Utf8("QUERY"),Utf8("id")) | +------------------------------------------------------------------+ | 42 | +------------------------------------------------------------------+ ``` Reference: [Spark SQL `parse_url`](https://spark.apache.org/docs/latest/api/sql/index.html#parse_url). ### `position`[​](#position "Direct link to position") Alias of [`strpos`](#strpos). ### `repeat`[​](#repeat "Direct link to repeat") Returns a string consisting of the input string repeated `n` times. ``` repeat(str, n) ``` #### Arguments[​](#arguments-67 "Direct link to Arguments") * **str**: String expression. * **n**: Number of repetitions. #### Example[​](#example-33 "Direct link to Example") ``` > select repeat('data', 3); +-------------------------------+ | repeat(Utf8("data"),Int64(3)) | +-------------------------------+ | datadatadata | +-------------------------------+ ``` ### `replace`[​](#replace "Direct link to replace") Replaces all occurrences of `substr` in `str` with `replacement`. ``` replace(str, substr, replacement) ``` #### Arguments[​](#arguments-68 "Direct link to Arguments") * **str**: String expression. * **substr**: Substring to replace. * **replacement**: Replacement string. #### Example[​](#example-34 "Direct link to Example") ``` > select replace('ABabbaBA', 'ab', 'cd'); +-------------------------------------------------+ | replace(Utf8("ABabbaBA"),Utf8("ab"),Utf8("cd")) | +-------------------------------------------------+ | ABcdbaBA | +-------------------------------------------------+ ``` ### `reverse`[​](#reverse "Direct link to reverse") Returns the string with the character order reversed. ``` reverse(str) ``` #### Arguments[​](#arguments-69 "Direct link to Arguments") * **str**: String expression. #### Example[​](#example-35 "Direct link to Example") ``` > select reverse('datafusion'); +-----------------------------+ | reverse(Utf8("datafusion")) | +-----------------------------+ | noisufatad | +-----------------------------+ ``` ### `right`[​](#right "Direct link to right") Returns the last `n` characters from the right side of the string. ``` right(str, n) ``` #### Arguments[​](#arguments-70 "Direct link to Arguments") * **str**: String expression. * **n**: Number of characters to return. #### Example[​](#example-36 "Direct link to Example") ``` > select right('datafusion', 6); +------------------------------------+ | right(Utf8("datafusion"),Int64(6)) | +------------------------------------+ | fusion | +------------------------------------+ ``` Related function: [`left`](#left) ### `rpad`[​](#rpad "Direct link to rpad") Pads the right side of the string with another string until the result reaches the specified length. If the padding string is omitted, a space is used. ``` rpad(str, n[, padding_str]) ``` #### Arguments[​](#arguments-71 "Direct link to Arguments") * **str**: String expression. * **n**: Target length. * **padding\_str**: Optional string to pad with. #### Example[​](#example-37 "Direct link to Example") ``` > select rpad('datafusion', 20, '_-'); +-----------------------------------------------+ | rpad(Utf8("datafusion"),Int64(20),Utf8("_-")) | +-----------------------------------------------+ | datafusion_-_-_-_-_- | +-----------------------------------------------+ ``` Related function: [`lpad`](#lpad) ### `rtrim`[​](#rtrim "Direct link to rtrim") Removes the longest string containing only characters in `trim_str` from the end of `str`. If `trim_str` is omitted, whitespace is removed. ``` rtrim(str[, trim_str]) ``` #### Arguments[​](#arguments-72 "Direct link to Arguments") * **str**: String expression. * **trim\_str**: Optional string of characters to trim. Defaults to whitespace. #### Example[​](#example-38 "Direct link to Example") ``` > select rtrim(' datafusion '); +-------------------------------+ | rtrim(Utf8(" datafusion ")) | +-------------------------------+ | datafusion | +-------------------------------+ > select rtrim('___datafusion___', '_'); +-------------------------------------------+ | rtrim(Utf8("___datafusion___"),Utf8("_")) | +-------------------------------------------+ | ___datafusion | +-------------------------------------------+ ``` Alternative syntax: `trim(TRAILING trim_str FROM str)` Related functions: [`btrim`](#btrim), [`ltrim`](#ltrim) ### `split_part`[​](#split_part "Direct link to split_part") Splits the string on the specified delimiter and returns the substring at the given position (1-based). ``` split_part(str, delimiter, pos) ``` #### Arguments[​](#arguments-73 "Direct link to Arguments") * **str**: String expression. * **delimiter**: Delimiter string. * **pos**: Position of the part to return (1-based). #### Example[​](#example-39 "Direct link to Example") ``` > select split_part('1.2.3.4.5', '.', 3); +--------------------------------------------------+ | split_part(Utf8("1.2.3.4.5"),Utf8("."),Int64(3)) | +--------------------------------------------------+ | 3 | +--------------------------------------------------+ ``` ### `starts_with`[​](#starts_with "Direct link to starts_with") Returns true if `str` starts with the substring `substr`. ``` starts_with(str, substr) ``` #### Arguments[​](#arguments-74 "Direct link to Arguments") * **str**: String expression. * **substr**: Substring to test for. #### Example[​](#example-40 "Direct link to Example") ``` > select starts_with('datafusion','data'); +----------------------------------------------+ | starts_with(Utf8("datafusion"),Utf8("data")) | +----------------------------------------------+ | true | +----------------------------------------------+ ``` ### `strpos`[​](#strpos "Direct link to strpos") Returns the position (1-based) of the first occurrence of `substr` in `str`. Returns 0 if not found. ``` strpos(str, substr) ``` #### Arguments[​](#arguments-75 "Direct link to Arguments") * **str**: String expression. * **substr**: Substring to search for. #### Example[​](#example-41 "Direct link to Example") ``` > select strpos('datafusion', 'fus'); +----------------------------------------+ | strpos(Utf8("datafusion"),Utf8("fus")) | +----------------------------------------+ | 5 | +----------------------------------------+ ``` Alternative syntax: `position(substr in origstr)` Aliases: `instr`, `position` ### `substr`[​](#substr "Direct link to substr") Extracts a substring from `str`, starting at `start_pos` for `length` characters. If `length` is omitted, returns the rest of the string. ``` substr(str, start_pos[, length]) ``` #### Arguments[​](#arguments-76 "Direct link to Arguments") * **str**: String expression. * **start\_pos**: Start position (1-based). * **length**: Optional number of characters to extract. #### Example[​](#example-42 "Direct link to Example") ``` > select substr('datafusion', 5, 3); +----------------------------------------------+ | substr(Utf8("datafusion"),Int64(5),Int64(3)) | +----------------------------------------------+ | fus | +----------------------------------------------+ ``` Alternative syntax: `substring(str from start_pos for length)` Aliases: `substring` ### `substr_index`[​](#substr_index "Direct link to substr_index") Returns the substring from `str` before or after a specified number of occurrences of the delimiter `delim`. If `count` is positive, returns everything to the left of the final delimiter (counting from the left). If `count` is negative, returns everything to the right of the final delimiter (counting from the right). ``` substr_index(str, delim, count) ``` #### Arguments[​](#arguments-77 "Direct link to Arguments") * **str**: String expression. * **delim**: Delimiter string. * **count**: Number of occurrences (positive or negative). #### Example[​](#example-43 "Direct link to Example") ``` > select substr_index('www.apache.org', '.', 1); +---------------------------------------------------------+ | substr_index(Utf8("www.apache.org"),Utf8("."),Int64(1)) | +---------------------------------------------------------+ | www | +---------------------------------------------------------+ > select substr_index('www.apache.org', '.', -1); +----------------------------------------------------------+ | substr_index(Utf8("www.apache.org"),Utf8("."),Int64(-1)) | +----------------------------------------------------------+ | org | +----------------------------------------------------------+ ``` Aliases: `substring_index` ### `substring`[​](#substring "Direct link to substring") Alias of [`substr`](#substr). ### `substring_index`[​](#substring_index "Direct link to substring_index") Alias of [`substr_index`](#substr_index). ### `to_hex`[​](#to_hex "Direct link to to_hex") Converts an integer to its hexadecimal string representation. ``` to_hex(int) ``` #### Arguments[​](#arguments-78 "Direct link to Arguments") * **int**: Integer expression. #### Example[​](#example-44 "Direct link to Example") ``` > select to_hex(12345689); +-------------------------+ | to_hex(Int64(12345689)) | +-------------------------+ | bc6159 | +-------------------------+ ``` ### `translate`[​](#translate "Direct link to translate") Replaces each character in `str` that matches a character in `chars` with the corresponding character in `translation`. If `translation` is shorter than `chars`, extra characters are removed. ``` translate(str, chars, translation) ``` #### Arguments[​](#arguments-79 "Direct link to Arguments") * **str**: String expression. * **chars**: Characters to translate. * **translation**: Replacement characters. #### Example[​](#example-45 "Direct link to Example") ``` > select translate('twice', 'wic', 'her'); +--------------------------------------------------+ | translate(Utf8("twice"),Utf8("wic"),Utf8("her")) | +--------------------------------------------------+ | there | +--------------------------------------------------+ ``` ### `trim`[​](#trim "Direct link to trim") Alias of [`btrim`](#btrim). ### `upper`[​](#upper "Direct link to upper") Converts all characters in the string to upper case. ``` upper(str) ``` #### Arguments[​](#arguments-80 "Direct link to Arguments") * **str**: String expression. #### Example[​](#example-46 "Direct link to Example") ``` > select upper('dataFusion'); +---------------------------+ | upper(Utf8("dataFusion")) | +---------------------------+ | DATAFUSION | +---------------------------+ ``` Related functions: [`initcap`](#initcap), [`lower`](#lower) ### `uuid`[​](#uuid "Direct link to uuid") Returns a [UUID v4](https://en.wikipedia.org/wiki/Universally_unique_identifier#Version_4_\(random\)) string value that is unique per row. ``` uuid() ``` #### Example[​](#example-47 "Direct link to Example") ``` > select uuid(); +--------------------------------------+ | uuid() | +--------------------------------------+ | 6ec17ef8-1934-41cc-8d59-d0c8f9eea1f0 | +--------------------------------------+ ``` *** ## Binary String Functions[​](#binary-string-functions "Direct link to Binary String Functions") Binary string functions help encode and decode binary data, such as base64 and hexadecimal conversions. These are useful for working with encoded data or binary blobs. * [bit\_get](#bit_get) * [bit\_count](#bit_count) * [bitmap\_count](#bitmap_count) ### `bit_get`[​](#bit_get "Direct link to bit_get") Returns the bit (0 or 1) at the specified zero-based position when counting from the least-significant bit of an integral or binary expression, matching Spark SQL semantics. ``` bit_get(value, position) ``` #### Arguments[​](#arguments-81 "Direct link to Arguments") * **value**: Integer or binary expression whose bits are inspected. * **position**: Zero-based index of the bit to return. Must be non-negative. #### Example[​](#example-48 "Direct link to Example") ``` > select bit_get(11, 2) as bit; +-----+ | bit | +-----+ | 0 | +-----+ ``` Reference: [Spark SQL `bit_get`](https://spark.apache.org/docs/latest/api/sql/index.html#bit_get). ### `bit_count`[​](#bit_count "Direct link to bit_count") Counts the number of set bits in an integral or binary expression. Useful for quick popcount operations on bitmaps or packed flags, aligned with Spark SQL behavior. ``` bit_count(value) ``` #### Arguments[​](#arguments-82 "Direct link to Arguments") * **value**: Integer or binary expression. #### Example[​](#example-49 "Direct link to Example") ``` > select bit_count(255) as popcnt; +--------+ | popcnt | +--------+ | 8 | +--------+ ``` Reference: [Spark SQL `bit_count`](https://spark.apache.org/docs/latest/api/sql/index.html#bit_count). ### `bitmap_count`[​](#bitmap_count "Direct link to bitmap_count") Returns the number of set bits in a binary bitmap produced by functions such as `bitmap_construct_agg`, mirroring the Spark SQL implementation. ``` bitmap_count(bitmap) ``` #### Arguments[​](#arguments-83 "Direct link to Arguments") * **bitmap**: Binary expression representing a bitmap. #### Example[​](#example-50 "Direct link to Example") ``` > select bitmap_count(x'0F') as popcnt; +--------+ | popcnt | +--------+ | 4 | +--------+ ``` Reference: [Spark SQL `bitmap_count`](https://spark.apache.org/docs/latest/api/sql/index.html#bitmap_count). ## Regular Expression Functions[​](#regular-expression-functions "Direct link to Regular Expression Functions") Regular expression functions help match, extract, and replace patterns in strings. Spice.ai uses a PCRE-like regular expression syntax. Spice supports the following regular expressions: * [regexp\_like](#regexp_like) * [regexp\_match](#regexp_match) * [regexp\_replace](#regexp_replace) * [regexp\_count](#regexp_count) * [regexp\_instr](#regexp_instr) ### `regexp_like`[​](#regexp_like "Direct link to regexp_like") Returns true if a regular expression has at least one match in a string, false otherwise. ``` regexp_like(str, regexp[, flags]) ``` #### Arguments[​](#arguments-84 "Direct link to Arguments") * **str**: String expression to operate on. Can be a constant, column, or function, and any combination of operators. * **regexp**: Regular expression to operate on. Can be a constant, column, or function, and any combination of operators. * **flags**: Optional regular expression flags that control the behavior of the regular expression. The following flags are supported: * **i**: case-insensitive: letters match both upper and lower case * **m**: multi-line mode: ^ and $ match begin/end of line * **s**: allow . to match \n * **R**: enables CRLF mode: when multi-line mode is enabled, \r\n is used * **U**: swap the meaning of x\* and x\*? #### Example[​](#example-51 "Direct link to Example") ``` > select regexp_like('Köln', '[a-zA-Z]ö[a-zA-Z]{2}'); +--------------------------------------------------------+ | regexp_like(Utf8("Köln"),Utf8("[a-zA-Z]ö[a-zA-Z]{2}")) | +--------------------------------------------------------+ | true | +--------------------------------------------------------+ > SELECT regexp_like('aBc', '(b|d)', 'i'); +--------------------------------------------------+ | regexp_like(Utf8("aBc"),Utf8("(b|d)"),Utf8("i")) | +--------------------------------------------------+ | true | +--------------------------------------------------+ ``` ### `regexp_match`[​](#regexp_match "Direct link to regexp_match") Returns the first regular expression matches in a string. ``` regexp_match(str, regexp[, flags]) ``` #### Arguments[​](#arguments-85 "Direct link to Arguments") * **str**: String expression to operate on. Can be a constant, column, or function, and any combination of operators. * **regexp**: Regular expression to match against. Can be a constant, column, or function. * **flags**: Optional regular expression flags that control the behavior of the regular expression. The following flags are supported: * **i**: case-insensitive: letters match both upper and lower case * **m**: multi-line mode: ^ and $ match begin/end of line * **s**: allow . to match \n * **R**: enables CRLF mode: when multi-line mode is enabled, \r\n is used * **U**: swap the meaning of x\* and x\*? #### Example[​](#example-52 "Direct link to Example") ``` > select regexp_match('Köln', '[a-zA-Z]ö[a-zA-Z]{2}'); +---------------------------------------------------------+ | regexp_match(Utf8("Köln"),Utf8("[a-zA-Z]ö[a-zA-Z]{2}")) | +---------------------------------------------------------+ | [Köln] | +---------------------------------------------------------+ SELECT regexp_match('aBc', '(b|d)', 'i'); +---------------------------------------------------+ | regexp_match(Utf8("aBc"),Utf8("(b|d)"),Utf8("i")) | +---------------------------------------------------+ | [B] | +---------------------------------------------------+ ``` ### `regexp_replace`[​](#regexp_replace "Direct link to regexp_replace") Replaces substrings in a string that match a regular expression. ``` regexp_replace(str, regexp, replacement[, flags]) ``` #### Arguments[​](#arguments-86 "Direct link to Arguments") * **str**: String expression to operate on. Can be a constant, column, or function, and any combination of operators. * **regexp**: Regular expression to match against. Can be a constant, column, or function. * **replacement**: Replacement string expression to operate on. Can be a constant, column, or function, and any combination of operators. * **flags**: Optional regular expression flags that control the behavior of the regular expression. The following flags are supported: * **g**: (global) Search globally and don’t return after the first match * **i**: case-insensitive: letters match both upper and lower case * **m**: multi-line mode: ^ and $ match begin/end of line * **s**: allow . to match \n * **R**: enables CRLF mode: when multi-line mode is enabled, \r\n is used * **U**: swap the meaning of x\* and x\*? #### Example[​](#example-53 "Direct link to Example") ``` > select regexp_replace('foobarbaz', 'b(..)', 'X\\1Y', 'g'); +------------------------------------------------------------------------+ | regexp_replace(Utf8("foobarbaz"),Utf8("b(..)"),Utf8("X\1Y"),Utf8("g")) | +------------------------------------------------------------------------+ | fooXarYXazY | +------------------------------------------------------------------------+ SELECT regexp_replace('aBc', '(b|d)', 'Ab\\1a', 'i'); +-------------------------------------------------------------------+ | regexp_replace(Utf8("aBc"),Utf8("(b|d)"),Utf8("Ab\1a"),Utf8("i")) | +-------------------------------------------------------------------+ | aAbBac | +-------------------------------------------------------------------+ ``` ### `regexp_count`[​](#regexp_count "Direct link to regexp_count") Returns the number of matches that a regular expression has in a string. ``` regexp_count(str, regexp[, start, flags]) ``` #### Arguments[​](#arguments-87 "Direct link to Arguments") * **str**: String expression to operate on. Can be a constant, column, or function, and any combination of operators. * **regexp**: Regular expression to operate on. Can be a constant, column, or function, and any combination of operators. * **start**: Optional start position (the first position is 1) to search for the regular expression. Can be a constant, column, or function. * **flags**: Optional regular expression flags that control the behavior of the regular expression. The following flags are supported: * **i**: case-insensitive: letters match both upper and lower case * **m**: multi-line mode: ^ and $ match begin/end of line * **s**: allow . to match \n * **R**: enables CRLF mode: when multi-line mode is enabled, \r\n is used * **U**: swap the meaning of x\* and x\*? #### Example[​](#example-54 "Direct link to Example") ``` > select regexp_count('abcAbAbc', 'abc', 2, 'i'); +---------------------------------------------------------------+ | regexp_count(Utf8("abcAbAbc"),Utf8("abc"),Int64(2),Utf8("i")) | +---------------------------------------------------------------+ | 1 | +---------------------------------------------------------------+ ``` ### `regexp_instr`[​](#regexp_instr "Direct link to regexp_instr") Returns the position in a string where the specified occurrence of a POSIX regular expression is located. ``` regexp_instr(str, regexp[, start[, N[, flags[, subexpr]]]]) ``` #### Arguments[​](#arguments-88 "Direct link to Arguments") * **str**: String expression to operate on. Can be a constant, column, or function, and any combination of operators. * **regexp**: Regular expression to operate on. Can be a constant, column, or function, and any combination of operators. * **start**: Optional start position (the first position is 1) to search for the regular expression. Can be a constant, column, or function. Defaults to 1. * **N**: Optional. The N-th occurrence of pattern to find. Defaults to 1 (first match). Can be a constant, column, or function. * **flags**: Optional regular expression flags that control the behavior of the regular expression. The following flags are supported: * **i**: case-insensitive: letters match both upper and lower case * **m**: multi-line mode: ^ and $ match begin/end of line * **s**: allow . to match \n * **R**: enables CRLF mode: when multi-line mode is enabled, \r\n is used * **U**: swap the meaning of x\* and x\*? * **subexpr**: Optional. Specifies which capture group (subexpression) to return the position for. Defaults to 0, which returns the position of the entire match. #### Example[​](#example-55 "Direct link to Example") ``` > SELECT regexp_instr('ABCDEF', 'C(.)(..)'); +---------------------------------------------------------------+ | regexp_instr(Utf8("ABCDEF"),Utf8("C(.)(..)")) | +---------------------------------------------------------------+ | 3 | +---------------------------------------------------------------+ ``` ## Time and Date Functions[​](#time-and-date-functions "Direct link to Time and Date Functions") Time and date functions help extract, format, and manipulate temporal data. Functions include `current_date`, `now`, `date_part`, `date_trunc`, and various conversion functions. These are essential for time series analysis and working with timestamps. * [current\_date](#current_date) * [current\_time](#current_time) * [current\_timestamp](#current_timestamp) * [date\_bin](#date_bin) * [date\_add](#date_add) * [date\_sub](#date_sub) * [last\_day](#last_day) * [next\_day](#next_day) * [date\_format](#date_format) * [date\_part](#date_part) * [date\_trunc](#date_trunc) * [datepart](#datepart) * [datetrunc](#datetrunc) * [from\_unixtime](#from_unixtime) * [make\_date](#make_date) * [now](#now) * [to\_char](#to_char) * [to\_date](#to_date) * [to\_local\_time](#to_local_time) * [to\_timestamp](#to_timestamp) * [to\_timestamp\_micros](#to_timestamp_micros) * [to\_timestamp\_millis](#to_timestamp_millis) * [to\_timestamp\_nanos](#to_timestamp_nanos) * [to\_timestamp\_seconds](#to_timestamp_seconds) * [to\_unixtime](#to_unixtime) * [today](#today) ### `current_date`[​](#current_date "Direct link to current_date") Returns the current UTC date. The `current_date()` return value is determined at query time and will return the same date, no matter when in the query plan the function executes. ``` current_date() ``` #### Aliases[​](#aliases "Direct link to Aliases") * today ### `current_time`[​](#current_time "Direct link to current_time") Returns the current UTC time. The `current_time()` return value is determined at query time and will return the same time, no matter when in the query plan the function executes. ``` current_time() ``` ### `current_timestamp`[​](#current_timestamp "Direct link to current_timestamp") *Alias of [now](#now).* ### `date_bin`[​](#date_bin "Direct link to date_bin") Calculates time intervals and returns the start of the interval nearest to the specified timestamp. Use `date_bin` to downsample time series data by grouping rows into time-based "bins" or "windows" and applying an aggregate or selector function to each window. For example, if you "bin" or "window" data into 15 minute intervals, an input timestamp of `2023-01-01T18:18:18Z` will be updated to the start time of the 15 minute bin it is in: `2023-01-01T18:15:00Z`. ``` date_bin(interval, expression, origin-timestamp) ``` #### Arguments[​](#arguments-89 "Direct link to Arguments") * **interval**: Bin interval. * **expression**: Time expression to operate on. Can be a constant, column, or function. * **origin-timestamp**: Optional. Starting point used to determine bin boundaries. If not specified defaults 1970-01-01T00:00:00Z (the UNIX epoch in UTC). The following intervals are supported: * nanoseconds * microseconds * milliseconds * seconds * minutes * hours * days * weeks * months * years * century #### Example[​](#example-56 "Direct link to Example") ``` -- Bin the timestamp into 1 day intervals > SELECT date_bin(interval '1 day', time) as bin FROM VALUES ('2023-01-01T18:18:18Z'), ('2023-01-03T19:00:03Z') t(time); +---------------------+ | bin | +---------------------+ | 2023-01-01T00:00:00 | | 2023-01-03T00:00:00 | +---------------------+ 2 row(s) fetched. -- Bin the timestamp into 1 day intervals starting at 3AM on 2023-01-01 > SELECT date_bin(interval '1 day', time, '2023-01-01T03:00:00') as bin FROM VALUES ('2023-01-01T18:18:18Z'), ('2023-01-03T19:00:03Z') t(time); +---------------------+ | bin | +---------------------+ | 2023-01-01T03:00:00 | | 2023-01-03T03:00:00 | +---------------------+ 2 row(s) fetched. ``` ### `date_add`[​](#date_add "Direct link to date_add") Adds a number of days to a DATE or TIMESTAMP expression, matching Spark SQL semantics. Negative offsets move backwards in time. ``` date_add(start_date, num_days) ``` #### Arguments[​](#arguments-90 "Direct link to Arguments") * **start\_date**: DATE or TIMESTAMP expression. * **num\_days**: Integer number of days to add. #### Example[​](#example-57 "Direct link to Example") ``` > select date_add(date '2024-02-27', 3); +---------------------------------+ | date_add(Date32("2024-02-27"),Int64(3)) | +---------------------------------+ | 2024-03-01 | +---------------------------------+ ``` Reference: [Spark SQL `date_add`](https://spark.apache.org/docs/latest/api/sql/index.html#date_add). ### `date_sub`[​](#date_sub "Direct link to date_sub") Subtracts a number of days from a DATE or TIMESTAMP expression using Spark-compatible behavior. ``` date_sub(start_date, num_days) ``` #### Arguments[​](#arguments-91 "Direct link to Arguments") * **start\_date**: DATE or TIMESTAMP expression. * **num\_days**: Integer number of days to subtract. #### Example[​](#example-58 "Direct link to Example") ``` > select date_sub(date '2024-03-05', 7); +---------------------------------+ | date_sub(Date32("2024-03-05"),Int64(7)) | +---------------------------------+ | 2024-02-27 | +---------------------------------+ ``` Reference: [Spark SQL `date_sub`](https://spark.apache.org/docs/latest/api/sql/index.html#date_sub). ### `last_day`[​](#last_day "Direct link to last_day") Returns the last day of the month that contains the input date or timestamp, matching Spark SQL semantics. ``` last_day(expression) ``` #### Arguments[​](#arguments-92 "Direct link to Arguments") * **expression**: DATE or TIMESTAMP expression. #### Example[​](#example-59 "Direct link to Example") ``` > select last_day(date '2024-02-14'); +----------------------------------+ | last_day(Date32("2024-02-14")) | +----------------------------------+ | 2024-02-29 | +----------------------------------+ ``` Reference: [Spark SQL `last_day`](https://spark.apache.org/docs/latest/api/sql/index.html#last_day). ### `next_day`[​](#next_day "Direct link to next_day") Returns the first date after `start_date` that matches the requested day of week. Valid day names include full names (e.g., `Monday`) or abbreviations such as `Mon`, matching Spark SQL behavior. ``` next_day(start_date, day_of_week) ``` #### Arguments[​](#arguments-93 "Direct link to Arguments") * **start\_date**: DATE or TIMESTAMP expression. * **day\_of\_week**: String literal naming the target weekday. #### Example[​](#example-60 "Direct link to Example") ``` > select next_day(date '2024-02-14', 'FRI'); +--------------------------------------------------+ | next_day(Date32("2024-02-14"),Utf8("FRI")) | +--------------------------------------------------+ | 2024-02-16 | +--------------------------------------------------+ ``` Reference: [Spark SQL `next_day`](https://spark.apache.org/docs/latest/api/sql/index.html#next_day). ### `date_format`[​](#date_format "Direct link to date_format") *Alias of [to\_char](#to_char).* ### `date_part`[​](#date_part "Direct link to date_part") Returns the specified part of the date as an integer. ``` date_part(part, expression) ``` #### Arguments[​](#arguments-94 "Direct link to Arguments") * **part**: Part of the date to return. The following date parts are supported: * year * quarter (emits value in inclusive range \[1, 4] based on which quartile of the year the date is in) * month * week (week of the year) * day (day of the month) * hour * minute * second * millisecond * microsecond * nanosecond * dow (day of the week where Sunday is 0) * doy (day of the year) * epoch (seconds since Unix epoch) * isodow (day of the week where Monday is 0) * **expression**: Time expression to operate on. Can be a constant, column, or function. #### Alternative Syntax[​](#alternative-syntax "Direct link to Alternative Syntax") ``` extract(field FROM source) ``` #### Aliases[​](#aliases-1 "Direct link to Aliases") * datepart ### `date_trunc`[​](#date_trunc "Direct link to date_trunc") Truncates a timestamp value to a specified precision. ``` date_trunc(precision, expression) ``` #### Arguments[​](#arguments-95 "Direct link to Arguments") * **precision**: Time precision to truncate to. The following precisions are supported: * year / YEAR * quarter / QUARTER * month / MONTH * week / WEEK * day / DAY * hour / HOUR * minute / MINUTE * second / SECOND * millisecond / MILLISECOND * microsecond / MICROSECOND * **expression**: Time expression to operate on. Can be a constant, column, or function. #### Aliases[​](#aliases-2 "Direct link to Aliases") * datetrunc ### `datepart`[​](#datepart "Direct link to datepart") *Alias of [date\_part](#date_part).* ### `datetrunc`[​](#datetrunc "Direct link to datetrunc") *Alias of [date\_trunc](#date_trunc).* ### `from_unixtime`[​](#from_unixtime "Direct link to from_unixtime") Converts an integer to RFC3339 timestamp format (`YYYY-MM-DDT00:00:00.000000000Z`). Integers and unsigned integers are interpreted as seconds since the unix epoch (`1970-01-01T00:00:00Z`) return the corresponding timestamp. ``` from_unixtime(expression[, timezone]) ``` #### Arguments[​](#arguments-96 "Direct link to Arguments") * **expression**: The expression to operate on. Can be a constant, column, or function, and any combination of operators. * **timezone**: Optional timezone to use when converting the integer to a timestamp. If not provided, the default timezone is UTC. #### Example[​](#example-61 "Direct link to Example") ``` > select from_unixtime(1599572549, 'America/New_York'); +-----------------------------------------------------------+ | from_unixtime(Int64(1599572549),Utf8("America/New_York")) | +-----------------------------------------------------------+ | 2020-09-08T09:42:29-04:00 | +-----------------------------------------------------------+ ``` ### `make_date`[​](#make_date "Direct link to make_date") Make a date from year/month/day component parts. ``` make_date(year, month, day) ``` #### Arguments[​](#arguments-97 "Direct link to Arguments") * **year**: Year to use when making the date. Can be a constant, column or function, and any combination of arithmetic operators. * **month**: Month to use when making the date. Can be a constant, column or function, and any combination of arithmetic operators. * **day**: Day to use when making the date. Can be a constant, column or function, and any combination of arithmetic operators. #### Example[​](#example-62 "Direct link to Example") ``` > select make_date(2023, 1, 31); +-------------------------------------------+ | make_date(Int64(2023),Int64(1),Int64(31)) | +-------------------------------------------+ | 2023-01-31 | +-------------------------------------------+ > select make_date('2023', '01', '31'); +-----------------------------------------------+ | make_date(Utf8("2023"),Utf8("01"),Utf8("31")) | +-----------------------------------------------+ | 2023-01-31 | +-----------------------------------------------+ ``` ### `now`[​](#now "Direct link to now") Returns the current UTC timestamp. The `now()` return value is determined at query time and will return the same timestamp, no matter when in the query plan the function executes. ``` now() ``` #### Aliases[​](#aliases-3 "Direct link to Aliases") * current\_timestamp ### `to_char`[​](#to_char "Direct link to to_char") Returns a string representation of a date, time, timestamp or duration based on a [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html). Unlike the PostgreSQL equivalent of this function numerical formatting is not supported. ``` to_char(expression, format) ``` #### Arguments[​](#arguments-98 "Direct link to Arguments") * **expression**: Expression to operate on. Can be a constant, column, or function that results in a date, time, timestamp or duration. * **format**: A [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html) string to use to convert the expression. * **day**: Day to use when making the date. Can be a constant, column or function, and any combination of arithmetic operators. #### Example[​](#example-63 "Direct link to Example") ``` > select to_char('2023-03-01'::date, '%d-%m-%Y'); +----------------------------------------------+ | to_char(Utf8("2023-03-01"),Utf8("%d-%m-%Y")) | +----------------------------------------------+ | 01-03-2023 | +----------------------------------------------+ ``` #### Aliases[​](#aliases-4 "Direct link to Aliases") * date\_format ### `to_date`[​](#to_date "Direct link to to_date") Converts a value to a date (`YYYY-MM-DD`). Supports strings, integer and double types as input. Strings are parsed as YYYY-MM-DD (e.g. '2023-07-20') if no [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html)s are provided. Integers and doubles are interpreted as days since the unix epoch (`1970-01-01T00:00:00Z`). Returns the corresponding date. Note: `to_date` returns Date32, which represents its values as the number of days since unix epoch(`1970-01-01`) stored as signed 32 bit value. The largest supported date value is `9999-12-31`. ``` to_date('2017-05-31', '%Y-%m-%d') ``` #### Arguments[​](#arguments-99 "Direct link to Arguments") * **expression**: String expression to operate on. Can be a constant, column, or function, and any combination of operators. * **format\_n**: Optional [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html) strings to use to parse the expression. Formats will be tried in the order they appear with the first successful one being returned. If none of the formats successfully parse the expression an error will be returned. #### Example[​](#example-64 "Direct link to Example") ``` > select to_date('2023-01-31'); +-------------------------------+ | to_date(Utf8("2023-01-31")) | +-------------------------------+ | 2023-01-31 | +-------------------------------+ > select to_date('2023/01/31', '%Y-%m-%d', '%Y/%m/%d'); +---------------------------------------------------------------------+ | to_date(Utf8("2023/01/31"),Utf8("%Y-%m-%d"),Utf8("%Y/%m/%d")) | +---------------------------------------------------------------------+ | 2023-01-31 | +---------------------------------------------------------------------+ ``` ### `to_local_time`[​](#to_local_time "Direct link to to_local_time") Converts a timestamp with a timezone to a timestamp without a timezone (with no offset or timezone information). This function handles daylight saving time changes. ``` to_local_time(expression) ``` #### Arguments[​](#arguments-100 "Direct link to Arguments") * **expression**: Time expression to operate on. Can be a constant, column, or function. #### Example[​](#example-65 "Direct link to Example") ``` > SELECT to_local_time('2024-04-01T00:00:20Z'::timestamp); +---------------------------------------------+ | to_local_time(Utf8("2024-04-01T00:00:20Z")) | +---------------------------------------------+ | 2024-04-01T00:00:20 | +---------------------------------------------+ > SELECT to_local_time('2024-04-01T00:00:20Z'::timestamp AT TIME ZONE 'Europe/Brussels'); +---------------------------------------------+ | to_local_time(Utf8("2024-04-01T00:00:20Z")) | +---------------------------------------------+ | 2024-04-01T00:00:20 | +---------------------------------------------+ > SELECT time, arrow_typeof(time) as type, to_local_time(time) as to_local_time, arrow_typeof(to_local_time(time)) as to_local_time_type FROM ( SELECT '2024-04-01T00:00:20Z'::timestamp AT TIME ZONE 'Europe/Brussels' AS time ); +---------------------------+------------------------------------------------+---------------------+-----------------------------+ | time | type | to_local_time | to_local_time_type | +---------------------------+------------------------------------------------+---------------------+-----------------------------+ | 2024-04-01T00:00:20+02:00 | Timestamp(Nanosecond, Some("Europe/Brussels")) | 2024-04-01T00:00:20 | Timestamp(Nanosecond, None) | +---------------------------+------------------------------------------------+---------------------+-----------------------------+ # combine `to_local_time()` with `date_bin()` to bin on boundaries in the timezone rather # than UTC boundaries > SELECT date_bin(interval '1 day', to_local_time('2024-04-01T00:00:20Z'::timestamp AT TIME ZONE 'Europe/Brussels')) AS date_bin; +---------------------+ | date_bin | +---------------------+ | 2024-04-01T00:00:00 | +---------------------+ > SELECT date_bin(interval '1 day', to_local_time('2024-04-01T00:00:20Z'::timestamp AT TIME ZONE 'Europe/Brussels')) AT TIME ZONE 'Europe/Brussels' AS date_bin_with_timezone; +---------------------------+ | date_bin_with_timezone | +---------------------------+ | 2024-04-01T00:00:00+02:00 | +---------------------------+ ``` ### `to_timestamp`[​](#to_timestamp "Direct link to to_timestamp") Converts a value to a timestamp (`YYYY-MM-DDT00:00:00Z`). Supports strings, integer, unsigned integer, and double types as input. Strings are parsed as RFC3339 (e.g. '2023-07-20T05:44:00') if no \[Chrono formats] are provided. Integers, unsigned integers, and doubles are interpreted as seconds since the unix epoch (`1970-01-01T00:00:00Z`). Returns the corresponding timestamp. Note: `to_timestamp` returns `Timestamp(Nanosecond)`. The supported range for integer input is between `-9223372037` and `9223372036`. Supported range for string input is between `1677-09-21T00:12:44.0` and `2262-04-11T23:47:16.0`. Please use `to_timestamp_seconds` for the input outside of supported bounds. ``` to_timestamp(expression[, ..., format_n]) ``` #### Arguments[​](#arguments-101 "Direct link to Arguments") * **expression**: Expression to operate on. Can be a constant, column, or function, and any combination of arithmetic operators. * **format\_n**: Optional [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html) strings to use to parse the expression. Formats will be tried in the order they appear with the first successful one being returned. If none of the formats successfully parse the expression an error will be returned. #### Example[​](#example-66 "Direct link to Example") ``` > select to_timestamp('2023-01-31T09:26:56.123456789-05:00'); +-----------------------------------------------------------+ | to_timestamp(Utf8("2023-01-31T09:26:56.123456789-05:00")) | +-----------------------------------------------------------+ | 2023-01-31T14:26:56.123456789 | +-----------------------------------------------------------+ > select to_timestamp('03:59:00.123456789 05-17-2023', '%c', '%+', '%H:%M:%S%.f %m-%d-%Y'); +--------------------------------------------------------------------------------------------------------+ | to_timestamp(Utf8("03:59:00.123456789 05-17-2023"),Utf8("%c"),Utf8("%+"),Utf8("%H:%M:%S%.f %m-%d-%Y")) | +--------------------------------------------------------------------------------------------------------+ | 2023-05-17T03:59:00.123456789 | +--------------------------------------------------------------------------------------------------------+ ``` ### `to_timestamp_micros`[​](#to_timestamp_micros "Direct link to to_timestamp_micros") Converts a value to a timestamp (`YYYY-MM-DDT00:00:00.000000Z`). Supports strings, integer, and unsigned integer types as input. Strings are parsed as RFC3339 (e.g. '2023-07-20T05:44:00') if no [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html)s are provided. Integers and unsigned integers are interpreted as microseconds since the unix epoch (`1970-01-01T00:00:00Z`) Returns the corresponding timestamp. ``` to_timestamp_micros(expression[, ..., format_n]) ``` #### Arguments[​](#arguments-102 "Direct link to Arguments") * **expression**: Expression to operate on. Can be a constant, column, or function, and any combination of arithmetic operators. * **format\_n**: Optional [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html) strings to use to parse the expression. Formats will be tried in the order they appear with the first successful one being returned. If none of the formats successfully parse the expression an error will be returned. #### Example[​](#example-67 "Direct link to Example") ``` > select to_timestamp_micros('2023-01-31T09:26:56.123456789-05:00'); +------------------------------------------------------------------+ | to_timestamp_micros(Utf8("2023-01-31T09:26:56.123456789-05:00")) | +------------------------------------------------------------------+ | 2023-01-31T14:26:56.123456 | +------------------------------------------------------------------+ > select to_timestamp_micros('03:59:00.123456789 05-17-2023', '%c', '%+', '%H:%M:%S%.f %m-%d-%Y'); +---------------------------------------------------------------------------------------------------------------+ | to_timestamp_micros(Utf8("03:59:00.123456789 05-17-2023"),Utf8("%c"),Utf8("%+"),Utf8("%H:%M:%S%.f %m-%d-%Y")) | +---------------------------------------------------------------------------------------------------------------+ | 2023-05-17T03:59:00.123456 | +---------------------------------------------------------------------------------------------------------------+ ``` ### `to_timestamp_millis`[​](#to_timestamp_millis "Direct link to to_timestamp_millis") Converts a value to a timestamp (`YYYY-MM-DDT00:00:00.000Z`). Supports strings, integer, and unsigned integer types as input. Strings are parsed as RFC3339 (e.g. '2023-07-20T05:44:00') if no [Chrono formats](https://docs.rs/chrono/latest/chrono/format/strftime/index.html) are provided. Integers and unsigned integers are interpreted as milliseconds since the unix epoch (`1970-01-01T00:00:00Z`). Returns the corresponding timestamp. ``` to_timestamp_millis(expression[, ..., format_n]) ``` #### Arguments[​](#arguments-103 "Direct link to Arguments") * **expression**: Expression to operate on. Can be a constant, column, or function, and any combination of arithmetic operators. * **format\_n**: Optional [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html) strings to use to parse the expression. Formats will be tried in the order they appear with the first successful one being returned. If none of the formats successfully parse the expression an error will be returned. #### Example[​](#example-68 "Direct link to Example") ``` > select to_timestamp_millis('2023-01-31T09:26:56.123456789-05:00'); +------------------------------------------------------------------+ | to_timestamp_millis(Utf8("2023-01-31T09:26:56.123456789-05:00")) | +------------------------------------------------------------------+ | 2023-01-31T14:26:56.123 | +------------------------------------------------------------------+ > select to_timestamp_millis('03:59:00.123456789 05-17-2023', '%c', '%+', '%H:%M:%S%.f %m-%d-%Y'); +---------------------------------------------------------------------------------------------------------------+ | to_timestamp_millis(Utf8("03:59:00.123456789 05-17-2023"),Utf8("%c"),Utf8("%+"),Utf8("%H:%M:%S%.f %m-%d-%Y")) | +---------------------------------------------------------------------------------------------------------------+ | 2023-05-17T03:59:00.123 | +---------------------------------------------------------------------------------------------------------------+ ``` ### `to_timestamp_nanos`[​](#to_timestamp_nanos "Direct link to to_timestamp_nanos") Converts a value to a timestamp (`YYYY-MM-DDT00:00:00.000000000Z`). Supports strings, integer, and unsigned integer types as input. Strings are parsed as RFC3339 (e.g. '2023-07-20T05:44:00') if no [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html)s are provided. Integers and unsigned integers are interpreted as nanoseconds since the unix epoch (`1970-01-01T00:00:00Z`). Returns the corresponding timestamp. ``` to_timestamp_nanos(expression[, ..., format_n]) ``` #### Arguments[​](#arguments-104 "Direct link to Arguments") * **expression**: Expression to operate on. Can be a constant, column, or function, and any combination of arithmetic operators. * **format\_n**: Optional [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html) strings to use to parse the expression. Formats will be tried in the order they appear with the first successful one being returned. If none of the formats successfully parse the expression an error will be returned. #### Example[​](#example-69 "Direct link to Example") ``` > select to_timestamp_nanos('2023-01-31T09:26:56.123456789-05:00'); +-----------------------------------------------------------------+ | to_timestamp_nanos(Utf8("2023-01-31T09:26:56.123456789-05:00")) | +-----------------------------------------------------------------+ | 2023-01-31T14:26:56.123456789 | +-----------------------------------------------------------------+ > select to_timestamp_nanos('03:59:00.123456789 05-17-2023', '%c', '%+', '%H:%M:%S%.f %m-%d-%Y'); +--------------------------------------------------------------------------------------------------------------+ | to_timestamp_nanos(Utf8("03:59:00.123456789 05-17-2023"),Utf8("%c"),Utf8("%+"),Utf8("%H:%M:%S%.f %m-%d-%Y")) | +--------------------------------------------------------------------------------------------------------------+ | 2023-05-17T03:59:00.123456789 | +---------------------------------------------------------------------------------------------------------------+ ``` ### `to_timestamp_seconds`[​](#to_timestamp_seconds "Direct link to to_timestamp_seconds") Converts a value to a timestamp (`YYYY-MM-DDT00:00:00.000Z`). Supports strings, integer, and unsigned integer types as input. Strings are parsed as RFC3339 (e.g. '2023-07-20T05:44:00') if no [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html)s are provided. Integers and unsigned integers are interpreted as seconds since the unix epoch (`1970-01-01T00:00:00Z`). Returns the corresponding timestamp. ``` to_timestamp_seconds(expression[, ..., format_n]) ``` #### Arguments[​](#arguments-105 "Direct link to Arguments") * **expression**: Expression to operate on. Can be a constant, column, or function, and any combination of arithmetic operators. * **format\_n**: Optional [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html) strings to use to parse the expression. Formats will be tried in the order they appear with the first successful one being returned. If none of the formats successfully parse the expression an error will be returned. #### Example[​](#example-70 "Direct link to Example") ``` > select to_timestamp_seconds('2023-01-31T09:26:56.123456789-05:00'); +-------------------------------------------------------------------+ | to_timestamp_seconds(Utf8("2023-01-31T09:26:56.123456789-05:00")) | +-------------------------------------------------------------------+ | 2023-01-31T14:26:56 | +-------------------------------------------------------------------+ > select to_timestamp_seconds('03:59:00.123456789 05-17-2023', '%c', '%+', '%H:%M:%S%.f %m-%d-%Y'); +----------------------------------------------------------------------------------------------------------------+ | to_timestamp_seconds(Utf8("03:59:00.123456789 05-17-2023"),Utf8("%c"),Utf8("%+"),Utf8("%H:%M:%S%.f %m-%d-%Y")) | +----------------------------------------------------------------------------------------------------------------+ | 2023-05-17T03:59:00 | +----------------------------------------------------------------------------------------------------------------+ ``` ### `to_unixtime`[​](#to_unixtime "Direct link to to_unixtime") Converts a value to seconds since the unix epoch (`1970-01-01T00:00:00Z`). Supports strings, dates, timestamps and double types as input. Strings are parsed as RFC3339 (e.g. '2023-07-20T05:44:00') if no [Chrono formats](https://docs.rs/chrono/latest/chrono/format/strftime/index.html) are provided. ``` to_unixtime(expression[, ..., format_n]) ``` #### Arguments[​](#arguments-106 "Direct link to Arguments") * **expression**: Expression to operate on. Can be a constant, column, or function, and any combination of arithmetic operators. * **format\_n**: Optional [Chrono format](https://docs.rs/chrono/latest/chrono/format/strftime/index.html) strings to use to parse the expression. Formats will be tried in the order they appear with the first successful one being returned. If none of the formats successfully parse the expression an error will be returned. #### Example[​](#example-71 "Direct link to Example") ``` > select to_unixtime('2020-09-08T12:00:00+00:00'); +------------------------------------------------+ | to_unixtime(Utf8("2020-09-08T12:00:00+00:00")) | +------------------------------------------------+ | 1599566400 | +------------------------------------------------+ > select to_unixtime('01-14-2023 01:01:30+05:30', '%q', '%d-%m-%Y %H/%M/%S', '%+', '%m-%d-%Y %H:%M:%S%#z'); +-----------------------------------------------------------------------------------------------------------------------------+ | to_unixtime(Utf8("01-14-2023 01:01:30+05:30"),Utf8("%q"),Utf8("%d-%m-%Y %H/%M/%S"),Utf8("%+"),Utf8("%m-%d-%Y %H:%M:%S%#z")) | +-----------------------------------------------------------------------------------------------------------------------------+ | 1673638290 | +-----------------------------------------------------------------------------------------------------------------------------+ ``` ### `today`[​](#today "Direct link to today") *Alias of [current\_date](#current_date).* ## Array Functions[​](#array-functions "Direct link to Array Functions") Array functions in Spice.ai SQL help construct, transform, and query array data types. These functions operate on array expressions, which can be constants, columns, or results of other functions. The implementation closely follows the PostgreSQL dialect. The following array functions are supported: * [array](#array) * [array\_any\_value](#array_any_value) * [array\_append](#array_append) * [array\_cat](#array_cat) * [array\_concat](#array_concat) * [array\_contains](#array_contains) * [array\_dims](#array_dims) * [array\_distance](#array_distance) * [array\_distinct](#array_distinct) * [array\_element](#array_element) * [array\_except](#array_except) * [array\_has](#array_has) * [array\_has\_all](#array_has_all) * [array\_has\_any](#array_has_any) * [array\_intersect](#array_intersect) * [array\_length](#array_length) * [array\_max](#array_max) * [array\_min](#array_min) * [array\_ndims](#array_ndims) * [array\_pop\_back](#array_pop_back) * [array\_pop\_front](#array_pop_front) * [array\_position](#array_position) * [array\_positions](#array_positions) * [array\_prepend](#array_prepend) * [array\_remove](#array_remove) * [array\_remove\_n](#array_remove_n) * [array\_remove\_all](#array_remove_all) * [array\_repeat](#array_repeat) * [array\_replace](#array_replace) * [array\_replace\_n](#array_replace_n) * [array\_replace\_all](#array_replace_all) * [array\_resize](#array_resize) * [array\_reverse](#array_reverse) * [array\_slice](#array_slice) * [array\_sort](#array_sort) * [array\_to\_string](#array_to_string) * [array\_union](#array_union) * [arrays\_zip](#arrays_zip) * [cardinality](#cardinality) * [empty](#empty) * [flatten](#flatten) * [make\_array](#make_array) * [range](#range) * [generate\_series](#generate_series) * [string\_to\_array](#string_to_array) ### `array`[​](#array "Direct link to array") Constructs an array from the provided expressions using Spark-compatible semantics. Inputs are evaluated left to right, cast to a common element type, and collected into a single Arrow list value without removing duplicates or nulls. ``` array(expression[, ..., expression_n]) ``` #### Arguments[​](#arguments-107 "Direct link to Arguments") * **expression**: Value to include in the array. Expressions must be implicitly castable to a shared element type. * **expression\_n**: Additional expressions to append to the array. #### Example[​](#example-72 "Direct link to Example") ``` > select array(1, 2, 3); +-----------------------------------+ | array(Int64(1),Int64(2),Int64(3)) | +-----------------------------------+ | [1, 2, 3] | +-----------------------------------+ ``` Reference: [Spark SQL `array`](https://spark.apache.org/docs/latest/api/sql/index.html#array). ### `array_any_value`[​](#array_any_value "Direct link to array_any_value") Returns the first non-null element in the array. If all elements are null, returns null. ``` array_any_value(array) ``` #### Arguments[​](#arguments-108 "Direct link to Arguments") * **array**: Array expression. Can be a constant, column, or function, and any combination of array operators. #### Example[​](#example-73 "Direct link to Example") ``` > select array_any_value([NULL, 1, 2, 3]); +-------------------------------+ | array_any_value(List([NULL,1,2,3])) | +-------------------------------------+ | 1 | +-------------------------------------+ ``` #### Aliases[​](#aliases-5 "Direct link to Aliases") * list\_any\_value ### `array_append`[​](#array_append "Direct link to array_append") Appends an element to the end of an array and returns the new array. ``` array_append(array, element) ``` #### Arguments[​](#arguments-109 "Direct link to Arguments") * **array**: Array expression. Can be a constant, column, or function, and any combination of array operators. * **element**: Element to append to the array. #### Example[​](#example-74 "Direct link to Example") ``` > select array_append([1, 2, 3], 4); +--------------------------------------+ | array_append(List([1,2,3]),Int64(4)) | +--------------------------------------+ | [1, 2, 3, 4] | +--------------------------------------+ ``` #### Aliases[​](#aliases-6 "Direct link to Aliases") * list\_append * array\_push\_back * list\_push\_back ### `array_cat`[​](#array_cat "Direct link to array_cat") Alias of [`array_concat`](#array_concat). ### `array_concat`[​](#array_concat "Direct link to array_concat") Concatenates two or more arrays into a single array. ``` array_concat(array[, ..., array_n]) ``` #### Arguments[​](#arguments-110 "Direct link to Arguments") * **array**: Array expression. Can be a constant, column, or function, and any combination of array operators. * **array\_n**: Additional array expressions to concatenate. #### Example[​](#example-75 "Direct link to Example") ``` > select array_concat([1, 2], [3, 4], [5, 6]); +---------------------------------------------------+ | array_concat(List([1,2]),List([3,4]),List([5,6])) | +---------------------------------------------------+ | [1, 2, 3, 4, 5, 6] | +---------------------------------------------------+ ``` #### Aliases[​](#aliases-7 "Direct link to Aliases") * array\_cat * list\_concat * list\_cat ### `array_contains`[​](#array_contains "Direct link to array_contains") Returns true if the array contains the specified element. ``` array_contains(array, element) ``` #### Arguments[​](#arguments-111 "Direct link to Arguments") * **array**: Array expression. Can be a constant, column, or function, and any combination of array operators. * **element**: Element to search for in the array. #### Example[​](#example-76 "Direct link to Example") ``` > select array_contains([1, 2, 3], 2); +----------------------------------------+ | array_contains(List([1,2,3]),Int64(2)) | +----------------------------------------+ | true | +----------------------------------------+ ``` **Note**: For array-to-array containment operations, use the [`@>` operator](/docs/next/reference/sql/operators#op_arr_contains). ### `array_dims`[​](#array_dims "Direct link to array_dims") Returns an array of the array's dimensions. For a 2D array, returns the number of rows and columns. ``` array_dims(array) ``` #### Arguments[​](#arguments-112 "Direct link to Arguments") * **array**: Array expression. Can be a constant, column, or function, and any combination of array operators. #### Example[​](#example-77 "Direct link to Example") ``` > select array_dims([[1, 2, 3], [4, 5, 6]]); +---------------------------------+ | array_dims(List([1,2,3,4,5,6])) | +---------------------------------+ | [2, 3] | +---------------------------------+ ``` #### Aliases[​](#aliases-8 "Direct link to Aliases") * list\_dims ### `array_distance`[​](#array_distance "Direct link to array_distance") Returns the Euclidean distance between two input arrays of equal length. ``` array_distance(array1, array2) ``` #### Arguments[​](#arguments-113 "Direct link to Arguments") * **array1**: Array expression. Can be a constant, column, or function, and any combination of array operators. * **array2**: Array expression. Can be a constant, column, or function, and any combination of array operators. #### Example[​](#example-78 "Direct link to Example") ``` > select array_distance([1, 2], [1, 4]); +------------------------------------+ | array_distance(List([1,2], [1,4])) | +------------------------------------+ | 2.0 | +------------------------------------+ ``` #### Aliases[​](#aliases-9 "Direct link to Aliases") * list\_distance ### `array_distinct`[​](#array_distinct "Direct link to array_distinct") Returns a new array with duplicate elements removed, preserving the order of first occurrence. ``` array_distinct(array) ``` #### Arguments[​](#arguments-114 "Direct link to Arguments") * **array**: Array expression. Can be a constant, column, or function, and any combination of array operators. #### Example[​](#example-79 "Direct link to Example") ``` > select array_distinct([1, 3, 2, 3, 1, 2, 4]); +---------------------------------+ | array_distinct(List([1,2,3,4])) | +---------------------------------+ | [1, 3, 2, 4] | +---------------------------------+ ``` #### Aliases[​](#aliases-10 "Direct link to Aliases") * list\_distinct ### `array_element`[​](#array_element "Direct link to array_element") Extracts the element at the specified index from the array. Indexing is 1-based. ``` array_element(array, index) ``` #### Arguments[​](#arguments-115 "Direct link to Arguments") * **array**: Array expression. Can be a constant, column, or function, and any combination of array operators. * **index**: Index to extract the element from the array (1-based). #### Example[​](#example-80 "Direct link to Example") ``` > select array_element([1, 2, 3, 4], 3); +-----------------------------------------+ | array_element(List([1,2,3,4]),Int64(3)) | +-----------------------------------------+ | 3 | +-----------------------------------------+ ``` #### Aliases[​](#aliases-11 "Direct link to Aliases") * array\_extract * list\_element * list\_extract ### `array_except`[​](#array_except "Direct link to array_except") Returns an array containing elements in `array1` that are not in `array2`, preserving first-occurrence order and without duplicates. ``` array_except(array1, array2) ``` Alias: `list_except`. ### `array_has`[​](#array_has "Direct link to array_has") Returns `true` if the array contains the specified element. ``` array_has(array, element) ``` Aliases: `array_contains`, `list_has`. ### `array_has_all`[​](#array_has_all "Direct link to array_has_all") Returns `true` if every element of `sub_array` is present in `array`. ``` array_has_all(array, sub_array) ``` Alias: `list_has_all`. ### `array_has_any`[​](#array_has_any "Direct link to array_has_any") Returns `true` if `array` and `sub_array` share at least one element. ``` array_has_any(array, sub_array) ``` Aliases: `list_has_any`, `arrays_overlap`. ### `array_intersect`[​](#array_intersect "Direct link to array_intersect") Returns an array of elements present in both input arrays, deduplicated. ``` array_intersect(array1, array2) ``` Alias: `list_intersect`. ### `array_length`[​](#array_length "Direct link to array_length") Returns the length of the array at the given (optional) dimension. Dimension defaults to 1. ``` array_length(array[, dimension]) ``` Alias: `list_length`. ### `array_max`[​](#array_max "Direct link to array_max") Returns the maximum element of the array, ignoring `NULL`s. ``` array_max(array) ``` Alias: `list_max`. ### `array_min`[​](#array_min "Direct link to array_min") Returns the minimum element of the array, ignoring `NULL`s. ``` array_min(array) ``` Alias: `list_min`. ### `array_ndims`[​](#array_ndims "Direct link to array_ndims") Returns the number of dimensions of the array. ``` array_ndims(array) ``` Alias: `list_ndims`. ### `array_pop_back`[​](#array_pop_back "Direct link to array_pop_back") Returns the array with the last element removed. ``` array_pop_back(array) ``` Alias: `list_pop_back`. ### `array_pop_front`[​](#array_pop_front "Direct link to array_pop_front") Returns the array with the first element removed. ``` array_pop_front(array) ``` Alias: `list_pop_front`. ### `array_position`[​](#array_position "Direct link to array_position") Returns the 1-based position of the first occurrence of `element` in `array`, or `NULL` if not found. An optional `from_index` starts the search at a later position. ``` array_position(array, element[, from_index]) ``` Aliases: `list_position`, `array_indexof`, `list_indexof`. ### `array_positions`[​](#array_positions "Direct link to array_positions") Returns a 1-based array of all positions where `element` occurs in `array`. ``` array_positions(array, element) ``` Alias: `list_positions`. ### `array_prepend`[​](#array_prepend "Direct link to array_prepend") Prepends an element to the beginning of an array. ``` array_prepend(element, array) ``` Aliases: `list_prepend`, `array_push_front`, `list_push_front`. ### `array_remove`[​](#array_remove "Direct link to array_remove") Returns the array with the first occurrence of `element` removed. ``` array_remove(array, element) ``` Alias: `list_remove`. ### `array_remove_n`[​](#array_remove_n "Direct link to array_remove_n") Returns the array with the first `max` occurrences of `element` removed. ``` array_remove_n(array, element, max) ``` Alias: `list_remove_n`. ### `array_remove_all`[​](#array_remove_all "Direct link to array_remove_all") Returns the array with all occurrences of `element` removed. ``` array_remove_all(array, element) ``` Alias: `list_remove_all`. ### `array_repeat`[​](#array_repeat "Direct link to array_repeat") Returns an array containing `element` repeated `count` times. ``` array_repeat(element, count) ``` Alias: `list_repeat`. ### `array_replace`[​](#array_replace "Direct link to array_replace") Replaces the first occurrence of `from` with `to` in `array`. ``` array_replace(array, from, to) ``` Alias: `list_replace`. ### `array_replace_n`[​](#array_replace_n "Direct link to array_replace_n") Replaces the first `max` occurrences of `from` with `to` in `array`. ``` array_replace_n(array, from, to, max) ``` Alias: `list_replace_n`. ### `array_replace_all`[​](#array_replace_all "Direct link to array_replace_all") Replaces every occurrence of `from` with `to` in `array`. ``` array_replace_all(array, from, to) ``` Alias: `list_replace_all`. ### `array_resize`[​](#array_resize "Direct link to array_resize") Resizes `array` to the given length, padding with `value` (or `NULL` if omitted) when growing. ``` array_resize(array, size[, value]) ``` Alias: `list_resize`. ### `array_reverse`[​](#array_reverse "Direct link to array_reverse") Returns the array with elements in reverse order. ``` array_reverse(array) ``` Alias: `list_reverse`. ### `array_slice`[​](#array_slice "Direct link to array_slice") Returns a slice of the array from `begin` to `end` (1-based, inclusive). Negative indices count from the end. ``` array_slice(array, begin, end[, stride]) ``` Alias: `list_slice`. ### `array_sort`[​](#array_sort "Direct link to array_sort") Returns `array` sorted in ascending order (default). Optional arguments control sort direction (`ASC`/`DESC`) and null placement (`NULLS FIRST`/`NULLS LAST`). ``` array_sort(array[, desc[, nulls_first]]) ``` Alias: `list_sort`. ### `array_to_string`[​](#array_to_string "Direct link to array_to_string") Concatenates array elements into a single string using the given delimiter. Optional `null_string` replaces `NULL` elements. ``` array_to_string(array, delimiter[, null_string]) ``` Aliases: `list_to_string`, `array_join`, `list_join`. ### `array_union`[​](#array_union "Direct link to array_union") Returns the set-union of two arrays, deduplicated. ``` array_union(array1, array2) ``` Alias: `list_union`. ### `arrays_zip`[​](#arrays_zip "Direct link to arrays_zip") Merges the given arrays element-wise into an array of structs. Shorter arrays are padded with `NULL`s. ``` arrays_zip(array1[, array2, ...]) ``` Alias: `list_zip`. ### `cardinality`[​](#cardinality "Direct link to cardinality") Returns the total number of elements in an array (including nested elements) or the number of entries in a map. ``` cardinality(array_or_map) ``` ### `empty`[​](#empty "Direct link to empty") Returns `true` if the array has length 0 (or is `NULL`). ``` empty(array) ``` Aliases: `array_empty`, `list_empty`. ### `flatten`[​](#flatten "Direct link to flatten") Flattens a nested array into a single-level array. ``` flatten(array) ``` ### `make_array`[​](#make_array "Direct link to make_array") Constructs an array (Arrow list) from the given expressions. SQL `[expr1, expr2, ...]` literal syntax compiles to this function. ``` make_array(expression1[, ..., expression_n]) ``` Alias: `make_list`. ### `range`[​](#range "Direct link to range") Generates a numeric or date range as an array, half-open on the upper bound. When the step is omitted, the default is `1`. ``` range(start, stop[, step]) range(stop) ``` For dates, `step` is an interval literal, e.g. `interval '1 day'`. Use [`generate_series`](#generate_series) for the inclusive-upper-bound variant. ### `generate_series`[​](#generate_series "Direct link to generate_series") Like [`range`](#range), but the upper bound is inclusive. ``` generate_series(start, stop[, step]) ``` ### `string_to_array`[​](#string_to_array "Direct link to string_to_array") Splits a string into an array of substrings using the given delimiter. An optional `null_string` turns matching substrings into `NULL`s. ``` string_to_array(str, delimiter[, null_string]) ``` Alias: `string_to_list`. *** ## Struct Functions[​](#struct-functions "Direct link to Struct Functions") Struct functions help construct and access structured data types (Arrow structs). These are useful for working with nested or composite data. * [struct](#struct) * [named\_struct](#named_struct) * [get\_field](#get_field) ### `struct`[​](#struct "Direct link to struct") Constructs an anonymous Arrow struct from the given values. Field names default to `c0`, `c1`, ... in the order provided. ``` struct(expression1[, ..., expression_n]) ``` #### Example[​](#example-81 "Direct link to Example") ``` > SELECT struct(1, 'spice', true); +-----------------------------------------------------+ | struct(Int64(1),Utf8("spice"),Boolean(true)) | +-----------------------------------------------------+ | {c0: 1, c1: spice, c2: true} | +-----------------------------------------------------+ ``` ### `named_struct`[​](#named_struct "Direct link to named_struct") Constructs an Arrow struct from alternating field-name / field-value pairs. ``` named_struct(name1, expression1[, name2, expression2, ...]) ``` #### Example[​](#example-82 "Direct link to Example") ``` > SELECT named_struct('id', 1, 'label', 'spice'); +--------------------------------------------------------+ | named_struct(Utf8("id"),Int64(1),Utf8("label"),Utf8("spice")) | +--------------------------------------------------------+ | {id: 1, label: spice} | +--------------------------------------------------------+ ``` ### `get_field`[​](#get_field "Direct link to get_field") Extracts a field by name from a struct or map. `struct.field` and `struct['field']` sugar invoke this function. ``` get_field(expression, field_name) ``` ## Map Functions[​](#map-functions "Direct link to Map Functions") Map functions help construct and query key-value data structures. These are useful for semi-structured or JSON-like data. * [map](#map) * [map\_keys](#map_keys) * [map\_values](#map_values) * [map\_entries](#map_entries) * [map\_extract](#map_extract) ### `map`[​](#map "Direct link to map") Constructs an Arrow map from alternating key/value arguments, or from two arrays (one of keys, one of values). ``` map(key1, value1[, key2, value2, ...]) map(keys_array, values_array) ``` #### Example[​](#example-83 "Direct link to Example") ``` > SELECT map('a', 1, 'b', 2); +-------------------------------------------------------+ | map(Utf8("a"),Int64(1),Utf8("b"),Int64(2)) | +-------------------------------------------------------+ | {a: 1, b: 2} | +-------------------------------------------------------+ ``` ### `map_keys`[​](#map_keys "Direct link to map_keys") Returns the keys of a map as an array. ``` map_keys(map) ``` ### `map_values`[​](#map_values "Direct link to map_values") Returns the values of a map as an array. ``` map_values(map) ``` ### `map_entries`[​](#map_entries "Direct link to map_entries") Returns the entries of a map as an array of structs `[{key, value}, ...]`. ``` map_entries(map) ``` ### `map_extract`[​](#map_extract "Direct link to map_extract") Looks up a key in a map and returns the associated value, or `NULL` if the key is absent. ``` map_extract(map, key) ``` Alias: `element_at`. ## Hashing Functions[​](#hashing-functions "Direct link to Hashing Functions") Hashing functions compute cryptographic hashes and checksums for data integrity, fingerprinting, and security applications. Binary digest output is returned as a `Binary` (bytes) array; use `encode(..., 'hex')` to render as hex. * [digest](#digest) * [md5](#md5) * [sha224](#sha224) * [sha256](#sha256) * [sha384](#sha384) * [sha512](#sha512) ### `digest`[​](#digest "Direct link to digest") Computes the digest of the input using the named hash algorithm. Supported algorithms: `'md5'`, `'sha224'`, `'sha256'`, `'sha384'`, `'sha512'`, `'blake2s'`, `'blake2b'`, `'blake3'`. ``` digest(expression, algorithm) ``` #### Example[​](#example-84 "Direct link to Example") ``` > SELECT encode(digest('spice.ai', 'sha256'), 'hex'); ``` ### `md5`[​](#md5 "Direct link to md5") Computes the MD5 128-bit hash of a string and returns the result as a lowercase hex string. ``` md5(expression) ``` ### `sha224`[​](#sha224 "Direct link to sha224") Computes the SHA-224 hash and returns a binary digest. ``` sha224(expression) ``` ### `sha256`[​](#sha256 "Direct link to sha256") Computes the SHA-256 hash and returns a binary digest. ``` sha256(expression) ``` ### `sha384`[​](#sha384 "Direct link to sha384") Computes the SHA-384 hash and returns a binary digest. ``` sha384(expression) ``` ### `sha512`[​](#sha512 "Direct link to sha512") Computes the SHA-512 hash and returns a binary digest. ``` sha512(expression) ``` ## Encoding Functions[​](#encoding-functions "Direct link to Encoding Functions") Binary encoding utilities for converting between binary data and text representations. * [encode](#encode) * [decode](#decode) ### `encode`[​](#encode "Direct link to encode") Encodes a string or binary value using the specified encoding. Supported encodings: `'hex'`, `'base64'`. ``` encode(expression, encoding) ``` #### Example[​](#example-85 "Direct link to Example") ``` > SELECT encode('spice', 'base64'); +--------------------------------------+ | encode(Utf8("spice"),Utf8("base64")) | +--------------------------------------+ | c3BpY2U= | +--------------------------------------+ ``` ### `decode`[​](#decode "Direct link to decode") Decodes text back to binary using the specified encoding. Supported encodings: `'hex'`, `'base64'`. ``` decode(expression, encoding) ``` ## Union Functions[​](#union-functions "Direct link to Union Functions") Union functions help work with union (variant) data types. * [union\_extract](#union_extract) * [union\_tag](#union_tag) ### `union_extract`[​](#union_extract "Direct link to union_extract") Extracts the value of a named member from a union, returning `NULL` if the union's active member doesn't match. ``` union_extract(expression, field_name) ``` ### `union_tag`[​](#union_tag "Direct link to union_tag") Returns the name of the active member of a union value as a string. ``` union_tag(expression) ``` ## Metadata Functions[​](#metadata-functions "Direct link to Metadata Functions") PostgreSQL-compatible functions for reading table and column comments from registered datasets. Comments originate from the dataset's source (for connectors that surface `COMMENT ON TABLE` / `COMMENT ON COLUMN` metadata, such as PostgreSQL, MySQL, Snowflake, and Databricks) or from `description` metadata attached to the dataset schema. * [obj\_description](#obj_description) * [col\_description](#col_description) ### `obj_description`[​](#obj_description "Direct link to obj_description") Returns the comment attached to a registered table, or `NULL` if the table has no comment. ``` obj_description(table_identifier) obj_description(table_identifier, catalog_name) obj_description(schema_name, table_name) obj_description(catalog_name, schema_name, table_name) ``` #### Arguments[​](#arguments-116 "Direct link to Arguments") * **table\_identifier**: Either a string holding a possibly-qualified table name (`'table'`, `'schema.table'`, or `'catalog.schema.table'`) or an integer table OID. Unqualified names are resolved against the session's default catalog and schema. * **catalog\_name**: When supplied as the second argument with `'pg_class'`, the call is treated as PostgreSQL-style `obj_description(oid, 'pg_class')`; any other value returns `NULL`. * **schema\_name**, **table\_name**: Explicit schema and table parts. The three-argument form additionally takes the catalog as the first argument. #### Example[​](#example-86 "Direct link to Example") ``` > SELECT obj_description('public.taxi_trips'); +----------------------------------------+ | obj_description(Utf8("public.taxi_trips")) | +----------------------------------------+ | NYC yellow-cab trip records | +----------------------------------------+ ``` ### `col_description`[​](#col_description "Direct link to col_description") Returns the comment attached to a column on a registered table, or `NULL` if no comment exists. ``` col_description(table_identifier, column) col_description(catalog_name, schema_name, table_name, column) ``` #### Arguments[​](#arguments-117 "Direct link to Arguments") * **table\_identifier**: Possibly-qualified table name (string) or table OID (integer), resolved the same way as in [`obj_description`](#obj_description). * **column**: Either a column name (string) or a 1-based ordinal position (integer). * **catalog\_name**, **schema\_name**, **table\_name**: Explicit catalog, schema, and table parts when the four-argument form is used. #### Example[​](#example-87 "Direct link to Example") ``` > SELECT col_description('public.taxi_trips', 'fare_amount'); +---------------------------------------------------------+ | col_description(Utf8("public.taxi_trips"),Utf8("fare_amount")) | +---------------------------------------------------------+ | Total fare in USD, excluding tip | +---------------------------------------------------------+ > SELECT col_description('public.taxi_trips', 3); +--------------------------------------------------+ | col_description(Utf8("public.taxi_trips"),Int64(3)) | +--------------------------------------------------+ | Pickup datetime in source-local time | +--------------------------------------------------+ ``` *** ## Other Functions[​](#other-functions "Direct link to Other Functions") Additional scalar functions include type casting, type inspection, and version reporting. * [arrow\_cast](#arrow_cast) * [arrow\_try\_cast](#arrow_try_cast) * [arrow\_typeof](#arrow_typeof) * [arrow\_metadata](#arrow_metadata) * [version](#version) * [ai](#ai-and-embed) * [embed](#ai-and-embed) * [bucket](#bucket) * [truncate](#truncate) ### `arrow_cast`[​](#arrow_cast "Direct link to arrow_cast") Casts an expression to a specific Arrow data type. Use this function when you need precise control over the target Arrow type, such as specifying timestamp precision. ``` arrow_cast(expression, arrow_type) ``` #### Arguments[​](#arguments-118 "Direct link to Arguments") * **expression**: The value to cast. * **arrow\_type**: A string specifying the target Arrow type (e.g., `'Int32'`, `'Utf8'`, `'Timestamp(Second, None)'`). #### Example[​](#example-88 "Direct link to Example") ``` > SELECT arrow_cast(now(), 'Timestamp(Second, None)') AS now_seconds; +---------------------+ | now_seconds | +---------------------+ | 2024-01-15T10:30:45 | +---------------------+ > SELECT arrow_cast('123', 'Int64') AS num; +-----+ | num | +-----+ | 123 | +-----+ ``` See [Data Types Reference](/docs/next/reference/datatypes) for supported Arrow types. ### `arrow_try_cast`[​](#arrow_try_cast "Direct link to arrow_try_cast") Like [`arrow_cast`](#arrow_cast) but returns `NULL` instead of erroring when the cast fails. ``` arrow_try_cast(expression, arrow_type) ``` ### `arrow_typeof`[​](#arrow_typeof "Direct link to arrow_typeof") Returns the Arrow data type of the given expression as a string. ``` arrow_typeof(expression) ``` #### Arguments[​](#arguments-119 "Direct link to Arguments") * **expression**: Any SQL expression. #### Example[​](#example-89 "Direct link to Example") ``` > SELECT arrow_typeof(1); +------------------------+ | arrow_typeof(Int64(1)) | +------------------------+ | Int64 | +------------------------+ > SELECT arrow_typeof(now()); +-------------------------------+ | arrow_typeof(now()) | +-------------------------------+ | Timestamp(Nanosecond, None) | +-------------------------------+ > SELECT arrow_typeof(interval '1 month'); +------------------------------+ | arrow_typeof(...) | +------------------------------+ | Interval(MonthDayNano) | +------------------------------+ ``` ### `arrow_metadata`[​](#arrow_metadata "Direct link to arrow_metadata") Returns the Arrow schema metadata associated with an expression as a map of key/value strings. Useful for inspecting field-level metadata (units, comments, logical type hints) attached during ingest. ``` arrow_metadata(expression) ``` ### `version`[​](#version "Direct link to version") Returns the underlying DataFusion runtime version string. ``` version() ``` ### `ai` and `embed`[​](#ai-and-embed "Direct link to ai-and-embed") See [AI Functions](/docs/next/reference/sql/ai) for `ai()` (LLM text generation) and `embed()` (vector embedding generation). ### `bucket`[​](#bucket "Direct link to bucket") Assigns a deterministic bucket identifier for a value by hashing the input and projecting it into a fixed number of buckets. Helpful for `partition_by` expressions and for co-locating related rows during acceleration refreshes. ``` bucket(num_buckets, value) ``` #### Arguments[​](#arguments-120 "Direct link to Arguments") * **num\_buckets**: Positive integer literal indicating how many buckets to distribute values across. Must be in the range `[1, 1_000_000]`. The literal's integer type (`Int8` … `Int64`, `UInt8` … `UInt64`) determines the return type. * **value**: Expression to hash. Accepts strings, numbers, and other scalar types supported by the query engine. #### Return Type[​](#return-type "Direct link to Return Type") Returns an integer in the range `[0, num_buckets - 1]`, matching the integer type of `num_buckets`. The same input value always maps to the same bucket for a given `num_buckets` (the hash uses a fixed seed, so buckets are stable across processes and runtime restarts). #### Example[​](#example-90 "Direct link to Example") ``` -- Partition account IDs into 100 stable buckets SELECT account_id, bucket(100, account_id) AS account_bucket FROM accounts; ``` In `spicepod.yaml`, use the function directly inside `partition_by` to build file-based accelerations: ``` datasets: - name: my_table acceleration: enabled: true engine: duckdb mode: file partition_by: - bucket(100, account_id) ``` ### `truncate`[​](#truncate "Direct link to truncate") Iceberg-style truncate transform. For numeric values, rounds down to the nearest multiple of `width`. For strings and binary, returns the first `width` characters/bytes. Useful for partitioning by wide numeric ranges or string prefixes. ``` truncate(width, value) ``` #### Arguments[​](#arguments-121 "Direct link to Arguments") * **width**: Positive `Int64` literal that defines the bucket size or, for strings/binary, the number of leading units to retain. Maximum: `i64::MAX / 2`. * **value**: Expression to truncate. Accepts: * Signed integers: `Int8`, `Int16`, `Int32`, `Int64` * Unsigned integers: `UInt8`, `UInt16`, `UInt32`, `UInt64` * Decimals: `Decimal128`, `Decimal256` * Strings: `Utf8` * Binary: `Binary` #### Return Type[​](#return-type-1 "Direct link to Return Type") Returns the same type as `value`. For numbers, the result is the largest multiple of `width` that is less than or equal to `value`. For strings/binary, the result is the first `width` characters/bytes. #### Example[​](#example-91 "Direct link to Example") ``` -- Numeric: floor-bucket integers into ranges of 10 SELECT truncate(10, 101) AS truncated_id; -- returns 100 -- Truncate event timestamps to the start of each hour (3600 seconds) SELECT truncate(3600, extract(epoch FROM event_time)) AS hour_start FROM events; -- String: keep the first 2 characters (e.g., country prefix) SELECT truncate(2, 'United Kingdom'); -- returns 'Un' ``` *** Spice.ai aims for compatibility with PostgreSQL, but some functions or behaviors may differ depending on the underlying engine version. --- # Search in SQL This section documents search capabilities in Spice SQL, including vector search, full-text search, and lexical filtering methods. These features help retrieve relevant data using semantic similarity, keyword matching, and pattern-based filtering. ## Table of Contents[​](#table-of-contents "Direct link to Table of Contents") * [Table of Contents](#table-of-contents) * [Vector Search (`vector_search`)](#vector-search-vector_search) * [Usage](#usage) * [Example](#example) * [Multi-Query (Late-Interaction) Form](#multi-query-late-interaction-form) * [Full-Text Search (`text_search`)](#full-text-search-text_search) * [Usage](#usage-1) * [Example](#example-1) * [Reciprocal Rank Fusion (`rrf`)](#reciprocal-rank-fusion-rrf) * [Usage](#usage-2) * [Examples](#examples) * [Reranking (`rerank`)](#reranking-rerank) * [Lexical Search: LIKE, =, and Regex](#lexical-search-like--and-regex) * [LIKE (Pattern Matching)](#like-pattern-matching) * [= (Keyword/Exact Match)](#-keywordexact-match) * [Regex Filtering](#regex-filtering) * [Example](#example-2) *** ## Vector Search (`vector_search`)[​](#vector-search-vector_search "Direct link to vector-search-vector_search") Vector search retrieves records by semantic similarity using embeddings. It is ideal for finding related content even when exact keywords differ. ### Usage[​](#usage "Direct link to Usage") ``` SELECT id, score FROM vector_search(table, 'search query') ORDER BY score DESC LIMIT 5; ``` * `table`: Dataset name (required) * `query`: Search text, or an array of strings for [multi-query](#multi-query-late-interaction-form) search (required) * `column`: Column name (optional if only one embedding column; required when the table has multiple embedded columns) * `limit`: Maximum results (optional). When omitted, the engine-defined maximum is used. * `include_score`: Include relevance scores (optional, default `TRUE`) * `distance_metric`: Similarity metric used to rank candidate vectors (optional, named argument). Supported values: `'cosine'` (default) and `'l2'` (negated Euclidean distance). `'dot'` is parsed but not yet wired through the scan path. * `rank_weight`: Per-query ranking weight (optional, named argument). Only meaningful when `vector_search` is passed as a subquery to [`rrf`](#reciprocal-rank-fusion-rrf). #### Filter Pushdown[​](#filter-pushdown "Direct link to Filter Pushdown") `WHERE` predicates on base table columns (e.g., `created_at`, `product_category`) are pushed down as **pre-filters** — they are applied before the similarity ranking, so only matching rows are scored and returned. This means results reflect the top-K *within the filtered set*, not the top-K of the entire table filtered afterward. Predicates on computed columns like `score` are applied as post-filters after ranking. #### Example[​](#example "Direct link to Example") ``` -- Filters on created_at are pushed down before ranking SELECT review_id, rating, customer_id, body, score FROM vector_search(reviews, 'issues with same day shipping', 1500) WHERE created_at >= to_unixtime(now() - INTERVAL '7 days') ORDER BY score DESC LIMIT 2; ``` To override the similarity metric, pass `distance_metric` as a named argument: ``` SELECT id, body, score FROM vector_search(reviews, 'issues with shipping', distance_metric => 'l2') ORDER BY score DESC LIMIT 10; ``` See [Vector-Based Search](/docs/next/features/search/vector-search) for configuration and advanced usage. ### Multi-Query (Late-Interaction) Form[​](#multi-query-late-interaction-form "Direct link to Multi-Query (Late-Interaction) Form") When the target column is a [multi-vector column](/docs/next/features/search/multi-vector), `vector_search` also accepts an array of query strings. Each query is embedded independently and the per-row score is `Σ_q max_e cos(q, e)` — ColBERT-style late interaction. Passing an array to a scalar or chunked column returns an error. At most 32 query strings are accepted per call. ``` SELECT product_id, name, score FROM vector_search(products, ['hiking', 'waterproof', 'lightweight'], tags) ORDER BY score DESC LIMIT 10; ``` *** ## Full-Text Search (`text_search`)[​](#full-text-search-text_search "Direct link to full-text-search-text_search") Full-text search uses BM25 scoring to retrieve records matching keywords in indexed columns. ### Usage[​](#usage-1 "Direct link to Usage") ``` SELECT id, score FROM text_search(table, 'search terms', col) ORDER BY score DESC LIMIT 5; ``` * `table`: Dataset name (required) * `query`: Keyword or phrase (required) * `column`: Column to search (optional if the table has a single full-text index; required when multiple columns are indexed) * `limit`: Maximum results (optional). Defaults to 1000, which is the maximum supported. * `include_score`: Include relevance scores (optional, default `TRUE`) * `rank_weight`: Per-query ranking weight (optional, named argument). Only meaningful when `text_search` is passed as a subquery to [`rrf`](#reciprocal-rank-fusion-rrf). By default, `text_search` retrieves up to 1000 results. To request fewer, specify a smaller `limit`. #### Filter Pushdown[​](#full-text-filter-pushdown "Direct link to Filter Pushdown") With the built-in Tantivy engine, `WHERE` predicates on columns carried in the full-text index are pushed into the index scan as **pre-filters** — they are applied before the top-K limit, so the results are the top-K *within the filtered set* rather than the filtered remainder of an unfiltered top-K. A predicate the index cannot apply is left to Spice's query engine above the scan, and one the index applies only approximately is re-checked there, so results are the same either way; only how many rows survive the limit changes. ``` -- state and additions are applied inside the index, before the limit of 5 SELECT id, title, score FROM text_search(doc.pulls, 'search keywords', body, 5) WHERE state = 'open' AND additions > 100; ``` Filterable columns are the dataset's primary key (or [`full_text_search.row_id`](/docs/next/reference/spicepod/datasets#columnsfull_text_searchrow_id)) and any column declared with [`metadata.vectors`](/docs/next/reference/spicepod/datasets#columnsmetadatavectors), whose values are carried into the index alongside the searched text. Columns of a type the index cannot represent — dates and timestamps — are skipped, with a warning logged at startup naming the column. Predicates that push down: `=`, `!=`, `<`, `<=`, `>`, `>=`, `BETWEEN`, `IN`, a prefix `LIKE 'x%'` on a string column, and `AND` / `OR` / `NOT` combinations of them. Predicates that do not, and are applied above the scan instead: anything on the searched text column itself (it is tokenized), on a floating-point column, or on a binary column; a case-insensitive or negated `LIKE`; an `IN` list containing `NULL`; and any comparison whose operands are not a column and a literal. The Tantivy warm tier used with [`engine: elasticsearch`](/docs/next/features/search/full-text#warm-tier) is built without those extra columns — its schema is the primary key and `_score` alone — so only primary-key predicates push into it. #### Example[​](#example-1 "Direct link to Example") ``` SELECT id, title, score FROM text_search(doc.pulls, 'search keywords', body) ORDER BY score DESC LIMIT 5; ``` See [Full-Text Search](/docs/next/features/search/full-text) for configuration and details. *** ## Reciprocal Rank Fusion (`rrf`)[​](#reciprocal-rank-fusion-rrf "Direct link to reciprocal-rank-fusion-rrf") Reciprocal Rank Fusion (RRF) combines results from multiple search queries to improve relevance by merging rankings from different search methods. Advanced features include per-query ranking weights, recency boosting, and flexible decay functions. ### Usage[​](#usage-2 "Direct link to Usage") `rrf` is variadic and takes two or more search UDTF calls as arguments. Named parameters provide advanced control over ranking, recency, and fusion behavior. info The `rrf` function automatically adds a `fused_score` column to the result set, which contains the combined relevance score from all input search queries. Results are sorted by `fused_score DESC` by default when no explicit `ORDER BY` clause is specified. ``` SELECT id, content, fused_score FROM rrf( vector_search(table, 'search query', rank_weight => 20), text_search(table, 'search terms', column), join_key => 'id', -- explicit join key for performance k => 60.0 -- smoothing parameter ) ORDER BY fused_score DESC LIMIT 10; ``` **Arguments:** Note that `rank_weight` is specified as the last argument to either a `text_search` or `vector_search` UDTF call (as shown above). All other arguments can be specified in any order after the search calls (within an `rrf` invocation). | Parameter | Type | Required | Description | | ------------------- | ---------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | | `query_1` | Search UDTF call | Yes | First search query (e.g., `vector_search`, `text_search`) | | `query_2` | Search UDTF call | Yes | Second search query. `rrf` requires at least two subqueries. | | `...` | Search UDTF call | No | Additional search queries (variadic) | | `join_key` | String | No | Column name to use for joining subquery results. If omitted, the primary key is inferred from the underlying tables; otherwise rows are auto-hashed. | | `k` | Float | No | Smoothing parameter for RRF scoring (default: 60.0) | | `limit` | Integer | No | Upper bound on the fused result set. Also propagated as a default limit to any nested search subquery that does not specify its own. | | `time_column` | String | No | Column name containing timestamps for recency boosting | | `recency_decay` | String | No | Decay function: 'linear' or 'exponential' (default: 'exponential') | | `decay_constant` | Float | No | Decay rate for exponential decay (default: 0.01) | | `decay_scale_secs` | Float | No | Time scale in seconds for decay (default: 86400) | | `decay_window_secs` | Float | No | Window size for linear decay in seconds (default: 86400) | | `rank_weight` | Float | No | Per-query ranking weight (**specified within the individual search subquery call**) | #### Filter Pushdown[​](#filter-pushdown-1 "Direct link to Filter Pushdown") `WHERE` predicates on base table columns (e.g., `review_date`, `product_category`) are pushed down into each nested search subquery as **pre-filters** — they are applied before ranking and fusion, so each subquery only considers matching rows. This means the fused results reflect the top-K *within the filtered set*, not a post-filtered slice of unfiltered rankings. Predicates on computed columns like `fused_score` are applied as post-filters after fusion. ``` -- review_date and product_category are pushed into each vector_search before ranking SELECT review_id, review_headline FROM rrf( vector_search(amazon_reviews, 'cannot exit the app', rank_weight => 20), vector_search(amazon_reviews, 'app not working', rank_weight => 10), join_key => 'review_id', k => 60.0 ) WHERE review_date > '2015-06-15' AND product_category = 'Mobile_Apps' LIMIT 10; ``` #### Examples[​](#examples "Direct link to Examples") **Basic Hybrid Search:** ``` -- Combine vector and text search for enhanced relevance SELECT id, title, content, fused_score FROM rrf( vector_search(documents, 'machine learning algorithms'), text_search(documents, 'neural networks deep learning', content), join_key => 'id' -- explicit join key for performance ) WHERE fused_score > 0.01 ORDER BY fused_score DESC LIMIT 5; ``` **Weighted Ranking:** ``` -- Boost semantic search over exact text matching SELECT fused_score, title, content FROM rrf( text_search(posts, 'artificial intelligence', rank_weight => 50.0), vector_search(posts, 'AI machine learning', rank_weight => 200.0) ) ORDER BY fused_score DESC LIMIT 10; ``` **Recency-Boosted Search:** ``` -- Exponential decay favoring recent content SELECT fused_score, title, created_at FROM rrf( text_search(news, 'breaking news'), vector_search(news, 'latest updates'), time_column => 'created_at', recency_decay => 'exponential', decay_constant => 0.05, decay_scale_secs => 3600 -- 1 hour scale ) ORDER BY fused_score DESC LIMIT 10; ``` **Linear Decay:** ``` -- Linear decay over 24 hours SELECT fused_score, content FROM rrf( text_search(posts, 'trending'), vector_search(posts, 'viral popular'), time_column => 'created_at', recency_decay => 'linear', decay_window_secs => 86400 ) ORDER BY fused_score DESC; ``` **How RRF works:** * Each input query is ranked independently by score * Rankings are combined using the formula: `RRF Score = Σ(rank_weight / (k + rank))` * Documents appearing in multiple result sets receive higher scores * The `k` parameter controls ranking sensitivity (lower = more sensitive to rank position) **Advanced query tuning**: * **Rank weighting**: Individual queries can be weighted using `rank_weight` parameter * **Recency boosting**: When `time_column` is specified, scores are multiplied by a decay factor * **Exponential decay**: `e^(-decay_constant * age_in_units)` where age is in `decay_scale_secs` * **Linear decay**: `max(0, 1 - (age_in_units / decay_window_secs))` * **Auto-join**: When no `join_key` is specified, `rrf` infers the primary key from the underlying tables; if none is available, rows are joined by an auto-generated row identifier *** ## Reranking (`rerank`)[​](#reranking-rerank "Direct link to reranking-rerank") Reranking reorders candidate results using a dedicated reranker model or an LLM-as-reranker for improved relevance. The input can be any search UDTF (`vector_search`, `text_search`, `rrf`) or a plain table. ### Usage[​](#usage-3 "Direct link to Usage") ``` SELECT * FROM rerank( , document => 'column_name', model => 'reranker_name', limit => 10 ) ``` **Arguments:** | Parameter | Type | Required | Description | | ----------------- | ------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | | `input` | Table or UDTF | Yes | Input rows to rerank. Can be a search UDTF call (`vector_search`, `text_search`, `rrf`) or a table name. | | `model` | String | Yes | Name of a registered reranker or chat model. | | `document` | String | Yes | Column containing the text to send to the reranker for scoring. | | `query` | String | No | Query string for relevance scoring. Auto-extracted from nested search UDTFs when omitted; required for bare-table inputs. | | `limit` | Integer | No | Maximum number of results to return. | | `strategy` | String | No | LLM reranking strategy: `'listwise'` (default) or `'pointwise'`. Only applies when the model resolves to a chat model. | | `prompt_template` | String | No | Custom prompt template for LLM-as-reranker. Use `{query}` and `{document}` placeholders. Only applies when the model resolves to a chat model. | #### Query Auto-Propagation[​](#query-auto-propagation "Direct link to Query Auto-Propagation") When the input is a search UDTF (`vector_search`, `text_search`, or `rrf` wrapping search UDTFs), the query string is automatically extracted from the nested call. Single-string, `make_array(...)`, and `ARRAY[...]` query forms are all supported. For multi-query inputs, the first query string is used. For bare-table inputs, `query` must be provided explicitly. #### Examples[​](#examples-1 "Direct link to Examples") **Rerank hybrid search results:** ``` SELECT * FROM rerank( rrf( vector_search(docs, 'delta lake time travel', limit => 50), text_search(docs, 'delta lake time travel', limit => 50) ), document => 'content', model => 'cohere_rr', limit => 10 ); ``` **Rerank a plain table with an explicit query:** ``` SELECT * FROM rerank( tickets, query => 'auth failures', document => 'body', model => 'voyage_rr', limit => 5 ); ``` **LLM-as-reranker with custom prompt:** ``` SELECT * FROM rerank( vector_search(kb, 'onboarding checklist', limit => 40), document => 'content', model => 'gpt_mini', strategy => 'pointwise', prompt_template => 'Rate 0-1: is this useful for a new hire?\nQuery: {query}\nDoc: {document}', limit => 10 ); ``` See [Reranking](/docs/next/features/search/rerank) for configuration, provider setup, and additional examples. *** ## Lexical Search: LIKE, =, and Regex[​](#lexical-search-like--and-regex "Direct link to Lexical Search: LIKE, =, and Regex") Spice SQL supports traditional filtering for exact and pattern-based matches: ### LIKE (Pattern Matching)[​](#like-pattern-matching "Direct link to LIKE (Pattern Matching)") ``` SELECT * FROM my_table WHERE column LIKE '%substring%'; ``` * `%` matches any sequence of characters. * `_` matches a single character. ### = (Keyword/Exact Match)[​](#-keywordexact-match "Direct link to = (Keyword/Exact Match)") ``` SELECT * FROM my_table WHERE column = 'exact value'; ``` Returns rows where the column exactly matches the value. ### Regex Filtering[​](#regex-filtering "Direct link to Regex Filtering") Spice SQL supports the PostgreSQL regex operators `~` (match), `~*` (case-insensitive match), `!~` (not match), and `!~*` (case-insensitive not match) — see [Operators](/docs/next/reference/sql/operators#op_re_match). Alternatively, use scalar functions such as `regexp_like`, `regexp_match`, and `regexp_replace`. For details and examples, see the [Scalar Functions documentation](/docs/next/reference/sql/scalar_functions#regular-expression-functions). #### Example[​](#example-2 "Direct link to Example") ``` SELECT * FROM my_table WHERE column ~ '^spice.*ai$'; -- Or, equivalently: SELECT * FROM my_table WHERE regexp_like(column, '^spice.*ai$'); ``` *** For more on hybrid and advanced search, see [Search Functionality](/docs/next/features/search) and [Vector-Based Search](/docs/next/features/search/vector-search) --- # SELECT info Spice is built on [Apache DataFusion](https://datafusion.apache.org/) and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. ## SELECT syntax[​](#select-syntax "Direct link to SELECT syntax") The queries in Spice scan data from tables and return 0 or more rows. Spice follows PostgreSQL conventions for identifier handling: unquoted identifiers (table and column names) are normalized to lowercase. To reference a table or column with uppercase or mixed-case characters, wrap the identifier in double quotes. ``` -- These are equivalent (both reference the lowercase table name) SELECT * FROM lineitem; SELECT * FROM LINEITEM; -- Double quotes preserve the exact casing SELECT * FROM "LINEITEM"; ``` See [dataset `name` configuration](/docs/next/reference/spicepod/datasets#name) for how to set a case-sensitive dataset name in the Spicepod manifest. Spice supports the following syntax for queries: \[ [WITH](#with-clause) with\_query \[, ...] ]
[SELECT](#select-clause) \[ ALL | DISTINCT ] select\_expr \[, ...]
\[ [FROM](#from-clause) from\_item \[, ...] ]
\[ [JOIN](#join-clause) join\_item \[, ...] ]
\[ [WHERE](#where-clause) condition ]
\[ [GROUP BY](#group-by-clause) grouping\_element \[, ...] ]
\[ [HAVING](#having-clause) condition]
\[ [QUALIFY](#qualify-clause) condition ]
\[ [UNION](#union-clause) \[ ALL | select ] ] \[ [ORDER BY](#order-by-clause) expression \[ ASC | DESC ]\[, ...] ]
\[ [LIMIT](#limit-clause) count ]
\[ [EXCLUDE | EXCEPT](#exclude-except-replace-and-ilike-clauses) ] ### Window Functions (OVER Clause)[​](#window-functions-over-clause "Direct link to Window Functions (OVER Clause)") Window functions perform calculations across a set of rows related to the current row. Use the `OVER` clause to define the window: ``` SELECT employee_id, salary, ROW_NUMBER() OVER (ORDER BY salary DESC) AS salary_rank, SUM(salary) OVER (PARTITION BY dept_id) AS dept_total FROM employees; ``` The `OVER` clause supports: * `PARTITION BY`: Divides rows into groups * `ORDER BY`: Defines row ordering within each partition * Frame specifications: `ROWS BETWEEN ... AND ...` ### WITH clause[​](#with-clause "Direct link to WITH clause") A WITH clause assigns names to subqueries so they can be referenced by name. ``` WITH x AS (SELECT a, MAX(b) AS b FROM t GROUP BY a) SELECT a, b FROM x; ``` ### SELECT clause[​](#select-clause "Direct link to SELECT clause") The `SELECT` clause is used to select data from a database by defining the colummns it returns. Each `select_expr` in the SELECT list can be an expression or wildcards. Example: ``` SELECT a, b, a + b FROM table; ``` The `DISTINCT` quantifier can be added to make the query return all distinct rows. By default `ALL` will be used, which returns all the rows. ``` SELECT DISTINCT person, age FROM employees; ``` ### FROM clause[​](#from-clause "Direct link to FROM clause") The `FROM` clause is used to specify which table to select data from. Example: ``` SELECT t.a FROM table AS t; ``` ### WHERE clause[​](#where-clause "Direct link to WHERE clause") The `WHERE` clause is used define the conditions to filter the query results. Example: ``` SELECT a FROM table WHERE a > 10; ``` ### JOIN clause[​](#join-clause "Direct link to JOIN clause") Spice supports `INNER JOIN`, `LEFT OUTER JOIN`, `RIGHT OUTER JOIN`, `FULL OUTER JOIN`, `NATURAL JOIN` and `CROSS JOIN`. The following examples are based on this table: ``` select * from x; +----------+----------+ | column_1 | column_2 | +----------+----------+ | 1 | 2 | +----------+----------+ ``` #### INNER JOIN[​](#inner-join "Direct link to INNER JOIN") The keywords `JOIN` or `INNER JOIN` define a join that only shows rows where there is a match in both tables. ``` select * from x inner join x y ON x.column_1 = y.column_1; +----------+----------+----------+----------+ | column_1 | column_2 | column_1 | column_2 | +----------+----------+----------+----------+ | 1 | 2 | 1 | 2 | +----------+----------+----------+----------+ ``` #### LEFT OUTER JOIN[​](#left-outer-join "Direct link to LEFT OUTER JOIN") The keywords `LEFT JOIN` or `LEFT OUTER JOIN` define a join that includes all rows from the left table even if there is not a match in the right table. When there is no match, null values are produced for the right side of the join. ``` select * from x left join x y ON x.column_1 = y.column_2; +----------+----------+----------+----------+ | column_1 | column_2 | column_1 | column_2 | +----------+----------+----------+----------+ | 1 | 2 | | | +----------+----------+----------+----------+ ``` #### RIGHT OUTER JOIN[​](#right-outer-join "Direct link to RIGHT OUTER JOIN") The keywords `RIGHT JOIN` or `RIGHT OUTER JOIN` define a join that includes all rows from the right table even if there is not a match in the left table. When there is no match, null values are produced for the left side of the join. ``` select * from x right join x y ON x.column_1 = y.column_2; +----------+----------+----------+----------+ | column_1 | column_2 | column_1 | column_2 | +----------+----------+----------+----------+ | | | 1 | 2 | +----------+----------+----------+----------+ ``` #### FULL OUTER JOIN[​](#full-outer-join "Direct link to FULL OUTER JOIN") The keywords `FULL JOIN` or `FULL OUTER JOIN` define a join that is effectively a union of a `LEFT OUTER JOIN` and `RIGHT OUTER JOIN`. It will show all rows from the left and right side of the join and will produce null values on either side of the join where there is not a match. ``` select * from x full outer join x y ON x.column_1 = y.column_2; +----------+----------+----------+----------+ | column_1 | column_2 | column_1 | column_2 | +----------+----------+----------+----------+ | 1 | 2 | | | | | | 1 | 2 | +----------+----------+----------+----------+ ``` #### NATURAL JOIN[​](#natural-join "Direct link to NATURAL JOIN") A natural join defines an inner join based on common column names found between the input tables. When no common column names are found, it behaves like a cross join. ``` select * from x natural join x y; +----------+----------+ | column_1 | column_2 | +----------+----------+ | 1 | 2 | +----------+----------+ ``` #### CROSS JOIN[​](#cross-join "Direct link to CROSS JOIN") A cross join produces a cartesian product that matches every row in the left side of the join with every row in the right side of the join. ``` select * from x cross join x y; +----------+----------+----------+----------+ | column_1 | column_2 | column_1 | column_2 | +----------+----------+----------+----------+ | 1 | 2 | 1 | 2 | +----------+----------+----------+----------+ ``` ### GROUP BY clause[​](#group-by-clause "Direct link to GROUP BY clause") The `GROUP BY` clause groups together input rows that have the same value into summary rows. `GROUP BY` is typically used with aggregrate functions (`COUNT()`, `MAX()`, `SUM()`), but if no aggregate functions are included, the query with a `GROUP BY` clause is the same as `SELECT DISTINCT`. Example: ``` SELECT a, b, MAX(c) FROM table GROUP BY a, b; ``` Some aggregation functions accept optional ordering requirement, such as `ARRAY_AGG`. If a requirement is given, aggregation is calculated in the order of the requirement. Example: ``` SELECT a, b, ARRAY_AGG(c, ORDER BY d) FROM table GROUP BY a, b; ``` #### `GROUP BY ALL`[​](#group-by-all "Direct link to group-by-all") Use GROUP BY ALL to group by every column in the SELECT list that isn’t inside an aggregate function. This keeps the column definitions in one place, simplifies the query, and prevents bugs by keeping the SELECT granularity aligned with the GROUP BY granularity (e.g., preventing unintended duplication). Example: ``` SELECT a, b, MAX(c) FROM table GROUP BY ALL; ``` ### HAVING clause[​](#having-clause "Direct link to HAVING clause") The `HAVING` clause can be used with `GROUP BY` to eliminate groups that don't satisfy the condition given. Example: ``` SELECT a, b, MAX(c) FROM table GROUP BY a, b HAVING MAX(c) > 10; ``` ### QUALIFY clause[​](#qualify-clause "Direct link to QUALIFY clause") The `QUALIFY` clause filters the results of window functions. It is evaluated after window functions are computed, similar to how `HAVING` filters results after `GROUP BY`. Example: ``` SELECT employee_id, dept_id, salary, ROW_NUMBER() OVER (PARTITION BY dept_id ORDER BY salary DESC) AS rank FROM employees QUALIFY rank <= 3; ``` This query returns only the top 3 highest-paid employees in each department. ### UNION clause[​](#union-clause "Direct link to UNION clause") The `UNION` clause combines the results of two or more `SELECT` statements. By default `UNION` removes duplicates. To include duplicates, use `UNION ALL`. Example: ``` SELECT a, b, c FROM table1 UNION ALL SELECT a, b, c FROM table2; ``` ### ORDER BY clause[​](#order-by-clause "Direct link to ORDER BY clause") Orders the results by the referenced expression. By default it uses ascending order (`ASC`). This order can be changed to descending by adding `DESC` after the order-by expressions. Examples: ``` SELECT age, person FROM table ORDER BY age; SELECT age, person FROM table ORDER BY age DESC; SELECT age, person FROM table ORDER BY age, person DESC; ``` #### `ORDER BY ALL`[​](#order-by-all "Direct link to order-by-all") Order from left to right (by age, then by person) in ascending order: ``` SELECT age, person FROM table ORDER BY ALL; ``` ### LIMIT clause[​](#limit-clause "Direct link to LIMIT clause") Limits the number of rows to be a maximum of `count` rows. `count` should be a non-negative integer. Example: ``` SELECT age, person FROM table LIMIT 10; ``` ### EXCLUDE, EXCEPT, REPLACE, and ILIKE clauses[​](#exclude-except-replace-and-ilike-clauses "Direct link to EXCLUDE, EXCEPT, REPLACE, and ILIKE clauses") Spice supports the following wildcard modifiers on `SELECT *`: * `EXCLUDE (col1, col2, ...)` / `EXCEPT (col1, col2, ...)` — omit the named columns. * `REPLACE (expr AS col, ...)` — substitute the named columns with a new expression. * `ILIKE 'pattern'` — emit only columns whose names match the case-insensitive pattern. `RENAME` is parsed but not yet implemented. Example selecting all columns except for `age` and `person`: ``` SELECT * EXCEPT(age, person) FROM table; ``` ``` SELECT * EXCLUDE(age, person) FROM table; ``` Example replacing a column's value while keeping all other columns: ``` SELECT * REPLACE (upper(name) AS name) FROM customers; ``` Example selecting all columns whose names contain "date": ``` SELECT * ILIKE '%date%' FROM events; ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` SELECT name, age FROM employees WHERE age > 30 ORDER BY age DESC; ``` --- # Subqueries info Spice is built on [Apache DataFusion](https://datafusion.apache.org/) and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. A subquery, also known as an inner query or nested query, is a query inside another query. Subqueries can appear in the `SELECT`, `FROM`, `WHERE`, and `HAVING` clauses. The examples below reference these sample tables: The examples below are based on the following tables. ``` SELECT * FROM x; +----------+----------+ | column_1 | column_2 | +----------+----------+ | 1 | 2 | +----------+----------+ | 2 | 4 | +----------+----------+ ``` ``` SELECT * FROM y; +--------+--------+ | number | string | +--------+--------+ | 1 | one | +--------+--------+ | 2 | two | +--------+--------+ | 3 | three | +--------+--------+ | 4 | four | +--------+--------+ ``` ## Subquery operators[​](#subquery-operators "Direct link to Subquery operators") * [\[ NOT \] EXISTS](#-not--exists) * [\[ NOT \] IN](#-not--in) ### \[ NOT ] EXISTS[​](#-not--exists "Direct link to \[ NOT ] EXISTS") The `EXISTS` operator returns rows for which a *[correlated subquery](#correlated-subqueries)* produces one or more matches. The `NOT EXISTS` operator returns rows for which the *correlated subquery* produces zero matches. Only *correlated subquery* are supported. ``` [NOT] EXISTS (subquery) ``` ### \[ NOT ] IN[​](#-not--in "Direct link to \[ NOT ] IN") The IN operator returns rows that match any value produced by a *[correlated subquery](#correlated-subqueries)* or listed values. The NOT IN operator returns rows that do not match any of these values. ``` expression [NOT] IN (subquery|list-literal) ``` #### Examples[​](#examples "Direct link to Examples") ``` SELECT * FROM x WHERE column_1 IN (1,3); +----------+----------+ | column_1 | column_2 | +----------+----------+ | 1 | 2 | +----------+----------+ ``` ``` SELECT * FROM x WHERE column_1 NOT IN (1,3); +----------+----------+ | column_1 | column_2 | +----------+----------+ | 2 | 4 | +----------+----------+ ``` ## SELECT clause subqueries[​](#select-clause-subqueries "Direct link to SELECT clause subqueries") `SELECT` clause subqueries use values returned from the inner query as part of the outer query's `SELECT` list. The `SELECT` clause only supports [scalar subqueries](#scalar-subqueries) that return a single value per execution of the inner query. The returned value can be unique per row. ``` SELECT [expression1[, expression2, ..., expressionN],] () ``` **Note**: `SELECT` clause subqueries can be used as an alternative to `JOIN` operations. ### Example[​](#example "Direct link to Example") ``` SELECT column_1, ( SELECT first_value(string) FROM y WHERE number = x.column_1 ) AS "numeric string" FROM x; +----------+----------------+ | column_1 | numeric string | +----------+----------------+ | 1 | one | | 2 | two | +----------+----------------+ ``` ## FROM clause subqueries[​](#from-clause-subqueries "Direct link to FROM clause subqueries") A subquery in the `FROM` clause produces a result set that is then referenced by the outer query. ``` SELECT expression1[, expression2, ..., expressionN] FROM () ``` ### Example[​](#example-1 "Direct link to Example") The following query returns the average of maximum values per room. The inner query returns the maximum value for each field from each room. The outer query uses the results of the inner query and returns the average maximum value for each field. ``` SELECT column_2 FROM ( SELECT * FROM x WHERE column_1 > 1 ); +----------+ | column_2 | +----------+ | 4 | +----------+ ``` ## WHERE clause subqueries[​](#where-clause-subqueries "Direct link to WHERE clause subqueries") A subquery in the `WHERE` clause compares an expression to the subquery result, returning *true* or *false*. Rows that evaluate to *false* or NULL are filtered from the final result. Both correlated and non-correlated subqueries are supported in `WHERE` clause subqueries, as well as scalar and non-scalar subqueries (depending on the operator). ``` SELECT expression1[, expression2, ..., expressionN] FROM WHERE expression operator () ``` **Note:** `WHERE` clause subqueries can be used as an alternative to `JOIN` operations. ### Examples[​](#examples-1 "Direct link to Examples") #### `WHERE` clause with scalar subquery[​](#where-clause-with-scalar-subquery "Direct link to where-clause-with-scalar-subquery") The following query returns all rows with `column_2` values above the average of all `number` values in `y`. ``` SELECT * FROM x WHERE column_2 > ( SELECT AVG(number) FROM y ); +----------+----------+ | column_1 | column_2 | +----------+----------+ | 2 | 4 | +----------+----------+ ``` #### `WHERE` clause with non-scalar subquery[​](#where-clause-with-non-scalar-subquery "Direct link to where-clause-with-non-scalar-subquery") Non-scalar subqueries must use the `[NOT] IN` or `[NOT] EXISTS` operators and can only return a single column. The values in the returned column are evaluated as a list. The following query returns all rows with `column_2` values in table `x` that are in the list of numbers with string lengths greater than three from table `y`. ``` SELECT * FROM x WHERE column_2 IN ( SELECT number FROM y WHERE length(string) > 3 ); +----------+----------+ | column_1 | column_2 | +----------+----------+ | 2 | 4 | +----------+----------+ ``` ### `WHERE` clause with correlated subquery[​](#where-clause-with-correlated-subquery "Direct link to where-clause-with-correlated-subquery") The following query returns rows with `column_2` values from table `x` greater than the average `string` value length from table `y`. The subquery in the `WHERE` clause uses the `column_1` value from the outer query to return the average `string` value length for that specific value. ``` SELECT * FROM x WHERE column_2 > ( SELECT AVG(length(string)) FROM y WHERE number = x.column_1 ); +----------+----------+ | column_1 | column_2 | +----------+----------+ | 2 | 4 | +----------+----------+ ``` ## HAVING clause subqueries[​](#having-clause-subqueries "Direct link to HAVING clause subqueries") A subquery in the `HAVING` clause compares an expression using aggregate functions to the subquery result and returns *true* or *false*. Rows that evaluate to *false* are excluded. Both correlated and non-correlated subqueries are possible, as well as scalar and non-scalar subqueries (depending on the operator). ``` SELECT aggregate_expression1[, aggregate_expression2, ..., aggregate_expressionN] FROM WHERE GROUP BY column_expression1[, column_expression2, ..., column_expressionN] HAVING expression operator () ``` ### Examples[​](#examples-2 "Direct link to Examples") The following query calculates the averages of even and odd numbers in table `y` and returns the averages that are equal to the maximum value of `column_1` in table `x`. #### `HAVING` clause with a scalar subquery[​](#having-clause-with-a-scalar-subquery "Direct link to having-clause-with-a-scalar-subquery") ``` SELECT AVG(number) AS avg, (number % 2 = 0) AS even FROM y GROUP BY even HAVING avg = ( SELECT MAX(column_1) FROM x ); +-------+--------+ | avg | even | +-------+--------+ | 2 | false | +-------+--------+ ``` #### `HAVING` clause with a non-scalar subquery[​](#having-clause-with-a-non-scalar-subquery "Direct link to having-clause-with-a-non-scalar-subquery") Non-scalar subqueries must use the `[NOT] IN` or `[NOT] EXISTS` operators and can only return a single column. The values in the returned column are evaluated as a list. The following query calculates the averages of even and odd numbers in table `y` and returns the averages that are in `column_1` of table `x`. ``` SELECT AVG(number) AS avg, (number % 2 = 0) AS even FROM y GROUP BY even HAVING avg IN ( SELECT column_1 FROM x ); +-------+--------+ | avg | even | +-------+--------+ | 2 | false | +-------+--------+ ``` ## Subquery categories[​](#subquery-categories "Direct link to Subquery categories") Subqueries can be categorized as one or more of the following based on the behavior of the subquery: * [correlated](#correlated-subqueries) or [non-correlated](#non-correlated-subqueries) * [scalar](#scalar-subqueries) or [non-scalar](#non-scalar-subqueries) ### Correlated subqueries[​](#correlated-subqueries "Direct link to Correlated subqueries") A **correlated** subquery depends on the values of the current row processed by the outer query. Spice uses DataFusion execution engine, which rewrites correlated subqueries into joins to improve performance. Correlated subqueries are typically less performant than non-correlated subqueries. ### Non-correlated subqueries[​](#non-correlated-subqueries "Direct link to Non-correlated subqueries") A **non-correlated** subquery does not depend on the outer query. The inner query runs first and passes its result to the outer query. ### Scalar subqueries[​](#scalar-subqueries "Direct link to Scalar subqueries") A **scalar** subquery returns exactly one value (one column of one row). If no rows match, the subquery returns NULL. ### Non-scalar subqueries[​](#non-scalar-subqueries "Direct link to Non-scalar subqueries") A **non-scalar** subquery can return 0, 1, or more rows, each potentially containing one or more columns. If no rows qualify, it returns zero rows. If there are no values for a particular column, it returns NULL for that column. --- # Spice.ai Open Source System Requirements This document outlines the system requirements for running Spice.ai Open Source. Ensure your environment meets these requirements to achieve optimal performance and stability. Note that resource requirements, particularly memory and storage, are highly dependent on workload and data. Specific recommendations are provided later in this document. ## Operating Systems and Architectures[​](#operating-systems-and-architectures "Direct link to Operating Systems and Architectures") Spice.ai supports the following operating systems and architectures: * **Linux (x86\_64)**: Linux 5.10 or later, Intel 64-bit processor * **Linux (ARM64)**: Linux 5.10 or later, ARM 64-bit processor * **Apple macOS (ARM64)**: macOS 14 "Sonoma" or later, Apple 64-bit processor (M-series) * **Microsoft Windows (x64)**: Windows 11 or later, Intel 64-bit processor ## Linux Distribution Support[​](#linux-distribution-support "Direct link to Linux Distribution Support") Spice.ai supports the following Linux distributions: * **Ubuntu Server LTS**: 22.04 "Jammy Jellyfish" or later * **Debian**: 12 "Bookworm" or later * **Amazon Linux**: 2 or later ## Server or Instance Hardware Requirements[​](#server-or-instance-hardware-requirements "Direct link to Server or Instance Hardware Requirements") Recommended minimum hardware specifications: * **CPU**: Quad-core processor * **Memory**: 8 GB RAM * **Storage**: 50 GB available disk space ## Network Requirements[​](#network-requirements "Direct link to Network Requirements") * **Internet Connection**: Required for downloading dependencies and updates. * **Remote Data Sources**: A low-latency, high-bandwidth connection (1-10 Gbps+) is recommended for accessing remote data, model, tools, and embedding providers. ### Port Requirements[​](#port-requirements "Direct link to Port Requirements") The following ports are used: | Description | Port | Required | | ----------------------------------- | ----- | ------------------------------------------------- | | HTTP / HTTPS (if TLS is configured) | 8090 | Yes | | Metrics Endpoint | 9090 | No (disabled by default, enable with `--metrics`) | | Arrow Flight / ADBC/ODBC/JDBC | 50051 | Yes | ## Kubernetes Requirements[​](#kubernetes-requirements "Direct link to Kubernetes Requirements") When deploying Spice.ai on Kubernetes, it is important to configure your pod specifications appropriately to ensure optimal performance. Below are the recommended minimum configurations for CPU and memory requests: ### PodSpec Configuration[​](#podspec-configuration "Direct link to PodSpec Configuration") ``` apiVersion: v1 kind: Pod metadata: name: spice-ai-pod spec: containers: - name: spice-ai-container image: spiceai/spiceai:latest resources: requests: memory: '8Gi' cpu: '4' limits: memory: '16Gi' # Set higher than request for burst capacity # Do not set CPU limits - see recommendations below ``` CPU Limits Avoid setting CPU limits for Spice pods. CPU limits can cause [throttling](https://home.robusta.dev/blog/stop-using-cpu-limits) even when CPU is available, leading to degraded query performance and increased latency. Instead, set appropriate CPU requests to guarantee scheduling and allow pods to burst when needed. For more details, see [Kubernetes CPU requests and limits](https://www.datadoghq.com/blog/kubernetes-cpu-requests-limits/). ## Resource Requirements Based on Workload and Data[​](#resource-requirements-based-on-workload-and-data "Direct link to Resource Requirements Based on Workload and Data") Spice resource requirements, particularly memory, are highly dependent on workload and data. The following table provides memory recommendations based on the dataset size and `refresh_mode`: | Refresh Mode | Memory Recommendation | | ---------------------- | --------------------- | | `refresh_mode: full` | 2.5x the dataset size | | `refresh_mode: append` | 1.5x the dataset size | See [Memory Management and Best Practices](/docs/next/reference/memory) for a detailed guide on memory considerations. ## Additional Considerations[​](#additional-considerations "Direct link to Additional Considerations") * **Security**: Regularly update your operating system and software dependencies to the latest versions to ensure security and stability. * **Backups**: Implement a regular backup strategy for your database and configuration files. * **Monitoring**: Ensure proper observability tools are in place for monitoring system performance and resource utilization. --- # Task History The Spice runtime stores information about completed tasks in the `spice.runtime.task_history` table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. A span is a unit of trace data that encapsulates the details of a task's execution, including its duration, inputs, and outputs. Spans enable hierarchical tracing by grouping tasks under a parent span, which provides a view of task dependencies and the overall execution flow. For configuration options, see the [`runtime.task_history` reference](/docs/next/reference/spicepod/runtime#runtimetask_history). ## Configuration[​](#configuration "Direct link to Configuration") Task history is enabled by default and retains records for 8 hours. To adjust retention and other settings, configure the `runtime.task_history` section in `spicepod.yaml`: ``` runtime: task_history: enabled: true captured_output: none retention_period: 8h retention_check_interval: 15m ``` * **`enabled`**: Enable or disable task history. Defaults to `true`. * **`captured_output`**: Level of output captured. Defaults to `none`. * **`captured_context`**: How much of the AI and search task payload is stored. Defaults to `truncated`. See [Captured context](#captured-context). * **`retention_period`**: How long records are retained. Defaults to `8h`. Longer retention periods increase memory usage. * **`retention_check_interval`**: How often old records are checked for removal. Defaults to `15m`. For the full list of parameters, see the [`runtime.task_history` reference](/docs/next/reference/spicepod/runtime#runtimetask_history). ## Captured Context[​](#captured-context "Direct link to Captured Context") Tasks that carry user or model content — `ai_chat`, `ai_completion`, `responses`, `text_embed`, `search`, `nsql`, `scheduled_worker`, and every `tool_use::*` task — store that content in the `input` and `captured_output` columns. `captured_context` controls how much of it is written: ``` runtime: task_history: captured_context: redacted ``` | Value | Behavior | | --------------------- | ---------------------------------------------------------------------------------------------- | | `truncated` (default) | Payloads longer than 4096 characters are cut at that point and suffixed with `...[truncated]`. | | `redacted` | The payload is replaced with `[redacted]`. | | `full` | The payload is stored in full. | Values are case-sensitive; an unrecognized value fails at load. Other task types (for example `sql_query` and `accelerated_refresh`) are unaffected by this setting. note `captured_context` shapes prompt, tool, and search payloads only. Whether task output is recorded at all is controlled separately by `captured_output`, which defaults to `none`. ## Querying Task History[​](#querying-task-history "Direct link to Querying Task History") Task history is queryable as a standard SQL table at `runtime.task_history` (or `spice.runtime.task_history`). To retrieve all recorded tasks, run: ``` SELECT * FROM runtime.task_history; ``` This query can be issued through any Spice SQL interface, including the [HTTP API](/docs/next/api/HTTP/post-sql), [Arrow Flight SQL](/docs/next/api/arrow-flight-sql), or the [Spice SQL REPL](/docs/next/cli/reference/sql). ## Persisting Task History[​](#persisting-task-history "Direct link to Persisting Task History") Task history is stored in-memory and subject to the configured `retention_period`. To persist task history beyond the retention window, set up a [worker](/docs/next/reference/spicepod/workers) with a cron schedule that periodically writes records to an external dataset. For example, to back up task history to an Iceberg table every 10 minutes: ``` datasets: - from: glue:team_app.task_history name: task_history_sink mode: read_write params: glue_auth: key glue_region: us-east-1 glue_key: ${secrets:AWS_GLUE_ACCESS_KEY} glue_secret: ${secrets:AWS_GLUE_SECRET_ACCESS_KEY} workers: - name: backup-task-history cron: '*/10 * * * *' # every 10 minutes sql: | INSERT INTO task_history_sink SELECT * FROM runtime.task_history WHERE start_time >= NOW() - INTERVAL '10' MINUTE; ``` This approach writes recent task history records to a durable store on a regular schedule, ensuring data is available for later analysis even after the in-memory retention window expires. ## Table Schema[​](#table-schema "Direct link to Table Schema") ``` describe runtime.task_history; ``` Output ``` +-----------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------+ | column_name | data_type | is_nullable | +-----------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------+ | trace_id | Utf8 | NO | | span_id | Utf8 | NO | | parent_span_id | Utf8 | YES | | task | Utf8 | NO | | input | Utf8 | NO | | captured_output | Utf8 | YES | | start_time | Timestamp(Nanosecond, None) | NO | | end_time | Timestamp(Nanosecond, None) | NO | | execution_duration_ms | Float64 | NO | | error_message | Utf8 | YES | | labels | Map(Field { name: "entries", data_type: Struct([Field { name: "keys", data_type: Utf8, nullable: false, dict_id: 0, dict_is_ordered: false, metadata: {} }, Field { name: "values", data_type: Utf8, nullable: false, dict_id: 0, dict_is_ordered: false, metadata: {} }]), nullable: false, dict_id: 0, dict_is_ordered: false, metadata: {} }, false) | NO | +-----------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------+ ``` * **trace\_id**: Unique identifier for the entire trace this task is part of. * **span\_id**: Unique identifier for this specific task within the trace. * **parent\_span\_id**: Identifier of the parent task, if any. * **task**: Name or description of the task being performed (e.g., `sql_query`). * **input**: Input data or parameters for the task. * **captured\_output**: Output or result of the task, if available. * **start\_time**: Time when the task started. * **end\_time**: Time when the task ended. * **execution\_duration\_ms**: Duration of the task execution in milliseconds. * **error\_message**: Error message if the task failed, otherwise null. * **labels**: Key-value pairs for additional metadata or attributes associated with the task. ## Retrieve all tasks within a specific timeframe[​](#retrieve-all-tasks-within-a-specific-timeframe "Direct link to Retrieve all tasks within a specific timeframe") ``` SELECT trace_id, span_id, task, start_time, end_time, execution_duration_ms, error_message FROM spice.runtime.task_history WHERE start_time >= NOW() - INTERVAL '10 MINUTES' AND end_time <= NOW(); ``` Example output: ``` +----------------------------------+------------------+---------------------+----------------------------+----------------------------+-----------------------+---------------------------------------------------------------------------------------------+ | trace_id | span_id | task | start_time | end_time | execution_duration_ms | error_message | +----------------------------------+------------------+---------------------+----------------------------+----------------------------+-----------------------+---------------------------------------------------------------------------------------------+ | 687e0970f8c49d19c5a08764ea2d4dc1 | f4f52ed29db8b151 | text_embed | 2024-11-25T05:39:37.444749 | 2024-11-25T05:39:53.577195 | 16132.446000000002 | | | 687e0970f8c49d19c5a08764ea2d4dc1 | e47b17bd9fd9fe37 | accelerated_refresh | 2024-11-25T05:39:31.112504 | 2024-11-25T05:39:53.579933 | 22467.429 | | | 1e881188e5fd252b26adb8a8d838efb8 | 532b0019ad778094 | sql_query | 2024-11-25T05:40:38.864982 | 2024-11-25T05:40:38.871090 | 6.108 | | | 2ee1c700b450034bb6c2da3de2e2386c | 235dafed1e7d8c02 | sql_query | 2024-11-25T05:39:38.249113 | 2024-11-25T05:39:39.387258 | 1138.145 | | | 20e75df9ea77ba1c8cb99a2632cdd091 | d07551cd172ffa80 | sql_query | 2024-11-25T05:39:39.458135 | 2024-11-25T05:39:39.482181 | 24.046000000000003 | | | ca1d470b12191726b61d825df6f2ce2a | 65597a0bc0a4fde3 | sql_query | 2024-11-25T05:39:39.675726 | 2024-11-25T05:39:39.822479 | 146.753 | | | ac5abd8bfec7e5aa7c19fc84772c55f1 | 316622ac359e3c00 | sql_query | 2024-11-25T05:39:39.872946 | 2024-11-25T05:39:39.872994 | 0.048 | This feature is not implemented: The context currently only supports a single SQL statement | | 1c640298e248ba297a12b1e3b59fffc7 | 031c3a25dc56d8e9 | sql_query | 2024-11-25T05:39:40.467032 | 2024-11-25T05:39:40.486156 | 19.124 | | | 2c4d9abee740ced8ae423e0eb4fcff6b | a324b699b8bcf338 | sql_query | 2024-11-25T05:39:40.525506 | 2024-11-25T05:39:40.525526 | 0.02 | This feature is not implemented: The context currently only supports a single SQL statement | | e5ed7f7a98e62f493ef8af2e0cd7734e | e84c30862a546bb5 | sql_query | 2024-11-25T05:39:40.560891 | 2024-11-25T05:39:40.560911 | 0.02 | This feature is not implemented: The context currently only supports a single SQL statement | | d471f83092a95bde8663438cda74627f | 3dd9c4d4ebff4cb9 | sql_query | 2024-11-25T05:39:40.600892 | 2024-11-25T05:39:40.647092 | 46.199999999999996 | | | 701874d7282dd47791e7519b343a9694 | 5dacf75c4537ee0e | accelerated_refresh | 2024-11-25T05:39:30.452534 | 2024-11-25T05:39:30.452900 | 0.366 | | | 2e6b672a49a8cd5f0862a760661dc846 | f813941e0699e783 | accelerated_refresh | 2024-11-25T05:39:30.848425 | 2024-11-25T05:39:30.857242 | 8.817 | | | 18d76b6389898cc5253a49294607477d | cc0d06a4e69cbcd5 | health | 2024-11-25T05:39:30.451626 | 2024-11-25T05:39:31.563876 | 1112.25 | | | c75af81360e8962639faa64e6804b830 | 1ea2c95b243a5717 | accelerated_refresh | 2024-11-25T05:39:31.036470 | 2024-11-25T05:39:31.607845 | 571.375 | | | 817d88778e91322640414263779ce7f1 | 513a58d83f0416a7 | accelerated_refresh | 2024-11-25T05:39:30.998455 | 2024-11-25T05:39:32.076359 | 1077.904 | | | 3c507ee30211e6fab7d8a2eaf686e451 | d9be117925fb6d42 | accelerated_refresh | 2024-11-25T05:39:31.061851 | 2024-11-25T05:39:32.078412 | 1016.561 | | | aa6010405a12a14b6afaf76e9fabedb8 | 1f50a2b177003c54 | accelerated_refresh | 2024-11-25T05:39:30.933543 | 2024-11-25T05:39:32.476197 | 1542.654 | | | 3c75d16b6b4b8da98c551d115e1c049c | 9a16dc065a95236a | sql_query | 2024-11-25T05:42:27.386754 | 2024-11-25T05:42:27.386859 | 0.10500000000000001 | SQL error: ParserError("Expected: an SQL statement, found: ELECT") | +----------------------------------+------------------+---------------------+----------------------------+----------------------------+-----------------------+---------------------------------------------------------------------------------------------+ ``` ## Retrieve the most recent error messages[​](#retrieve-the-most-recent-error-messages "Direct link to Retrieve the most recent error messages") ``` SELECT trace_id, task, error_message, SUBSTRING(input, 1, 100) AS input_preview, start_time FROM spice.runtime.task_history WHERE error_message IS NOT NULL ORDER BY start_time DESC LIMIT 5; ``` ``` +----------------------------------+-----------+---------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------+----------------------------+ | trace_id | task | error_message | input_preview | start_time | +----------------------------------+-----------+---------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------+----------------------------+ | 352539e75fdb1a3d5fc3b48bfd4b4bae | sql_query | Error during planning: Invalid function 'date'. | SELECT DATE(start_time) AS task_date, COUNT(*) AS task_count | 2024-11-25T06:17:40.573970 | | | | Did you mean 'tanh'? | FROM spice.runtime.task_history | | | | | | GROUP B | | | f6672d562ad97dde0bb4db428723461f | sql_query | This feature is not implemented: The context currently only supports a single SQL statement | with ssales as (select c_last_name ,c_first_name ,s_store_name ,ca_state ,s_ | 2024-11-25T06:06:39.800900 | | b16fa36e5a2f7f119fc3834875f6bdee | sql_query | This feature is not implemented: The context currently only supports a single SQL statement | with frequent_ss_items as (select substr(i_item_desc,1,30) itemdesc,i_item_sk item_sk,d_date soldda | 2024-11-25T06:06:39.760532 | | 8d126ea506a374c0c6239c11ee5cbe5a | sql_query | This feature is not implemented: The context currently only supports a single SQL statement | with cross_items as (select i_item_sk ss_item_sk from item, (select iss.i_brand_id brand_id | 2024-11-25T06:06:39.118422 | | 198a9dcc496f0435cff69de61cc07874 | sql_query | This feature is not implemented: The context currently only supports a single SQL statement | with ssales as (select c_last_name ,c_first_name ,s_store_name ,ca_state ,s_ | 2024-11-25T06:04:11.317419 | +----------------------------------+-----------+---------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------+----------------------------+ ``` ## Summarize number of tasks by type[​](#summarize-number-of-tasks-by-type "Direct link to Summarize number of tasks by type") ``` SELECT task, COUNT(*) AS task_count, AVG(execution_duration_ms) AS avg_duration_ms FROM spice.runtime.task_history GROUP BY task ORDER BY task_count DESC; ``` Example output: ``` +-------------------------------+------------+---------------------+ | task | task_count | avg_duration_ms | +-------------------------------+------------+---------------------+ | sql_query | 65 | 55.10198461538462 | | accelerated_refresh | 27 | 749.1187407407407 | | ai_completion | 9 | 5026.337888888888 | | tool_use::list_datasets | 4 | 0.16899999999999998 | | text_embed | 4 | 3341.08975 | | ai_chat | 4 | 7151.03675 | | search | 3 | 384.376 | | tool_use::search | 3 | 385.0406666666667 | | tool_use::get_readiness | 1 | 0.12999999999999998 | | tool_use::sample_data | 1 | 2.275 | | health | 1 | 661.0169999999999 | +-------------------------------+------------+---------------------+ ``` ## Identify the longest-running tasks[​](#identify-the-longest-running-tasks "Direct link to Identify the longest-running tasks") ``` SELECT task, trace_id, parent_span_id, execution_duration_ms, labels FROM spice.runtime.task_history ORDER BY execution_duration_ms DESC LIMIT 10; ``` Example output: ``` +---------------------+----------------------------------+------------------+-----------------------+------------------------------------------------------------------------------------------------+ | task | trace_id | parent_span_id | execution_duration_ms | labels | +---------------------+----------------------------------+------------------+-----------------------+------------------------------------------------------------------------------------------------+ | accelerated_refresh | d9c38c7e58a02ec939240385a4a25a04 | | 1093711.474 | {sql: SELECT * FROM react.issues} | | ai_chat | 7a6427313880942316bf3018cd23a198 | | 17202.836000000003 | {model: gpt-4o} | | ai_completion | 7a6427313880942316bf3018cd23a198 | 59b1fd88c8397e3f | 17202.475 | {model: gpt-4o, total_tokens: 2673, prompt_tokens: 1807, completion_tokens: 866, stream: true} | | accelerated_refresh | 96758c1132164204a68e1a7234a06cda | | 15660.023000000001 | {sql: SELECT * FROM react.docs} | | text_embed | 96758c1132164204a68e1a7234a06cda | 109c489b24602356 | 12406.787 | {outputs_produced: 2086} | | ai_chat | b2a69503a1b83215603ead321eea6f61 | | 6445.162 | {model: gpt-4o} | | ai_completion | b2a69503a1b83215603ead321eea6f61 | 95411c59fc9c8cb8 | 6444.1990000000005 | {prompt_tokens: 1454, stream: true, total_tokens: 1484, model: gpt-4o, completion_tokens: 30} | | ai_completion | b2a69503a1b83215603ead321eea6f61 | 95411c59fc9c8cb8 | 5608.6359999999995 | {prompt_tokens: 1529, total_tokens: 1559, model: gpt-4o, completion_tokens: 30, stream: true} | | text_embed | 65880ecfc884a41555ac4d21ceef9aef | | 5143.494000000001 | {outputs_produced: 1} | | text_embed | f51e5e9d4e26de31a2f7d5e9286dd8f4 | | 4769.832 | {outputs_produced: 1} | +---------------------+----------------------------------+------------------+-----------------------+------------------------------------------------------------------------------------------------+ ``` ## Retrieve details of all tasks associated with a specific trace[​](#retrieve-details-of-all-tasks-associated-with-a-specific-trace "Direct link to Retrieve details of all tasks associated with a specific trace") ``` SELECT task, trace_id, parent_span_id, span_id, execution_duration_ms, start_time, end_time, SUBSTRING(input, 1, 100) AS input_preview, SUBSTRING(captured_output, 1, 100) AS output_preview, error_message, labels FROM spice.runtime.task_history WHERE trace_id = 'b2a69503a1b83215603ead321eea6f61' ORDER BY start_time; ``` Example output: ``` +-------------------------------+----------------------------------+------------------+------------------+-----------------------+----------------------------+----------------------------+------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------+ | task | trace_id | parent_span_id | span_id | execution_duration_ms | start_time | end_time | input_preview | output_preview | error_message | labels | +-------------------------------+----------------------------------+------------------+------------------+-----------------------+----------------------------+----------------------------+------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------+ | ai_chat | b2a69503a1b83215603ead321eea6f61 | | 95411c59fc9c8cb8 | 6445.162 | 2024-11-25T06:03:31.197980 | 2024-11-25T06:03:37.643142 | {"messages":[{"role":"user","content":"how to install react"},{"role":"user","content":"top 3 recent | It seems that the dataset containing the recent React issues is currently being refreshed and is not | | {model: gpt-4o} | | tool_use::list_datasets | b2a69503a1b83215603ead321eea6f61 | 95411c59fc9c8cb8 | aa5ba649da3f0581 | 0.367 | 2024-11-25T06:03:31.198139 | 2024-11-25T06:03:31.198506 | | [{"can_search_documents":true,"description":"React.js documentation and reference, from https://reac | | {tool: list_datasets} | | ai_completion | b2a69503a1b83215603ead321eea6f61 | 95411c59fc9c8cb8 | 8bd67a43da1b4312 | 6444.1990000000005 | 2024-11-25T06:03:31.198906 | 2024-11-25T06:03:37.643105 | {"messages":[{"role":"assistant","tool_calls":[{"id":"initial_list_datasets","type":"function","func | | | {prompt_tokens: 1454, stream: true, total_tokens: 1484, model: gpt-4o, completion_tokens: 30} | | tool_use::search | b2a69503a1b83215603ead321eea6f61 | 95411c59fc9c8cb8 | 49b25ff51580d4fa | 140.337 | 2024-11-25T06:03:31.893913 | 2024-11-25T06:03:32.034250 | {"text":"how to install react","datasets":["spice.react.issues"],"limit":3} | | Error occurred interacting with datafusion: Failed to execute query: External error: Acceleration not ready; loading initial data for react.issues | {tool: search} | | search | b2a69503a1b83215603ead321eea6f61 | 49b25ff51580d4fa | 9feb8c9a54079cbd | 140.24200000000002 | 2024-11-25T06:03:31.894 | 2024-11-25T06:03:32.034242 | how to install react | | Error occurred interacting with datafusion: Failed to execute query: External error: Acceleration not ready; loading initial data for react.issues | {limit: 3, tables: spice.react.issues} | | text_embed | b2a69503a1b83215603ead321eea6f61 | 9feb8c9a54079cbd | 7f18fd8d96a1baeb | 122.771 | 2024-11-25T06:03:31.894072 | 2024-11-25T06:03:32.016843 | "how to install react" | | | {outputs_produced: 1} | | sql_query | b2a69503a1b83215603ead321eea6f61 | 9feb8c9a54079cbd | 025f6b1e5cd502a6 | 16.892 | 2024-11-25T06:03:32.017320 | 2024-11-25T06:03:32.034212 | WITH ranked_docs as ( | | Failed to execute query: External error: Acceleration not ready; loading initial data for react.issues | {error_code: QueryExecutionError, protocol: Internal, query_execution_duration_ms: 8.20325, datasets: spice.react.issues, rows_produced: 0} | | | | | | | | | SELECT id, dist, offset FROM ( | | | | | | | | | | | | SELECT | | | | | | | | | | | | | | | | | ai_completion | b2a69503a1b83215603ead321eea6f61 | 95411c59fc9c8cb8 | 1b2a3273a1dc4cdf | 5608.6359999999995 | 2024-11-25T06:03:32.034451 | 2024-11-25T06:03:37.643087 | {"messages":[{"role":"assistant","tool_calls":[{"id":"initial_list_datasets","type":"function","func | | | {prompt_tokens: 1529, total_tokens: 1559, model: gpt-4o, completion_tokens: 30, stream: true} | | tool_use::sample_data | b2a69503a1b83215603ead321eea6f61 | 95411c59fc9c8cb8 | 8bd10431fb18df87 | 2.275 | 2024-11-25T06:03:33.039583 | 2024-11-25T06:03:33.041858 | TopNSample({"dataset":"spice.react.issues","limit":3,"order_by":"created_at DESC"}) | | | {sample_method: top_n_sample, tool: top_n_sample} | | sql_query | b2a69503a1b83215603ead321eea6f61 | 8bd10431fb18df87 | f1b2e06225aa2ba3 | 2.193 | 2024-11-25T06:03:33.039625 | 2024-11-25T06:03:33.041818 | SELECT * FROM spice.react.issues ORDER BY created_at DESC LIMIT 3 | | Failed to execute query: External error: Acceleration not ready; loading initial data for react.issues | {query_execution_duration_ms: 1.5519999, datasets: spice.react.issues, protocol: Internal, error_code: QueryExecutionError, rows_produced: 0} | | ai_completion | b2a69503a1b83215603ead321eea6f61 | 95411c59fc9c8cb8 | b649e4bbe8aac7a3 | 4601.02 | 2024-11-25T06:03:33.042037 | 2024-11-25T06:03:37.643057 | {"messages":[{"role":"assistant","tool_calls":[{"id":"initial_list_datasets","type":"function","func | | | {prompt_tokens: 1599, completion_tokens: 11, model: gpt-4o, stream: true, total_tokens: 1610} | | tool_use::get_readiness | b2a69503a1b83215603ead321eea6f61 | 95411c59fc9c8cb8 | f40f0b4cd608de8b | 0.12999999999999998 | 2024-11-25T06:03:33.499867 | 2024-11-25T06:03:33.499997 | | {"dataset:call_center":"Ready","dataset:catalog_page":"Ready","dataset:catalog_returns":"Ready","dat | | {tool: get_readiness} | | ai_completion | b2a69503a1b83215603ead321eea6f61 | 95411c59fc9c8cb8 | e018a56742b5064f | 4142.789 | 2024-11-25T06:03:33.500219 | 2024-11-25T06:03:37.643008 | {"messages":[{"role":"assistant","tool_calls":[{"id":"initial_list_datasets","type":"function","func | | | {stream: true, completion_tokens: 254, prompt_tokens: 1920, model: gpt-4o, total_tokens: 2174} | +-------------------------------+----------------------------------+------------------+------------------+-----------------------+----------------------------+----------------------------+------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------+ ``` ## Retrieve details of most recent chat query[​](#retrieve-details-of-most-recent-chat-query "Direct link to Retrieve details of most recent chat query") ``` SELECT task, trace_id, execution_duration_ms, start_time, SUBSTRING(input, 1, 100) AS input_preview, SUBSTRING(captured_output, 1, 100) AS output_preview, error_message, labels FROM spice.runtime.task_history WHERE trace_id = ( SELECT trace_id FROM spice.runtime.task_history WHERE task = 'ai_chat' ORDER BY start_time DESC LIMIT 1 ) ORDER BY start_time; ``` ``` -----------------+------------------------------------------------------------------------------------------------------+---------------+--------------------------------------------------------------------------------------------------------------+ | task | trace_id | execution_duration_ms | start_time | input_preview | output_preview | error_message | labels | +-------------------------------+----------------------------------+-----------------------+----------------------------+------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------+---------------+--------------------------------------------------------------------------------------------------------------+ | ai_chat | bb2de94b6f575b6c001f39cfded8bff4 | 5665.4169999999995 | 2024-11-25T07:02:43.240005 | {"messages":[{"role":"user","content":"how to install react"},{"role":"user","content":"top 3 recent | Here are the three most recent issues related to React, along with their summaries and links: | | {model: gpt-4o} | | | | | | | | | | | | | | | | 1. ** | | | | tool_use::list_datasets | bb2de94b6f575b6c001f39cfded8bff4 | 0.157 | 2024-11-25T07:02:43.240048 | | [{"can_search_documents":true,"description":"React.js documentation and reference, from https://reac | | {tool: list_datasets} | | ai_completion | bb2de94b6f575b6c001f39cfded8bff4 | 5664.964 | 2024-11-25T07:02:43.240416 | {"messages":[{"role":"assistant","tool_calls":[{"id":"initial_list_datasets","type":"function","func | | | {prompt_tokens: 2688, total_tokens: 2716, completion_tokens: 28, model: gpt-4o, stream: true} | | tool_use::search | bb2de94b6f575b6c001f39cfded8bff4 | 710.9609999999999 | 2024-11-25T07:02:44.344935 | {"text":"recent issues","datasets":["spice.react.issues"],"limit":3} | | | {tool: search} | | search | bb2de94b6f575b6c001f39cfded8bff4 | 710.842 | 2024-11-25T07:02:44.345018 | recent issues | {Full { catalog: "spice", schema: "react", table: "issues" }: VectorSearchTableResult { data: [Recor | | {limit: 3, tables: spice.react.issues} | | text_embed | bb2de94b6f575b6c001f39cfded8bff4 | 562.453 | 2024-11-25T07:02:44.345072 | "recent issues" | | | {outputs_produced: 1} | | sql_query | bb2de94b6f575b6c001f39cfded8bff4 | 147.672 | 2024-11-25T07:02:44.908148 | WITH ranked_docs as ( | [{"title_chunk":"app:lintVitalReleaseBug issu","id":"I_kwDOAJy2Ks47fqsI","title":"app:lintVitalRelea | | {datasets: spice.react.issues, protocol: Internal, query_execution_duration_ms: 139.69496, rows_produced: 3} | | | | | | SELECT id, dist, offset FROM ( | | | | | | | | | SELECT | | | | | | | | | | | | | | ai_completion | bb2de94b6f575b6c001f39cfded8bff4 | 3849.3140000000003 | 2024-11-25T07:02:45.055985 | {"messages":[{"role":"assistant","tool_calls":[{"id":"initial_list_datasets","type":"function","func | | | {model: gpt-4o, total_tokens: 3240, stream: true, completion_tokens: 271, prompt_tokens: 2969} | +-------------------------------+----------------------------------+-----------------------+----------------------------+------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------+---------------+--------------------------------------------------------------------------------------------------------------+ ``` ## Retrieve Recent Queries for Specific Dataset[​](#retrieve-recent-queries-for-specific-dataset "Direct link to Retrieve Recent Queries for Specific Dataset") ``` SELECT task, start_time, execution_duration_ms, SUBSTRING(input, 1, 100) AS input_preview, error_message, labels FROM spice.runtime.task_history WHERE 'catalog_sales' = ANY(string_to_array(labels['datasets'], ',')) ORDER BY start_time DESC LIMIT 5; ``` Example output: ``` +-----------+----------------------------+-----------------------+------------------------------------------------------------------------------------------------------+---------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | task | start_time | execution_duration_ms | input_preview | error_message | labels | +-----------+----------------------------+-----------------------+------------------------------------------------------------------------------------------------------+---------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | sql_query | 2024-11-25T07:15:31.278018 | 19.424 | with ss as ( select i_manufact_id,sum(ss_ext_sales_price) total_sales from store_sales | | {accelerated: true, protocol: FlightSQL, query_execution_duration_ms: 18.765831, rows_produced: 15, datasets: store_sales,customer_address,item,date_dim,web_sales,catalog_sales} | | sql_query | 2024-11-25T07:15:31.233147 | 1.661 | select sum(cs_ext_discount_amt) as "excess discount amount" from catalog_sales ,item ,dat | | {query_execution_duration_ms: 1.380667, accelerated: true, datasets: date_dim,catalog_sales,item, rows_produced: 1, protocol: FlightSQL} | | sql_query | 2024-11-25T07:15:30.889905 | 30.101 | select i_item_id ,i_item_desc ,s_store_id ,s_store_name ,stddev_samp(ss_quantit | | {query_execution_duration_ms: 29.511086, datasets: store_sales,store,item,date_dim,store_returns,catalog_sales, rows_produced: 0, protocol: FlightSQL, accelerated: true} | | sql_query | 2024-11-25T07:15:30.658339 | 24.942 | select i_item_id, avg(cs_quantity) agg1, avg(cs_list_price) agg2, avg(cs_co | | {query_execution_duration_ms: 24.620039, protocol: FlightSQL, datasets: item,date_dim,catalog_sales,promotion,customer_demographics, accelerated: true, rows_produced: 73} | | sql_query | 2024-11-25T07:15:30.574257 | 40.908 | select i_item_id ,i_item_desc ,s_store_id ,s_store_name ,min(ss_net_profit) as store_sales_prof | | {protocol: FlightSQL, rows_produced: 0, accelerated: true, datasets: store_returns,date_dim,store,store_sales,item,catalog_sales, query_execution_duration_ms: 40.29496} | +-----------+----------------------------+-----------------------+------------------------------------------------------------------------------------------------------+---------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ ``` ## SQL Query Plan Capture[​](#sql-query-plan-capture "Direct link to SQL Query Plan Capture") Spice supports automated SQL query plan capture through `EXPLAIN` or `EXPLAIN ANALYZE` statements. This feature stores query plans in the task history, providing detailed execution information for analysis and debugging. Query plans are captured asynchronously after query completion to avoid blocking query execution. ### Configuration[​](#configuration-1 "Direct link to Configuration") Configure SQL query plan capture using `runtime.task_history` parameters: ``` runtime: task_history: captured_plan: explain analyze min_sql_duration: 5s min_plan_duration: 10s ``` #### Parameters[​](#parameters "Direct link to Parameters") * **`captured_plan`**: Determines which plan type is captured. Options: * `none` (default): No plans are captured * `explain`: Captures the logical query plan without executing it * `explain analyze`: Captures the query plan with actual execution metrics * **`min_sql_duration`**: Minimum query execution duration before a plan is captured. Plans are captured only for queries that exceed this threshold. * **`min_plan_duration`**: Minimum plan execution duration before a plan is captured. This threshold applies to the execution time of the `EXPLAIN` or `EXPLAIN ANALYZE` operation itself. ### Captured Output[​](#captured-output "Direct link to Captured Output") Query plans are stored in the `captured_output` column of the task history table. This column contains the result of the `EXPLAIN` or `EXPLAIN ANALYZE` statement generated for the original query. ### Query Examples[​](#query-examples "Direct link to Query Examples") Retrieve query plans for slow queries: ``` SELECT start_time, execution_duration_ms, SUBSTRING(input, 1, 80) AS query_preview, captured_output FROM spice.runtime.task_history WHERE task = 'sql_query' AND captured_output IS NOT NULL AND execution_duration_ms > 5000 ORDER BY start_time DESC LIMIT 5; ``` Find queries with specific execution patterns: ``` SELECT start_time, execution_duration_ms, captured_output FROM spice.runtime.task_history WHERE task = 'sql_query' AND captured_output LIKE '%TableScan%' AND execution_duration_ms > 1000 ORDER BY execution_duration_ms DESC; ``` --- # SDKs Spice provides official SDKs for programmatic access to datasets and AI features. Use these SDKs to integrate Spice into applications written in Python, Go, Rust, and other languages. ## [📄️Dotnet SDK](/docs/next/sdks/dotnet) [Connect to Spice using the Dotnet SDK](/docs/next/sdks/dotnet) ## [📄️Go SDK](/docs/next/sdks/golang) [Connect to Spice using the Go SDK](/docs/next/sdks/golang) ## [📄️Java SDK](/docs/next/sdks/java) [Connect to Spice using the Java SDK](/docs/next/sdks/java) ## [📄️JavaScript SDK](/docs/next/sdks/javascript) [Connect to Spice using the JavaScript SDK](/docs/next/sdks/javascript) ## [📄️Python SDK](/docs/next/sdks/python) [Connect to Spice using the Python SDK](/docs/next/sdks/python) ## [📄️Rust SDK](/docs/next/sdks/rust) [Connect to Spice using the Rust SDK](/docs/next/sdks/rust) --- # Dotnet SDK ## Dotnet SDK for Spice.ai[​](#dotnet-sdk-for-spiceai "Direct link to Dotnet SDK for Spice.ai") [github.com/spiceai/spice-dotnet](https://github.com/spiceai/spice-dotnet) ### Install[​](#install "Direct link to Install") ``` dotnet add package spiceai ``` ### Connect to Spice Runtime[​](#connect-to-spice-runtime "Direct link to Connect to Spice Runtime") Create a `SpiceClient` using default configuration: ``` using Spice; var client = new SpiceClientBuilder().Build(); var data = await client.Query( "SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10" ); ``` Or pass a custom flight address: ``` using Spice; var client = new SpiceClientBuilder() .WithFlightAddress("http://my_remote_spice_instance:50051") .Build(); var data = await client.Query( "SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10" ); ``` ### Parameterized Queries[​](#parameterized-queries "Direct link to Parameterized Queries") Use `Query()` with a `Dictionary` for parameterized queries to prevent SQL injection and improve performance (v1.1.0+): ``` using Spice; var client = new SpiceClientBuilder().Build(); var parameters = new Dictionary { { "min_price", 10.0 }, { "category", "electronics" } }; var data = await client.Query( "SELECT * FROM products WHERE price > :min_price AND category = :category", parameters ); ``` For more details, see [Parameterized Queries](/docs/next/features/query-federation/parameterized-queries). --- # Go SDK ## Go SDK for Spice.ai[​](#go-sdk-for-spiceai "Direct link to Go SDK for Spice.ai") [github.com/spiceai/gospice](https://github.com/spiceai/gospice) ### Install[​](#install "Direct link to Install") ``` go get github.com/spiceai/gospice/v8@latest ``` ### Connect to Spice Runtime[​](#connect-to-spice-runtime "Direct link to Connect to Spice Runtime") Import the package: ``` import "github.com/spiceai/gospice/v8" ``` Create a `SpiceClient` using default configuration: ``` spice := gospice.NewSpiceClient() defer spice.Close() if err := spice.Init(); err != nil { panic(fmt.Errorf("error initializing SpiceClient: %w", err)) } ``` Or pass a custom flight address: ``` if err := spice.Init( gospice.WithFlightAddress("grpc://my_remote_spice_instance:50051"), ); err != nil { panic(fmt.Errorf("error initializing SpiceClient: %w", err)) } ``` ### Execute a Query[​](#execute-a-query "Direct link to Execute a Query") Execute a query and get back an Apache Arrow RecordReader: ``` reader, err := spice.Sql( context.Background(), "SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10", ) if err != nil { panic(fmt.Errorf("error querying: %w", err)) } defer reader.Release() for reader.Next() { record := reader.RecordBatch() defer record.Release() fmt.Println(record) } ``` ### Parameterized Queries[​](#parameterized-queries "Direct link to Parameterized Queries") Use `SqlWithParams()` for parameterized queries to prevent SQL injection and improve performance: ``` reader, err := spice.SqlWithParams( context.Background(), "SELECT * FROM customers WHERE c_custkey > $1 LIMIT $2", 100, // $1: customer key threshold 10, // $2: limit ) if err != nil { panic(fmt.Errorf("error querying: %w", err)) } defer reader.Release() ``` Multiple parameter types are supported with automatic type inference: ``` reader, err := spice.SqlWithParams( ctx, "SELECT * FROM orders WHERE customer_id = $1 AND order_date > $2 AND total > $3", 42, // int "2024-01-01", // string 100.50, // float64 ) ``` For explicit type control, use typed parameter constructors: ``` import "github.com/apache/arrow-go/v18/arrow" reader, err := spice.SqlWithParams( ctx, "SELECT * FROM financial WHERE amount >= $1", gospice.Decimal128Param(amountBytes, 19, 4), ) ``` For more details, see [Parameterized Queries](/docs/next/features/query-federation/parameterized-queries). ### Health Checks[​](#health-checks "Direct link to Health Checks") Check if the Spice instance is healthy: ``` ctx := context.Background() // Unauthenticated health check if spice.IsSpiceHealthy(ctx) { fmt.Println("Spice is healthy") } // Authenticated readiness check (requires API key for Spice Cloud) if spice.IsSpiceReady(ctx) { fmt.Println("Spice is ready") } ``` ### Connection to Spice Cloud[​](#connection-to-spice-cloud "Direct link to Connection to Spice Cloud") Connect to Spice Cloud with an API key: ``` spice := gospice.NewSpiceClient() defer spice.Close() if err := spice.Init( gospice.WithApiKey(os.Getenv("SPICE_API_KEY")), gospice.WithSpiceCloudAddress(), ); err != nil { panic(fmt.Errorf("error initializing SpiceClient: %w", err)) } ``` --- # Java SDK ## Java SDK for Spice.ai[​](#java-sdk-for-spiceai "Direct link to Java SDK for Spice.ai") [github.com/spiceai/spice-java](https://github.com/spiceai/spice-java) ### Installation[​](#installation "Direct link to Installation") * Maven * Gradle Add the following dependency: ``` ai.spice spiceai 0.6.0 compile ``` Add the following dependency: ``` implementation 'ai.spice:spiceai:0.6.0' ``` ### Connect to Spice runtime[​](#connect-to-spice-runtime "Direct link to Connect to Spice runtime") Create a `SpiceClient` using default configuration. Requires local Spice OSS running: [follow the quickstart](https://github.com/spiceai/spiceai?tab=readme-ov-file#%EF%B8%8F-quickstart-local-machine) ``` import org.apache.arrow.flight.FlightStream; import ai.spice.SpiceClient; public class App { public static void main( String[] args ) { try { SpiceClient client = SpiceClient.builder() .build(); FlightStream res = client.query("SELECT \"VendorID\", \"tpep_pickup_datetime\", \"fare_amount\" FROM taxi_trips LIMIT 10"); while (res.next()) { System.out.println(res.getRoot().contentToTSVString()); } } catch (Exception e) { System.err.println("An unexpected error occurred: " + e.getMessage()); } } } ``` Or pass custom flight address: ``` SpiceClient client = SpiceClient.builder() .withFlightAddress(new URI("grpc://my_remote_spice_instance:50051")) .build(); ``` ### Connection retry[​](#connection-retry "Direct link to Connection retry") The `SpiceClient` implements connection retry mechanism (3 attempts by default). The number of attempts can be configured with `withMaxRetries`: ``` SpiceClient client = SpiceClient.builder() .withMaxRetries(5) // Setting to 0 will disable retries .build(); ``` Retries are performed for connection and system internal errors. It is the SDK user's responsibility to properly handle other errors, for example RESOURCE\_EXHAUSTED (HTTP 429). ### Parameterized Queries[​](#parameterized-queries "Direct link to Parameterized Queries") The SDK supports parameterized queries using ADBC (v0.5.0+). Use `queryWithParams()` for queries with user input to prevent SQL injection: ``` import org.apache.arrow.vector.VectorSchemaRoot; import org.apache.arrow.vector.ipc.ArrowReader; import ai.spice.SpiceClient; public class Example { public static void main(String[] args) { try (SpiceClient client = SpiceClient.builder().build()) { // Query with automatic type inference ArrowReader reader = client.queryWithParams( "SELECT * FROM taxi_trips WHERE trip_distance > $1 LIMIT 10", 5.0); // Double is inferred as Float64 while (reader.loadNextBatch()) { VectorSchemaRoot root = reader.getVectorSchemaRoot(); System.out.println(root.contentToTSVString()); } reader.close(); } catch (Exception e) { System.err.println("Error: " + e.getMessage()); } } } ``` For explicit type control, use the `Param` class: ``` import ai.spice.Param; ArrowReader reader = client.queryWithParams( "SELECT * FROM orders WHERE order_id = $1 AND amount >= $2", Param.int64(12345), Param.decimal128(new BigDecimal("99.99"), 10, 2)); ``` For more details, see [Parameterized Queries](/docs/next/features/query-federation/parameterized-queries). ### Memory Configuration[​](#memory-configuration "Direct link to Memory Configuration") The `SpiceClient` uses an Arrow `RootAllocator` for managing off-heap memory. By default, it uses all available memory. You can configure the memory limit using megabytes: ``` SpiceClient client = SpiceClient.builder() .withArrowMemoryLimitMB(1024) // 1GB limit .build(); ``` ### Spice.ai Runtime commands[​](#spiceai-runtime-commands "Direct link to Spice.ai Runtime commands") #### Accelerated dataset refresh[​](#accelerated-dataset-refresh "Direct link to Accelerated dataset refresh") Use `refresh` method to perform [Accelerated Dataset](/docs/next/components/data-accelerators) refresh. See full [dataset refresh example](https://github.com/spiceai/spice-java/blob/trunk/src/main/java/ai/spice/example/ExampleDatasetRefreshSpiceOSS.java). ``` SpiceClient client = SpiceClient.builder() .. .build(); client.refresh("taxi_trips") ``` --- # JavaScript SDK ## JavaScript SDK for Spice.ai[​](#javascript-sdk-for-spiceai "Direct link to JavaScript SDK for Spice.ai") [github.com/spiceai/spice.js](https://github.com/spiceai/spice.js) Parameterized Queries Parameterized queries are supported via the `sql()` method with a `parameters` option. See [Parameterized Queries](/docs/next/features/query-federation/parameterized-queries) for more information. ### Install[​](#install "Direct link to Install") * npm * yarn * pnpm ``` npm i @spiceai/spice ``` ``` yarn add @spiceai/spice ``` ``` pnpm add @spiceai/spice ``` ### Connect to spice runtime[​](#connect-to-spice-runtime "Direct link to Connect to spice runtime") Create a `SpiceClient` using default configuration: ``` import { SpiceClient } from '@spiceai/spice'; const main = async () => { const spiceClient = new SpiceClient(); const table = await spiceClient.sql( 'SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10;' ); console.table(table.toArray()); }; main(); ``` Or pass custom flight address: ``` const spiceClient = new SpiceClient({ flightUrl: 'my_remote_spice_instance:50051' }); ``` --- # Python SDK ## Python SDK for Spice.ai[​](#python-sdk-for-spiceai "Direct link to Python SDK for Spice.ai") [github.com/spiceai/spicepy](https://github.com/spiceai/spicepy) Parameterized Queries Parameterized queries are supported via the `query_with_params()` method. Install with `pip install spicepy[params]` for ADBC support. See [Parameterized Queries](/docs/next/features/query-federation/parameterized-queries) and [ADBC](/docs/next/api/adbc) for more information. ### Install[​](#install "Direct link to Install") ``` pip install git+https://github.com/spiceai/spicepy@v3.1.0 ``` ### Connect to a local Spice Runtime[​](#connect-to-a-local-spice-runtime "Direct link to Connect to a local Spice Runtime") By default, the Python SDK will connect to a locally running Spice Runtime: ``` from spicepy import Client client = Client() data = client.query( 'SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10;', timeout=5*60 ) pd = data.read_pandas() ``` ### Connect to Spice Cloud[​](#connect-to-spice-cloud "Direct link to Connect to Spice Cloud") To connect to Spice Cloud, specify the required Flight and HTTP URLs to connect to Spice Cloud: ``` from spicepy import Client from spicepy.config import ( DEFAULT_FLIGHT_URL, DEFAULT_HTTP_URL ) client = Client( api_key="", flight_url=DEFAULT_FLIGHT_URL, http_url=DEFAULT_HTTP_URL ) data = client.query( 'SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10;', timeout=5*60 ) pd = data.read_pandas() ``` ### Connect to a remote Spice Runtime[​](#connect-to-a-remote-spice-runtime "Direct link to Connect to a remote Spice Runtime") By specifying a custom Flight and HTTP URL, the Python SDK can connect to a remote Spice Runtime - for example, a centralised Spice Runtime instance. Example code: ``` from spicepy import Client client = Client( flight_url="grpc://your-remote-spice-runtime-host:50051", http_url="http://your-remote-spice-runtime-host:8090" ) data = client.query( 'SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10;', timeout=5*60 ) pd = data.read_pandas() ``` ### Refresh an accelerated dataset[​](#refresh-an-accelerated-dataset "Direct link to Refresh an accelerated dataset") The SDK supports refreshing an accelerated dataset with the `Client.refresh_dataset()` method. Refresh a dataset by calling this method on an instance of an SDK client: ``` from spicepy import Client, RefreshOpts client = Client() client.refresh_dataset("taxi_trips", None) # refresh with no refresh options client.refresh_dataset("taxi_trips", RefreshOpts(refresh_sql="SELECT * FROM taxi_trips LIMIT 10")) # refresh with overridden refresh SQL # RefreshOpts support all refresh parameters RefreshOpts( refresh_sql="SELECT * FROM taxi_trips LIMIT 10", refresh_mode="full", refresh_jitter_max="1m" ) ``` --- # Rust SDK ## Rust SDK for Spice.ai[​](#rust-sdk-for-spiceai "Direct link to Rust SDK for Spice.ai") [github.com/spiceai/spice-rs](https://github.com/spiceai/spice-rs) ### Install[​](#install "Direct link to Install") ``` cargo add spiceai ``` ### Connect to Spice Runtime[​](#connect-to-spice-runtime "Direct link to Connect to Spice Runtime") Create a `SpiceClient` using default configuration: ``` use spiceai::ClientBuilder; #[tokio::main] async fn main() { let client = ClientBuilder::new() .build() .await .unwrap(); let data = client.query( "SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10" ).await; } ``` Or pass a custom flight address: ``` use spiceai::ClientBuilder; #[tokio::main] async fn main() { let client = ClientBuilder::new() .flight_url("http://my_remote_spice_instance:50051") .build() .await .unwrap(); let data = client.query( "SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10" ).await; } ``` ### Parameterized Queries[​](#parameterized-queries "Direct link to Parameterized Queries") Use `query_with_params()` for parameterized queries to prevent SQL injection and improve performance (v3.0.0+): ``` use spiceai::ClientBuilder; use arrow::array::{Float64Array, StringArray}; use arrow::record_batch::RecordBatch; use std::sync::Arc; #[tokio::main] async fn main() { let client = ClientBuilder::new() .build() .await .unwrap(); // Create parameters as a RecordBatch let min_price = Float64Array::from(vec![10.0]); let category = StringArray::from(vec!["electronics"]); let params = RecordBatch::try_from_iter(vec![ ("1", Arc::new(min_price) as _), ("2", Arc::new(category) as _), ]).unwrap(); let data = client.query_with_params( "SELECT * FROM products WHERE price > $1 AND category = $2", Some(params) ).await; } ``` For more details, see [Parameterized Queries](/docs/next/features/query-federation/parameterized-queries). --- # Troubleshooting Spice Spice provides a number of methods to support debugging Runtime operations, including capturing verbose logs and reviewing task history in SQL queries or AI completions. For hands-on examples, see the [Spice.ai Cookbook](https://github.com/spiceai/cookbook) for recipes demonstrating common troubleshooting scenarios. ## Common Issues[​](#common-issues "Direct link to Common Issues") ### `spice` command not found after installation[​](#spice-command-not-found-after-installation "Direct link to spice-command-not-found-after-installation") The Spice binary directory is not in your `PATH`. Add it: ``` # Add to your shell profile (.zshrc, .bashrc, etc.) for persistence export PATH="$PATH:$HOME/.spice/bin" ``` Verify with `spice version`. ### Dataset fails to load or refresh[​](#dataset-fails-to-load-or-refresh "Direct link to Dataset fails to load or refresh") Check the runtime output for error messages. Common causes: * **Incorrect credentials**: Verify connection parameters (`pg_host`, `pg_user`, `pg_pass`, etc.) and that secrets are configured correctly. Use `--verbose` to see secret resolution logs. * **Network connectivity**: Ensure the Spice runtime can reach the data source. Test with `curl` or `psql` from the same host. * **Schema mismatch**: If the source schema changed since the last refresh (for example, columns were added, removed, or types changed), the accelerated table will fail to update. Spice infers the dataset schema at startup and intentionally blocks runtime schema evolution to protect against unintentional or breaking changes. Restart the runtime so Spice re-infers the schema from the source. See [Schema Inference](/docs/next/components/data-connectors#schema-inference) for details. Review the `task_history` table for detailed error messages: ``` SELECT task, error_message FROM runtime.task_history WHERE error_message IS NOT NULL ORDER BY start_time DESC LIMIT 5; ``` ### Slow query performance[​](#slow-query-performance "Direct link to Slow query performance") * **Check if acceleration is enabled**: Unaccelerated datasets query the remote source directly, adding network latency. Add `acceleration: enabled: true` to the dataset configuration. * **Review the query plan**: Run `EXPLAIN` before the query to verify it executes against the local accelerator and not the remote source. * **Check cache status**: For repeated queries, verify caching is active by inspecting the `Results-Cache-Status` HTTP header. A `MISS` on repeated identical queries may indicate a low `item_ttl`. ### AI chat returns incorrect or empty results[​](#ai-chat-returns-incorrect-or-empty-results "Direct link to AI chat returns incorrect or empty results") * **Verify model deployment**: Check the runtime logs for `Model [name] deployed, ready for inferencing`. If the model failed to load, review error messages in the logs. * **Check tools configuration**: Ensure `tools: auto` is set in the model params so the model can discover and query datasets. * **Use `spice trace ai_chat`**: Inspect the trace output to see which tools the model called and whether SQL queries succeeded or failed. ### Container OOM-killed despite a configured memory limit[​](#container-oom-killed-despite-a-configured-memory-limit "Direct link to Container OOM-killed despite a configured memory limit") `runtime.query.memory_limit` bounds the query execution pool, not the process. Accelerator caches, serialization buffers, embedded engine pools, and allocator retention sit outside it, so a process can be killed while the query pool still reports unused capacity. * **Compare the pools against actual memory**: if `process_resident_memory_bytes` is far above `query_memory_pool_used_bytes`, the memory is off-pool and lowering the query limit will not recover it. * **Count the per-dataset baseline**: accelerator caches are allocated per dataset with a floor that does not shrink with the container, so the idle footprint grows with dataset count regardless of data size. * **Check whether the limit was set explicitly**: when unset, the runtime derives the query limit with the per-dataset reservations subtracted. An explicit value opts out of that and must leave room for the baseline itself. * **Do not read a flat, high memory graph as a leak**: caches fill to their ceilings and allocators retain freed pages. What matters is whether it plateaus, and how far below the limit. * **Read the startup warnings before load testing**: the runtime warns at startup when the configured caches cannot fit beside the query pool, which detects an undersized deployment in seconds rather than days. * **Lowering the query limit is often the wrong lever**: bounding `runtime.query.max_concurrent_queries` reduces the peak directly, whereas lowering `runtime.query.memory_limit` shrinks each query's budget without reducing how many run at once. See [Managing Memory Usage](/docs/next/reference/memory) for the sizing model and validation guidance. ### Port conflicts on startup[​](#port-conflicts-on-startup "Direct link to Port conflicts on startup") If Spice fails to bind to its default ports (`8090` for HTTP, `50051` for Flight, `9090` for metrics), another process is already using that port. Override the ports: ``` spiced --http 0.0.0.0:3000 --flight 0.0.0.0:50052 --metrics 0.0.0.0:9091 ``` ## Verbose Logging[​](#verbose-logging "Direct link to Verbose Logging") Running `spiced` with `--verbose` produces immediate debug logs for diagnosing issues in real time. Running `spiced` with `--very-verbose` captures trace logs, which are useful for examining function outputs and other low-level details. The verbosity flags are also available for `spice run`, providing consistent behavior during local testing or production operation. Example `--verbose` output: ``` 2025-02-07T00:24:46.590576Z DEBUG runtime::secrets: Found secret replacement: Store name: secrets, Key: OPENAI_KEY, Span: 0..21 2025-02-07T00:24:46.592381Z DEBUG runtime::accelerated_table::refresh: Starting scheduled refresh 2025-02-07T00:24:46.592664Z DEBUG runtime::accelerated_table::refresh_task: Loading data for dataset runtime.task_history 2025-02-07T00:24:46.592844Z DEBUG runtime::accelerated_table: [retention] Evicting data for runtime.task_history where start_time < 2025-02-06T16:24:46+00:00... 2025-02-07T00:24:46.592859Z DEBUG runtime::accelerated_table: [retention] Expr BinaryExpr(BinaryExpr { left: Cast(Cast { expr: Column(Column { relation: None, name: "start_time" }), data_type: Timestamp(Nanosecond, None) }), op: Lt, right: Literal(TimestampNanosecond(1738859086592808104, None)) }) 2025-02-07T00:24:46.595460Z DEBUG runtime::accelerated_table::refresh_task: Loaded 0 rows for dataset runtime.task_history in 2ms. 2025-02-07T00:24:46.595591Z DEBUG runtime::accelerated_table::refresh_task_runner: Refresh task successfully completed for dataset runtime.task_history 2025-02-07T00:24:46.595623Z DEBUG runtime::accelerated_table::refresh: Received refresh task completion callback: Ok(()) 2025-02-07T00:24:46.599033Z DEBUG runtime::model::wrapper: Ignoring unknown default key: api_key 2025-02-07T00:24:46.599526Z DEBUG runtime::embeddings::table: Column 'content' does not have needed embeddings in base table. Will augment with model xl_embed. 2025-02-07T00:24:46.601038Z DEBUG runtime::accelerated_table: [retention] Evicted 0 records for runtime.task_history 2025-02-07T00:24:46.790058Z DEBUG runtime::datafusion: Creating accelerated table Dataset { <... dataset configuration ...> } 2025-02-07T00:24:46.869373Z INFO runtime::init::dataset: Dataset emails registered (*****), acceleration (duckdb), results cache enabled. 2025-02-07T00:24:46.870623Z DEBUG runtime::accelerated_table::refresh: Starting scheduled refresh 2025-02-07T00:24:46.870857Z INFO runtime::accelerated_table::refresh_task: Loading data for dataset emails 2025-02-07T00:24:49.087869Z INFO runtime::init::model: Model [openai] deployed, ready for inferencing 2025-02-07T00:24:52.196121Z INFO runtime::accelerated_table::refresh_task: Loaded 10 rows (1.77 MiB) for dataset emails in 5s 325ms. 2025-02-07T00:24:52.196303Z DEBUG runtime::accelerated_table::refresh_task_runner: Refresh task successfully completed for dataset emails 2025-02-07T00:24:52.196344Z DEBUG runtime::accelerated_table::refresh: Received refresh task completion callback: Ok(()) ``` Example `--very-verbose` output: ``` 2025-02-07T00:26:16.153386Z TRACE runtime::embeddings::table: For `EmbeddingTable`, additional embedding columns to compute: ["content"] 2025-02-07T00:26:16.433905Z TRACE runtime::embeddings::execution_plan: Of embedding columns: ["content"], only need to create embeddings for columns: ["content"] 2025-02-07T00:26:16.433990Z TRACE runtime::embeddings::execution_plan: Embedding column 'content' with model xl_embed 2025-02-07T00:26:16.434713Z DEBUG datafusion_table_providers::duckdb::write: Deleting all data from table. 2025-02-07T00:26:18.368600Z DEBUG hyper_util::client::legacy::pool: reuse idle connection for ("https", api.openai.com) 2025-02-07T00:26:18.369335Z DEBUG hyper_util::client::legacy::pool: pooling idle connection for ("https", api.openai.com) 2025-02-07T00:26:18.369652Z INFO runtime::init::model: Model [openai] deployed, ready for inferencing 2025-02-07T00:26:21.311004Z DEBUG hyper_util::client::legacy::pool: pooling idle connection for ("https", api.openai.com) 2025-02-07T00:26:21.426867Z TRACE runtime::embeddings::execution_plan: Successfully embedded column 'content' with chunking 2025-02-07T00:26:21.426981Z TRACE runtime::accelerated_table::refresh_task: [refresh] Received 10 rows for dataset: emails ``` For more information, view the [tracing documentation](/docs/next/cli/tracing) ## Use `spice trace` for task tracing[​](#use-spice-trace-for-task-tracing "Direct link to use-spice-trace-for-task-tracing") Use the `spice trace ai_chat` command to inspect processes involved in generating AI chat responses. The `spice trace` command supports tracing any task type, like `spice trace sql_query`: ``` [d06e1fb508e009eb] ( 1.84ms) sql_query ``` This step is helpful for reviewing any tool usage or tasks invoked during AI completions, providing more information on the steps an AI took to arrive at a result. For example, running a chat using the `taxi_trips` dataset: ``` Using model: openai chat> When was the last taxi trip completed? The last taxi trip was completed on January 12, 2024, at 12:49:45 PM. Time: 4.24s (first token 3.85s). Tokens: 826. Prompt: 793. Completion: 33 (85.50/s). ``` To inspect the last chat, run `spice trace ai_chat`, producing an AI chat trace output: ``` [8153c6563c7f9d88] ( 4234.38ms) ai_chat ├── [8656eaacb6c7a57d] ( 0.13ms) tool_use::list_datasets ├── [3873d8257d8ea30c] ( 4233.47ms) ai_completion ├── [02f2def1712f1473] ( 1.11ms) tool_use::sql │ └── [1e4e5f4e79e74e5e] ( 0.91ms) sql_query ├── [8d1db4d4db80c021] ( 3185.46ms) ai_completion ├── [0c7b421812ca8180] ( 0.17ms) tool_use::table_schema ├── [a49553aca9f19384] ( 2166.24ms) ai_completion ├── [ec69ce81b3d71b1a] ( 3.38ms) tool_use::sql │ └── [24769e9b068656ed] ( 3.33ms) sql_query └── [6ff16c04ecf6f6ff] ( 961.39ms) ai_completion ``` In this example, the trace logs show that the model attempted an SQL query twice - getting the first query incorrect, due to a syntax issue in the SQL the model generated. Before running the second query, it retrieved the table schema - which it used for the second query to successfully retrieve the last taxi trip time. For more information, view the [`spice trace` documentation](/docs/next/cli/reference/trace). ## Reviewing the Task History[​](#reviewing-the-task-history "Direct link to Reviewing the Task History") Query the `task_history` table to review completed tasks handled by the Runtime. Results in this table can include tasks for accelerator refresh, SQL queries, text embedding, AI calls, and more. Reviewing this table can provide information on SQL query issues or other processes that may produce errors. Example `task_history` query: ``` select start_time, end_time, task, captured_output, error_message from runtime.task_history; ``` The `task_history` table also includes start and end times, including execution duration and any error messages during the operation. An example `task_history` output with a failed SQL query: ``` +-------------------------------+-------------------------------+---------------------+-----------------+-------------------------------------------------------------------+ | start_time | end_time | task | captured_output | error_message | +-------------------------------+-------------------------------+---------------------+-----------------+-------------------------------------------------------------------+ | 2025-02-07T00:29:13.429351004 | 2025-02-07T00:29:13.432404760 | accelerated_refresh | | | | 2025-02-07T00:29:13.429022167 | 2025-02-07T00:29:13.432472389 | accelerated_refresh | | | | 2025-02-07T00:29:19.313382657 | 2025-02-07T00:29:19.313648021 | sql_query | | Error during planning: table 'spice.public.not_a_table' not found | +-------------------------------+-------------------------------+---------------------+-----------------+-------------------------------------------------------------------+ ``` For more information, view the [task history documentation](/docs/next/reference/task_history) ## Logging Additional Captured Output[​](#logging-additional-captured-output "Direct link to Logging Additional Captured Output") Set `runtime.task_history.captured_output` to `truncated` to store summarized SQL query information in the `captured_output` `task_history` column. Enable captured output by running `spiced` with `--set-runtime task_history.captured_output=truncated` or adjusting the Spicepod parameter `runtime.task_history.captured_output`. Example Spicepod excerpt: ``` runtime: task_history: captured_output: truncated ``` Example captured output: ``` +-------------------------------+-------------------------------+---------------------+-------------------------------+---------------+ | start_time | end_time | task | captured_output | error_message | +-------------------------------+-------------------------------+---------------------+-------------------------------+---------------+ | 2025-02-07T00:17:41.999469156 | 2025-02-07T00:17:42.002922183 | accelerated_refresh | | | | 2025-02-07T00:17:42.007874330 | 2025-02-07T00:17:44.512541448 | health | | | | 2025-02-07T00:17:44.510484956 | 2025-02-07T00:17:48.889947970 | text_embed | | | | 2025-02-07T00:17:42.278785968 | 2025-02-07T00:17:48.913729643 | accelerated_refresh | | | | 2025-02-07T00:17:54.717312222 | 2025-02-07T00:17:54.728507220 | sql_query | [{"subject":"Hello, world!"}] | | +-------------------------------+-------------------------------+---------------------+-------------------------------+---------------+ ``` ## Capturing SQL Query Plans in Task History[​](#capturing-sql-query-plans-in-task-history "Direct link to Capturing SQL Query Plans in Task History") Configure Spice to automatically capture SQL query plans (`EXPLAIN` or `EXPLAIN ANALYZE`) in the task history. This feature captures query execution information asynchronously, storing results in the `captured_output` column for later analysis without impacting query performance. Configure plan capture using the `runtime.task_history` settings: ``` runtime: task_history: captured_plan: explain analyze min_sql_duration: 5s min_plan_duration: 10s ``` ### Configuration Parameters[​](#configuration-parameters "Direct link to Configuration Parameters") * **`captured_plan`**: Determines the type of plan captured: * `none` (default): No plans are captured * `explain`: Captures logical and physical query plans * `explain analyze`: Captures plans with actual execution metrics * **`min_sql_duration`**: Minimum query duration before plan capture. Only queries exceeding this threshold generate a plan. * **`min_plan_duration`**: Minimum plan execution duration before storage. Plans that execute faster than this threshold are discarded. ### Query Examples[​](#query-examples "Direct link to Query Examples") Review captured plans for slow queries: ``` SELECT start_time, execution_duration_ms, SUBSTRING(input, 1, 80) AS query_preview, captured_output FROM spice.runtime.task_history WHERE task = 'sql_query' AND captured_output IS NOT NULL ORDER BY execution_duration_ms DESC LIMIT 5; ``` Plans are captured asynchronously after query completion, ensuring query execution proceeds without blocking. For more details, view the [task history documentation](/docs/next/reference/task_history#sql-query-plan-capture). ## SQL Explain Plans[​](#sql-explain-plans "Direct link to SQL Explain Plans") Spice supports generating `EXPLAIN` plans, which can be used to debug SQL that may not be producing the correct result. An explain plan in Spice can provide information about the data sources the SQL will be executed against, including the re-written SQL that will be executed against that data source (if applicable). For example, an explain plan could be used to debug the SQL that is generated and executed against separate federated sources during an SQL query which joins their results (like a PostgreSQL database and a MySQL database). Generate an explain plan by adding the `EXPLAIN` keyword before the query: ``` explain select * from taxi_trips; ``` ``` +---------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | plan_type | plan | +---------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | logical_plan | BytesProcessedNode | | | Federated | | | Projection: taxi_trips.VendorID, taxi_trips.tpep_pickup_datetime, taxi_trips.tpep_dropoff_datetime, taxi_trips.passenger_count, taxi_trips.trip_distance, taxi_trips.RatecodeID, taxi_trips.store_and_fwd_flag, taxi_trips.PULocationID, taxi_trips.DOLocationID, taxi_trips.payment_type, taxi_trips.fare_amount, taxi_trips.extra, taxi_trips.mta_tax, taxi_trips.tip_amount, taxi_trips.tolls_amount, taxi_trips.improvement_surcharge, taxi_trips.total_amount, taxi_trips.congestion_surcharge, taxi_trips.Airport_fee | | | TableScan: taxi_trips projection=[VendorID, tpep_pickup_datetime, tpep_dropoff_datetime, passenger_count, trip_distance, RatecodeID, store_and_fwd_flag, PULocationID, DOLocationID, payment_type, fare_amount, extra, mta_tax, tip_amount, tolls_amount, improvement_surcharge, total_amount, congestion_surcharge, Airport_fee] | | physical_plan | BytesProcessedExec | | | SchemaCastScanExec | | | VirtualExecutionPlan name=duckdb compute_context=:memory: sql=SELECT taxi_trips.VendorID, taxi_trips.tpep_pickup_datetime, taxi_trips.tpep_dropoff_datetime, taxi_trips.passenger_count, taxi_trips.trip_distance, taxi_trips.RatecodeID, taxi_trips.store_and_fwd_flag, taxi_trips.PULocationID, taxi_trips.DOLocationID, taxi_trips.payment_type, taxi_trips.fare_amount, taxi_trips.extra, taxi_trips.mta_tax, taxi_trips.tip_amount, taxi_trips.tolls_amount, taxi_trips.improvement_surcharge, taxi_trips.total_amount, taxi_trips.congestion_surcharge, taxi_trips.Airport_fee FROM taxi_trips rewritten_sql=SELECT "taxi_trips"."VendorID", "taxi_trips"."tpep_pickup_datetime", "taxi_trips"."tpep_dropoff_datetime", "taxi_trips"."passenger_count", "taxi_trips"."trip_distance", "taxi_trips"."RatecodeID", "taxi_trips"."store_and_fwd_flag", "taxi_trips"."PULocationID", "taxi_trips"."DOLocationID", "taxi_trips"."payment_type", "taxi_trips"."fare_amount", "taxi_trips"."extra", "taxi_trips"."mta_tax", "taxi_trips"."tip_amount", "taxi_trips"."tolls_amount", "taxi_trips"."improvement_surcharge", "taxi_trips"."total_amount", "taxi_trips"."congestion_surcharge", "taxi_trips"."Airport_fee" FROM "taxi_trips" | | | | +---------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ ``` ## Debugging Sandbox Container[​](#debugging-sandbox-container "Direct link to Debugging Sandbox Container") If you try to open a shell in a Spice container the way you would with most images, it fails: **On your machine:** ``` $ kubectl exec -it my-spicepod -- bash error: exec: "bash": executable file not found in $PATH ``` The error comes from *inside* the container: `kubectl` reached the pod successfully, then found no `bash` there to execute. This is expected, not a broken container. The sections below explain why, and give three ways to debug depending on what you need to look at. ### Why there is no shell[​](#why-there-is-no-shell "Direct link to Why there is no shell") The published Spice image is built `FROM scratch` — an empty base image with no Linux userspace at all. It contains only: * the `spiced` binary * the shared libraries `spiced` links against * CA certificates and timezone data There is no `bash`, no `sh`, not even `ls`, and no package manager — so nothing can be installed into a running container either. `spiced` also runs as the unprivileged user `65534` (`nobody`), whose login shell is set to `/usr/sbin/nologin`. This is deliberate: an image with no shell and no packages has a far smaller attack surface and far fewer CVEs to patch. The trade-off is that debugging needs the techniques below instead of `kubectl exec ... -- bash`. For background on how this image is built, see the [Docker Sandbox Guide](/docs/deployment/docker/sandbox). ### Choosing an approach[​](#choosing-an-approach "Direct link to Choosing an approach") | What you need to do | Use | Requirements | | ------------------------------------------------------------ | ----------------------------------------------------------------------- | ---------------------------------------------------------------- | | Run SQL queries against the running runtime | [SQL REPL](#sql-repl) | None — works out of the box | | Inspect files, processes, or the network of a Kubernetes pod | [Ephemeral container](#debug-kubernetes-pods-with-ephemeral-containers) | Kubernetes v1.25+, and permission to create ephemeral containers | | Get an interactive shell against a local Docker container | [busybox shell](#debugging-with-a-shell) | Docker, and the container started with an extra volume | Start with the SQL REPL if the question is about data or queries. Reach for an ephemeral container when the question is about the environment — config files, mounted volumes, DNS, or connectivity. ### Where commands run[​](#where-commands-run "Direct link to Where commands run") Debugging a container means moving between machines, and the same command can behave differently — or fail — depending on where it is typed. Three contexts appear below, and **every command block states which one it belongs to**: | Context | Where it actually executes | How you get there | | ------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------- | | **Your machine** | Your own workstation or CI runner — wherever `kubectl` and `docker` are installed. These commands talk *to* the cluster or Docker daemon, not inside it. | The default; no extra step | | **Inside the Spice container** | The sandbox container running `spiced`. No shell here, so only `spiced` itself, or a mounted binary such as `busybox`, can be run. | `kubectl exec` / `docker exec` | | **Inside the debug container** | A *separate* temporary container in the same pod, with its own filesystem and its own tools. Spice's files are **not** at their usual paths here. | `kubectl debug` drops you straight in | Two consequences that catch people out: * `kubectl` and `docker` commands are typed on **your machine**, but the part after `--` (or after the container name) executes **inside the container**. In `kubectl exec -it my-pod -- spiced --repl`, `kubectl` runs locally while `spiced --repl` runs in the container — which is why `localhost` in that command means the container's own localhost, not your machine's. * `kubectl debug` does **not** put you inside the Spice container. It puts you in a new debug container beside it. That is why inspecting Spice's files from there needs the `/proc/1/root` prefix described below. ### SQL REPL[​](#sql-repl "Direct link to SQL REPL") The REPL needs no shell, because `spiced` is itself the binary being executed. **On your machine** — the `spiced --repl` part after the container name executes **inside the Spice container**: ``` # Docker docker exec -it spiced --repl # Kubernetes kubectl exec -it -- spiced --repl ``` Because `spiced --repl` runs inside the container, it connects to that container's own `http://localhost:50051` Flight endpoint — attaching to the runtime already serving there. The interactive SQL prompt that follows is therefore executing queries **inside the deployment**, not locally. This is the quickest way to check whether a dataset loaded, inspect a schema, or reproduce a slow query from inside the deployment. note Running `spiced --repl` on **your machine** instead is a different thing entirely: it would try to reach a runtime on your own `localhost:50051`. To query a remote runtime from your workstation without `kubectl exec`, use `spice sql --endpoint ` (it accepts `http://`, `https://`, `grpc://`, and `grpc+tls://`), pointing at an address the runtime is reachable on — for example one published by `kubectl port-forward`. ### Debug Kubernetes Pods with Ephemeral Containers[​](#debug-kubernetes-pods-with-ephemeral-containers "Direct link to Debug Kubernetes Pods with Ephemeral Containers") An [ephemeral container](https://kubernetes.io/docs/concepts/workloads/pods/ephemeral-containers/) is a temporary extra container that Kubernetes adds to a pod that is already running. Because it brings its own image, it can supply all the tools the Spice image lacks — while the Spice process keeps running untouched. The pod is **not** restarted and no configuration is changed. This is the recommended way to debug Spice on Kubernetes. #### 1. Find the pod and container names[​](#1-find-the-pod-and-container-names "Direct link to 1. Find the pod and container names") **On your machine:** ``` # List pods in the namespace kubectl get pods -n # List the container names inside the pod kubectl get pod -n -o jsonpath='{.spec.containers[*].name}' ``` The Helm chart names the Spice container `spiceai`. Run the second command rather than assuming, since a custom manifest may name it something else. #### 2. Start the debug container[​](#2-start-the-debug-container "Direct link to 2. Start the debug container") **On your machine:** ``` kubectl debug -it my-spicepod-name \ -n my-spicepod-ns \ --image=ubuntu:24.04 \ --target=spiceai \ --profile=sysadmin ``` This command returns with an interactive prompt, and **from here on you are inside the debug container** — a different filesystem from both your machine and the Spice container. What each flag does: | Flag | Meaning | | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `-it` | Attach an interactive terminal, as with `kubectl exec -it`. | | `--image` | The image providing the debugging tools, and therefore the only source of tools available once inside. `ubuntu:24.04` gives a familiar shell and `apt` to install more; any image works. | | `--target` | The container in the pod to attach to, from step 1. This is what makes the Spice process and its filesystem visible — omit it and you get an isolated container that sees nothing of the runtime. | | `--profile=sysadmin` | Applies the `sysadmin` [static debugging profile](https://kubernetes.io/docs/tasks/debug/debug-application/debug-running-pod/#static-profile), which grants the elevated capabilities needed to inspect another container's files and processes. | #### 3. Inspect the runtime[​](#3-inspect-the-runtime "Direct link to 3. Inspect the runtime") The debug container has **its own** filesystem — from `ubuntu:24.04`, not from Spice — so `/app` here is empty, and Spice's files are not where you might expect. Because `--target` shares the process namespace, `spiced` is visible as PID 1, and Linux exposes any process's root filesystem at `/proc//root`. So everything belonging to Spice is reachable under **`/proc/1/root`**: **Inside the debug container** (reading the Spice container's files): ``` # The Spicepod definition the runtime actually loaded cat /proc/1/root/app/spicepod.yaml # The runtime's working directory, including acceleration files ls -l /proc/1/root/app # Confirm which process is PID 1 and how it was invoked cat /proc/1/cmdline | tr '\0' ' ' ``` The `/proc/1/root` prefix is what makes the difference: `ls /app` reads the *debug* container's empty directory, while `ls /proc/1/root/app` reads the *Spice* container's real working directory. Dropping the prefix is the most common source of confusion — it does not error, it just shows you the wrong container's filesystem. Networking is the exception to that rule. All containers in a pod share a single network namespace, so network tools run **inside the debug container** already see exactly what the runtime sees — no `/proc/1/root` prefix applies, and none is needed. Note that the tools available are whatever the `--image` you chose provides, not what Spice provides. `ubuntu:24.04` is a minimal image, so install what you need first — this is safe, because it modifies only the throwaway debug container: **Inside the debug container:** ``` # Install network tools into the debug container (not into Spice) apt update && apt install -y curl dnsutils iproute2 # Resolve and reach a data source exactly as the runtime would dig my-postgres.my-namespace.svc.cluster.local curl -v telnet://my-postgres.my-namespace.svc.cluster.local:5432 # The runtime's own listening sockets: HTTP 8090, metrics 9090, Flight 50051 ss -ltnp ``` Images purpose-built for network debugging, such as `nicolaka/netshoot`, bundle these tools and skip the install step. When finished, type `exit` to leave the debug container and return to **your machine**. note Ephemeral containers require Kubernetes v1.25 or later, and creating one requires permission on the `pods/ephemeralcontainers` subresource — a `Forbidden` error means the account lacks it, not that the command is wrong. An ephemeral container cannot be removed from a pod once added; it remains in the pod spec until the pod is replaced. Some hardened clusters also reject `--profile=sysadmin`; if so, try `--profile=general` (fewer capabilities, so cross-container file inspection may not work). ### Debugging with a shell[​](#debugging-with-a-shell "Direct link to Debugging with a shell") For local Docker debugging, a shell can be added from outside the image. [busybox](https://busybox.net/) ships a single *statically compiled* binary containing dozens of standard tools — because it depends on no system libraries, it runs inside the sandbox image even though that image has no userspace of its own. Mounting it as a volume supplies a shell without rebuilding the image. This technique is for Docker. On Kubernetes, use an [ephemeral container](#debug-kubernetes-pods-with-ephemeral-containers) instead — it needs no volume and no restart. Note that the volume must be mounted when the container is **created**, so this requires starting a new container rather than attaching to a running one. **On your machine** (the Docker host) — all four commands: ``` # Create a volume for the busybox binary docker volume create busybox # Copy the busybox binary to the volume docker run --rm -v busybox:/data busybox:stable-musl sh -c "mkdir -p /data && cp /bin/busybox /data/busybox" # Run the Spice.ai container with the busybox binary mounted, ensure that any other volumes are mounted as well (i.e. for spicepod) docker run -v busybox:/busy -v :/app/spicepod -d --name spiceai-debug spiceai/spiceai:latest # Exec into the container — the shell that follows runs INSIDE the Spice container docker exec -it spiceai-debug /busy/busybox sh ``` That last command hands you a shell **inside the Spice container** itself — unlike `kubectl debug`, there is no separate debug container here, so paths such as `/app` are already Spice's own and need no `/proc/1/root` prefix. Because the busybox tools are not on `PATH`, every command in that shell must be prefixed with `/busy/busybox`: **Inside the Spice container:** ``` /busy/busybox ls -l /app /busy/busybox cat /app/spicepod.yaml ``` --- # Spice.ai Use Cases Spice supports a range of use cases across data infrastructure, search, and AI. Each use case below describes a specific scenario with architecture guidance, configuration examples, and links to relevant cookbook recipes. For hands-on examples, see the [Spice.ai Cookbook](https://github.com/spiceai/cookbook). ## Data Federation, Acceleration, and SQL Query[​](#data-federation-acceleration-and-sql-query "Direct link to Data Federation, Acceleration, and SQL Query") * [**Reverse-ETL**](/docs/next/use-cases/data/reverse-etl): Serve data from warehouses and data lakes to operational systems, applications, and dashboards, eliminating complex pipelines. * [**ETL-free Workflows and Data Migrations**](/docs/next/use-cases/data/etl-free-workflows): Enable data migrations and workflows without ETL federating legacy and modern systems for faster time-to-market and lower operational overhead. * [**Database CDN**](/docs/next/use-cases/data/database-cdn): Locally replicate working sets of data for operational applications, caching dynamic data for high performance, low-latency, and resilience. * [**Data Mesh**](/docs/next/use-cases/data/data-mesh): Unified data access across disparate sources with acceleration. * [**Object-Store Native Database**](/docs/next/use-cases/data/object-store-data-engine): Federates, accelerates, and queries object-store data for real-time data access without centralized warehouses. ## Caching[​](#caching "Direct link to Caching") * [**Write-Through Cache**](/docs/next/use-cases/caching/write-through-cache): Write data through Spice to both a local accelerator and the upstream source, keeping both layers consistent. * [**Read-Through Cache**](/docs/next/use-cases/caching/read-through-cache): Fetch data from the upstream source on cache miss, with stale-while-revalidate and stale-if-error semantics. * [**SQL/Database Cache**](/docs/next/use-cases/caching/sql-database-cache): Cache SQL database tables locally with acceleration and cache SQL query results in memory. * [**S3 Cache**](/docs/next/use-cases/caching/s3-cache): Cache S3 and object store data locally with smart refresh skip for unchanged files. * [**HTTP Cache**](/docs/next/use-cases/caching/http-cache): Cache HTTP API responses locally with request filtering, TTL, and stale-while-revalidate support. ## Search and Retrieval[​](#search-and-retrieval "Direct link to Search and Retrieval") * [**Enterprise Search**](/docs/next/use-cases/search/enterprise-search): Semantic and full-text-search search with hybrid vector and keyword capabilities. * [**Object-Store Native Search**](/docs/next/use-cases/search/object-store-search-engine): Enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. * [**Simplifying Real-Time Data Collection and Search**](/docs/next/use-cases/search/data-collection-and-search): Processes streaming and static data with integrated search for real-time insights in health-tech, focusing on application logic. ## Retrieval-Augmented-Generation (RAG)[​](#retrieval-augmented-generation-rag "Direct link to Retrieval-Augmented-Generation (RAG)") * [**RAG for Contextual Applications**](/docs/next/use-cases/rag/applications): Combines structured and unstructured data for context-rich AI outputs in SaaS chatbots, improving user interactions. * [**RAG for AI-Powered Reporting**](/docs/next/use-cases/rag/reporting): Generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. ## AI Applications and Agents[​](#ai-applications-and-agents "Direct link to AI Applications and Agents") * [**Real-Time Decision-Making for Intelligent Applications**](/docs/next/use-cases/ai/real-time-decision-making): Powers instant, context-aware decisions for security applications by grounding AI in federated, low-latency datasets. * [**Edge-Enabled AI Applications and Agents**](/docs/next/use-cases/ai/edge-ai): Deploys AI applications across cloud and edge for low-latency decisions in security IoT use cases. * [**Tool-Augmented AI with Model Context Protocol Server**](/docs/next/use-cases/ai/federated-mcp-server): Extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. * [**Agentic AI Applications and Agents**](/docs/next/use-cases/ai/agentic-apps): Builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. * [**Multi-Tenant AI Agents**](/docs/next/use-cases/ai/multi-tenant-agents): Deploy AI agents across many SaaS tenants with strict isolation and no per-tenant ETL pipelines. --- # AI Applications and Agents ## [📄️Real-Time Decision-Making](/docs/next/use-cases/ai/real-time-decision-making) [Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets.](/docs/next/use-cases/ai/real-time-decision-making) ## [📄️Edge AI](/docs/next/use-cases/ai/edge-ai) [Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases.](/docs/next/use-cases/ai/edge-ai) ## [📄️MCP Server](/docs/next/use-cases/ai/tool-calling-ai) [Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality.](/docs/next/use-cases/ai/tool-calling-ai) ## [📄️Federated MCP Client](/docs/next/use-cases/ai/federated-mcp-server) [Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, improving threat analysis.](/docs/next/use-cases/ai/federated-mcp-server) ## [📄️Object-Store AI Engine](/docs/next/use-cases/ai/object-store-ai-engine) [Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights.](/docs/next/use-cases/ai/object-store-ai-engine) ## [📄️Agentic AI Apps](/docs/next/use-cases/ai/agentic-apps) [Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making.](/docs/next/use-cases/ai/agentic-apps) ## [📄️Multi-Tenant AI Agents](/docs/next/use-cases/ai/multi-tenant-agents) [Deploy AI agents across many SaaS tenants with strict isolation and no per-tenant ETL pipelines.](/docs/next/use-cases/ai/multi-tenant-agents) --- # Agentic AI Applications and Agents Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making to improve user experiences and operational efficiency. Unlike generic AI agent frameworks (e.g., AutoGen, CrewAI) that often lack direct integration with enterprise data or real-time capabilities, Spice.ai combines federated data access, hybrid search, and AI inference to deliver autonomous agents that operate with enterprise-grade governance and low-latency performance. This makes it ideal for SaaS platforms requiring dynamic, context-driven automation in customer-facing workflows. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Federated Data Access**: Queries disparate data sources (e.g., Databricks, PostgreSQL, cloud storage) in real time, providing agents with a unified view of customer data, unlike siloed frameworks that limit context. * **Hybrid Search**: Integrates vector similarity search (VSS) for semantic understanding (e.g., user intent from support tickets) with keyword/BM25 search for precise data retrieval, enabling agents to make informed decisions. * **AI Gateway**: Powers agents with large language models (LLMs), supporting hosted (e.g., OpenAI) and local (e.g., Llama) models for privacy and cost efficiency, optimized for real-time inference in SaaS environments. * **Governance and Observability**: Uses Databricks Unity Catalog for compliance and provides end-to-end visibility into agent performance and data flows, ensuring reliability and auditability, unlike generic agent platforms. ## Example[​](#example "Direct link to Example") A SaaS customer success platform deploys a Spice.ai-powered agent to automate ticket resolution by querying real-time user data from PostgreSQL, historical support interactions from Databricks, and unstructured knowledge base articles via hybrid search. The agent uses LLM inference to generate personalized responses and escalate critical issues, reducing resolution time compared to manual processes or generic AI agents lacking deep data integration. The [Federated SQL Query recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) and [Vector-Based Search documentation](/docs/next/features/search) guide implementation of data access and search for agentic workflows. ## Benefits[​](#benefits "Direct link to Benefits") * **Automation**: Context-aware agents reduce manual effort, improving efficiency in SaaS customer workflows. * **Responsiveness**: Real-time data and inference enable rapid, accurate decision-making, improving user satisfaction. * **Compliance**: Governed data access ensures alignment with SaaS regulatory requirements, fostering trust. ### Learn More[​](#learn-more "Direct link to Learn More") * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Vector and Hybrid Search**: [Documentation](/docs/next/features/search) and [Searching GitHub Files Recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md). * **AI Gateway**: [Documentation](/docs/next/features/large-language-models) and [Running Llama3 Locally Recipe](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md). * **Semantic Model**: [Documentation](/docs/next/features/semantic-model). * **Observability**: [Documentation](/docs/next/features/observability). --- # Edge-Enabled AI Applications and Agents Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases, ensuring rapid threat detection and response in distributed environments. Unlike cloud-centric AI platforms (e.g., AWS SageMaker, Google Vertex AI) that rely on constant connectivity and introduce latency, Spice.ai’s edge data materialization and local model inference enable real-time, resilient AI operations. This is critical for security applications requiring immediate action and functionality in low-connectivity scenarios, outperforming cloud-dependent solutions in speed and reliability. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Edge Acceleration**: Materializes data (e.g., sensor logs, event streams) at the edge for fast, local queries, minimizing latency compared to cloud-only inference, essential for security IoT responsiveness. * **Unified Queries**: Accesses cloud and edge data sources (e.g., Databricks, on-premises sensors) via federated SQL, simplifying distributed architectures for security deployments. * **Local Models**: Deploys lightweight AI models (e.g., NVIDIA NIM, OSS Llama) at the edge, reducing costs and ensuring data privacy, critical for sensitive security data in regulated environments. * **Resilience**: Maintains edge functionality during network disruptions, ensuring continuous threat monitoring and response, unlike cloud-reliant platforms vulnerable to outages. ## Example[​](#example "Direct link to Example") A security IoT system processes real-time sensor data from edge devices to detect unauthorized access attempts in a corporate facility, using local AI models to prioritize alerts without cloud dependency. This ensures instant threat detection and response, even during network outages, outperforming cloud-dependent systems that suffer from latency and connectivity issues. The [Running Llama3 Locally recipe](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md) demonstrates edge model deployment for such scenarios. ## Benefits[​](#benefits "Direct link to Benefits") * **Low Latency**: Edge processing delivers instant threat detection, critical for security IoT applications. * **Reliability**: Offline capabilities ensure continuous operation in distributed environments, improving system uptime. * **Privacy**: Local inference protects sensitive security data, aligning with compliance requirements in regulated industries. ### Learn More[​](#learn-more "Direct link to Learn More") * **AI Gateway**: [Documentation](/docs/next/features/large-language-models) and [Running Llama3 Locally Recipe](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md). * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md). --- # Federated MCP Client for Distributed Tool Ecosystems Spice.ai federates external Model Context Protocol (MCP) servers over Server-Sent Events (SSE) for scalable, tool-driven AI applications in security, improving threat analysis with distributed tool ecosystems. Unlike centralized AI orchestration platforms (e.g., Apache Airflow for AI workflows) that introduce complexity and latency, Spice.ai's federated MCP client approach supports modular, scalable tool integration with integrated data and AI capabilities. This makes it ideal for security applications requiring distributed, real-time threat intelligence, outperforming platforms with rigid, centralized architectures. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **MCP Federation**: Unifies outputs from multiple external MCP servers (e.g., threat intelligence, anomaly detection tools), enabling complex, distributed workflows without centralized orchestration, critical for security’s dynamic threat landscape. * **SSE Connectivity**: Simplifies integration with remote tools via Server-Sent Events, reducing networking overhead compared to custom API integrations, ensuring efficient communication in distributed security systems. * **Hybrid Access**: Combines MCP tool outputs with federated SQL queries and vector search (e.g., for log analysis or threat pattern matching), delivering comprehensive insights that surpass tool-only platforms. * **Scalability**: Distributes tool execution across cloud and edge environments, optimizing performance and resilience for global security operations. ## Example[​](#example "Direct link to Example") A security operations center federates external MCP servers for threat intelligence, behavioral analysis, and log parsing, delivering AI-driven threat reports enriched with real-time network data from Databricks. This streamlines incident response by unifying disparate tools, reducing response times compared to fragmented integrations, and enhances threat detection accuracy. The [Federated SQL Query recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) provides patterns adaptable to MCP federation workflows. ## Benefits[​](#benefits "Direct link to Benefits") * **Modularity**: Integrates diverse security tools into a cohesive workflow, improving flexibility and adaptability. * **Scalability**: Supports distributed architectures for global security operations, handling high-volume threat data. * **Efficiency**: Reduces integration complexity with standardized SSE connectivity, speeding up deployment. ### Learn More[​](#learn-more "Direct link to Learn More") * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Vector and Hybrid Search**: [Documentation](/docs/next/features/search) and [Searching GitHub Files Recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md). --- # Multi-Tenant AI Agents Spice serves as the data substrate for multi-tenant AI agents, giving each tenant real-time access to its own data sources with configurable isolation and no per-tenant ETL pipelines. A lightweight Rust runtime can be deployed once for all tenants, once per tenant, or in a hybrid topology, and every deployment exposes the same SQL, search, and AI inference APIs over HTTP, Arrow Flight, FlightSQL, ODBC, JDBC, and ADBC. Pipeline-per-integration architectures collapse at scale: every tenant brings its own schema, refresh cadence, and ownership, and ETL orchestration becomes the bottleneck. Spice federates queries across 40+ data sources and accelerates them locally, so adding a tenant is a Spicepod configuration change rather than a new pipeline. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Zero-ETL federation**: Query PostgreSQL, Snowflake, Databricks, S3, Kafka, and 25+ other sources through a single SQL interface. Onboarding a tenant is a dataset declaration, not a pipeline build. * **Configurable isolation**: Choose logical, config-level, or runtime-level tenant boundaries. Isolation becomes a deployment property rather than something enforced in application code. * **Sandboxed runtimes**: The Spice runtime is lightweight enough to deploy one instance per tenant or per agent, giving each agent its own sources, acceleration layers, and secrets. * **Local acceleration with CDC**: Materialize tenant working sets into Arrow, DuckDB, SQLite, or PostgreSQL accelerators and keep them current with change data capture. * **Declarative configuration**: The [`spicepod.yaml`](/docs/next/reference/spicepod) manifest defines datasets, models, secrets, and acceleration behavior, so tenant topology is version-controlled and auditable. ## Deployment Patterns[​](#deployment-patterns "Direct link to Deployment Patterns") Four patterns trade off operational simplicity against isolation strength. Most SaaS workloads converge on a hybrid configuration. ### Pattern 1: Query-Time Tenant Isolation[​](#pattern-1-query-time-tenant-isolation "Direct link to Pattern 1: Query-Time Tenant Isolation") One runtime serves many tenants. Datasets reference tenant-partitioned tables, and the application includes a tenant filter in every query. ``` version: v1 kind: Spicepod name: saas-shared datasets: - from: postgres:public.events name: events params: pg_host: db.shared.internal pg_port: '5432' pg_db: app pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true engine: arrow refresh_check_interval: 1s ``` This is the simplest and cheapest pattern to operate, but isolation is logical: correctness depends on every query path applying the tenant filter. **View-based variant**: Move tenant filtering into the Spicepod using [views](/docs/next/reference/spicepod/views), so agents query a tenant-specific view rather than constructing the filter at runtime. ``` datasets: - name: events from: postgres:public.events views: - name: view_tenant_abc sql: "select * from events where tenant='tenant_abc'" - name: view_tenant_xyz sql: "select * from events where tenant='tenant_xyz'" ``` ### Pattern 2: Config-Level Tenant Isolation[​](#pattern-2-config-level-tenant-isolation "Direct link to Pattern 2: Config-Level Tenant Isolation") One runtime, but each tenant gets its own dataset entry. The manifest is typically generated from tenant-onboarding metadata. ``` version: v1 kind: Spicepod name: saas-many-datasets datasets: - from: postgres:tenant_abc.events name: tenant_abc_events params: pg_host: db.pool.internal pg_db: app pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} - from: postgres:tenant_xyz.events name: tenant_xyz_events params: pg_host: db.pool.internal pg_db: app pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} ``` Tenant boundaries are explicit in configuration and easier to audit, but the manifest grows with tenant count and reload time and memory planning become operational concerns at large scale. ### Pattern 3: Runtime-Level Tenant Isolation[​](#pattern-3-runtime-level-tenant-isolation "Direct link to Pattern 3: Runtime-Level Tenant Isolation") Run a dedicated Spicepod per tenant and route requests using tenant context from auth or session claims. Each runtime has an independent manifest, compute envelope, and cache. ``` version: v1 kind: Spicepod name: tenant-abc datasets: - from: postgres:public.events name: events params: pg_host: tenant-abc-db.internal pg_db: app pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true engine: duckdb mode: file refresh_check_interval: 1s params: duckdb_file: /var/lib/spice/tenant-abc.db ``` This gives the clearest isolation model and per-tenant operational control, at the cost of higher operational overhead as tenant count grows. It fits regulated workloads that require strict data boundaries. ### Pattern 4: Hybrid Isolation[​](#pattern-4-hybrid-isolation "Direct link to Pattern 4: Hybrid Isolation") Large tenants run on dedicated Spicepods while the long tail shares a partitioned deployment. A tenant-aware router decides placement and clients query a single logical interface. ``` # router config (conceptual) tenants: - id: enterprise-abc spicepod: spicepod-tenant-abc - id: enterprise-xyz spicepod: spicepod-tenant-xyz default: spicepod: spicepod-shared ``` Hybrid isolation keeps the query interface stable while allowing tenant placement by policy, isolating high-load or regulated tenants on dedicated runtimes and using shared capacity for the long tail. It is the recommended starting point for variable SaaS workloads. ## Choosing a Pattern[​](#choosing-a-pattern "Direct link to Choosing a Pattern") The right shape depends on isolation requirements, tenant count, and query volume per tenant. | Requirement | Recommended Pattern | | ---------------------------------------------------- | ------------------- | | Small tenant count, uniform workload | Pattern 1 or 1-view | | Explicit config-level boundaries, auditable manifest | Pattern 2 | | Regulated tenants, strict per-tenant SLOs | Pattern 3 | | Power-law tenant distribution | Pattern 4 (hybrid) | ## Benefits[​](#benefits "Direct link to Benefits") * **Zero-ETL onboarding**: Add a tenant by declaring a dataset, not by building a pipeline. * **Deployment-level isolation**: Match isolation strength to business and compliance requirements without changing application code. * **Consistent interface**: Agents query SQL, search, and inference APIs the same way regardless of how tenants are partitioned underneath. ## Learn More[​](#learn-more "Direct link to Learn More") * [Multi-Tenancy for AI Agents without the Pipelines](https://spice.ai/blog/multi-tenancy-for-ai-agents-without-pipelines) — the engineering deep dive behind these patterns. * [Spicepod reference](/docs/next/reference/spicepod) and [datasets reference](/docs/next/reference/spicepod/datasets). * [Views](/docs/next/reference/spicepod/views) for declarative tenant filtering. * [Data Acceleration](/docs/next/features/data-acceleration) and [Change Data Capture](/docs/next/features/data-acceleration/data-refresh). * [Federated SQL Query recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). --- # Object-Store Based SQL Query, Search, and LLM Inference Engine Spice.ai enables SQL queries, hybrid search, and large language model (LLM) inference on object-store data for security applications, delivering real-time insights from distributed data sources with minimal infrastructure overhead. Unlike traditional data platforms (e.g., Snowflake, BigQuery) or separate search and AI frameworks (e.g., Elasticsearch, LangChain) that require complex data pipelines and centralized storage, Spice.ai integrates SQL querying, vector/keyword search, and LLM inference directly on object stores (e.g., S3, Azure Blob). This unified approach reduces latency, simplifies architecture, and ensures compliance for security applications, outperforming fragmented solutions that demand extensive data movement and integration. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Object-Store SQL Queries**: Executes federated SQL queries directly on object-store data (e.g., S3, Databricks Delta Lake) alongside other sources (e.g., PostgreSQL), eliminating the need for data ingestion into centralized warehouses, reducing costs and complexity. * **Hybrid Search**: Combines vector similarity search (VSS) for semantic analysis of unstructured data (e.g., security logs, threat reports) with keyword/BM25 search for precise retrieval, delivering context-aware results critical for security investigations. * **AI Gateway for LLM Inference**: Integrates LLMs (hosted like OpenAI or local like Llama) to process query and search results, generating actionable insights (e.g., threat summaries) with low-latency inference, optimized for object-store data access. * **Performance and Compliance**: Materializes hot datasets using Change Data Capture (CDC) for low-latency access and uses Databricks Unity Catalog for governance, ensuring compliance with security regulations (e.g., GDPR, SOC 2) unlike generic platforms. ## Example[​](#example "Direct link to Example") A security operations platform uses Spice.ai to query object-store data (e.g., S3-stored network logs), perform hybrid search to identify threat patterns (semantic via VSS, precise via BM25), and run LLM inference to generate real-time threat intelligence reports. This unified workflow detects anomalies in minutes without moving data to a centralized warehouse, outperforming traditional platforms requiring complex ETL pipelines and separate AI tools. The [Vector-Based Search documentation](/docs/next/features/search) and [Federated SQL Query recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) provide guidance for implementing search and query workflows on object stores. ## Benefits[​](#benefits "Direct link to Benefits") * **Efficiency**: Direct querying and inference on object stores eliminate data movement, reducing infrastructure costs and complexity. * **Real-Time Insights**: Low-latency search and LLM inference deliver rapid threat detection, critical for security applications. * **Compliance**: Governed data access ensures alignment with security regulations, increasing trust and auditability. ### Learn More[​](#learn-more "Direct link to Learn More") * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Vector and Hybrid Search**: [Documentation](/docs/next/features/search) and [Searching GitHub Files Recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md). * **AI Gateway**: [Documentation](/docs/next/features/large-language-models) and [Running Llama3 Locally Recipe](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md). * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) --- # Real-Time Decision-Making for Intelligent Applications Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. Unlike batch-processing analytics platforms (e.g., traditional data warehouses), Spice.ai delivers real-time decisions by unifying disparate data sources and accelerating access, outpacing siloed pipelines that introduce delays and complexity. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Federated SQL Queries**: Unifies disparate sources (e.g., PostgreSQL, Databricks, Snowflake) in a single SQL interface, eliminating the need for custom connectors and reducing integration overhead compared to point-to-point solutions. * **Data Acceleration**: Materializes hot datasets near applications using CDC, achieving sub-second latency, a critical advantage over cloud-only solutions with higher latency. * **AI Gateway**: Integrate AI into your applications with Spice.ai’s AI Gateway. It supports hosted models like OpenAI and Anthropic and local models such as OSS Llama and NVIDIA NIM. Fine-tuning and model distillation are simplified, helping faster cycles of development and deployment. * **Observability**: Provides end-to-end visibility into data and AI workflows, enabling rapid debugging and performance optimization, unlike fragmented analytics tools. ## Example[​](#example "Direct link to Example") A ride-sharing app optimizes driver assignments in milliseconds by combining real-time driver locations (from Kafka), user preferences (from PostgreSQL), and traffic conditions (from external APIs). This delivers faster, more accurate assignments than competitors relying on delayed batch updates, improving user satisfaction and operational efficiency. Developers can implement this using the Federated SQL Query recipe, which demonstrates querying across multiple sources with optimized push-down techniques. ## Benefits[​](#benefits "Direct link to Benefits") * **Speed**: Sub-second decision-making enhances user experience in high-stakes applications. * **Scalability**: Handles growing data volumes and user concurrency without infrastructure overhaul. * **Flexibility**: Adapts to diverse data sources and AI models, future-proofing application stacks. ### Learn More[​](#learn-more "Direct link to Learn More") * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md) for an example. * **AI Gateway**: [Documentation](/docs/next/features/large-language-models) and [Running Llama3 Locally Recipe](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md) for details. * **Vector Search**: [Documentation](/docs/next/features/search) and [Searching GitHub Files Recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md). * **Semantic Model**: [Documentation](/docs/next/features/semantic-model). * **Observability**: [Documentation](/docs/next/features/observability). --- # Tool-Augmented AI with Model Context Protocol Server Spice.ai extends AI with custom tools via Model Context Protocol (MCP) server in financial services (finserv), integrating domain-specific APIs for enhanced functionality and precise decision-making in high-stakes financial applications. Unlike generic tool-calling frameworks (e.g., OpenAI Functions) that rely on external APIs and lack deep data integration, Spice.ai's stdio-based MCP servers provide low-latency, secure tool interactions with federated data access. This makes it ideal for finserv applications requiring compliance, real-time performance, and tailored AI capabilities, outperforming platforms with limited domain-specific tool integration. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Internal MCP Hosting**: Runs stdio-based MCP servers within the Spice runtime, reducing latency and external dependencies compared to API-only frameworks, critical for finserv’s high-speed transaction and risk assessment needs. * **Tool Integration**: Extends large language models (LLMs) with custom tools (e.g., risk assessment APIs, portfolio optimization engines), offering flexibility for domain-specific tasks like fraud detection or trading analysis. * **Data Synergy**: Combines MCP tool outputs with federated SQL queries (e.g., transaction data from Databricks, PostgreSQL) for richer, context-aware AI responses, surpassing standalone tool frameworks that lack data integration. * **Security**: Internal execution ensures data privacy, aligning with finserv’s stringent compliance requirements for handling sensitive financial data. ## Example[​](#example "Direct link to Example") A finserv platform uses Spice.ai as an MCP server to integrate a risk assessment tool, enabling an LLM to generate real-time fraud alerts by combining transaction data (via federated queries from Databricks) with the tool's risk scoring outputs. This delivers faster, more accurate fraud detection than generic AI agents reliant on external APIs, reducing financial losses and improving regulatory compliance. The [Spice.ai MCP documentation](/docs/next/features/large-language-models/mcp) provides detailed setup guidance for implementing MCP server workflows. ## Benefits[​](#benefits "Direct link to Benefits") * **Customization**: Tailors AI capabilities to finserv-specific needs with custom tools, improving decision-making accuracy. * **Performance**: Internal tool execution minimizes latency, enabling real-time fraud detection and risk analysis. * **Compliance**: Protects sensitive financial data with secure, internal processing, meeting regulatory standards. ### Learn More[​](#learn-more "Direct link to Learn More") * **MCP Documentation**: [Documentation](/docs/next/features/large-language-models/mcp). * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **AI Gateway**: [Documentation](/docs/next/features/large-language-models) and [Running Llama3 Locally Recipe](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md). --- # Caching Spice.ai provides two complementary caching layers. Dataset acceleration materializes upstream data into a fast local engine. The SQL results cache stores query output in memory so that identical queries return instantly without re-executing against the accelerator. Both layers are configurable independently and work together to minimize latency and upstream load. ## [📄️Write-Through Cache](/docs/next/use-cases/caching/write-through-cache) [Use Spice.ai as a write-through cache that writes data to both a local accelerator and the upstream data source.](/docs/next/use-cases/caching/write-through-cache) ## [📄️Read-Through Cache](/docs/next/use-cases/caching/read-through-cache) [Use Spice.ai as a read-through cache with the SQL results cache for federated data sources and HTTP APIs.](/docs/next/use-cases/caching/read-through-cache) ## [📄️SQL/Database Cache](/docs/next/use-cases/caching/sql-database-cache) [Use Spice.ai to cache SQL database tables and query results locally for low-latency access and reduced load on upstream databases.](/docs/next/use-cases/caching/sql-database-cache) ## [📄️S3 Cache](/docs/next/use-cases/caching/s3-cache) [Use Spice.ai to cache S3 and object store data locally for fast, repeatable SQL queries without re-reading from remote storage.](/docs/next/use-cases/caching/s3-cache) ## [📄️HTTP Cache](/docs/next/use-cases/caching/http-cache) [Use Spice.ai to cache HTTP API responses locally, reducing API call frequency and providing fast, SQL-queryable access to external API data.](/docs/next/use-cases/caching/http-cache) --- # HTTP Cache Spice.ai caches HTTP API responses at two layers: a connector-level HTTP response cache that respects upstream `Cache-Control` headers, and dataset-level acceleration with `refresh_mode: caching` that stores API responses in a local accelerator with configurable TTL and stale-while-revalidate semantics. This pattern is useful for applications that query external REST APIs repeatedly — for example, caching search results from a third-party API so that the same query served to multiple users does not trigger redundant upstream requests. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Two-Layer Caching**: The HTTP connector caches raw responses based on upstream `Cache-Control: max-age` headers. Dataset-level caching adds TTL, stale-while-revalidate, and stale-if-error on top, providing application-controlled cache behavior independent of upstream headers. * **SQL-Queryable API Data**: Cached HTTP responses are stored as structured tables, queryable with standard SQL including filters, joins, and aggregations. * **Request Filtering**: Controls which request paths, query parameters, and body content are cacheable via `allowed_request_paths`, `request_query_filters`, and `request_body_filters`, preventing unbounded cache growth. * **Durable Cache**: File-backed accelerators (Cayenne, DuckDB, SQLite) persist cached responses to disk for fast restarts without re-fetching. ## Example[​](#example "Direct link to Example") ### Accelerated HTTP Dataset with Caching Mode[​](#accelerated-http-dataset-with-caching-mode "Direct link to Accelerated HTTP Dataset with Caching Mode") Cache JSON responses from a REST API with stale-while-revalidate: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_cache params: file_format: json allowed_request_paths: '/search/shows,/shows/*' request_query_filters: enabled max_request_query_length: 1024 acceleration: enabled: true refresh_mode: caching engine: cayenne mode: file params: caching_ttl: 30s caching_stale_while_revalidate_ttl: 2m caching_stale_if_error: enabled ``` Query the cached API data over SQL: ``` SELECT * FROM tv_shows_cache WHERE request_path = '/search/shows' AND request_query = 'q=breaking+bad'; ``` The first query triggers a fetch from the upstream API. Subsequent queries within 30 seconds are served from the local accelerator. After 30 seconds, the cached result is served immediately while Spice revalidates in the background. If the upstream API is unavailable, the stale cached result is served instead of an error. ### Request Filtering[​](#request-filtering "Direct link to Request Filtering") Limit which API paths and parameters are cached to prevent unbounded cache growth: ``` datasets: - from: https://api.example.com name: api_cache params: file_format: json allowed_request_paths: '/v1/search,/v1/items/*' request_query_filters: enabled max_request_query_length: 1024 request_body_filters: enabled max_request_body_bytes: 16384 acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 1m ``` Only requests matching the allowed paths with query and body sizes within the configured limits are cached. ## Query Results Cache[​](#query-results-cache "Direct link to Query Results Cache") The SQL results cache adds a third caching layer for HTTP API data. While the HTTP connector caches raw responses and `refresh_mode: caching` stores parsed data in the accelerator, the SQL results cache stores the output of executed SQL queries in memory so that identical queries return instantly without re-querying the accelerator. ``` datasets: - from: https://api.tvmaze.com name: tv_shows_cache params: file_format: json allowed_request_paths: '/search/shows,/shows/*' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: cayenne mode: file params: caching_ttl: 30s caching_stale_if_error: enabled runtime: caching: sql_results: enabled: true item_ttl: 10s ``` In this three-layer configuration: 1. The HTTP connector caches raw responses based on upstream `Cache-Control` headers. 2. The dataset accelerator caches parsed data with a 30-second TTL and stale-if-error. 3. The SQL results cache stores query output in memory for 10 seconds. Identical SQL queries within 10 seconds are served from memory. After 10 seconds, the query re-executes against the accelerator (which may still be serving from its dataset-level cache). The `Results-Cache-Status` response header indicates cache state: `HIT`, `MISS`, `BYPASS`, or `STALE`. Clients can bypass the results cache using the `Cache-Control: no-cache` header. warning Do not configure `stale_while_revalidate_ttl` on both the SQL results cache (`runtime.caching.sql_results`) and the dataset caching accelerator (`acceleration.params.caching_stale_while_revalidate_ttl`) for the same dataset. Use one or the other to avoid conflicting revalidation behavior. ## Benefits[​](#benefits "Direct link to Benefits") * **Reduced API Costs**: Three caching layers minimize redundant upstream HTTP requests. * **Low Latency**: Queries are served from the closest cache layer with data — memory, accelerator, or HTTP cache. * **Resilience**: Stale-if-error keeps the application functional when upstream APIs experience downtime or rate limiting. ### Learn More[​](#learn-more "Direct link to Learn More") * **Caching Refresh Mode**: [Documentation](/docs/next/features/data-acceleration/refresh-modes/caching) for detailed configuration, schema, and parameters. * **Caching**: [Documentation](/docs/next/features/caching) for SQL results cache configuration, Cache-Control directives, and response headers. * **HTTP(s) Data Connector**: [Documentation](/docs/next/components/data-connectors/https) for authentication, headers, and connector-specific parameters. * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) for acceleration engines and modes. --- # Read-Through Cache Spice.ai provides a read-through caching pattern through the SQL results cache. When a query is executed, the result is stored in an in-memory cache. Identical queries within the TTL window are served directly from memory without re-executing against the upstream data source. This works for both federated data sources (PostgreSQL, MySQL, S3, etc.) and HTTP API datasets. For HTTP-based datasets, Spice also supports a dataset-level `refresh_mode: caching` that fetches data from the upstream API on cache miss and stores it in the local accelerator. The SQL results cache operates on top of this, adding a fast in-memory layer for repeated SQL queries. ## Federated Data Sources[​](#federated-data-sources "Direct link to Federated Data Sources") For datasets without acceleration enabled, queries are federated directly to the upstream source. The SQL results cache stores the output of these queries in memory, so that identical queries within the TTL return instantly without a network round-trip to the source. This is effective for dashboards, reporting queries, or any workload where the same query is executed repeatedly within a short window against a remote database. ``` datasets: - from: postgres:production.public.customers name: customers runtime: caching: sql_results: enabled: true item_ttl: 30s stale_while_revalidate_ttl: 5m ``` The first execution of `SELECT * FROM customers WHERE region = 'us-west'` federates the query to PostgreSQL. The result is cached in memory. Identical queries within 30 seconds return from the cache (`HIT`). Between 30 seconds and 5 minutes 30 seconds, stale results are served immediately (`STALE`) while Spice re-executes the query against the upstream source in the background. After 5 minutes 30 seconds without access, the entry is evicted and the next query is a `MISS`. ### Configuration[​](#configuration "Direct link to Configuration") ``` runtime: caching: sql_results: enabled: true # Default: true max_size: 256MiB # Default: 128MiB item_ttl: 30s # Default: 1s stale_while_revalidate_ttl: 5m # Default: 0s (disabled) eviction_policy: lru # lru (default) or tiny_lfu cache_key_type: plan # plan (default) or sql encoding: zstd # none (default) or zstd ``` | Parameter | Default | Description | | ---------------------------- | ------- | --------------------------------------------------------------------------------------------------------------- | | `item_ttl` | `1s` | Duration a cached entry is considered fresh. | | `stale_while_revalidate_ttl` | `0s` | Grace period to serve stale entries while re-executing the query in the background. | | `eviction_policy` | `lru` | Cache replacement policy. `tiny_lfu` provides higher hit rates for skewed access patterns. | | `cache_key_type` | `plan` | `plan` matches semantically equivalent queries. `sql` matches only identical SQL strings (faster but stricter). | | `encoding` | `none` | `zstd` compresses cached results to fit more entries in memory. | ### Cache-Control[​](#cache-control "Direct link to Cache-Control") Clients can control cache behavior per-request using the `Cache-Control` header (HTTP/Flight API) or the `--cache-control` flag (Spice SQL REPL): ``` # Bypass cache and fetch fresh from upstream curl -H "cache-control: no-cache" -XPOST http://localhost:8090/v1/sql \ -d "SELECT * FROM customers WHERE region = 'us-west'" # Only return cached results; error on cache miss curl -H "cache-control: only-if-cached" -XPOST http://localhost:8090/v1/sql \ -d "SELECT * FROM customers WHERE region = 'us-west'" # Accept stale results up to 60 seconds old curl -H "cache-control: max-stale=60" -XPOST http://localhost:8090/v1/sql \ -d "SELECT * FROM customers WHERE region = 'us-west'" ``` The `Results-Cache-Status` response header indicates cache state: `HIT`, `MISS`, `BYPASS`, or `STALE`. ## HTTP Data Sources[​](#http-data-sources "Direct link to HTTP Data Sources") For HTTP-based datasets, Spice provides dataset-level caching with `refresh_mode: caching`. On a cache miss, Spice fetches data from the upstream API, returns it to the caller, and stores it in the local accelerator. The SQL results cache adds an in-memory layer on top, caching the output of SQL queries against the accelerated data. ``` datasets: - from: https://api.tvmaze.com name: tv_search_cache params: file_format: json allowed_request_paths: '/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: cayenne mode: file params: caching_ttl: 30s caching_stale_if_error: enabled runtime: caching: sql_results: enabled: true item_ttl: 10s ``` In this configuration: 1. The first query fetches from the upstream API and caches the response in the accelerator. 2. The query result is also stored in the in-memory SQL results cache. 3. Identical SQL queries within 10 seconds are served from memory without touching the accelerator. 4. After 10 seconds, the query re-executes against the accelerator (still serving from the dataset cache if within the 30-second `caching_ttl`). ### HTTP Dataset Cache Parameters[​](#http-dataset-cache-parameters "Direct link to HTTP Dataset Cache Parameters") | Parameter | Default | Description | | ------------------------------------ | ---------- | ------------------------------------------------------------------------------------------------- | | `caching_ttl` | `0s` | Duration a cached entry in the accelerator is considered fresh. | | `caching_stale_while_revalidate_ttl` | `0s` | Duration after TTL expiry during which stale data is served while revalidating in the background. | | `caching_stale_if_error` | `disabled` | When `enabled`, serves stale cached data if the upstream fetch fails. | warning Do not configure `stale_while_revalidate_ttl` on both the SQL results cache (`runtime.caching.sql_results`) and the dataset caching accelerator (`acceleration.params.caching_stale_while_revalidate_ttl`) for the same dataset. Use one or the other to avoid conflicting revalidation behavior. ## Benefits[​](#benefits "Direct link to Benefits") * **Read-Through for Any Source**: The SQL results cache provides read-through semantics for any data source — federated databases, accelerated datasets, or HTTP APIs — with no application code changes. * **Reduced Upstream Load**: Repeated queries are served from the in-memory cache, protecting upstream databases and APIs from read amplification. * **Stale-While-Revalidate**: Expired entries are served immediately while Spice refreshes data in the background, avoiding latency spikes. * **Resilience**: For HTTP sources, `stale_if_error` keeps the application functional during upstream outages. ### Learn More[​](#learn-more "Direct link to Learn More") * **Caching**: [Documentation](/docs/next/features/caching) for SQL results cache configuration, Cache-Control directives, stale-while-revalidate behavior, and response headers. * **Caching Refresh Mode**: [Documentation](/docs/next/features/data-acceleration/refresh-modes/caching) for HTTP dataset-level caching configuration. * **HTTP(s) Data Connector**: [Documentation](/docs/next/components/data-connectors/https) for HTTP-specific parameters like `allowed_request_paths` and request filters. * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) for acceleration engines and modes. --- # S3 Cache Spice.ai caches S3 and object store data by accelerating remote datasets into a local engine. Instead of scanning remote Parquet, CSV, or JSON files on every query, Spice materializes the data locally and refreshes it on a configurable schedule. For single-file datasets, Spice tracks the object's metadata (size, last modified, ETag) and skips refresh when the file has not changed, reducing S3 API costs. This pattern is useful for analytics workloads over object store data — for example, querying a Parquet dataset in S3 repeatedly throughout the day without incurring the latency and cost of a full scan on each query. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Local Acceleration**: Materializes S3 data into a fast local engine (Cayenne, DuckDB, SQLite, Arrow) so queries run at local speed instead of scanning remote storage. * **Smart Refresh Skip**: For single-file S3 datasets, Spice checks the object's metadata before refreshing. If the file has not changed (same size, last modified timestamp, and version/ETag), the refresh is skipped entirely. * **Multiple File Formats**: Supports Parquet, CSV, JSON, and other formats stored in S3 or S3-compatible systems (MinIO, Cloudflare R2). * **Folder-Level Datasets**: Point a dataset at an S3 folder to load all files within it as a single table, with periodic refresh to pick up new files. ## Example[​](#example "Direct link to Example") ### Accelerated S3 Dataset[​](#accelerated-s3-dataset "Direct link to Accelerated S3 Dataset") Cache a Parquet dataset from S3 with periodic refresh: ``` datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet acceleration: enabled: true engine: cayenne refresh_check_interval: 10m ``` Queries against `taxi_trips` run against the local accelerator. Every 10 minutes, Spice checks S3 for changes and refreshes the accelerated copy if the data has changed. ### Private Bucket with Authentication[​](#private-bucket-with-authentication "Direct link to Private Bucket with Authentication") ``` datasets: - from: s3://my-private-bucket/events/ name: events params: file_format: parquet s3_auth: key s3_key: ${secrets:AWS_ACCESS_KEY_ID} s3_secret: ${secrets:AWS_SECRET_ACCESS_KEY} s3_region: us-west-2 acceleration: enabled: true engine: duckdb mode: file refresh_check_interval: 5m ``` Using `mode: file` persists the accelerated data to disk, so the cache survives Spice restarts without re-reading from S3. ## Query Results Cache[​](#query-results-cache "Direct link to Query Results Cache") For analytics workloads that run the same queries repeatedly against S3 data — such as dashboards or scheduled reports — the SQL results cache stores query output in memory. Identical queries within the TTL window return instantly without re-executing against the accelerator. ``` datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet acceleration: enabled: true engine: cayenne refresh_check_interval: 10m runtime: caching: sql_results: enabled: true item_ttl: 30s stale_while_revalidate_ttl: 5m eviction_policy: lru ``` With this configuration, the first execution of a query runs against the accelerated S3 data and the result is cached in memory. Identical queries within 30 seconds are served from the cache. Between 30 seconds and 5 minutes 30 seconds, stale results are served immediately while Spice re-executes the query in the background. The `Results-Cache-Status` response header indicates cache state: `HIT`, `MISS`, `BYPASS`, or `STALE`. Clients can use `Cache-Control: no-cache` to bypass the cache and force a fresh query execution. ## Benefits[​](#benefits "Direct link to Benefits") * **Reduced S3 Costs**: Fewer GET requests and less data transfer. Smart refresh skip avoids unnecessary reads for unchanged files. * **Fast Queries**: Local acceleration delivers sub-millisecond to low-millisecond query times instead of seconds-long remote scans. The results cache eliminates accelerator query overhead for repeated queries. * **Cold Start Resilience**: File-backed accelerators persist cached data across restarts. ### Learn More[​](#learn-more "Direct link to Learn More") * **S3 Data Connector**: [Documentation](/docs/next/components/data-connectors/s3) for authentication, configuration, and supported formats. * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [Data Refresh](/docs/next/features/data-acceleration/data-refresh). * **Caching**: [Documentation](/docs/next/features/caching) for SQL results cache configuration, Cache-Control directives, and response headers. * **DuckDB Data Accelerator**: [Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README) for file-backed acceleration. --- # SQL and Database Cache Spice.ai caches SQL database data at two layers: dataset acceleration, which materializes upstream tables into a fast local engine, and SQL results caching, which stores query results in memory for repeated queries. Together, these layers reduce load on upstream databases and deliver sub-millisecond query performance. Dataset acceleration is suited for working sets that applications query repeatedly, while results caching handles identical SQL queries across short time windows — for example, a dashboard that multiple users hit with the same query within seconds. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Dataset Acceleration**: Materializes tables from PostgreSQL, MySQL, or other SQL connectors into a local accelerator (Cayenne, DuckDB, SQLite, Arrow), with configurable refresh intervals or CDC-based change tracking. * **SQL Results Cache**: Caches query results in memory with LRU eviction, configurable TTL, and stale-while-revalidate support. Enabled by default for HTTP and Arrow Flight SQL APIs. * **Reduced Upstream Load**: Repeated queries are served from the local cache or accelerator without hitting the upstream database, protecting production databases from read amplification. * **CDC Support**: Debezium-based change data capture keeps the accelerated copy in sync with the source database in near-real-time, without polling. ## Example[​](#example "Direct link to Example") ### Dataset Acceleration[​](#dataset-acceleration "Direct link to Dataset Acceleration") Accelerate a PostgreSQL table locally with periodic refresh: ``` datasets: - from: postgres:production.public.customers name: customers acceleration: enabled: true engine: cayenne refresh_check_interval: 1m ``` With CDC-based refresh using Debezium: ``` datasets: - from: debezium:cdc.public.orders name: orders acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes ``` ### SQL Results Cache[​](#sql-results-cache "Direct link to SQL Results Cache") Enable the SQL results cache with a longer TTL and stale-while-revalidate for dashboard workloads: ``` runtime: caching: sql_results: enabled: true max_size: 512MiB item_ttl: 30s stale_while_revalidate_ttl: 5m eviction_policy: lru ``` With this configuration, the first execution of a query runs against the accelerated table, and the result is cached in memory. Identical queries within 30 seconds are served from the in-memory cache. After 30 seconds, the cached result is served stale while Spice re-executes the query in the background. After 5 minutes 30 seconds without access, the entry is evicted. The `Results-Cache-Status` response header indicates cache state for each query: | Value | Meaning | | -------- | --------------------------------------------------------- | | `HIT` | Result served from the cache. | | `MISS` | No cached result; query executed against the accelerator. | | `STALE` | Stale result served while revalidating in the background. | | `BYPASS` | Cache bypassed by client request. | Clients can control cache behavior per-request using the `Cache-Control` header (HTTP API) or `--cache-control` flag (Spice SQL REPL). For example, `Cache-Control: no-cache` bypasses the cache for a single request while still caching the result for future queries. `Cache-Control: only-if-cached` returns only cached results, erroring on a miss. ## Benefits[​](#benefits "Direct link to Benefits") * **Sub-Millisecond Reads**: Accelerated data is colocated with the application, eliminating network round-trips to the source database. * **Production Database Protection**: Both caching layers absorb read traffic, keeping upstream database load predictable. * **Freshness Control**: Choose between periodic refresh, CDC, or query-level TTL depending on the consistency requirements. ### Learn More[​](#learn-more "Direct link to Learn More") * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [Data Refresh](/docs/next/features/data-acceleration/data-refresh). * **Caching**: [Documentation](/docs/next/features/caching) for SQL results cache configuration and parameters. * **CDC**: [Documentation](/docs/next/features/cdc) for Debezium-based change data capture. * **CQRS Cookbook**: [Recipe](https://github.com/spiceai/cookbook/tree/trunk/cqrs#readme) for command-query separation patterns with Spice. --- # Write-Through Cache Spice.ai functions as a write-through cache by accepting writes via SQL and propagating them to both the local accelerator and the upstream data source. Applications read from the fast local accelerator while writes flow through to the authoritative source, keeping both layers consistent without manual synchronization. This pattern is common in operational applications that need low-latency reads alongside durable writes — for example, an application that updates user preferences through Spice and expects the changes to be reflected in both the local cache and the upstream PostgreSQL database. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Transparent Write Path**: SQL `INSERT INTO` statements write to the upstream connector and the accelerator reflects the change on the next refresh cycle, requiring no application-level cache invalidation logic. * **Consistent Reads**: The local accelerator serves reads at sub-millisecond latency while the source of truth stays in sync through periodic or CDC-based refresh. * **Connector Flexibility**: Write-through works with any [write-capable connector](/docs/next/features/data-ingestion) (e.g., Apache Iceberg, AWS Glue) combined with any acceleration engine. * **Automatic Refresh**: CDC (`refresh_mode: changes`) or periodic refresh (`refresh_check_interval`) propagates upstream changes back to the accelerator, closing the read-after-write loop. ## Example[​](#example "Direct link to Example") An application writes new order records to an Iceberg table and reads from a locally accelerated copy. Writes go to the upstream Iceberg table, and the accelerator refreshes periodically to pick up the new rows. ``` datasets: - from: iceberg:my_catalog.orders name: orders access: read_write acceleration: enabled: true engine: cayenne refresh_check_interval: 10s ``` Insert data through Spice using standard SQL: ``` INSERT INTO orders (order_id, customer_id, total) VALUES (1001, 42, 99.95); ``` The write goes to the upstream Iceberg table. On the next refresh cycle (within 10 seconds in this configuration), the accelerator picks up the new row and subsequent reads return it at local-accelerator speed. ## Query Results Cache[​](#query-results-cache "Direct link to Query Results Cache") In addition to dataset acceleration, Spice provides an in-memory SQL results cache that stores the output of repeated queries. When applications issue the same `SELECT` against the accelerated dataset within the TTL window, the result is served directly from memory without re-executing the query against the accelerator. This is particularly effective in write-through scenarios where reads far outnumber writes — the accelerator absorbs the write-path overhead, while the results cache eliminates redundant read-path computation. ``` datasets: - from: iceberg:my_catalog.orders name: orders access: read_write acceleration: enabled: true engine: cayenne refresh_check_interval: 10s runtime: caching: sql_results: enabled: true item_ttl: 5s ``` With this configuration, identical queries within 5 seconds are served from the in-memory results cache. After the TTL expires, the next query re-executes against the accelerator and the result is cached again. The `Results-Cache-Status` response header indicates whether a query was served from the cache (`HIT`), executed fresh (`MISS`), or served stale during background revalidation (`STALE`). Clients can bypass the cache for a specific request using the `Cache-Control: no-cache` header or the `--cache-control no-cache` flag in the Spice SQL REPL. ## Benefits[​](#benefits "Direct link to Benefits") * **Low-Latency Reads**: The accelerator serves queries without network round-trips to the upstream source. The results cache adds a further in-memory layer for repeated queries. * **Durable Writes**: Data is persisted in the authoritative source, not just the cache. * **No Cache Invalidation Logic**: Refresh handles consistency automatically — no application code needed to synchronize the cache. ### Learn More[​](#learn-more "Direct link to Learn More") * **Data Ingestion**: [Documentation](/docs/next/features/data-ingestion) for write-capable connectors and SQL write syntax. * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README). * **Caching**: [Documentation](/docs/next/features/caching) for SQL results cache configuration, Cache-Control directives, and response headers. * **CDC**: [Documentation](/docs/next/features/cdc) for change data capture-based refresh. --- # Data Federation, Acceleration, and SQL Query ## [📄️Reverse-ETL](/docs/next/use-cases/data/reverse-etl) [Serve enriched data from warehouses and data lakes to operational systems and applications, eliminating complex pipelines.](/docs/next/use-cases/data/reverse-etl) ## [📄️ETL-free Workflows](/docs/next/use-cases/data/etl-free-workflows) [Federate legacy and modern data systems without ETL for faster migrations, lower overhead, and zero application downtime.](/docs/next/use-cases/data/etl-free-workflows) ## [📄️Data Mesh](/docs/next/use-cases/data/data-mesh) [Give domain teams decentralized, real-time data ownership and access across disparate sources through a unified SQL interface.](/docs/next/use-cases/data/data-mesh) ## [📄️Object-Store Data Federation](/docs/next/use-cases/data/object-store-data-engine) [Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses.](/docs/next/use-cases/data/object-store-data-engine) ## [📄️Resilience and Performance](/docs/next/use-cases/data/application-resilence-and-acceleration) [Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance.](/docs/next/use-cases/data/application-resilence-and-acceleration) ## [📄️Database CDN](/docs/next/use-cases/data/database-cdn) [Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience.](/docs/next/use-cases/data/database-cdn) --- # Application Resilience and Performance Optimization Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance for uninterrupted user experiences in high-traffic, customer-facing environments. Unlike traditional CDNs (e.g., Cloudflare, Akamai) focused on static content delivery, Spice.ai targets dynamic, operational data with real-time caching and materialization, enabling high-availability SaaS applications with minimal infrastructure. This approach delivers superior performance and uptime compared to cloud-dependent databases or generic caching solutions, addressing the critical needs of SaaS platforms for scalability and reliability. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Data Colocation**: Caches dynamic data (e.g., user sessions, application states) locally using Change Data Capture (CDC), reducing latency compared to remote database queries, critical for responsive SaaS applications. * **Resilience**: Maintains local data replicas to ensure availability during cloud outages or network disruptions, unlike cloud-dependent architectures that risk downtime in high-traffic scenarios. * **Scalability**: Optimizes data access for high concurrency, supporting thousands of simultaneous users, surpassing traditional database setups that struggle with scale in SaaS environments. * **Monitoring**: Provides end-to-end visibility into cache performance, data freshness, and system health, ensuring reliability and rapid debugging, unlike fragmented monitoring in generic caching tools. ## Example[​](#example "Direct link to Example") A SaaS project management platform caches user task data and project states locally, ensuring uninterrupted access to critical features during peak usage or cloud outages. This delivers a seamless user experience, unlike cloud-only platforms prone to latency spikes or downtime, improving user productivity and platform reliability. The [CQRS Cookbook](https://github.com/spiceai/cookbook/tree/trunk/cqrs#readme) illustrates colocation strategies for such use cases. ## Benefits[​](#benefits "Direct link to Benefits") * **Availability**: Local replicas ensure uptime, improving user trust and retention in SaaS applications. * **Performance**: Reduced latency improves responsiveness for customer-facing features, critical for user satisfaction. * **Scalability**: Supports high user concurrency, enabling growth without performance degradation. ### Learn More[​](#learn-more "Direct link to Learn More") * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md). * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Observability**: [Documentation](/docs/next/features/observability). --- # Data Mesh for Unified Data Access Spice supports data mesh architectures by giving domain teams decentralized, real-time data access through a unified SQL interface. Each team manages its own datasets while Spice federates and accelerates queries across all sources, removing the need for centralized data pipelines. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Federated SQL Queries**: Query disparate sources (PostgreSQL, Databricks, S3, on-premises systems) through a single SQL interface. Domain teams access their own data without relying on a central data team. * **Local Acceleration**: Materialize domain-specific datasets near applications using CDC-based refresh, delivering low-latency access without copying data into a central warehouse. * **Governance**: Integrates with Databricks Unity Catalog for role-based access control and credential vendoring, so teams maintain security and compliance without custom infrastructure. * **Observability**: End-to-end visibility into data flows and query performance across domains, simplifying monitoring and debugging. ## Example[​](#example "Direct link to Example") An organization runs multiple teams, each owning their data in separate systems — one team in PostgreSQL, another in Databricks, a third in S3. Spice federates all three sources and accelerates frequently queried datasets locally, so any application can query across domains with consistent performance. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - from: postgres:team_a.customers name: customers acceleration: enabled: true engine: duckdb - from: databricks:team_b.transactions name: transactions acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes - from: s3://team-c-data/reports/ name: reports params: file_format: parquet acceleration: enabled: true ``` This configuration federates customer data from PostgreSQL, transaction data from Databricks (with CDC refresh), and report data from S3, accelerating all three locally for unified access. The [Federated SQL Query recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) demonstrates unified data access patterns for such scenarios. ## Benefits[​](#benefits "Direct link to Benefits") * **Decentralization**: Teams own and manage their own data while applications query a single endpoint. * **Performance**: Local acceleration delivers consistent low-latency queries across all domains. * **Governance**: Centralized access control without centralized data infrastructure. ### Learn More[​](#learn-more "Direct link to Learn More") * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md). * **Observability**: [Documentation](/docs/next/features/observability). --- # Database CDN for Enhanced Performance Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience in high-traffic, customer-facing environments. Unlike traditional CDNs (e.g., Cloudflare, Akamai) designed for static content delivery, Spice.ai focuses on dynamic, operational data with real-time caching and materialization, delivering superior performance and uptime compared to cloud-dependent databases or generic caching solutions. This makes it ideal for SaaS platforms requiring low-latency data access and continuous availability to support seamless user experiences. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Dynamic Data Colocation**: Caches dynamic data (e.g., user sessions, application states) locally using Change Data Capture (CDC), reducing latency compared to remote database queries, critical for responsive SaaS applications. * **Resilience**: Maintains local data replicas to ensure availability during cloud outages or network disruptions, unlike cloud-dependent architectures that risk downtime in high-traffic scenarios. * **Scalability**: Optimizes data access for high concurrency, supporting thousands of simultaneous users, surpassing traditional database setups that struggle with scale in SaaS environments. * **Monitoring**: Provides end-to-end visibility into cache performance, data freshness, and system health through built-in observability, ensuring reliability and rapid debugging, unlike fragmented monitoring in generic caching tools. ## Example[​](#example "Direct link to Example") A SaaS customer relationship management (CRM) platform caches user interaction data and account states locally, ensuring uninterrupted access to critical features during peak usage or cloud outages. This delivers a seamless user experience, unlike cloud-only platforms prone to latency spikes or downtime, improving customer satisfaction and retention. The [CQRS Cookbook](https://github.com/spiceai/cookbook/tree/trunk/cqrs#readme) illustrates colocation strategies for such use cases. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - from: postgres:crm_db.user_sessions name: user_sessions acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes refresh_check_interval: 1m ``` This configuration accelerates the `user_sessions` table from a PostgreSQL database, storing it locally in DuckDB file mode for fast, resilient access. ## Benefits[​](#benefits "Direct link to Benefits") * **Availability**: Local replicas ensure uptime, improving user trust and retention in SaaS applications. * **Performance**: Reduced latency improves responsiveness for customer-facing features, critical for user engagement. * **Scalability**: Supports high user concurrency, enabling growth without performance degradation. ### Learn More[​](#learn-more "Direct link to Learn More") * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README). * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README). * **Observability**: [Documentation](/docs/next/features/observability). --- # ETL-free Workflows and Data Migrations Spice federates legacy and modern data systems without ETL, supporting data migrations and cross-system workflows with zero application downtime. Applications query a single SQL endpoint that spans both old and new systems, making migrations incremental and reversible. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Single Endpoint**: Applications connect to one SQL interface regardless of how many backend systems are involved. Migrating from one database to another requires no application code changes. * **No Data Movement**: Federated queries read from each source directly, avoiding the cost and risk of bulk data replication during migrations. * **Acceleration**: Frequently accessed data is materialized locally for low-latency queries, so performance stays consistent even while backend systems are being migrated. * **Observability**: Built-in monitoring tracks query performance and data freshness across all sources, giving visibility into migration progress. ## Example[​](#example "Direct link to Example") An application migrates from a legacy PostgreSQL database to a cloud-native system. During the transition, Spice federates both sources so the application queries a single endpoint. Data from the legacy system is accelerated locally to maintain performance while traffic shifts to the new backend. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - from: postgres:legacy_db.orders name: legacy_orders acceleration: enabled: true engine: duckdb - from: mysql:new_system.orders name: new_orders acceleration: enabled: true engine: duckdb ``` This configuration federates order data from both the legacy PostgreSQL database and the new MySQL system, accelerating both locally in DuckDB for consistent query performance. ## Benefits[​](#benefits "Direct link to Benefits") * **Zero Downtime**: Applications keep working during migrations — no staged cutovers or maintenance windows required. * **Incremental Migration**: Move tables one at a time while the application queries both systems transparently. * **Lower Risk**: No bulk data replication means fewer failure modes and easier rollback. ### Learn More[​](#learn-more "Direct link to Learn More") * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md). * **Observability**: [Documentation](/docs/next/features/observability). --- # Object-Store Data Engine Spice.ai federates, accelerates, and queries object-store data for financial services (finserv) applications, enabling real-time data access without the need for centralized data warehouses, streamlining workflows and ensuring compliance. Unlike traditional data platforms (e.g., Snowflake, BigQuery) that rely on costly data ingestion into centralized repositories, Spice.ai’s federated query engine and data acceleration capabilities operate directly on object stores (e.g., S3, Azure Blob) alongside other sources, reducing infrastructure overhead and latency. This makes it ideal for finserv applications demanding fast, secure, and compliant data access for operational workflows, outperforming solutions dependent on complex ETL pipelines. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Federated SQL Queries**: Executes SQL queries directly on object-store data (e.g., S3, Databricks Delta Lake) and other sources (e.g., PostgreSQL, on-premises systems) in a unified interface, eliminating data movement and simplifying workflows compared to centralized warehouse solutions. * **Data Acceleration**: Materializes frequently accessed datasets (e.g., transaction logs) from object stores using Change Data Capture (CDC), delivering low-latency access critical for finserv’s time-sensitive applications, surpassing cloud-only platforms with higher latency. * **Governance**: Integrates with Databricks Unity Catalog for role-based security and credential vendoring, ensuring compliance with finserv regulations (e.g., GDPR, SEC), unlike generic data federation tools lacking comprehensive governance. * **Observability**: Provides end-to-end visibility into data flows, query performance, and system health, enabling rapid debugging and optimization, reducing overhead compared to fragmented monitoring in traditional platforms. ## Example[​](#example "Direct link to Example") A finserv platform queries transaction logs stored in S3, combines them with real-time market data from PostgreSQL, and materializes high-frequency trading datasets for low-latency portfolio analysis. This approach avoids moving data to a centralized warehouse, reducing costs and ensuring compliance with regulatory requirements, unlike ETL-heavy platforms that introduce delays and complexity. The [Federated SQL Query recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) and [DuckDB Data Accelerator recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md) provide practical guidance for implementing federated queries and data materialization. ## Benefits[​](#benefits "Direct link to Benefits") * **Efficiency**: Direct querying on object stores eliminates data movement, reducing infrastructure costs and complexity in finserv. * **Performance**: Accelerated data access ensures rapid insights for high-speed trading and risk analysis. * **Compliance**: Governed data access aligns with strict financial regulations, ensuring security and auditability. ### Learn More[​](#learn-more "Direct link to Learn More") * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md). * **Observability**: [Documentation](/docs/next/features/observability). --- # Reverse-ETL for Operational Workflows Spice serves enriched data from warehouses and data lakes directly to operational systems and applications, eliminating the need for complex ETL pipelines. By federating queries across warehouse and operational data sources and accelerating results locally, Spice turns warehouse data into a live, queryable resource for applications. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Unified Query Layer**: Federated SQL queries span Databricks, Snowflake, PostgreSQL, and other sources through a single endpoint, removing the need for separate connectors or pipeline orchestration. * **Real-Time Sync**: Change Data Capture (CDC) keeps local replicas in sync with upstream sources, so operational systems always reflect the latest warehouse data without scheduled batch jobs. * **Local Acceleration**: Materializes working sets of warehouse data into fast local engines like DuckDB or SQLite, providing sub-millisecond query latency for application workloads. * **Governance**: Integrates with Databricks Unity Catalog for role-based access control and credential vendoring across federated sources. ## Example[​](#example "Direct link to Example") An application serves customer metrics from a Databricks lakehouse to an internal dashboard, keeping the data locally accelerated for fast queries and automatically refreshed as upstream data changes. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - from: databricks:catalog.schema.customer_metrics name: customer_metrics acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes refresh_check_interval: 10s ``` This configuration federates the `customer_metrics` table from Databricks and accelerates it locally in DuckDB. The `refresh_mode: changes` setting uses CDC to sync only changed rows, keeping operational data current. The [DuckDB Data Accelerator recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md) provides a practical guide to materializing datasets for such workflows. ## Benefits[​](#benefits "Direct link to Benefits") * **Simplicity**: No pipeline orchestration or ETL infrastructure to maintain — query warehouse data directly. * **Freshness**: CDC-based refresh keeps operational data current without manual sync. * **Performance**: Local acceleration delivers low-latency queries for application workloads. ### Learn More[​](#learn-more "Direct link to Learn More") * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md). * **Observability**: [Documentation](/docs/next/features/observability). --- # Spice for Retrieval-Augmented-Generation (RAG) Use Spice to access data across various data sources for Retrieval-Augmented-Generation (RAG). Spice enables developers to combine structured data via SQL queries and unstructured data through built-in vector similarity search. This combined data can then be fed to large language models (LLMs) through a native AI gateway, improving the models' ability to generate accurate and contextually relevant responses. For more details on using vector search, embeddings, and model providers, refer to the following documentation: * [Vector-Based Search](/docs/next/features/search/vector-search) * [Embedding Models](/docs/next/components/embeddings) * [Model Providers](/docs/next/components/models) --- # RAG for Contextual Applications Use Spice to access data across various data sources for Retrieval-Augmented-Generation (RAG). Spice enables developers to combine structured data via SQL queries and unstructured data through built-in vector similarity search. This combined data can then be fed to large language models (LLMs) through a native AI gateway, improving the models' ability to generate accurate and contextually relevant responses. ## Example Configuration[​](#example-configuration "Direct link to Example Configuration") The following `spicepod.yaml` configures a dataset with vector embeddings and an OpenAI model for RAG: ``` datasets: - from: s3://my-bucket/documents/ name: documents params: file_format: parquet columns: - name: content embeddings: - from: openai acceleration: enabled: true embeddings: - from: openai name: openai params: openai_api_key: ${ env:OPENAI_API_KEY } models: - from: openai:gpt-4o name: rag_model params: openai_api_key: ${ env:OPENAI_API_KEY } tools: auto ``` For more details on using vector search, embeddings, and model providers, refer to the following documentation: * [Vector-Based Search](/docs/next/features/search/vector-search) * [Embedding Models](/docs/next/components/embeddings) * [Model Providers](/docs/next/components/models) --- # Retrieval-Augmented Generation for AI-Powered Reporting Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech/healthcare, ensuring compliance and precision in regulated environments. Unlike traditional reporting tools (e.g., Tableau, Power BI) or generic RAG frameworks (e.g., LlamaIndex) that lack real-time data federation and comprehensive governance, Spice.ai combines enterprise-grade data access, vector search, and AI integration to deliver precise, real-time reports tailored to healthcare's stringent requirements, surpassing generic solutions in accuracy and trustworthiness. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Vector Similarity Search (VSS)**: Retrieves unstructured data (e.g., audit logs, regulatory guidelines) and embeddings efficiently, enabling semantic context for compliance-focused reports, critical for healthcare applications. * **AI Gateway**: Produces narrative, human-readable reports from complex datasets using large language models (LLMs), supporting hosted (e.g., OpenAI) and local (e.g., Llama) models for privacy and cost efficiency in healthcare. * **Federated Access**: Queries diverse data sources (e.g., Databricks, PostgreSQL, on-premises systems) in real time, providing a unified view for comprehensive reporting, unlike siloed reporting tools that limit data scope. * **Semantic Models**: Aligns AI-generated insights with enterprise data, ensuring compliance and accuracy, reducing risks of irrelevant outputs in regulated healthcare environments. ## Example[​](#example "Direct link to Example") A health-tech company generates real-time HIPAA compliance reports by combining structured patient data from PostgreSQL with unstructured audit logs and regulatory guidelines from cloud storage. The AI Gateway produces narrative summaries highlighting potential violations and actionable recommendations, delivered to compliance officers in minutes. This outperforms static BI dashboards and generic RAG tools lacking real-time integration and regulatory context, ensuring faster, more accurate compliance decisions. The [Vector-Based Search documentation](/docs/next/features/search/vector-search) provides guidance for implementing VSS in RAG workflows. ## Benefits[​](#benefits "Direct link to Benefits") * **Compliance**: Governed data access meets stringent healthcare regulatory standards, ensuring auditability and trust. * **Timeliness**: Real-time data federation delivers up-to-date reports for rapid compliance and operational decisions. * **Insightfulness**: AI-driven narratives provide context-rich, actionable insights, improving healthcare operational efficiency. ### Learn More[​](#learn-more "Direct link to Learn More") * **Vector Similarity Search**: [Documentation](/docs/next/features/search) and [Searching GitHub Files Recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md). * **AI Gateway**: [Documentation](/docs/next/features/large-language-models) and [Running Llama3 Locally Recipe](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md). * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Semantic Model**: [Documentation](/docs/next/features/semantic-model). --- # Search & Retrieval ## [📄️Real-Time Data Collection](/docs/next/use-cases/search/data-collection-and-search) [Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic.](/docs/next/use-cases/search/data-collection-and-search) ## [📄️Enterprise Search](/docs/next/use-cases/search/enterprise-search) [Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities.](/docs/next/use-cases/search/enterprise-search) ## [📄️Object-Store Search](/docs/next/use-cases/search/object-store-search-engine) [Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search.](/docs/next/use-cases/search/object-store-search-engine) --- # Simplifying Real-Time Data Collection and Search Spice.ai processes streaming and static data with integrated search capabilities for real-time insights, focusing on application logic and enabling rapid development of data-driven features. Unlike complex streaming platforms (e.g., Apache Flink) that require extensive infrastructure, Spice.ai unifies streaming and static data with vector and hybrid search for developers. This simplifies real-time data workflows and research-driven applications, minimizing setup and maintenance overhead compared to fragmented streaming and search solutions. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Streamlined Access**: Queries streaming (e.g., Kafka, Databricks Delta Live Tables) and static sources in a single SQL interface, reducing pipeline complexity compared to tools requiring separate stream and batch processing. * **Low Latency**: Materializes real-time datasets near applications using Change Data Capture (CDC), delivering faster insights than cloud-based streaming solutions with network latency. * **Hybrid Search**: Combines vector similarity search (VSS) for semantic research with keyword and BM25 scoring for precise data retrieval, enabling rich, context-aware applications unlike standalone streaming platforms. * **Observability**: Built-in monitoring simplifies debugging of data flows and search performance, providing end-to-end visibility absent in fragmented streaming and search tools. ## Example[​](#example "Direct link to Example") A gaming platform processes live player activity from Kafka streams to power in-game leaderboards and personalized challenges, while using hybrid search to enable players to research game strategies by querying unstructured content (e.g., guides, forums) and structured metadata. This unified approach bypasses the infrastructure overhead of separate streaming pipelines and search engines, improving player engagement and feature delivery. The [Searching GitHub Files recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md) demonstrates real-time data processing and search patterns adaptable to such use cases. ## Benefits[​](#benefits "Direct link to Benefits") * **Developer Focus**: Shifts effort from pipeline and search infrastructure management to application development. * **Responsiveness**: Delivers real-time insights and search results critical for user engagement. * **Versatility**: Integrates streaming data with semantic and precise search for diverse application needs. ### Learn More[​](#learn-more "Direct link to Learn More") * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md). * **Vector and Hybrid Search**: [Documentation](/docs/next/features/search) and [Searching GitHub Files Recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md). * **Observability**: [Documentation](/docs/next/features/observability). --- # Enterprise Search and Retrieval Spice.ai powers semantic and precise search for financial services (finserv) knowledge bases with hybrid vector and keyword capabilities, enabling rapid access to critical information for compliance and decision-making. Unlike standalone search engines (e.g., Elasticsearch, OpenSearch) that lack built-in data and AI integration, Spice.ai combines hybrid search (vector + keyword + BM25), federated data access, and AI workflows to deliver enterprise-grade search experiences tailored for finserv's regulatory and performance demands. This approach reduces data duplication costs and improves contextual relevance, making it ideal for high-stakes financial applications. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Hybrid Search**: Balances semantic (vector) and exact-match (keyword/BM25) search for optimal relevance, retrieving compliance documents or transaction records with precision, unlike single-mode search engines that compromise on flexibility. * **Federated Data Access**: Searches across Databricks, on-premises databases, and cloud storage without data duplication, minimizing storage costs and compliance risks compared to centralized search platforms. * **Performance**: Materializes datasets and uses query push-down optimizations for faster results than traditional enterprise search solutions, critical for finserv’s high-volume, time-sensitive queries. * **AI Integration**: Enhances search results with the AI Gateway for intelligent ranking and summarization, improving usability in complex finserv environments over basic search tools. ## Example[​](#example "Direct link to Example") A finserv firm enables traders to search a knowledge base combining structured transaction data from Databricks with unstructured compliance documents from cloud storage. Hybrid search retrieves relevant regulations semantically (via vector search) and precise contract terms (via BM25), streamlining compliance checks and reducing research time compared to siloed search tools. This improves operational efficiency and regulatory adherence. The [Searching GitHub Files recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md) and [Vector-Based Search documentation](/docs/next/features/search) provide implementation guidance for hybrid search workflows. ## Benefits[​](#benefits "Direct link to Benefits") * **Precision**: Hybrid search delivers highly relevant results for regulatory and operational needs, improving decision-making accuracy. * **Efficiency**: Federated access eliminates data duplication, reducing storage and management costs in finserv. * **Productivity**: AI-enhanced search accelerates information retrieval, enabling faster responses in fast-paced financial markets. ### Learn More[​](#learn-more "Direct link to Learn More") * **Vector and Hybrid Search**: [Documentation](/docs/next/features/search) and [Searching GitHub Files Recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md). * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **AI Gateway**: [Documentation](/docs/next/features/large-language-models) and [Running Llama3 Locally Recipe](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md). --- # Object-Store Native Search Engine Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search with real-time insights directly from distributed storage. Unlike standalone search engines (e.g., Elasticsearch, OpenSearch) that require data ingestion into centralized indexes, Spice.ai queries object-store native databases (e.g., S3, Azure Blob) with hybrid search (vector + keyword + BM25) and federated data access, eliminating data duplication and reducing infrastructure overhead. This makes it ideal for security applications needing fast, compliant, and context-aware search across vast, distributed datasets, outperforming traditional search platforms with complex ETL requirements. ## Why Spice.ai?[​](#why-spiceai "Direct link to Why Spice.ai?") * **Object-Store Native Search**: Executes hybrid search directly on object-store data (e.g., S3-stored logs, Databricks Delta Lake) without moving data to centralized indexes, reducing costs and complexity compared to traditional search engines. * **Hybrid Search**: Combines vector similarity search (VSS) for semantic analysis (e.g., threat intelligence reports) with keyword/BM25 search for precise retrieval (e.g., specific log entries), delivering comprehensive results for security investigations. * **Data Federation**: Integrates object-store data with other sources (e.g., PostgreSQL, on-premises systems) via federated SQL queries, providing a unified view without replication, unlike siloed search platforms. * **Performance and Compliance**: Materializes hot datasets using Change Data Capture (CDC) for low-latency access and integrates with Databricks Unity Catalog for governance, ensuring compliance with security regulations (e.g., GDPR, SOC 2). ## Example[​](#example "Direct link to Example") A security operations platform uses Spice.ai to search S3-stored network logs and threat intelligence data, combining semantic VSS for identifying emerging threat patterns with BM25 for precise log entry retrieval. This enables rapid incident analysis without moving data to a centralized index, reducing costs and ensuring compliance compared to ETL-dependent search engines. The [Vector-Based Search documentation](/docs/next/features/search) and [Searching GitHub Files recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md) provide guidance for implementing hybrid search on object stores. ## Benefits[​](#benefits "Direct link to Benefits") * **Efficiency**: Direct search on object stores eliminates data movement, reducing infrastructure costs and complexity in security applications. * **Precision**: Hybrid search delivers relevant, context-aware results for threat detection and analysis. * **Compliance**: Governed data access aligns with strict security regulations, ensuring trust and auditability. ### Learn More[​](#learn-more "Direct link to Learn More") * **Vector and Hybrid Search**: [Documentation](/docs/next/features/search) and [Searching GitHub Files Recipe](https://github.com/spiceai/cookbook/blob/trunk/search_github_files/README.md). * **Federated SQL Queries**: [Documentation](/docs/next/features/query-federation) and [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md). * **Data Acceleration**: [Documentation](/docs/next/features/data-acceleration) and [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md). * **Observability**: [Documentation](/docs/next/features/observability). --- ## A[​](#A "Direct link to A") * [Acknowledgements1](/docs/v1.10/tags/acknowledgements "Project acknowledgements, credits, and attributions.") * [ADBC1](/docs/v1.10/tags/adbc "Arrow Database Connectivity driver configuration and usage.") * [API4](/docs/v1.10/tags/api "HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference.") * [Arrow1](/docs/v1.10/tags/arrow "Apache Arrow columnar data format integration.") * [Arrow Flight SQL1](/docs/v1.10/tags/arrow-flight-sql "Arrow Flight SQL protocol implementation and configuration.") * [Auth1](/docs/v1.10/tags/auth "Authentication and authorization mechanisms.") * [Authentication1](/docs/v1.10/tags/authentication "User authentication methods and security protocols.") * [Azure2](/docs/v1.10/tags/azure "Microsoft Azure cloud services and integrations.") *** --- ## [Open Source Acknowledgements](/docs/v1.10/acknowledgements) Spice AI acknowledges the following open source projects for making this project possible: --- ## [ADBC: Arrow Database Connectivity](/docs/v1.10/api/adbc) ADBC API Documentation --- ## [ADBC: Arrow Database Connectivity](/docs/v1.10/api/adbc) ADBC API Documentation --- ## [ADBC: Arrow Database Connectivity](/docs/v1.10/api/adbc) ADBC API Documentation --- ## [Arrow Flight SQL API](/docs/v1.10/api/arrow-flight-sql) Query Spice using JDBC/ODBC/ADBC --- ## [Authentication](/docs/v1.10/api/auth) Authentication documentation --- ## [login](/docs/v1.10/cli/reference/login) Login to the Spice.ai Platform, or other services with sub-commands. --- ## [Azure BlobFS Data Connector](/docs/v1.10/components/data-connectors/abfs) Azure BlobFS Data Connector Documentation --- ## [Azure BlobFS Data Connector](/docs/v1.10/components/data-connectors/abfs) Azure BlobFS Data Connector Documentation --- ## [Caching](/docs/v1.10/features/caching) Learn how to use Spice in-memory caching --- ## [Catalog Connectors](/docs/v1.10/components/catalogs) --- ## [Cayenne Data Accelerator](/docs/v1.10/components/data-accelerators/cayenne) Cayenne Data Accelerator (Vortex) Documentation --- ## [Configuring Trace Levels](/docs/v1.10/cli/tracing) Configuring Spice.ai OSS trace output verbosity levels --- ## [Component Metrics](/docs/v1.10/features/observability/component_metrics) Learn how to enable optional component metrics. --- ## [Embedding Models](/docs/v1.10/components/embeddings) Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. --- ## [Language Model Overrides](/docs/v1.10/features/large-language-models/parameter_overrides) Learn how to override default LLM hyperparameters in Spice. --- ## [Cayenne Data Accelerator](/docs/v1.10/components/data-accelerators/cayenne) Cayenne Data Accelerator (Vortex) Documentation --- ## [Azure BlobFS Data Connector](/docs/v1.10/components/data-connectors/abfs) Azure BlobFS Data Connector Documentation --- ## [Delta Lake Data Connector](/docs/v1.10/components/data-connectors/delta-lake) Delta Lake Data Connector Documentation --- ## [Databricks Catalog Connector](/docs/v1.10/components/catalogs/databricks) Connect to a Databricks Unity Catalog provider. --- ## [Datasets](/docs/v1.10/reference/spicepod/datasets) Datasets YAML reference --- ## [Debezium Data Connector](/docs/v1.10/components/data-connectors/debezium) Debezium Data Connector Documentation --- ## [Troubleshooting Spice](/docs/v1.10/troubleshooting) Review and debug runtime tasks, logs, and diagnostic steps in Spice. --- ## [Databricks Data Connector](/docs/v1.10/components/data-connectors/databricks) Databricks Data Connector Documentation --- ## [Open Source Acknowledgements](/docs/v1.10/acknowledgements) Spice AI acknowledges the following open source projects for making this project possible: --- ## [Docker - Kubernetes](/docs/v1.10/deployment/docker) Running Spice.ai as Docker container --- ## [Docker - Kubernetes](/docs/v1.10/deployment/docker) Running Spice.ai as Docker container --- ## [DynamoDB Data Connector](/docs/v1.10/components/data-connectors/dynamodb) DynamoDB Data Connector Documentation --- ## [Azure OpenAI Embedding Models](/docs/v1.10/components/embeddings/azure) To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: --- ## [Evaluating Language Models](/docs/v1.10/features/large-language-models/evals) Learn how Spice evaluates, tracks, compares, and improves language model performance for specific tasks --- ## [Data Ingestion](/docs/v1.10/features/data-ingestion) Learn how to ingest data in Spice. --- ## [Data Connectors](/docs/v1.10/components/data-connectors) Learn how to use Data Connector to query external data. --- ## [Spicepods](/docs/v1.10/getting-started/spicepods) An introduction to Spicepods --- ## [GitHub Data Connector](/docs/v1.10/components/data-connectors/github) GitHub Data Connector Documentation --- ## [Glue Catalog Connector](/docs/v1.10/components/catalogs/glue) Connect to an AWS Glue Data Catalog. --- ## [Iceberg Catalog Connector](/docs/v1.10/components/catalogs/iceberg) Connect to an Iceberg catalog provider. --- ## [Memory Data Connector](/docs/v1.10/components/data-connectors/memory) Memory Data Connector Documentation --- ## [GitHub Data Connector](/docs/v1.10/components/data-connectors/github) GitHub Data Connector Documentation --- ## [Arrow Flight SQL API](/docs/v1.10/api/arrow-flight-sql) Query Spice using JDBC/ODBC/ADBC --- ## [Kafka Data Connector](/docs/v1.10/components/data-connectors/kafka) Kafka Data Connector Documentation --- ## [Docker - Kubernetes](/docs/v1.10/deployment/docker) Running Spice.ai as Docker container --- ## [Configuring Trace Levels](/docs/v1.10/cli/tracing) Configuring Spice.ai OSS trace output verbosity levels --- ## [login](/docs/v1.10/cli/reference/login) Login to the Spice.ai Platform, or other services with sub-commands. --- ## [YAML syntax for Spicepod manifests](/docs/v1.10/reference/spicepod) Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) --- ## [Model Context Protocol (MCP)](/docs/v1.10/features/large-language-models/mcp) Learn how to use the Model Context Protocol (MCP) with Spice. --- ## [Language Model Memory](/docs/v1.10/features/large-language-models/memory) Learn how to provide LLMs with memory --- ## [Embedding Models](/docs/v1.10/components/embeddings) Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. --- ## [MongoDB Data Connector](/docs/v1.10/components/data-connectors/mongodb) MongoDB Data Connector Documentation --- ## [MySQL Data Connector](/docs/v1.10/components/data-connectors/mysql) MySQL Data Connector Documentation --- ## [DynamoDB Data Connector](/docs/v1.10/components/data-connectors/dynamodb) DynamoDB Data Connector Documentation --- ## [Zipkin Integration](/docs/v1.10/monitoring/zipkin) Learn how to integrate Spice with Zipkin tracing. --- ## [Arrow Flight SQL API](/docs/v1.10/api/arrow-flight-sql) Query Spice using JDBC/ODBC/ADBC --- ## [Open Source Acknowledgements](/docs/v1.10/acknowledgements) Spice AI acknowledges the following open source projects for making this project possible: --- ## [Azure OpenAI Embedding Models](/docs/v1.10/components/embeddings/azure) To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: --- ## [Oracle Data Connector](/docs/v1.10/components/data-connectors/oracle) Oracle Data Connector Documentation --- ## [Language Model Overrides](/docs/v1.10/features/large-language-models/parameter_overrides) Learn how to override default LLM hyperparameters in Spice. --- ## [Catalog Connectors](/docs/v1.10/components/catalogs) --- ## [Language Model Overrides](/docs/v1.10/features/large-language-models/parameter_overrides) Learn how to override default LLM hyperparameters in Spice. --- ## [Cayenne Data Accelerator](/docs/v1.10/components/data-accelerators/cayenne) Cayenne Data Accelerator (Vortex) Documentation --- ## [Language Model Memory](/docs/v1.10/features/large-language-models/memory) Learn how to provide LLMs with memory --- ## [Datasets](/docs/v1.10/reference/spicepod/datasets) Datasets YAML reference --- ## [MySQL Data Connector](/docs/v1.10/components/data-connectors/mysql) MySQL Data Connector Documentation --- ## [Language Models Tools](/docs/v1.10/features/large-language-models/tools) Learn how LLMs interact with the Spice runtime. --- ## [Docker Sandbox Guide - v1.3.0](/docs/v1.10/deployment/docker/sandbox) Migrating to v1.3.0 --- ## [Embedding Models](/docs/v1.10/components/embeddings) Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. --- ## [Authentication](/docs/v1.10/api/auth) Authentication documentation --- ## [Docker - Kubernetes](/docs/v1.10/deployment/docker) Running Spice.ai as Docker container --- ## [Datasets](/docs/v1.10/reference/spicepod/datasets) Datasets YAML reference --- ## [Arrow Flight SQL API](/docs/v1.10/api/arrow-flight-sql) Query Spice using JDBC/ODBC/ADBC --- ## [Language Model Memory](/docs/v1.10/features/large-language-models/memory) Learn how to provide LLMs with memory --- ## [Configuring Trace Levels](/docs/v1.10/cli/tracing) Configuring Spice.ai OSS trace output verbosity levels --- ## [Troubleshooting Spice](/docs/v1.10/troubleshooting) Review and debug runtime tasks, logs, and diagnostic steps in Spice. --- ## [Unity Catalog Catalog Connector](/docs/v1.10/components/catalogs/unity-catalog) Connect to a Unity Catalog provider. --- ## [Views](/docs/v1.10/reference/spicepod/views) Views YAML reference --- ## [Cayenne Data Accelerator](/docs/v1.10/components/data-accelerators/cayenne) Cayenne Data Accelerator (Vortex) Documentation --- ## [Data Ingestion](/docs/v1.10/features/data-ingestion) Learn how to ingest data in Spice. --- ## [YAML syntax for Spicepod manifests](/docs/v1.10/reference/spicepod) Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) --- ## [Zipkin Integration](/docs/v1.10/monitoring/zipkin) Learn how to integrate Spice with Zipkin tracing. --- # Spice.ai Open Source **Spice** is a SQL query, search, and LLM-inference engine, written in Rust, for data-driven applications and AI agents. ![Spice.ai Open Source Data Query \& AI-Inference Compute Engine](/assets/images/spice.ai-compute-engine-new-c9b4c58d2c615acbe2537023202c09c9.png) Spice provides four industry-standard APIs in a lightweight, portable runtime (single \~140 MB binary): 1. **SQL Query & Search APIs**: Supports HTTP, Arrow Flight, Arrow Flight SQL, ODBC, JDBC, ADBC, and `vector_search` and `text_search` UDTFs. 2. **OpenAI-Compatible APIs**: Provides HTTP APIs for OpenAI SDK compatibility, local model serving (CUDA/Metal accelerated), and hosted model gateway. 3. **Iceberg Catalog REST APIs**: Offers a unified API for Iceberg Catalog. 4. **MCP HTTP+SSE APIs**: Enables integration with external tools via Model Context Protocol (MCP) using HTTP and Server-Sent Events (SSE). Goal 🎯 Developers can focus on building data apps and AI agents confidently, knowing they are grounded in data. Spice is primarily used for: * **Data Federation**: SQL query across any database, data warehouse, or data lake. [Learn More](https://spiceai.org/docs/features/query-federation). * **Data Materialization and Acceleration**: Materialize, accelerate, and cache database queries. [Read the MaterializedView interview - Building a CDN for Databases](https://materializedview.io/p/building-a-cdn-for-databases-spice-ai). * **Enterprise Search**: Keyword, vector, and full-text search with Tantivy-powered BM25 and vector similarity search for structured and unstructured data. [Learn More](https://github.com/spiceai/cookbook/blob/trunk/vectors/s3/README.md). * **AI Apps and Agents**: An AI-database powering retrieval-augmented generation (RAG) and intelligent agents. [Learn More](https://spiceai.org/docs/use-cases/rag). Watch 🎥 * [CMU Databases - Accelerating Data and AI with Spice.ai Open-Source](https://www.youtube.com/watch?v=tyM-ec1lKfU) * [Query Data using Spice, OpenAI, and MCP](https://www.youtube.com/watch?v=TFAu4qxjTPk\&list=PLesJrUXEx3U-dQul0PqLV3TGTdUmr3B6e\&index=8) * [Search with Amazon S3 Vectors](https://www.youtube.com/watch?v=QPbqPf5W36g) Spice is built on industry-leading technologies including [Apache DataFusion](https://datafusion.apache.org), Apache Arrow, Arrow Flight, SQLite, and DuckDB. If you want to build with DataFusion or DuckDB, Spice provides a simple, flexible, and production-ready engine. ![How Spice works](https://github.com/spiceai/spiceai/assets/80174/7d93ae32-d6d8-437b-88d3-d64fe089e4b7) Announcement 📣 Read the [Spice.ai 1.0-stable announcement](https://spiceai.org/blog/announcing-1.0-stable). ## Why Spice?[​](#why-spice "Direct link to Why Spice?") Spice simplifies building data-driven AI applications and agents by enabling fast, SQL-based querying, federation, and acceleration of data from multiple sources, grounding AI in real-time, reliable data. Co-locate datasets with apps and AI models to power AI feedback loops, enable RAG and search, and deliver low-latency data queries and AI inference with full control over cost and performance. ![Spice.ai](https://github.com/spiceai/spiceai/assets/80174/29e4421d-8942-4f2a-8397-e9d4fdeda36b) ### How is Spice different?[​](#how-is-spice-different "Direct link to How is Spice different?") 1. **AI-Native Runtime**: Combines data query and AI inference in a single engine for data-grounded, accurate AI. 2. **Application-Focused**: Designed for distributed deployment at the application or agent level, often as a 1:1 or 1 :N mapping, unlike centralized databases serving multiple apps. Multiple Spice instances can be deployed, even one per tenant or customer. 3. **Dual-Engine Acceleration**: Supports **OLAP** (Arrow/DuckDB) and **OLTP** (SQLite/PostgreSQL) engines at the dataset level for flexible performance across analytical and transactional workloads. 4. **Disaggregated Storage**: Separates compute from storage, co-locating materialized working datasets with applications, dashboards, or ML pipelines while accessing source data in its original storage. 5. **Edge to Cloud Native**: Deploy as a standalone instance, Kubernetes sidecar, microservice, or cluster across edge, on-prem, and public clouds. Chain multiple Spice instances for tier-optimized, distributed deployments. ## How does Spice compare?[​](#how-does-spice-compare "Direct link to How does Spice compare?") ### Data Query and Analytics[​](#data-query-and-analytics "Direct link to Data Query and Analytics") | Feature | **Spice** | Trino / Presto | Dremio | ClickHouse | Materialize | | -------------------------------- | -------------------------------------- | -------------------- | --------------------- | ------------------- | -------------------- | | **Primary Use-Case** | Data & AI apps/agents | Big data analytics | Interactive analytics | Real-time analytics | Real-time analytics | | **Primary Deployment Model** | Sidecar | Cluster | Cluster | Cluster | Cluster | | **Federated Query Support** | ✅ | ✅ | ✅ | ❌ | ❌ | | **Acceleration/Materialization** | ✅ (Arrow, SQLite, DuckDB, PostgreSQL) | Intermediate storage | Reflections (Iceberg) | Materialized views | ✅ (Real-time views) | | **Catalog Support** | ✅ (Iceberg, Unity Catalog, AWS Glue) | ✅ | ✅ | ❌ | ❌ | | **Query Result Caching** | ✅ | ✅ | ✅ | ✅ | Limited | | **Multi-Modal Acceleration** | ✅ (OLAP + OLTP) | ❌ | ❌ | ❌ | ❌ | | **Change Data Capture (CDC)** | ✅ (Debezium) | ❌ | ❌ | ❌ | ✅ (Debezium) | ### AI Apps and Agents[​](#ai-apps-and-agents "Direct link to AI Apps and Agents") | Feature | **Spice** | LangChain | LlamaIndex | AgentOps.ai | Ollama | | ----------------------------- | ------------------------------------ | ------------------ | ---------- | ---------------- | ----------------------------- | | **Primary Use-Case** | Data & AI apps | Agentic workflows | RAG apps | Agent operations | LLM apps | | **Programming Language** | Any language (HTTP interface) | JavaScript, Python | Python | Python | Any language (HTTP interface) | | **Unified Data + AI Runtime** | ✅ | ❌ | ❌ | ❌ | ❌ | | **Federated Data Query** | ✅ | ❌ | ❌ | ❌ | ❌ | | **Accelerated Data Access** | ✅ | ❌ | ❌ | ❌ | ❌ | | **Tools/Functions** | ✅ (MCP HTTP+SSE) | ✅ | ✅ | Limited | Limited | | **LLM Memory** | ✅ | ✅ | ❌ | ✅ | ❌ | | **Evaluations (Evals)** | ✅ | Limited | ❌ | Limited | ❌ | | **Search** | ✅ (Keyword, Vector, Full-Text) | ✅ | ✅ | Limited | Limited | | **Caching** | ✅ (Query and results caching) | Limited | ❌ | ❌ | ❌ | | **Embeddings** | ✅ (Built-in & pluggable models/DBs) | ✅ | ✅ | Limited | ❌ | ✅ = Fully supported
❌ = Not supported
Limited = Partial or restricted support ## Example Use-Cases[​](#example-use-cases "Direct link to Example Use-Cases") ### Data-Grounded Agentic AI Applications[​](#data-grounded-agentic-ai-applications "Direct link to Data-Grounded Agentic AI Applications") * **OpenAI-Compatible API**: Connect to hosted models (OpenAI, Anthropic, xAI) or deploy locally (Llama, NVIDIA NIM). [AI Gateway Recipe](https://github.com/spiceai/cookbook/blob/trunk/openai_sdk/README.md) * **Federated Data Access**: Query using SQL and NSQL (text-to-SQL) across databases, data warehouses, and data lakes with advanced query push-down. [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) * **Search and RAG**: Perform keyword, vector, and full-text search with Tantivy-powered BM25 and vector similarity search (VSS) integrated into SQL queries using `vector_search` and `text_search`. Supports multi-column vector search with reciprocal rank fusion. [Amazon S3 Vectors Recipe](https://github.com/spiceai/cookbook/blob/trunk/vectors/s3/README.md) * **LLM Memory and Observability**: Store and retrieve history and context for AI agents with visibility into data flows, model performance, and traces. [LLM Memory Recipe](https://github.com/spiceai/cookbook/blob/trunk/llm-memory/README.md) | [Observability Documentation](https://spiceai.org/docs/features/observability) ### Database CDN and Query Mesh[​](#database-cdn-and-query-mesh "Direct link to Database CDN and Query Mesh") * **Data Acceleration**: Co-locate materialized datasets in Arrow, SQLite, or DuckDB for sub-second queries. [DuckDB Data Accelerator Recipe](https://github.com/spiceai/cookbook/blob/trunk/duckdb/accelerator/README.md) * **Resiliency and Local Dataset Replication**: Maintain availability with local replicas of critical datasets. [Local Dataset Replication Recipe](https://github.com/spiceai/cookbook/blob/trunk/localpod/README.md) * **Responsive Dashboards**: Enable fast, real-time analytics for frontends and BI tools. [Sales BI Dashboard Demo](https://github.com/spiceai/cookbook/blob/trunk/sales-bi/README.md) * **Simplified Legacy Migration**: Unify legacy systems with modern infrastructure via federated SQL querying. [Federated SQL Query Recipe](https://github.com/spiceai/cookbook/blob/trunk/federation/README.md) ### Retrieval-Augmented Generation (RAG)[​](#retrieval-augmented-generation-rag "Direct link to Retrieval-Augmented Generation (RAG)") * **Unified Search with Vector Similarity**: Perform efficient vector similarity search across structured and unstructured data, with native support for Amazon S3 Vectors for petabyte-scale storage and querying. Supports distance metrics like cosine similarity, Euclidean distance, or dot product. [Amazon S3 Vectors Recipe](https://github.com/spiceai/cookbook/blob/trunk/vectors/s3/README.md) * **Semantic Knowledge Layer**: Define a semantic context model to enrich data for AI. [Semantic Model Documentation](/docs/features/semantic-model) * **Text-to-SQL**: Convert natural language queries into SQL using built-in NSQL and sampling tools. [Text-to-SQL Recipe](https://github.com/spiceai/cookbook/blob/trunk/text-to-sql/README.md) * **Model and Data Evaluations**: Assess model performance and data quality with integrated evaluation tools. [Language Model Evaluations Recipe](https://github.com/spiceai/cookbook/blob/trunk/evals/README.md) ## FAQ[​](#faq "Direct link to FAQ") * **Is Spice a cache?** No, but its data acceleration acts as an active cache, materialization, or prefetcher. Unlike traditional caches that fetch on miss, Spice prefetches and materializes filtered data on intervals, triggers, or via CDC. It also supports [results caching](https://spiceai.org/docs/features/caching). * **Is Spice a CDN for databases?** Yes, Spice enables shipping working datasets to where they're accessed most, like data-intensive applications or AI contexts, similar to a CDN. [Docs FAQ](/docs/faq) ### Watch a 30-second BI dashboard acceleration demo[​](#watch-a-30-second-bi-dashboard-acceleration-demo "Direct link to Watch a 30-second BI dashboard acceleration demo") []() See more demos on [YouTube](https://www.youtube.com/playlist?list=PLesJrUXEx3U9anekJvbjyyTm7r9A26ugK). ### Intelligent Applications and Agents[​](#intelligent-applications-and-agents "Direct link to Intelligent Applications and Agents") Spice enables developers to build data-grounded AI applications and agents by co-locating data and ML models with applications. Read more about the vision for [intelligent AI-driven applications](/docs/intelligent-applications). ### Connect with Us[​](#connect-with-us "Direct link to Connect with Us") * Build with Spice and share feedback at , [Slack](https://spice.ai/slack), [X](https://twitter.com/spice_ai), or [LinkedIn](https://www.linkedin.com/company/74148478). * [File an issue](https://github.com/spiceai/spiceai/issues/new) for bugs or issues. * Join our team ([We're hiring!](https://spice.ai/careers)). * Contribute code or documentation ([CONTRIBUTING.md](https://github.com/spiceai/spiceai/blob/trunk/CONTRIBUTING)). * ⭐️ Star the [Spice.ai repo](https://github.com/spiceai/spiceai) to show support! --- # Open Source Acknowledgements Spice AI acknowledges the following open source projects for making this project possible: ## Go Modules[​](#go-modules "Direct link to Go Modules") github.com/AzureAD/microsoft-authentication-library-for-go/apps, , MIT github.com/apache/arrow/go/v17, , Apache-2.0 github.com/cenkalti/backoff/v4, , MIT github.com/chzyer/readline, , MIT github.com/dustin/go-humanize, , MIT github.com/fsnotify/fsnotify, , BSD-3-Clause github.com/go-logr/logr, , Apache-2.0 github.com/go-logr/stdr, , Apache-2.0 github.com/go-viper/mapstructure/v2, , MIT github.com/gocarina/gocsv, , MIT github.com/goccy/go-json, , MIT github.com/golang-jwt/jwt/v5, , MIT github.com/google/flatbuffers/go, , Apache-2.0 github.com/google/uuid, , BSD-3-Clause github.com/hashicorp/go-cleanhttp, , MPL-2.0 github.com/hashicorp/go-retryablehttp, , MPL-2.0 github.com/joho/godotenv, , MIT github.com/klauspost/compress, , Apache-2.0 github.com/klauspost/compress/internal/snapref, , BSD-3-Clause github.com/klauspost/compress/zstd/internal/xxhash, , MIT github.com/kylelemons/godebug, , Apache-2.0 github.com/logrusorgru/aurora, , Unlicense github.com/manifoldco/promptui, , BSD-3-Clause github.com/mattn/go-runewidth, , MIT github.com/olekukonko/tablewriter, , MIT github.com/pelletier/go-toml/v2, , MIT github.com/peterh/liner, , MIT github.com/pierrec/lz4/v4, , BSD-3-Clause github.com/pkg/browser, , BSD-2-Clause github.com/rivo/uniseg, , MIT github.com/sagikazarmark/locafero, , MIT github.com/sourcegraph/conc, , MIT github.com/spf13/afero, , Apache-2.0 github.com/spf13/cast, , MIT github.com/spf13/cobra, , Apache-2.0 github.com/spf13/pflag, , BSD-3-Clause github.com/spf13/viper, , MIT github.com/spiceai/gospice/v7, , Apache-2.0 github.com/subosito/gotenv, , MIT github.com/zeebo/xxh3, , BSD-2-Clause go.opentelemetry.io/otel, , Apache-2.0 go.opentelemetry.io/otel/metric, , Apache-2.0 go.opentelemetry.io/otel/sdk, , Apache-2.0 go.opentelemetry.io/otel/trace, , Apache-2.0 go.yaml.in/yaml/v3, , MIT golang.org/x/exp, , BSD-3-Clause golang.org/x/mod/semver, , BSD-3-Clause golang.org/x/net, , BSD-3-Clause golang.org/x/sys, , BSD-3-Clause golang.org/x/text, , BSD-3-Clause golang.org/x/xerrors, , BSD-3-Clause google.golang.org/genproto/googleapis/rpc/status, , Apache-2.0 google.golang.org/grpc, , Apache-2.0 google.golang.org/protobuf, , BSD-3-Clause gopkg.in/yaml.v3, , MIT ## Rust Crates[​](#rust-crates "Direct link to Rust Crates") * aegis 0.9.3, MIT
* ahash 0.7.8, Apache-2.0 OR MIT
* ahash 0.8.12, Apache-2.0 OR MIT
* ansi\_term 0.12.1, MIT
* anyhow 1.0.100, Apache-2.0 OR MIT
* arrow 56.0.0, Apache-2.0
* arrow-array 56.0.0, Apache-2.0
* arrow-buffer 56.0.0, Apache-2.0
* arrow-cast 56.0.0, Apache-2.0
* arrow-csv 56.0.0, Apache-2.0
* arrow-flight 56.0.0, Apache-2.0
* arrow-ipc 56.0.0, Apache-2.0
* arrow-json 56.0.0, Apache-2.0
* arrow-odbc 18.1.2, MIT
* arrow-schema 56.0.0, Apache-2.0
* async-graphql 7.0.17, Apache-2.0 OR MIT
* async-graphql-axum 7.0.17, Apache-2.0 OR MIT
* async-openai 0.29.1, MIT
* async-stream 0.3.6, MIT
* async-trait 0.1.89, Apache-2.0 OR MIT
* aws-config 1.8.8, Apache-2.0
* aws-credential-types 1.2.8, Apache-2.0
* aws-runtime 1.5.12, Apache-2.0
* aws-sdk-bedrockruntime 1.109.0, Apache-2.0
* aws-sdk-cognitoidentity 1.87.0, Apache-2.0
* aws-sdk-cognitoidentityprovider 1.100.0, Apache-2.0
* aws-sdk-dynamodb 1.95.0, Apache-2.0
* aws-sdk-glue 1.124.0, Apache-2.0
* aws-sdk-s3 1.108.0, Apache-2.0
* aws-sdk-s3vectors 1.11.0, Apache-2.0
* aws-sdk-secretsmanager 1.90.0, Apache-2.0
* aws-sdk-sts 1.88.0, Apache-2.0
* aws-smithy-async 1.2.6, Apache-2.0
* aws-smithy-runtime 1.9.3, Apache-2.0
* aws-smithy-runtime-api 1.9.1, Apache-2.0
* aws-smithy-types 1.3.3, Apache-2.0
* axum 0.8.6, MIT
* axum-extra 0.10.3, MIT
* azure\_core 0.21.0, MIT
* azure\_core 0.28.0, MIT
* azure\_storage 0.21.0, MIT
* azure\_storage\_blobs 0.21.0, MIT
* backoff 0.4.0, Apache-2.0 OR MIT
* ballista 50.0.0, Apache-2.0
* ballista-core 50.0.0, Apache-2.0
* ballista-executor 50.0.0, Apache-2.0
* ballista-scheduler 50.0.0, Apache-2.0
* base64 0.13.1, Apache-2.0 OR MIT
* base64 0.21.7, Apache-2.0 OR MIT
* base64 0.22.1, Apache-2.0 OR MIT
* bb8 0.8.6, MIT
* bb8 0.9.0, MIT
* bb8-oracle 0.3.0, Apache-2.0 OR MIT
* bigdecimal 0.4.9, Apache-2.0 OR MIT
* bollard 0.18.1, Apache-2.0
* byte-unit 5.1.6, MIT
* bytes 1.10.1, MIT
* charset 0.1.5, Apache-2.0 OR MIT
* chrono 0.4.42, Apache-2.0 OR MIT
* chrono-tz 0.8.6, Apache-2.0 OR MIT
* chrono-tz 0.9.0, Apache-2.0 OR MIT
* chrono-tz 0.10.4, Apache-2.0 OR MIT
* clap 4.5.51, Apache-2.0 OR MIT
* clickhouse-rs 1.1.0-alpha.1, MIT
* criterion 0.5.1, Apache-2.0 OR MIT
* criterion 0.7.0, Apache-2.0 OR MIT
* croner 3.0.1, MIT
* csv 1.4.0, MIT OR Unlicense
* ctor 0.5.0, Apache-2.0 OR MIT
* ctrlc 3.5.1, Apache-2.0 OR MIT
* cudarc 0.12.2, Apache-2.0 OR MIT
* cudarc 0.13.9, Apache-2.0 OR MIT
* dashmap 6.1.0, MIT
* datafusion 50.3.0, Apache-2.0
* datafusion-catalog 50.3.0, Apache-2.0
* datafusion-common 50.3.0, Apache-2.0
* datafusion-datasource 50.3.0, Apache-2.0
* datafusion-execution 50.3.0, Apache-2.0
* datafusion-expr 50.3.0, Apache-2.0
* datafusion-federation 0.4.2, Apache-2.0
* datafusion-functions-json 0.50.0, Apache-2.0
* datafusion-physical-expr 50.3.0, Apache-2.0
* datafusion-physical-plan 50.3.0, Apache-2.0
* datafusion-proto 50.3.0, Apache-2.0
* datafusion-table-providers 0.1.0, Apache-2.0
* delta\_kernel 0.14.0, Apache-2.0
* derive\_builder 0.20.2, Apache-2.0 OR MIT
* dirs 6.0.0, Apache-2.0 OR MIT
* docx-rs 0.4.17, MIT
* dotenvy 0.15.7, MIT
* duckdb 1.3.2, MIT
* dyn-clone 1.0.20, Apache-2.0 OR MIT
* either 1.15.0, Apache-2.0 OR MIT
* env\_logger 0.11.8, Apache-2.0 OR MIT
* evalconverter 0.1.0, Apache-2.0
* fundu 2.0.1, MIT
* futures 0.3.31, Apache-2.0 OR MIT
* futures-util 0.3.31, Apache-2.0 OR MIT
* git2 0.20.2, Apache-2.0 OR MIT
* globset 0.4.18, MIT OR Unlicense
* governor 0.10.1, MIT
* graph-rs-sdk 2.0.1, MIT
* graphql-parser 0.4.1, Apache-2.0 OR MIT
* headers-accept 0.1.4, MIT
* hf-hub 0.4.3, Apache-2.0
* hostname 0.3.1, MIT
* hostname 0.4.1, MIT
* http 0.2.12, Apache-2.0 OR MIT
* http 1.3.1, Apache-2.0 OR MIT
* http-body-util 0.1.3, MIT
* humantime 2.3.0, Apache-2.0 OR MIT
* hyper 0.14.32, MIT
* hyper 1.7.0, MIT
* hyper-util 0.1.17, MIT
* iceberg 0.7.0, Apache-2.0
* iceberg-catalog-glue 0.7.0, Apache-2.0
* iceberg-catalog-rest 0.7.0, Apache-2.0
* iceberg-datafusion 0.7.0, Apache-2.0
* iceberg\_test\_utils 0.7.0, Apache-2.0
* imap 3.0.0-alpha.14, Apache-2.0 OR MIT
* indexmap 1.9.3, Apache-2.0 OR MIT
* indexmap 2.12.0, Apache-2.0 OR MIT
* indicatif 0.17.11, MIT
* indicatif 0.18.2, MIT
* insta 1.43.2, Apache-2.0
* itertools 0.10.5, Apache-2.0 OR MIT
* itertools 0.12.1, Apache-2.0 OR MIT
* itertools 0.13.0, Apache-2.0 OR MIT
* itertools 0.14.0, Apache-2.0 OR MIT
* jsonpath-rust 0.7.5, MIT
* jsonwebtoken 9.3.1, MIT
* keyring 3.6.3, Apache-2.0 OR MIT
* log 0.4.28, Apache-2.0 OR MIT
* logos 0.15.1, Apache-2.0 OR MIT
* mailparse 0.15.0, 0BSD
* mediatype 0.19.20, MIT
* mimalloc 0.1.48, MIT
* mistralrs 0.6.0, MIT
* mistralrs-core 0.6.0, MIT
* model2vec-rs 0.1.3, LICENSE
* moka 0.12.11, (Apache-2.0 OR MIT) AND Apache-2.0
* mongodb 3.3.0, Apache-2.0
* mysql\_async 0.35.1, Apache-2.0 OR MIT
* ndarray 0.15.6, Apache-2.0 OR MIT
* ndarray 0.16.1, Apache-2.0 OR MIT
* nix 0.30.1, MIT
* notify 8.2.0, CC0-1.0
* num\_cpus 1.17.0, Apache-2.0 OR MIT
* object\_store 0.12.4, Apache-2.0 OR MIT
* octocrab 0.47.0, Apache-2.0 OR MIT
* odbc-api 13.1.0, MIT
* once\_cell 1.21.3, Apache-2.0 OR MIT
* opendal 0.54.1, Apache-2.0
* opentelemetry 0.27.1, Apache-2.0
* opentelemetry 0.30.0, Apache-2.0
* opentelemetry-http 0.27.0, Apache-2.0
* opentelemetry-prometheus 0.27.0, Apache-2.0
* opentelemetry-proto 0.30.0, Apache-2.0
* opentelemetry-zipkin 0.27.0, Apache-2.0
* opentelemetry\_sdk 0.27.1, Apache-2.0
* opentelemetry\_sdk 0.30.0, Apache-2.0
* oracle 0.6.3, Apache-2.0 OR UPL-1.0
* parking\_lot 0.11.2, Apache-2.0 OR MIT
* parking\_lot 0.12.5, Apache-2.0 OR MIT
* parquet 56.0.0, Apache-2.0
* paste 1.0.15, Apache-2.0 OR MIT
* path-clean 1.0.1, Apache-2.0 OR MIT
* pdf-extract 0.8.0, MIT
* percent-encoding 2.3.2, Apache-2.0 OR MIT
* pin-project 1.1.10, Apache-2.0 OR MIT
* pkcs8 0.9.0, Apache-2.0 OR MIT
* pkcs8 0.10.2, Apache-2.0 OR MIT
* postcard 1.1.3, Apache-2.0 OR MIT
* prometheus 0.13.4, Apache-2.0
* prometheus-parse 0.2.5, Apache-2.0
* prost 0.11.9, Apache-2.0
* prost 0.12.6, Apache-2.0
* prost 0.13.5, Apache-2.0
* prost 0.14.1, Apache-2.0
* pulldown-cmark 0.12.2, MIT
* pulldown-cmark 0.13.0, MIT
* rand 0.7.3, Apache-2.0 OR MIT
* rand 0.8.5, Apache-2.0 OR MIT
* rand 0.9.2, Apache-2.0 OR MIT
* rayon 1.11.0, Apache-2.0 OR MIT
* rdkafka 0.38.0, MIT
* regex 1.12.2, Apache-2.0 OR MIT
* reqwest 0.12.24, Apache-2.0 OR MIT
* reqwest-eventsource 0.6.0, Apache-2.0 OR MIT
* rmcp 0.1.5, Apache-2.0 OR MIT
* rstest 0.25.0, Apache-2.0 OR MIT
* rstest 0.26.1, Apache-2.0 OR MIT
* rusqlite 0.31.0, MIT
* rustls 0.21.12, Apache-2.0 OR ISC OR MIT
* rustls 0.23.34, Apache-2.0 OR ISC OR MIT
* rustls-native-certs 0.6.3, Apache-2.0 OR ISC OR MIT
* rustls-native-certs 0.8.2, Apache-2.0 OR ISC OR MIT
* rustls-pemfile 1.0.4, Apache-2.0 OR ISC OR MIT
* rustls-pemfile 2.2.0, Apache-2.0 OR ISC OR MIT
* rustyline 17.0.2, MIT
* schemars 0.8.22, MIT
* schemars 0.9.0, MIT
* schemars 1.0.4, MIT
* scopeguard 1.2.0, Apache-2.0 OR MIT
* secrecy 0.10.3, Apache-2.0 OR MIT
* serde 1.0.228, Apache-2.0 OR MIT
* serde-value 0.7.0, MIT
* serde\_json 1.0.145, Apache-2.0 OR MIT
* serde\_yaml 0.9.34+deprecated, Apache-2.0 OR MIT
* sha2 0.10.9, Apache-2.0 OR MIT
* snafu 0.8.9, Apache-2.0 OR MIT
* snmalloc-rs 0.3.8, MIT
* snowflake-api 0.9.0, Apache-2.0
* spark-connect-rs 0.0.1-beta.4, Apache-2.0
* spiceai 3.1.0, Apache-2.0
* sqlparser 0.58.0, Apache-2.0
* ssh2 0.9.5, Apache-2.0 OR MIT
* strsim 0.10.0, MIT
* strsim 0.11.1, MIT
* suppaftp 5.4.0, Apache-2.0
* sysinfo 0.30.13, MIT
* sysinfo 0.35.2, MIT
* tantivy 0.24.2, MIT
* tempfile 3.23.0, Apache-2.0 OR MIT
* tera 1.20.0, MIT
* text-embeddings-backend 1.8.2,
* text-embeddings-backend-candle 1.8.2,
* text-embeddings-backend-core 1.8.2,
* text-embeddings-core 1.8.2,
* text-splitter 0.18.1, MIT
* tiberius 0.12.3, Apache-2.0 OR MIT
* tiktoken-rs 0.6.0, MIT
* tikv-jemallocator 0.6.1, Apache-2.0 OR MIT
* tokenizers 0.21.4, Apache-2.0
* tokio 1.48.0, MIT
* tokio-postgres 0.7.15, Apache-2.0 OR MIT
* tokio-rusqlite 0.5.1, MIT
* tokio-rustls 0.24.1, Apache-2.0 OR MIT
* tokio-rustls 0.26.4, Apache-2.0 OR MIT
* tokio-stream 0.1.17, MIT
* tokio-util 0.7.16, MIT
* tonic 0.13.1, MIT
* tonic-health 0.13.1, MIT
* tower 0.4.13, MIT
* tower 0.5.2, MIT
* tower-http 0.6.6, MIT
* tracing 0.1.41, MIT
* tracing-futures 0.2.5, MIT
* tracing-log 0.2.0, MIT
* tracing-opentelemetry 0.28.0, MIT
* tracing-subscriber 0.3.20, MIT
* tracing-test 0.2.5, MIT
* tract-core 0.22.0, Apache-2.0 OR MIT
* tract-onnx 0.22.0, Apache-2.0 OR MIT
* trust-dns-resolver 0.23.2, Apache-2.0 OR MIT
* turso 0.3.0, MIT
* twox-hash 2.1.2, MIT
* url 2.5.7, Apache-2.0 OR MIT
* utoipa 5.4.0, Apache-2.0 OR MIT
* utoipa-swagger-ui 9.0.2, Apache-2.0 OR MIT
* uuid 0.8.2, Apache-2.0 OR MIT
* uuid 1.18.1, Apache-2.0 OR MIT
* vortex 0.54.0, Apache-2.0
* vortex-datafusion 0.54.0, Apache-2.0
* winver 1.0.0, MIT
* x509-certificate 0.23.1, MPL-2.0
* zip 0.6.6, MIT
* zip 1.1.4, MIT
* zip 3.0.0, MIT
--- # API ## [📄️Overview](/docs/v1.10/api/overview) [Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools.](/docs/v1.10/api/overview) ## [📄️ADBC](/docs/v1.10/api/adbc) [ADBC API Documentation](/docs/v1.10/api/adbc) ## [📄️JDBC](/docs/v1.10/api/jdbc) [JDBC API Documentation](/docs/v1.10/api/jdbc) ## [📄️TLS](/docs/v1.10/api/tls) [Encryption in transit with TLS documentation](/docs/v1.10/api/tls) ## [📄️Authentication](/docs/v1.10/api/auth) [Authentication documentation](/docs/v1.10/api/auth) ## [📄️Arrow Flight SQL](/docs/v1.10/api/arrow-flight-sql) [Query Spice using JDBC/ODBC/ADBC](/docs/v1.10/api/arrow-flight-sql) ## [📄️ODBC](/docs/v1.10/api/odbc) [ODBC API Documentation](/docs/v1.10/api/odbc) ## [🗃HTTP](/docs/v1.10/api/HTTP/runtime) [25 items](/docs/v1.10/api/HTTP/runtime) --- # ADBC: Arrow Database Connectivity [ADBC](https://arrow.apache.org/adbc) is a set of APIs and libraries for Arrow-native access to databases. Spice supports ADBC clients using the [FlightSQL driver](https://arrow.apache.org/adbc/current/driver/flight_sql.html). ## Quickstart[​](#quickstart "Direct link to Quickstart") Get started with ADBC using Python. ### Installation[​](#installation "Direct link to Installation") Start a Python environment. ``` python ``` Install the ADBC driver manager, FlightSQL driver, and PyArrow. ``` pip install adbc_driver_manager adbc_driver_flightsql pyarrow ``` ### Create a connection to Spice over ADBC[​](#create-a-connection-to-spice-over-adbc "Direct link to Create a connection to Spice over ADBC") ``` >>> import adbc_driver_flightsql.dbapi >>> conn = adbc_driver_flightsql.dbapi.connect('grpc://localhost:50051') ``` ### Create a cursor[​](#create-a-cursor "Direct link to Create a cursor") ``` >>> cursor = conn.cursor() ``` ### Executing a query[​](#executing-a-query "Direct link to Executing a query") DBAPI interface: ``` >>> cursor.execute("SELECT 1, 2.0, 'Hello, world!'") >>> cursor.fetchone() (1, 2.0, 'Hello, world!') >>> cursor.execute("SHOW TABLES") >>> cursor.fetchall() [('spice', 'public', 'messages', 'BASE TABLE'), ('spice', 'runtime', 'task_history', 'BASE TABLE'), ('spice', 'information_schema', 'tables', 'VIEW'), ('spice', 'information_schema', 'views', 'VIEW'), ('spice', 'information_schema', 'columns', 'VIEW'), ('spice', 'information_schema', 'df_settings', 'VIEW'), ('spice', 'information_schema', 'schemata', 'VIEW')] ``` Arrow: ``` >>> cursor.execute("SELECT 1, 2.0, 'Hello, world!'") >>> cursor.fetch_arrow_table() pyarrow.Table 1: int64 2.0: double 'Hello, world!': string ---- 1: [[1]] 2.0: [[2]] 'Hello, world!': [["Hello, world!"]] ``` ## Parameterized Queries[​](#parameterized-queries "Direct link to Parameterized Queries") Spice supports parameterized queries when using ADBC clients. Parameterized queries help prevent SQL injection and improve code clarity by separating query logic from data values. The following example demonstrates how to use parameterized queries with the Python ADBC FlightSQL driver: ``` from adbc_driver_flightsql import DatabaseOptions from adbc_driver_flightsql.dbapi import connect with connect( "grpc://127.0.0.1:50051", ) as conn: with conn.cursor() as cur: cur.execute("SELECT $1 + 1 AS the_answer", parameters=(41,)) table = cur.fetch_arrow_table() print(table) cur.execute("SELECT 1 AS one") table = cur.fetch_arrow_table() print(table) conn.close() ``` --- # Arrow Flight SQL API [Arrow Flight SQL](https://arrow.apache.org/docs/format/FlightSql.html) is a protocol for interacting with SQL databases using the Arrow in-memory format and the Flight RPC framework. Spice implements the Flight SQL protocol, enabling querying of the datasets configured in Spice via tools that support connecting via one of the Arrow Flight SQL drivers, such as [DBeaver](https://dbeaver.io), [Tableau](https://www.tableau.com/), or [Power BI](https://www.microsoft.com/en-us/power-platform/products/power-bi). ![arrow flight and spice](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/0a8bc474-03c3-4c1c-8003-d250cd52b300/public) ## Authentication[​](#authentication "Direct link to Authentication") API Key authentication is supported for the Arrow Flight SQL endpoint. For more details, see [API Key Authentication](/docs/api/auth). --- # Authentication Spice supports adding optional authentication to its API endpoints via configurable API keys. Use the `auth` section as a child to `runtime` to provide the API keys. Multiple API keys can be specified, and any of the keys can be used to authenticate requests. ``` runtime: auth: api-key: enabled: true keys: - ${ secrets:api_key } # Use the secret replacement syntax to load the API key from a secret store - 1234567890 # Or specify the API key directly ``` To learn more about secrets, see [Secret Stores](/docs/components/secret-stores). info The API key authentication is applied on startup and changes will not take effect until the runtime is restarted. ## HTTP[​](#http "Direct link to HTTP") For HTTP routes, the API key is expected to be included in the `X-API-Key` header. ``` > curl -i "http://localhost:8090/v1/sql" -H "X-API-Key: 1234567890" -d 'SELECT 1' HTTP/1.1 200 OK content-type: text/plain; charset=utf-8 x-cache: Miss from spiceai content-length: 16 date: Fri, 08 Nov 2024 07:14:24 GMT [{"Int64(1)":1}] ``` The `/health` and `/v1/ready` endpoints are not protected and can be accessed without an API key. ## Flight SQL[​](#flight-sql "Direct link to Flight SQL") For the Flight SQL endpoint, the API key is expected to be included in the `Authorization` header as a Bearer token, i.e. `Authorization: Bearer ${ api_key }`. ## Spice CLI[​](#spice-cli "Direct link to Spice CLI") When API key authentication is enabled, the Spice CLI can connect to the runtime by specifying the `--api-key` argument. ``` spice sql --api-key 1234567890 spice status --api-key 1234567890 spice refresh taxi_trips --api-key 1234567890 # etc. ``` --- # Generate Package ``` POST /v1/packages/generate ``` This endpoint generates a zip package from a specified GitHub source. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 500 Package generated successfully Invalid request parameters Internal server error --- # List Catalogs ``` GET /v1/catalogs ``` List Catalogs ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of catalogs Internal server error occurred while processing catalogs --- # Get Iceberg API config ``` GET /v1/iceberg/config ``` This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. ## Responses[​](#responses "Direct link to Responses") * 200 API configuration retrieved successfully --- # List Datasets ``` GET /v1/datasets ``` This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, and additional filters can be applied using query parameters. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of datasets Internal server error occurred while processing datasets --- # List Iceberg namespaces ``` GET /v1/iceberg/namespaces ``` This endpoint retrieves namespaces available in the Iceberg catalog. If a `parent` namespace is provided, it will list the child namespaces under the specified parent. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 404 * 500 Namespaces retrieved successfully Bad request Namespace not found Internal server error --- # ML Prediction ``` GET /v1/models/:name/predict ``` Make a ML prediction using a specific model. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 500 Prediction made successfully Invalid request to the model Internal server error occurred during prediction --- # List Models ``` GET /v1/models ``` List all models, both machine learning and language models, available in the runtime. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of models in JSON format Internal server error occurred while processing models --- # List Spicepods ``` GET /v1/spicepods ``` Get a list of spicepods and their details. In CSV format, it will return a summarised form. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of spicepods Internal server error --- # Check Runtime Status ``` GET /v1/status ``` Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 List of connection statuses Error converting to CSV --- # Check Namespace exists ``` HEAD /v1/iceberg/namespaces/:namespace ``` This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. ## Responses[​](#responses "Direct link to Responses") * 200 * 404 Namespace exists Namespace does not exist --- # List Evals ``` GET /v1/evals ``` Return all evals available to run in the runtime. ## Responses[​](#responses "Direct link to Responses") * 200 All evals available in the Spice runtime --- # Send message to MCP server ``` POST /v1/mcp/sse ``` Send message to the MCP endoint, for a given session. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 202 * 404 * 413 * 500 Message accepted. Response will stream via SSE. Session not found. No active session for the given `session_id`. Payload too large. Maximum allowed size is 4MB. Internal server error. An unexpected issue occurred. --- # Establish an MCP SSE Connection ``` GET /v1/mcp/sse ``` Initiates a Server-Sent Events (SSE) connection using the Model Context Protocol (MCP) to interact with Spice tools. Once connected, clients can send messages via `POST /v1/mcp/sse` and receive responses through this SSE stream. --- # Update Refresh SQL ``` PATCH /v1/datasets/:name/acceleration ``` Update the refresh SQL for a dataset's acceleration. This endpoint allows for updating the `refresh_sql` parameter for a dataset's acceleration at runtime. The change is **temporary** and will revert to the `spicepod.yml` definition at the next runtime restart. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 500 The refresh SQL was updated successfully. The specified dataset was not found An internal server error occurred while updating the refresh SQL --- # Run Tool ``` POST /v1/tools/:name ``` The request body and JSON response formats match the tool’s specification. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 500 Tool Specific response, in JSON format Tool not found Error occured whilst calling the tool --- # Batch ML Predictions ``` POST /v1/predict ``` Perform a batch of ML predictions, using multiple models, in one request. This is useful for ensembling or A/B testing different models. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 500 Batch predictions completed successfully Internal server error occurred during batch prediction --- # Create Chat Completion ``` POST /v1/chat/completions ``` Creates a model response for the given chat conversation. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 500 Chat completion generated successfully The specified model was not found An internal server error occurred while processing the chat completion --- # Refresh Dataset ``` POST /v1/datasets/:name/acceleration/refresh ``` Trigger an on-demand refresh for an accelerated dataset. This endpoint triggers an on-demand refresh for an accelerated dataset. The refresh only applies to `full` and `append` refresh modes (not `changes` mode). ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 201 * 400 * 404 * 500 Dataset refresh triggered successfully Acceleration not enabled for the dataset Dataset not found Internal server error occurred while processing refresh --- # Create Embeddings ``` POST /v1/embeddings ``` Creates an embedding vector representing the input text. Get a vector representation of a given input that can be easily consumed by machine learning models and algorithms. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 500 Embedding created successfully Model not found Internal server error --- # Run Eval ``` POST /v1/evals/:name ``` Evaluate a model against a eval spice specification ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 Evaluation run successfully --- # Text-to-SQL (NSQL) ``` POST /v1/nsql ``` Generate and optionally execute a natural-language text-to-SQL (NSQL) query. This endpoint generates a SQL query using a natural language query (NSQL) and optionally executes it. The SQL query is generated by the specified model and executed if the `Accept` header is not set to `application/sql`. When `stream` is true, the response is streamed as Server-Sent Events (SSE). ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 500 SQL query executed successfully Invalid request parameters Internal server error --- # Search ``` POST /v1/search ``` Perform a vector similarity search (VSS) operation on a dataset. The search operation will return the most relevant matches based on cosine similarity with the input `text`. The datasets queries should have an embedding column, and the appropriate embedding model loaded. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 500 Search completed successfully Invalid request parameters Internal server error --- # SQL Query ``` POST /v1/sql ``` Execute a SQL query and return the results. This endpoint allows users to execute SQL queries directly from an HTTP request. The SQL query is sent as plain text in the request body. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 500 SQL query executed successfully Invalid SQL query or malformed input Internal server error --- # Check Readiness ``` GET /v1/ready ``` Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message "ready". If not, it returns a 503 status with the message "not ready". The behavior for when an accelerated dataset is considered ready is configurable via the `ready_state` parameter. See [Data refresh](https://spiceai.org/docs/components/data-accelerators/data-refresh#ready-state) for more details. ### Readiness Probe[​](#readiness-probe "Direct link to Readiness Probe") In production deployments, the /v1/ready endpoint can be used as a readiness probe for a Spice deployment to ensure traffic is routed to the Spice runtime only after all datasets have finished loading. Example Kubernetes readiness probe: ``` readinessProbe: httpGet: path: /v1/ready port: 8090 ``` ## Responses[​](#responses "Direct link to Responses") * 200 * 503 Service is ready Service is not ready --- Version: 1.9.2 # runtime The spiced runtime ### License --- # JDBC: Java Database Connectivity [JDBC](https://docs.oracle.com/javase/tutorial/jdbc/basics/index.html) (Java Database Connectivity) is a standard API for connecting to and interacting with databases. Spice supports JDBC clients through a JDBC driver implementation based on the [Flight SQL](https://arrow.apache.org/docs/format/FlightSql.html) protocol. This enables any JDBC-compatible application to connect to Spice, execute queries, and retrieve data. ## Download and install the Flight SQL JDBC driver[​](#download-and-install-the-flight-sql-jdbc-driver "Direct link to Download and install the Flight SQL JDBC driver") ### Download the Flight SQL JDBC driver[​](#download-the-flight-sql-jdbc-driver "Direct link to Download the Flight SQL JDBC driver") * Find the appropriate [Flight SQL JDBC driver](https://central.sonatype.com/artifact/org.apache.arrow/flight-sql-jdbc-driver/versions) version. * Click **Browse** next to the version you want to download * Click the `flight-sql-jdbc-driver-XX.XX.XX.jar` file (with only the `.jar` file extension) from the list of files to download the driver jar file ### Add the driver to your application[​](#add-the-driver-to-your-application "Direct link to Add the driver to your application") Follow the instructions specific to your application for adding a custom JDBC driver. Examples: **Tableau**: * Windows: `C:\Program Files\Tableau\Drivers` * Mac: `~/Library/Tableau/Drivers` * Linux: `/opt/tableau/tableau_driver/jdbc` * Start or restart Tableau [Full instruction](/docs/v1.10/clients/tableau) **JetBrains DataGrip**: * In Database Explorer menu, select "+" and choose "Driver" * Follow the steps to add the JDBC `.jar` file [Full instruction](/docs/v1.10/clients/jetbrains-datagrip) **DBeaver**: * In the DBeaver application menu bar, open the "Database" menu and choose: "Driver Manager" * Click the "New" button and follow instructions to add JDBC `.jar` file. [Full instruction](/docs/v1.10/clients/dbeaver) ## Configure JDBC connection[​](#configure-jdbc-connection "Direct link to Configure JDBC connection") 1. Use the following configuration settings: * **URL**: `jdbc:arrow-flight-sql://{host}:{port}` * **Dialect**: `PostgreSQL` For example: ![](/img/tableau/tableau-jdbc-conn.png) 1. **Ensure Spice is running** 2. Click **Connect** info Spice has [TLS support](/docs/v1.10/api/tls). For testing or non-production use cases for Spice without TLS, the following JDBC connection URL will bypass TLS `jdbc:arrow-flight-sql://{host}:{port}?useEncryption=false&disableCertificateVerification=true`. ### Authentication[​](#authentication "Direct link to Authentication") If [API Key authentication](/docs/v1.10/api/auth) is enabled, the API key can be provided in the JDBC connection URL as a query parameter: `jdbc:arrow-flight-sql://{host}:{port}?user=&password=` Replace `` with the API key value. The `user` and `password` parameters are required by the JDBC driver, but only the `password` parameter is used for the API key. ## Execute Test Query[​](#execute-test-query "Direct link to Execute Test Query") In the configured application, run a sample query, such as `SELECT * FROM taxi_trips;` ![Query Results](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/0e9f3c0f-2e03-47f9-8d5e-65e078d7e900/public "Query Results") ## Parameterized Queries[​](#parameterized-queries "Direct link to Parameterized Queries") Spice supports parameterized queries with JDBC. Parameterized queries help prevent SQL injection and improve code clarity by separating query logic from data values. --- # ODBC: Open Database Connectivity [ODBC](https://learn.microsoft.com/en-us/sql/odbc/microsoft-open-database-connectivity-odbc) (Open Database Connectivity) is a low-level, high-performance interface that is designed specifically for relational data stores as a standard way to connect to, and interact with a database. Spice supports ODBC clients through an ODBC driver implementation based on the [Flight SQL](https://arrow.apache.org/docs/format/FlightSql.html) protocol. This enables ODBC-compatible applications to connect to Spice, execute queries, and retrieve data. Limitations 1. ODBC support is currently in alpha, and not all functionality is supported 2. The Arrow Flight SQL ODBC driver is not available for 32-bit Windows versions 3. The Arrow Flight SQL ODBC driver is not supported on the Apple ARM architecture ## Install and configure the Flight SQL ODBC driver[​](#install-and-configure-the-flight-sql-odbc-driver "Direct link to Install and configure the Flight SQL ODBC driver") ### Download and install the Flight SQL ODBC driver[​](#download-and-install-the-flight-sql-odbc-driver "Direct link to Download and install the Flight SQL ODBC driver") * Download and install the driver from the [ODBC driver download page](https://www.dremio.com/drivers/odbc/) * [Windows instructions](https://docs.dremio.com/current/sonar/client-applications/drivers/arrow-flight-sql-odbc-driver/#downloading-and-installing-on-windows) * [Linux](https://docs.dremio.com/current/sonar/client-applications/drivers/arrow-flight-sql-odbc-driver/#downloading-and-installing-on-linux) * [macOS instructions](https://docs.dremio.com/current/sonar/client-applications/drivers/arrow-flight-sql-odbc-driver/#downloading-and-installing-on-macos) ### Configure Flight SQL ODBC driver[​](#configure-flight-sql-odbc-driver "Direct link to Configure Flight SQL ODBC driver") * Windows * Linux * macOS - Open **Start Menu** -> **Windows Administrative Tools** -> click **ODBC Data Sources (64-bit)** - In the **ODBC Data Source Administrator (64-bit)** dialog, click **System DSN** ![ODBC Data Source Administrator](/img/odbc/spice-odbc-windows-config.png) * Select **Arrow Flight SQL ODBC DSN** and click **Configure** * Specify Spice.ai OSS runtime `HOST`, `PORT`, in the `UseEncryption` field, specify one of these values: * `true`, if [Spice is configured for encrypted communication (TLS)](https://docs.spiceai.org/api/tls) * `false`, otherwise * For descriptions of all the parameters, see [ODBC Connection Parameters](#odbc-connection-parameters). - Ensure that `unixODBC` is installed. To verify whether `unixODBC` is installed, execute the following commands: ``` which odbcinst which isql ``` * Copy the content of the `odbc.ini` and `odbcinst.ini` from the `/opt/arrow-flight-sql-odbc-driver/conf` and paste into your system `/etc/odbc.ini` and `/etc/odbcinst.ini` files * Edit `odbc.ini`: specify Spice.ai OSS runtime `HOST`, `PORT`, in the `UseEncryption` field, specify one of these values: * `true`, if [Spice is configured for encrypted communication (TLS)](https://docs.spiceai.org/api/tls) * `false`, otherwise * For descriptions of all the parameters, see [ODBC Connection Parameters](#odbc-connection-parameters) * Run this command to verify `unixODBC` configuration ``` odbcinst -j ``` * Ensure that [ODBC Manager](http://www.odbcmanager.net/) is installed. * Launch **ODBC Manager** -> **System DSN** page, select **Arrow Flight SQL ODBC DSN** and click **Configure**. * Specify Spice.ai OSS runtime `HOST`, `PORT`, in the `UseEncryption` field, specify one of these values: * `true`, if [Spice is configured for encrypted communication (TLS)](https://docs.spiceai.org/api/tls) * `false`, otherwise ![ODBC Data Source Administrator](/img/odbc/spice-odbc-macos-config.png) * For descriptions of all the parameters, see [ODBC Connection Parameters](#odbc-connection-parameters). ### ODBC Connection Parameters[​](#odbc-connection-parameters "Direct link to ODBC Connection Parameters") | Name | Type | Description | | ------------------------------ | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | host | string | The IP address or hostname for the Spice runtime. | | port | integer | The Spice runtime Arrow Flight endpoint port number | | useEncryption | integer | Configures the driver to use an SSL-encrypted connection. Accepted values: `true` (default) - The client communicates with the Spice runtime only using SSL encryption and `false` - SSL encryption is disabled. | | disableCertificateVerification | integer | Specifies whether the driver should verify the host certificate against the trust store. Default is `false` | | useSystemTrustStore | integer | Controls whether to use a CA certificate from the system's trust store, or from a specified .pem file. If `true` - The driver verifies the connection using a certificate in the system trust store. IF `false` - The driver verifies the connection using the .pem file specified by the `trustedCerts` parameter. `true` on Windows and macOS, `false` on Linux by default | | trustedCerts | string | The full path of the .pem file containing certificates trusted by a CA, for the purpose of verifying the server. If this option is not set, then the driver defaults to using the trusted CA certificates .pem file installed by the driver. | note The ODBC driver for Arrow Flight SQL does not support password-protected `.pem/.crt` files or multiple `.crt` certificates in a single `.pem/.crt` file. ## Execute Test Query[​](#execute-test-query "Direct link to Execute Test Query") Ensure Spice runtime has started: ``` spiced --flight 0.0.0.0:50051 ``` ``` 2024-10-06T20:06:50.084017Z INFO runtime::opentelemetry: Spice Runtime OpenTelemetry listening on 127.0.0.1:50052 2024-10-06T20:06:50.084015Z INFO runtime::flight: Spice Runtime Flight listening on 0.0.0.0:50051 2024-10-06T20:06:50.086948Z INFO runtime::http: Spice Runtime HTTP listening on 127.0.0.1:8090 2024-10-06T20:06:50.297512Z INFO runtime: Initialized results cache; max size: 128.00 MiB, item ttl: 1s 2024-10-06T20:06:50.308775Z INFO runtime: Tool [search] ready to use 2024-10-06T20:06:50.308803Z INFO runtime: Tool [table_schema] ready to use 2024-10-06T20:06:50.308814Z INFO runtime: Tool [sql] ready to use 2024-10-06T20:06:50.308829Z INFO runtime: Tool [list_datasets] ready to use 2024-10-06T20:06:50.921420Z INFO runtime: Dataset taxi_trips registered (s3://spiceai-demo-datasets/taxi_trips/2024/), results cache enabled. ``` Configure the client app to use `Arrow Flight SQL ODBC Driver`. ![Example ODBC Client Configuration](/img/odbc/spice-odbc-example-config.png) Run a sample query, such as ``` SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10; ``` ![Example Query Results](/img/odbc/spice-odbc-example-query.png) ## Parameterized Queries[​](#parameterized-queries "Direct link to Parameterized Queries") Spice supports parameterized queries with ODBC. Parameterized queries help prevent SQL injection and improve code clarity by separating query logic from data values. --- # API Overview Spice provides high-performance, industry-standard APIs: ### SQL Query APIs[​](#sql-query-apis "Direct link to SQL Query APIs") * **Arrow Flight** / **Arrow Flight SQL**: High-performance SQL query. * **ODBC**, **JDBC**, **ADBC**: Standard SQL interfaces for database clients and analytics tools. ### OpenAI-Compatible APIs[​](#openai-compatible-apis "Direct link to OpenAI-Compatible APIs") * **HTTP APIs**: Compatible the OpenAI SDK, AI SDK with local model serving (CUDA/Metal accelerated) and gateway to hosted models. ### Iceberg Catalog REST APIs[​](#iceberg-catalog-rest-apis "Direct link to Iceberg Catalog REST APIs") * **HTTP APIs**: Unified API for consuming Apache Iceberg catalogs in data lake architectures. ### MCP API[​](#mcp-api "Direct link to MCP API") * **HTTP APIs**: The Model Context Protocol (MCP) helps integrate external tools and services into the Spice runtime. MCP tools can be accessed via HTTP APIs for tool integration and orchestration. For details, see the [MCP documentation](/docs/features/large-language-models/mcp). note HTTP Streaming support for MCP is coming soon. --- # TLS: Transport Layer Security [Transport Layer Security](https://en.wikipedia.org/wiki/Transport_Layer_Security) (TLS) is a cryptographic protocol that secures communication over a network. TLS is the successor to deprecated Secure-Sockets-Layer (SSL). Learn how to configure Spice to use TLS for encryption in transit. ## Pre-requisites[​](#pre-requisites "Direct link to Pre-requisites") A valid TLS certificate and private key in [PEM](https://en.wikipedia.org/wiki/Privacy-Enhanced_Mail) format are required. To generate certificates for testing, follow the [TLS Cookbook](https://github.com/spiceai/cookbook/tree/trunk/tls). ## Enable TLS via command line arguments[​](#enable-tls-via-command-line-arguments "Direct link to Enable TLS via command line arguments") Use `--tls-enabled true` to enable TLS from the command line. The arguments `--tls-certificate-file` and `--tls-key-file` specify the paths to the certificate and private key files. ``` # Provide the TLS certicate and key PEM files to the Spice runtime spiced --tls-enabled true --tls-certificate-file /path/to/cert.pem --tls-key-file /path/to/key.pem ``` Alternatively, to pass PEM-encoded certificate and private key strings directly, use the `--tls-certificate` and `--tls-key` arguments. ``` # Provide the TLS certicate and key using PEM-encoded strings to the Spice runtime export TLS_CERT=$(cat /path/to/cert.pem) export TLS_KEY=$(cat /path/to/key.pem) spiced --tls-enabled true --tls-certificate "$TLS_CERT" --tls-key "$TLS_KEY" ``` When using the Spice CLI, arguments, including the TLS arguments, are passed to `spice run` automatically. ``` # Run Spice using the CLI and provide the TLS certicate and key as PEM files spice run -- --tls-enabled true --tls-certificate-file /path/to/cert.pem --tls-key-file /path/to/key.pem ``` Note that `--` is used to separate the `spice run` arguments from the Spice runtime arguments. ## Enable TLS via spicepod.yaml[​](#enable-tls-via-spicepodyaml "Direct link to Enable TLS via spicepod.yaml") Use the `tls` section as a child to `runtime` to provide the certificate and key files/strings. ``` runtime: tls: enabled: true # Using filesystem paths certificate_file: /path/to/cert.pem key_file: /path/to/key.pem ``` ``` runtime: tls: enabled: true # Specify the certificate and key directly certificate: | -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- key: | -----BEGIN PRIVATE KEY----- ... -----END PRIVATE KEY----- ``` ``` runtime: tls: enabled: true # Provide the certificate and key using secrets certificate: ${secrets:tls_cert} key: ${secrets:tls_key} ``` To learn more about secrets, see [Secret Stores](/docs/components/secret-stores). info Changes to TLS configuration are not applied at runtime and will only take effect on startup. ## Output[​](#output "Direct link to Output") When TLS is enabled, the runtime output will print the TLS certificate details. ``` INFO runtime: All endpoints secured with TLS using certificate: CN=spiced.localhost, OU=IT, O=Widgets, Inc., L=Seattle, S=Washington, C=US ``` ## Using the Spice CLI[​](#using-the-spice-cli "Direct link to Using the Spice CLI") When TLS is enabled in the runtime, the Spice CLI can be configured to connect to the runtime using TLS by specifying the `--tls-root-certificate-file` argument, providing the path to the root certificate file. ``` spice sql --tls-root-certificate-file /path/to/root.pem ``` --- # Spice.ai OSS CLI documentation The Spice CLI is a set of commands to create and manage Spicepods and interact with the Spice runtime. ## Install[​](#install "Direct link to Install") The Spice CLI can be installed by: * macOS, Linux, and WSL * Windows - Running `curl https://install.spiceai.org | /bin/bash` - Using `brew`: `brew install spiceai/spiceai/spice` - Downloading the binary from [GitHub Releases](https://github.com/spiceai/spiceai/releases) * Shell command ``` curl -L "https://install.spiceai.org/Install.ps1" -o Install.ps1 && PowerShell -ExecutionPolicy Bypass -File ./Install.ps1 ``` * Downloading the binary from [GitHub Releases](https://github.com/spiceai/spiceai/releases) The `spice` program will be added to the PATH automatically for **bash**, **fish**, and **zsh** shells. After installing the Spice CLI for the first time, ensure you've got the correct version by running `spice version`. The Runtime version is not expected to be shown, as the runtime will be downloaded and installed automatically upon first run. ## Getting started[​](#getting-started "Direct link to Getting started") For getting started with Spice using the Spice CLI, see the [Getting Started Guide](/docs/v1.10/getting-started). Use `spice help` for all commands and `spice [command] --help` for more information about a command. A typical command-line workflow might be as follows: ``` # Start the runtime spice run ``` Run new shell in the same folder: ``` # Init new app spice init spice_app # Add the Quickstart Spicepod spice add spiceai/quickstart ``` Common commands are: | Command | Description | | ------------- | ------------------------------------------------------------- | | spice add | Add Pod - adds a pod to the project | | spice run | Run Spice - starts the Spice runtime, installing if necessary | | spice version | Spice CLI version | | spice help | Help about any command | | spice upgrade | Upgrades the Spice CLI to the latest release | See [Spice CLI command reference](/docs/v1.10/reference) for the full list of available commands. ## Updating[​](#updating "Direct link to Updating") To update to latest CLI, run the upgrade command. ``` spice upgrade ``` note Upgrade command is supported from CLI v0.3.1. For version < 0.3.1 users have to re-run the [install](#install) script. ## Uninstall[​](#uninstall "Direct link to Uninstall") The Spice CLI is installed by default to `$HOME/.spice/bin/spice` and a line added to the shell config, such as `.zshrc` It can be uninstalled by deleting the `spice` binary and removing the PATH addition from the rc file. --- # Spice.ai OSS CLI command reference ## spice[​](#spice "Direct link to spice") ### Usage[​](#usage "Direct link to Usage") ``` spice [command] [--help] ``` ### Full Command Reference[​](#full-command-reference "Direct link to Full Command Reference") | Command | Description | | -------------------------------------------------- | ---------------------------------------------------------------------------- | | [add](/docs/v1.10/cli/reference/add) | Add Pod - adds a pod to the project | | [catalogs](/docs/v1.10/cli/reference/catalogs) | List [catalogs](/docs/v1.10/components/catalogs) loaded by the Spice runtime | | [completion](/docs/v1.10/cli/reference/completion) | Generate the autocompletion script for the specified shell | | [dataset](/docs/v1.10/cli/reference/dataset) | Dataset operations | | [datasets](/docs/v1.10/cli/reference/datasets) | Lists datasets loaded by the Spice runtime | | help | Help about any command | | [init](/docs/v1.10/cli/reference/init) | Initialize Pod - initializes a new pod in the project | | [login](/docs/v1.10/cli/reference/login) | Login to the Spice.ai Platform | | [models](/docs/v1.10/cli/reference/models) | Lists models loaded by the Spice runtime | | [pods](/docs/v1.10/cli/reference/pods) | Lists Spicepods loaded by the Spice runtime | | [refresh](/docs/v1.10/cli/reference/refresh) | Refreshes an accelerated dataset loaded by the Spice runtime | | [run](/docs/v1.10/cli/reference/run) | Run Spice - starts the Spice runtime, installing if necessary | | [search](/docs/v1.10/cli/reference/search) | Perform embeddings-based searches across | | [sql](/docs/v1.10/cli/reference/sql) | Start an interactive SQL query session against the Spice runtime | | [status](/docs/v1.10/cli/reference/status) | Spice runtime status | | [upgrade](/docs/v1.10/cli/reference/upgrade) | Upgrades the Spice CLI to the latest release | | [version](/docs/v1.10/cli/reference/version) | Spice CLI version | ### Command Flags[​](#command-flags "Direct link to Command Flags") All commands have a help flag **--help** or **-h** to print its usage documentation: * **--help** | **-h** : Print the help message --- # add Add a Spicepod to the project. ### Usage[​](#usage "Direct link to Usage") ``` spice add [spicerack slug] [flags] ``` * `spicerack slug`: The slug to the Spicepod on [spicerack.org](https://spicerack.org). #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message ### Examples[​](#examples "Direct link to Examples") Adding a Spicepod from Spicerack (like `spiceai/quickstart`): ``` > spice add spiceai/quickstart ``` **Directory Structure**: The command makes two main modifications to the directory structure: 1. It creates the `spicepods` directory in the project root if it does not exist. 2. It adds the Spicepod defined by the Spicerack Slug in the relative path in the `spicepods` directory. For this example, the command would create the directories `spicepods/spiceai` and `spicepods/spiceai/quickstart`, instantiating a Spicepod under the latter. More generally, the Spicepod is placed under `spicepods/[slug]`, where `slug` is the Spicerack slug associated with that Spicepod. After running the command, the directory structure looks like this: ``` ├── spicepods/ │ ├── spiceai/ │ ├── quickstart/ │ ├── spicepod.yaml ├── spicepod.yaml └── ... ``` Any other Spicepods added using `spice add` are placed in the `spicepods` directory. `spice add` also creates the appropriate Spicepod for the given Spicerack slug. For this example with `spiceai/quickstart`, the command creates the following the Spicepod under `spicepods/spiceai/quickstart`: ``` # File: ./spicepods/spiceai/quickstart/spicepod.yaml version: v1beta1 kind: Spicepod name: quickstart datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips description: taxi trips in s3 params: file_format: parquet acceleration: enabled: true ``` The `add` command also includes the above Spicepod as a dependency in the root `spicepod.yaml`, creating this file if it does not exist: ``` # File: ./spicepod.yaml version: v1 kind: Spicepod name: Spice AI quickstart dependencies: - spiceai/quickstart ``` --- # catalogs List [catalogs](/docs/v1.10/components/catalogs) currently loaded by the Spice runtime. ### Requirements[​](#requirements "Direct link to Requirements") * Spice runtime must be running ### Usage[​](#usage "Direct link to Usage") ``` > spice catalogs [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate ### Example[​](#example "Direct link to Example") ``` > spice catalogs FROM NAME spiceai spiceai databricks us_census_data ``` --- # chat Start an interactive or one-shot chat with a [model](/docs/v1.10/components/models) registered in the Spice runtime. ## Requirements[​](#requirements "Direct link to Requirements") * Spice runtime must be running * At least one model defined in `spicepod.yaml` and the model is ready ## Usage[​](#usage "Direct link to Usage") ### Interative Chat: Invoke the command without arguments to open a REPL[​](#interative-chat-invoke-the-command-without-arguments-to-open-a-repl "Direct link to Interative Chat: Invoke the command without arguments to open a REPL") ``` spice chat [flags] ``` ### One-shot Chat: Pass a single message as the argument to send a one-shot chat request and print the response[​](#one-shot-chat-pass-a-single-message-as-the-argument-to-send-a-one-shot-chat-request-and-print-the-response "Direct link to One-shot Chat: Pass a single message as the argument to send a one-shot chat request and print the response") ``` spice chat [flags] [] ``` ## Flags[​](#flags "Direct link to Flags") * `--cloud` Use a Spice Cloud instance for chat. Requires `--api-key`. * `--endpoint ` Specifies the remote Spice instance endpoint. Supports `http://`, `https://`, `grpc://`, or `grpc+tls://` schemes. For example, `--endpoint http://my-remote-host:8090` (HTTP) or `--endpoint grpc://my-remote-host:50051` (Arrow Flight/gRPC). * `--http-endpoint ` (Deprecated) Runtime HTTP endpoint. Default: `http://localhost:8090`. * `--model ` Target model for the chat request. When omitted, the CLI uses the single ready model or prompts for a choice if several models are ready. * `--temperature ` Model temperature used for chat request. Default: `1.0`. * `--user-agent ` Custom `User-Agent` header sent with every request. * `--responses` Direct all chats to the `/v1/responses` endpoint, which exposes configured models that support [OpenAI's Responses API](https://platform.openai.com/docs/api-reference/responses) and enables access to [OpenAI-hosted tools](https://platform.openai.com/docs/guides/tools). To learn more about Spice's support for OpenAI's Responses API, view the [OpenAI model provider documentation](/docs/v1.10/components/models/openai) or the [Azure OpenAI model provider documentation](/docs/v1.10/components/models/azure). ## Examples[​](#examples "Direct link to Examples") When exactly one model is **ready**, `spice chat` opens a REPL that uses that model automatically: ``` > spice chat Using model: openai chat> hello Hello! How can I assist you today? Time: 0.57s (first token 0.53s). Tokens: 18. Prompt: 8. Completion: 10 (325.04/s). ``` #### Remote and Cloud Examples[​](#remote-and-cloud-examples "Direct link to Remote and Cloud Examples") ``` # Chat with Spice Cloud spice chat --cloud --api-key --model # Chat with a remote spiced instance over HTTP spice chat --endpoint http://my-remote-host:8090 --model # Chat with a remote spiced instance over Arrow Flight SQL (gRPC) spice chat --endpoint grpc://my-remote-host:50051 --model ``` When multiple models are **ready**, the command prompts for a selection before starting the REPL: ``` > spice chat Use the arrow keys to navigate: ↓ ↑ → ← ? Select model: ▸ openai llama Using model: openai chat> hello Hello! How can I assist you today? Time: 0.55s (first token 0.43s). Tokens: 18. Prompt: 8. Completion: 10 (80.09/s). ``` Passing `--model` skips the prompt and directs the request to the specified model. The flag works both in REPL mode and in one‑shot mode: ``` # REPL spice chat --model openai chat> hello Hello! How can I assist you today? Time: 0.61s (first token 0.58s). Tokens: 18. Prompt: 8. Completion: 10 (285.90/s). ``` Single prompt: ``` # One‑shot spice chat --model openai "hello" Hello! How can I assist you today? Time: 1.10s (first token 0.80s). Tokens: 18. Prompt: 8. Completion: 10 (33.74/s). ``` --- # Completion Generate the autocompletion script for spice for the specified shell. See each sub-command's help for details on how to use the generated script. ### Usage[​](#usage "Direct link to Usage") ``` spice completion [command] ``` Available `command`s * bash * fish * powershell * zsh #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message --- # connect Connect to an app on the Spice.ai Cloud Platform. ### Requirements[​](#requirements "Direct link to Requirements") * Authenticated to the Spice.ai Cloud Platform via [`login`](/docs/v1.10/cli/reference/login) **Notes**: Instead of pulling from Spicerack and integrating those Spicepods locally like `spice add`, `spice connect` enables connection to an app on the Spice.ai Cloud Platform without any local modifications or updates necessary. ### Usage[​](#usage "Direct link to Usage") ``` spice connect [app] [flags] ``` * `app`: The app slug in the Spice.ai Cloud Platform (i.e. `spiceai/tpch`). #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message ### Examples[​](#examples "Direct link to Examples") ``` > spice connect spiceai/quickstart ``` ``` > spice connect spiceai/tpch ``` --- # dataset Configure a Spice dataset. ### Usage[​](#usage "Direct link to Usage") ``` spice dataset [command] ``` Available `command`s: * `configure`: Create/configure a dataset directly from the command-line, including customizing components such as whether to add acceleration to the connector. **Note**: In order to run `spice dataset configure`, there *must* be a `spicepod.yaml` file in the root of your project directory. To create this file, see [`spice init`](/docs/v1.10/cli/reference/init). #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message ### Example[​](#example "Direct link to Example") When running `spice dataset configure`, Spice will prompt for four inputs: 1. The name of the dataset, labelled by `(1)` below. 2. The description of the dataset, labelled by `(2)` below. 3. The source of the dataset, labelled by `(3)` below. Consult [Spice's supported data connectors](/docs/v1.10/components/data-connectors) to see possible values for this field. Note: Spice may prompt for a file format if necessary, as shown in the example below. 4. Whether or not to enable acceleration for this dataset, labelled by `(4)`. The default value for this input is `y`, enabling acceleration for this dataset. Learn more about acceleration in the [dataset acceleration reference](/docs/v1.10/components/data-accelerators). ``` > spice dataset configure dataset name: (spiceai) taxi-trips # (1) description: Taxi Trips in S3 # (2) from: s3://spiceai-demo-datasets/taxi_trips/2024/ # (3) file_format (parquet/csv) (parquet) parquet locally accelerate (y/n)? (y) y # (4) 2025/01/10 14:07:46 INFO Saved datasets/test/dataset.yaml ``` After execution, the directory structure looks like this for the above example: ``` ├── datasets │ ├── taxi-trips │ ├── dataset.yaml ├── spicepod.yaml └── ... ``` The datasets folder includes the datasets for your project configured by using `spice dataset configure` or added manually. The `dataset.yaml` file in `./datasets/taxi-trips` is configured as defined by the inputs provided to `spice dataset configure`. For this example, the `dataset.yaml` file looks as follows: ``` from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi-trips description: Taxi trips in s3 acceleration: - enabled: false ``` The command additionally updates the root `spicepod.yaml` file to include the configured dataset as a reference (`ref`). For this example, `spicepod.yaml` would include the following: ``` version: v1 kind: Spicepod name: Taxi Trips with Spice datasets: - ref: datasets/taxi-trips ``` To learn more about Spice datasets and Spicepods, visit the [Spice dataset reference](/docs/v1.10/reference/spicepod/datasets) and [Spicepod reference](/docs/v1.10/reference/spicepod). --- # datasets Lists datasets loaded by the Spice runtime ### Usage[​](#usage "Direct link to Usage") ``` spice datasets [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `-h`, `--help` help for datasets ### Examples:[​](#examples "Direct link to Examples:") ``` >>> spice datasets FROM NAME REPLICATION ACCELERATION STATUS PROPERTIES spice.ai/spiceai/quickstart/datasets/taxi_trips taxi_trips false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.customer tpch.customer false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.lineitem tpch.lineitem false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.nation tpch.nation false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.orders tpch.orders false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.part tpch.part false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.partsupp tpch.partsupp false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.region tpch.region false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.supplier tpch.supplier false false Ready map[] ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice datasets --tls-root-certificate-file /path/to/cert.pem FROM NAME REPLICATION ACCELERATION STATUS PROPERTIES spice.ai/spiceai/quickstart/datasets/taxi_trips taxi_trips false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.customer tpch.customer false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.lineitem tpch.lineitem false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.nation tpch.nation false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.orders tpch.orders false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.part tpch.part false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.partsupp tpch.partsupp false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.region tpch.region false false Ready map[] spice.ai/spiceai/tpch/datasets/tpch.supplier tpch.supplier false false Ready map[] ``` --- # init Initialize Spice app in the current working directory. ### Usage[​](#usage "Direct link to Usage") ``` spice init [app_name] ``` * `app_name`: The name of the app. If this is not provided, Spice will prompt for the name, defaulting to the name of the current working directory. #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message ### Examples[​](#examples "Direct link to Examples") ``` > spice init taxi-trips 2025/01/09 16:43:02 INFO Initialized taxi-trips/spicepod.yaml ``` The command creates a `spicepod.yaml` with basic initial metadata about the app. For this example, the `spicepod.yaml` file is initialized to the following: ``` # File: ./taxi-trips/spicepod.yaml version: v1 kind: Spicepod name: taxi-trips ``` If no app name is provided, Spice initializes the Spicepod in the current working directory. For example: ``` > spice init name: (taxi-trips)? 2025/01/09 16:48:37 INFO Initialized spicepod.yaml ``` After execution, the current working directory contains the file `spicepod.yaml` with the same configuration as the previous example: ``` # File: ./spicepod.yaml version: v1 kind: Spicepod name: taxi-trips ``` --- # install Download and install the latest version of the Spice runtime. ### Usage[​](#usage "Direct link to Usage") ``` spice install [flavor] [flags] ``` #### flavor[​](#flavor "Direct link to flavor") * \`\` Install the core runtime that only includes data components * `ai` Install the AI-enabled runtime with both data components and AI components #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message * `-c`, `--cpu` Install the CPU accelerated version of the AI runtime * `-f`, `--force` Force installation of the latest released runtime ### Examples[​](#examples "Direct link to Examples") ``` spice install ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` spice install ai --cpu ``` --- # login Login to the Spice.ai Platform, or other services with sub-commands. ### Usage[​](#usage "Direct link to Usage") ``` spice login [command] [flags] ``` ### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message * `-k`, `--key` string API key (for spice.ai) #### Available Commands[​](#available-commands "Direct link to Available Commands") * `abfs` Login to a Azure Storage Account * `databricks` Login to a Databricks instance * `delta-lake` Configure credentials to access a Delta Lake table * `dremio` Login to a Dremio instance * `postgres` Login to a Postgres instance * `s3` Login to an s3 storage * `sharepoint` Login to a Microsoft 365 sharepoint account * `snowflake` Login to a Snowflake warehouse * `spark` Login to a Spark Connect remote #### Examples[​](#examples "Direct link to Examples") ``` spice login ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` spice login --key ``` --- # models Lists models loaded by the Spice runtime ### Usage[​](#usage "Direct link to Usage") ``` spice models [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `-h`, `--help` help for models ### Examples[​](#examples "Direct link to Examples") ``` >>> spice models NAME FROM DATASETS STATUS modlz file:/Users/jeadie/Downloads/model.onnx [] Ready ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice models --tls-root-certificate-file /path/to/cert.pem NAME FROM DATASETS STATUS modlz file:/Users/jeadie/Downloads/model.onnx [] Ready ``` --- # pods Lists Spicepods loaded by the Spice runtime ### Usage[​](#usage "Direct link to Usage") ``` spice pods [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `-h`, `--help` help for pods ### Examples[​](#examples "Direct link to Examples") ``` >>> spice pods VERSION NAME DATASETSCOUNT MODELSCOUNT DEPENDENCIESCOUNT v1 demo 2 1 0 v1 another_pod 3 0 1 ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice pods --tls-root-certificate-file /path/to/cert.pem VERSION NAME DATASETSCOUNT MODELSCOUNT DEPENDENCIESCOUNT v1 demo 2 1 0 v1 another_pod 3 0 1 ``` --- # refresh Refreshes an accelerated dataset loaded by the Spice runtime ### Usage[​](#usage "Direct link to Usage") ``` spice refresh [dataset] [flags] ``` `dataset` - an accelerated dataset name #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `--refresh-sql` SQL used to refresh the dataset, see [Refresh SQL docs](/docs/v1.10/features/data-acceleration/data-refresh#refresh-sql). * `--refresh-mode` Refresh mode to use, see [Refresh Modes docs](/docs/v1.10/features/data-acceleration/data-refresh#refresh-modes). * `-h`, `--help` Print this help message ### Examples[​](#examples "Direct link to Examples") ``` >>> spice refresh taxi_trips --refresh-sql "SELECT * FROM taxi_trips WHERE trip_amount > 10.0" ``` Refreshing dataset taxi\_trips ... Dataset refresh triggered for taxi\_trips. ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice refresh taxi_trips --refresh-mode append ``` Refreshing dataset taxi\_trips with append mode... Dataset refresh triggered for taxi\_trips. ``` ``` --- # run Run Spice - starts the Spice runtime, installing if necessary. ### Usage[​](#usage "Direct link to Usage") ``` spice run [flags] spice run [flags] -- [spiced flags] ``` #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` Print this help message. * `--flight-endpoint` Configure runtime Flight endpoint. Defaults to `http://127.0.0.1:50051`. * `--http-endpoint` Configure runtime HTTP endpoint. Defaults to `http://127.0.0.1:8090`. * `--metrics-endpoint` Configure the runtime Prometheus metrics endpoint (disabled by default). * `--open-telemetry-endpoint` Configure runtime OpenTelemetry endpoint. Defaults to `http://127.0.0.1:50052`. * `--captured-outputs` Configure the captured output setting for task history. Defaults to `truncated`. #### Spiced Flags[​](#spiced-flags "Direct link to Spiced Flags") Flags that are passed to the `spiced` runtime directly. * `--http` Configure runtime HTTP address \[default: 127.0.0.1:8090] * `--flight` Configure runtime Flight address \[default: 127.0.0.1:50051] * `--open_telemetry` Configure runtime OpenTelemetry address \[default: 127.0.0.1:50052] * `--tls-enabled` Enable TLS * `--tls-certificate` The TLS PEM-encoded certificate * `--tls-certificate-file` Path to the TLS PEM-encoded certificate file * `--tls-key` The TLS PEM-encoded key * `--tls-key-file` Path to the TLS PEM-encoded key file * `--set-runtime` Override [runtime configuration](/docs/v1.10/reference/spicepod#runtime) with a name/value pair specified as `name=value`. Multiple overrides can be specified by using the flag multiple times. ### Examples[​](#examples "Direct link to Examples") #### `--set-runtime`[​](#--set-runtime "Direct link to --set-runtime") The `--set-runtime` flag allows overriding runtime configuration values. It can be specified multiple times to set multiple values. It is used like this: `--set-runtime name=value`. The Spicepod YAML equivalent of that is: ``` runtime: name: value ``` Examples: `--set-runtime task_history.captured_output=none`: ``` runtime: task_history: captured_output: none ``` `--set-runtime results_cache.enabled=false`: ``` runtime: results_cache: enabled: false ``` `--set-runtime runtime.tls.enabled=true --set-runtime runtime.tls.certificate_file=/path/to/cert.pem --set-runtime runtime.tls.key_file=/path/to/key.pem`: ``` runtime: tls: enabled: true certificate_file: /path/to/cert.pem key_file: /path/to/key.pem ``` #### No arguments[​](#no-arguments "Direct link to No arguments") ``` spice run ``` #### `--captured-outputs none`[​](#--captured-outputs-none "Direct link to --captured-outputs-none") ``` # Set task history captured outputs to none spice run -- --captured-outputs none ``` #### `--http`[​](#--http "Direct link to --http") ``` # Expose the HTTP server on all interfaces spice run -- --http 0.0.0.0:8090 ``` #### `--flight`[​](#--flight "Direct link to --flight") ``` # Expose the HTTP & Flight servers on all interfaces with TLS spice run -- --http 0.0.0.0:8090 --flight 0.0.0.0:50051 --tls-enabled true --tls-certificate-file /path/to/cert.pem --tls-key-file /path/to/key.pem ``` #### `--open_telemetry`[​](#--open_telemetry "Direct link to --open_telemetry") ``` # Run Spice with OpenTelemetry enabled spice run -- --open_telemetry 0.0.0.0:50052 ``` --- # search Performs embeddings-based searches across search configured datasets. Note: Search requires the `ai` feature to be installed. ### Usage[​](#usage "Direct link to Usage") ``` spice search [query] [flags] ``` `query` - a search query #### Flags[​](#flags "Direct link to Flags") * `--cloud` Use a Spice Cloud instance for search. Requires `--api-key`. * `--endpoint ` Specifies the remote Spice instance HTTP endpoint (e.g., `http://localhost:8090`). * `--limit` Limit number of search results. * `--model` Model to use for search. * `--http-endpoint ` (Deprecated) HTTP endpoint for search (default: `http://localhost:8090`). ### Examples[​](#examples "Direct link to Examples") ``` >>> spice search --limit 2 ``` #### Remote and Cloud Examples[​](#remote-and-cloud-examples "Direct link to Remote and Cloud Examples") ``` # Search with Spice Cloud spice search --cloud --api-key # Search with a remote spiced instance spice search --endpoint http://my-remote-host:8090 ``` ``` search> artificial intelligence Rank 1, Score: 20.6, Datasets [pdf] Undergraduate Texts in Mathematics Editors: F. W. Gehring P. R. Halmos · Advisory Board: C. DePrima I. Herstein J. Kiefer W. LeVeque Kai Lai Chung Elementary Probability Theory with Stochastic Processes Springer Science+Business Media, LLC ... Rank 2, Score: 17.8, Datasets [pdf] Forecasting at Scale Sean J. Taylor y Facebook, Menlo Park, California, United States sjt@fb.com and Benjamin Letham y Facebook, Menlo Park, California, United States bletham@fb.com Abstract Forecasting is a common data science... ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice search --model gpt-3 --limit 1 ``` ``` search> machine learning Rank 1, Score: 25.4, Datasets [pdf] Machine Learning Yearning by Andrew Ng Machine Learning Yearning is a technical book by Andrew Ng that provides practical advice on how to structure machine learning projects. ... ``` --- # sql Start an interactive SQL query session against the Spice runtime ### Usage[​](#usage "Direct link to Usage") ``` spice sql [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--cloud` Use a Spice Cloud instance as the SQL backend. Requires `--api-key`. * `--endpoint ` Specifies the remote Spice instance endpoint. Supports `http://`, `https://`, `grpc://`, or `grpc+tls://` schemes. If not provided, uses the local spiced runtime. * `--flight-endpoint ` (Deprecated) Specifies the remote Spice instance Flight endpoint (treated as gRPC endpoint). If not provided, uses the local spiced runtime. * `--http-endpoint ` (Deprecated) HTTP endpoint of Spice (default: `http://127.0.0.1:8090`). * `--tls-root-certificate-file ` The path to the root certificate file used to verify the Spice.ai runtime server certificate. * `-h`, `--help` Print this help message. ### Examples[​](#examples "Direct link to Examples") ``` $ spice sql Welcome to the Spice.ai SQL REPL! Type 'help' for help. show tables; -- list available tables sql> show tables +---------------+--------------------+---------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------------+---------------+------------+ | datafusion | public | tmp_view_test | VIEW | | datafusion | information_schema | tables | VIEW | | datafusion | information_schema | views | VIEW | | datafusion | information_schema | columns | VIEW | | datafusion | information_schema | df_settings | VIEW | +---------------+--------------------+---------------+------------+ ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` $ spice sql --tls-root-certificate-file /path/to/cert.pem Welcome to the Spice.ai SQL REPL! Type 'help' for help. ``` #### Remote and Cloud Examples[​](#remote-and-cloud-examples "Direct link to Remote and Cloud Examples") ``` # Connect to Spice Cloud spice sql --cloud --api-key # Connect to a remote spiced instance over HTTP spice sql --endpoint http://my-remote-host:8090 # Connect to a remote spiced instance over Arrow Flight SQL (gRPC) spice sql --endpoint grpc://my-remote-host:50051 show tables; -- list available tables sql> show tables +---------------+--------------------+---------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------------+---------------+------------+ | datafusion | public | tmp_view_test | VIEW | | datafusion | information_schema | tables | VIEW | | datafusion | information_schema | views | VIEW | | datafusion | information_schema | columns | VIEW | | datafusion | information_schema | df_settings | VIEW | +---------------+--------------------+---------------+------------+ ``` --- # status Spice runtime status ### Usage[​](#usage "Direct link to Usage") ``` spice status [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `--tls-root-certificate-file` The path to the root certificate file used to verify the Spice.ai runtime server certificate * `-h`, `--help` help for status ### Examples[​](#examples "Direct link to Examples") ``` >>> spice status NAME ENDPOINT STATUS http 127.0.0.1:8090 Ready flight 127.0.0.1:50051 Ready metrics N/A Disabled opentelemetry 127.0.0.1:50052 Ready ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` >>> spice status --tls-root-certificate-file /path/to/cert.pem NAME ENDPOINT STATUS http 127.0.0.1:8090 Ready flight 127.0.0.1:50051 Ready metrics N/A Disabled opentelemetry 127.0.0.1:50052 Ready ``` --- # trace Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the `runtime.task_history` table. ### Usage[​](#usage "Direct link to Usage") ``` spice trace [task] [flags] ``` `task` - The name of the task whose trace is requested. Supported tasks include: * `accelerated_refresh` * `ai_chat` * `ai_completion` * `sql_query` * `nsql` * `tool_use::search` * `tool_use::list_datasets` * `tool_use::sql` * `tool_use::table_schema` * `tool_use::sample_data` * `tool_use::sql_query` * `tool_use::memory` * `search` * `scheduled_worker` These tasks are from the `task` column in the Spice SQL `runtime.task_history` table. #### Flags[​](#flags "Direct link to Flags") * `--trace-id` Retrieve the trace with the given trace ID (the column `trace_id` from `runtime.task_history`). * `--id` Retrieve the trace with the given `id` label (i.e. the task has a valid `id` within the `labels` column of `runtime.task_history`). * `--api-key` Specify the API key for authentication. * `--include-output`: Include, as an additional column, the captured output to each span (i.e. the `captured_output` column from `runtime.task_history`). Note: If captured outputs are not being stored, this will return an empty row. * `--include-input`: Include, as an additional column, the input to each span (i.e. the `input` column from `runtime.task_history`). The latest trace for the task will be used if neither `--trace-id` nor `--id` is specified. ### Examples[​](#examples "Direct link to Examples") #### Retrieve the trace for the last text-to-SQL operation[​](#retrieve-the-trace-for-the-last-text-to-sql-operation "Direct link to Retrieve the trace for the last text-to-SQL operation") ``` spice trace nsql ``` #### Retrieve the trace for a specific task by ID[​](#retrieve-the-trace-for-a-specific-task-by-id "Direct link to Retrieve the trace for a specific task by ID") ``` spice trace ai_chat --id chatcmpl-At6ZmDE8iAYRPeuQLA0FLlWxGKNnM ``` #### Retrieve a trace by `trace-id`[​](#retrieve-a-trace-by-trace-id "Direct link to retrieve-a-trace-by-trace-id") ``` spice trace sql_query --trace-id d5c6f1eed9f27257 ``` ### Output Example[​](#output-example "Direct link to Output Example") ``` TREE STATUS DURATION TASK a97f52ccd7687e64 ✅ 673.14ms ai_chat ├── 4eebde7b04321803 ✅ 0.04ms tool_use::list_datasets └── 4c9049e1bf1c3500 ✅ 671.91ms ai_completion ``` This output represents a structured trace of executed tasks. ### Output Example (with `--include-output`)[​](#output-example-with---include-output "Direct link to output-example-with---include-output") ``` TREE STATUS DURATION TASK OUTPUT a97f52ccd7687e64 ✅ 673.14ms ai_chat The capital of New York is Albany. ├── 4eebde7b04321803 ✅ 0.04ms tool_use::list_datasets [] └── 4c9049e1bf1c3500 ✅ 671.91ms ai_completion [{"content":"The capital of New York is Albany.","refusal":null,"tool_calls":null,"role":"assistant","function_call":null,"audio":null}] ``` --- # upgrade Upgrades the Spice CLI & Runtime to the latest release ### Usage[​](#usage "Direct link to Usage") ``` spice upgrade [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` help for upgrade ### Examples[​](#examples "Direct link to Examples") ``` spice upgrade ``` --- # version Outputs the current version of the Spice CLI and runtime ### Usage[​](#usage "Direct link to Usage") ``` spice version [flags] ``` #### Flags[​](#flags "Direct link to Flags") * `-h`, `--help` help for version ### Sample output[​](#sample-output "Direct link to Sample output") **Upgrade available**: ``` > spice version 2024/12/17 22:42:11 INFO CLI version: v0.18.3-beta 2024/12/17 22:42:13 INFO Runtime version: v0.20.0-beta+models 2024/12/17 22:42:14 INFO CLI version v1.0.0-rc.4 is now available! To upgrade, run "spice upgrade". ``` Learn more about upgrading the Spice CLI and runtime using `spice upgrade` [here.](/docs/v1.10/cli/reference/upgrade) **Latest Version**: ``` > spice version CLI version: v1.0.0-rc.4 Runtime version: v1.0.0-rc.4+models ``` --- # Configuring Trace Levels Trace output verbosity is determined by the following sources, listed in order of precedence: 1. Verbosity flags (`-v`/`--verbose`, `-vv`/`--very-verbose`). If these flags are provided, they override all other settings. 2. The `SPICED_LOG` environment variable. This is used only if verbosity flags are not set 3. The `runtime.output_level` YAML configuration file. This is used only if neither verbosity flags nor the environment variable are set. ### Default[​](#default "Direct link to Default") The default trace level is `INFO`, suitable for general information about the system. ``` SPICED_LOG="task_history=INFO,spiced=INFO,runtime=INFO,secrets=INFO,data_components=INFO,cache=INFO,extensions=INFO,spice_cloud=INFO,llms=INFO,reqwest_retry::middleware=off,WARN" ``` The equivalent `runtime.output_level` configuration is `info`: ``` runtime: output_level: info ``` ### Enabling Debug Mode[​](#enabling-debug-mode "Direct link to Enabling Debug Mode") Use the `-v`/`--verbose` CLI flags to enable detailed logs, useful for debugging. ``` spice run -v spiced -v ``` Alternatively you can use `runtime.output_level` yaml configuration: ``` runtime: output_level: verbose ``` This sets `SPICED_LOG` to `DEBUG` level: ``` SPICED_LOG="task_history=DEBUG,spiced=DEBUG,runtime=DEBUG,secrets=DEBUG,data_components=DEBUG,cache=DEBUG,extensions=DEBUG,spice_cloud=DEBUG,llms=DEBUG,DEBUG" spice run ``` ### Enabling Trace Mode[​](#enabling-trace-mode "Direct link to Enabling Trace Mode") Use the `-vv`/`--very-verbose` CLI flag to enable the most detailed logs, typically for in-depth troubleshooting. ``` spice run -vv spiced -vv ``` Alternatively you can use `runtime.output_level` yaml configuration: ``` runtime: output_level: very_verbose ``` This sets `SPICED_LOG` to `TRACE` level: ``` SPICED_LOG="task_history=TRACE,spiced=TRACE,runtime=TRACE,secrets=TRACE,data_components=TRACE,cache=TRACE,extensions=TRACE,spice_cloud=TRACE,llms=TRACE,TRACE" spice run ``` ### Granular Configuration[​](#granular-configuration "Direct link to Granular Configuration") For specific component trace configuration, adjust the trace levels as needed: ``` SPICED_LOG="spiced=INFO,runtime=DEBUG,data_components=WARN,cache=WARN" spice run ``` --- # Clients and Tools ## [📄️DBeaver](/docs/v1.10/clients/dbeaver) [Configure DBeaver to query Spice via JDBC](/docs/v1.10/clients/dbeaver) ## [📄️JetBrains DataGrip](/docs/v1.10/clients/jetbrains-datagrip) [Configure JetBrains Datagrip to query Spice via JDBC](/docs/v1.10/clients/jetbrains-datagrip) ## [📄️Apache Superset](/docs/v1.10/clients/superset) [Use Apache Superset to query and visualize datasets loaded in Spice.](/docs/v1.10/clients/superset) ## [📄️Tableau](/docs/v1.10/clients/tableau) [Use Tableau to to access, visualise and analyse datasets loaded in Spice.](/docs/v1.10/clients/tableau) ## [📄️Microsoft Power BI](/docs/v1.10/clients/powerbi) [Use Microsoft Power BI to access, visualize and analyze Spice datasets.](/docs/v1.10/clients/powerbi) --- # DBeaver 1. Start the Spice runtime with a dataset loaded. Follow the [quickstart guide](/docs/v1.10/getting-started) to get started. 2. Download [DBeaver Community Edition](https://dbeaver.io). 3. Download the [Apache Arrow Flight SQL JDBC driver](https://search.maven.org/search?q=a:flight-sql-jdbc-driver) - choose the "jar" option. 4. Launch DBeaver 5. In the DBeaver application menu bar, open the "Database" menu and choose: "Driver Manager": ![Driver manager menu option](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/691d1f83-c1d0-4ad8-ec8d-d8f37ccc9d00/public "Driver manager menu option") 6. Click the "New" button on the right: ![Driver manager new button](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/5783d944-daae-4735-99e9-976f974bc100/public "Driver manager new button") 7. Add the JDBC jar file: 1. Click the "Libraries" tab 2. Click the: "Add File" button 3. Choose the "flight-sql-jdbc-driver-15.0.1.jar" jar file (the file downloaded in step 3 above) - and click "Open" ![Select jar file](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/19900f7a-f00f-473d-780e-4a28c2ecd800/public "Select jar file") 4. Close the Driver editor window with the blue "OK" button on the lower-right 8. Enter the driver settings: 1. Click the "Settings" tab 2. In the "Driver Name" field - enter: `Apache Arrow Flight SQL` 3. In the "URL Template" field - enter: `jdbc:arrow-flight-sql://{host}:{port}?useEncryption=false&disableCertificateVerification=true` * If [API key authentication](/docs/v1.10/api/auth) is enabled, the URL template should be: `jdbc:arrow-flight-sql://{host}:{port}?useEncryption=false&disableCertificateVerification=true&user=&password=` - where `` is the API key value 1. In the "Driver Type" drop-down box - choose: "SQLite" 2. Select "No authentication" * This should be selected even if API key authentication is enabled in the runtime, as the API key is supplied via the URL template above. 1. The driver manager "Edit Driver" window should look like this: ![Driver Manager completed](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/20348c42-117b-4763-80d2-6e615b23ae00/public "Driver Manager completed") 2. Click the blue "OK" button on the lower-right to save the driver 3. Close the "Driver Manager" window by clicking the blue "Close" button on the lower-right. 9. Create a new Database Connection: 1. In the DBeaver application menu bar, open the "Database" menu and choose: "New Database Connection": ![New Database Connection](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/acdf7251-4238-44ee-9639-0c557518da00/public "New Database Connection") 2. In the "Connect to a database" window - type: `Flight` in the search bar 3. Choose the `Apache Arrow Flight SQL` driver - the window should look like this: ![Connect to a database window](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/61cee5fe-dc75-4ac1-e558-eea3aff4c100/public "Connect to a database window") 4. Click the blue "Next >" button on the bottom of the window 5. On the next screen, the JDBC URL should be filled out already - just supply the Host (`localhost`) and Port (`50051`) values for the Spice runtime. The window should look like this: ![Connect to a database window 2](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/2a2b2fdc-00db-49d3-5359-059b12342b00/public "Connect to a database window 2") 6. Click the "Test Connection" button - the window should look like this: ![Test Connection results](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/a3fc5f5f-a39f-47ce-7955-4b384ec1ae00/public "Test Connection results") 7. Click the blue "OK" button to close the Connection test window 8. Click the "Connection details (name, type, ...)" button on the right 9. In the "General" section, enter: `Spice Runtime` for the "Connection name". It should look like this: ![Name the Database Connection](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/f6d04fe1-92a1-4082-d4ea-e9daacaca200/public) 10. Click the blue "Finish" button to save the connection 10. Run a query: 1. Right-click on the Database Connection on the left - choose: "SQL Editor", and then: "Open SQL Console" as shown here: ![Open SQL Console](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/642a5885-9e3f-4dd7-ef43-72bfce27bb00/public "Open SQL Console") 2. In the Console window - run a query - something like: `SELECT * FROM taxi_trips;` 3. Click the triangle button to execute the SQL statement - as shown below (or use keyboard shortcut: Ctrl+Enter): ![Execute SQL](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/2134e47b-a066-47e9-1d48-06352675f400/public "Execute SQL") 4. See the query results as shown in this screenshot: ![Query Results](https://imagedelivery.net/HyTs22ttunfIlvyd6vumhQ/0e9f3c0f-2e03-47f9-8d5e-65e078d7e900/public "Query Results") 5. DBeaver is now configured to query the Spice runtime using SQL! 🎉 --- # JetBrains DataGrip 1. Start the Spice runtime with a dataset loaded. Follow the [quickstart guide](/docs/v1.10/getting-started) to get started. 2. Download [JetBrains DataGrip](https://www.jetbrains.com/datagrip). 3. Download the [Apache Arrow Flight SQL JDBC driver](https://search.maven.org/search?q=a:flight-sql-jdbc-driver) - Select "Versions", tab click "Browse" on most recent version and then download `flight-sql-jdbc-driver-.jar`. 4. Launch DataGrip 5. In Database Explorer menu, select "+" and choose "Driver" ![Data Sources and Drivers menu option](/assets/images/datagrip-1-12175a5760ca0dfde19f28d8253f2263.png "Data Sources and Drivers menu option") 6. Add the JSBC jar file: 1. Click the "+" button in "Driver Files" selection 2. Click the "Custom JARs" button 3. Choose the `flight-sql-jdbc-driver-.jar` jar file (the file downloaded in step 3 above) - and click "Open" 4. Click the "Class:" selector 5. Select `org.apache.arrow.driver.jdbc.ArrowFlightJdbcDriver` ![Driver Class selector](/assets/images/datagrip-3-62e98abc87e0f4705a08f2921bf2dbbc.png "Driver Class selector") 7. Enter the driver settings: 1. In the "Name" field - enter: `Apache Arrow Flight SQL` 2. Add "URL Template" Default: `jdbc:arrow-flight-sql://{host}:{port}\?useEncryption=false&disableCertificateVerification=true` 3. Click "Ok" ![Driver creation window](/assets/images/datagrip-4-6082661b869e86a8d39e08c804546735.png "Driver creation window") 8. Create a new Database Connection: 1. In Database Explorer menu, select "+", choose "Data Source" > "Arrow Flight JDBC" 2. Set the host to `localhost` and the port to `50051` 3. In "Authentication" select "No auth" 4. Click "Test Connection" to verify ![New Data Source](/assets/images/datagrip-5-08d8b998640048850709308946762442.png "New Data Source") 10. Run a query: 1. Right-click on the connection in Database Explorer and choose "New" > "Query Console" ![Create new Query Console](/assets/images/datagrip-6-7ca8eafad45c1890065deb2048c3e625.png "Create new Query Console") 2. In the Console window - add a query - something like: `SELECT * FROM taxi_trips;` and click the triangle button to execute the SQL statement 3. See the query results: ![Query Results](/assets/images/datagrip-7-81710c69e680dd7e49aad30e951aa2f0.png "Query Results") DataGrip is now configured to query the Spice runtime using SQL! 🎉 --- # Microsoft Power BI Connector Use the instructions below to get started with the **[Spice.ai Power BI Connector](https://github.com/spiceai/powerbi-connector)**—an [ADBC](https://github.com/apache/arrow-adbc)-based connector that enables [Microsoft Power BI](https://www.microsoft.com/en-us/power-platform/products/power-bi) users to easily connect to and visualize data loaded in [Spice.ai Enterprise](https://spiceai.org/) and [Spice Cloud Platform](https://spice.ai/) instances. ## Manual Connector Installation[​](#manual-connector-installation "Direct link to Manual Connector Installation") ### Power BI Desktop[​](#power-bi-desktop "Direct link to Power BI Desktop") 1. Download the latest `spice_adbc.mez` file from the [releases page](https://github.com/spiceai/powerbi-connector/releases) 2. Copy to your Power BI `Custom Connectors` directory: `C:\Users\[USERNAME]\Documents\Microsoft Power BI Desktop\Custom Connectors` ``` Invoke-WebRequest -Uri "https://github.com/spiceai/powerbi-connector/releases/latest/download/spice_adbc.mez" -OutFile "C:\Users\[USERNAME]\Documents\Microsoft Power BI Desktop\Custom Connectors\spice_adbc.mez" ``` 3. [Enable Uncertified Connectors](https://learn.microsoft.com/en-us/power-bi/connect-data/desktop-connector-extensibility#custom-connectors) in Power BI Desktop settings and restart Power BI Desktop. ## Adding Spice as a Data Source[​](#adding-spice-as-a-data-source "Direct link to Adding Spice as a Data Source") 1. Open Power BI Desktop. 2. Click on `Get Data` → `More...`. 3. In the dialog, select `Spice.ai` connector. ![Spice.ai connector](/img/powerbi/powerbi-spice-connector.png) 4. Click `Connect`. 5. Enter the **ADBC (Arrow Flight SQL) Endpoint**: * For Spice Cloud Platform:
`grpc+tls://flight.spiceai.io:443`
*(Use the region-specific address if applicable.)* * For on-premises/self-hosted Spice.ai: * Without TLS (default): `grpc://:50051` * With TLS: `grpc+tls://:50051` ![Spice.ai Connection Dialog](/img/powerbi/powerbi-spice-connection-dlg.png) 6. Select the `Data Connectivity` mode: * **Import**: Data is loaded into Power BI, enabling extensive functionality but requiring periodic refreshes and sufficient local memory to accommodate the dataset. * **DirectQuery**: Queries are executed directly against Spice in real-time, providing fast performance even on large datasets by leveraging Spice's optimized query engine. 7. Click `OK`. 8. Select `Authentication` option: * **Anonymous**: Select for unauthenticated on-premises deployments. * **API Key**: Your Spice.ai API key for authentication (required for Spice Cloud). Follow the [guide](https://docs.spice.ai/portal/apps/api-keys) to obtain it from the Spice Cloud portal. ![Spice.ai Authentication](/img/powerbi/powerbi-spice-auth-dlg.png) 9. Click `Connect` to establish the connection. ## Working with Spice datasets[​](#working-with-spice-datasets "Direct link to Working with Spice datasets") After establishing a connection, Spice datasets appear under their respective schemas, with the default schema being `spice.public`. When writing native queries, use the `PostgreSQL` dialect, as Spice is built on this standard. ![Spice PowerBI Example](/img/powerbi/powerbi-spice-example.png) ## Supported Data Types[​](#supported-data-types "Direct link to Supported Data Types") The following Apache Arrow / DataFusion SQL types are supported. Other types will result in a `Unable to understand the type for column` error. Please [report an issue](https://github.com/spiceai/powerbi-connector/issues) if support for additional types is required. | Arrow Type | DataFusion SQL Type | Power Query M Type | | ----------------------------------------------------------- | ------------------- | ------------------ | | Boolean | BOOLEAN | Logical | | Int16 | SMALLINT | Int16 | | Int32 | INTEGER | Int32 | | Int64 | BIGINT | Int64 | | Float32 | REAL | Single | | Float64 | DOUBLE | Double | | Decimal128 / Decimal256 | DECIMAL | Decimal | | Utf8 | VARCHAR | Text | | Date32 / Date64 | DATE | Date | | Time32 / Time64 | TIME | Time | | Timestamp | TIMESTAMP | DateTime | | List / LargeList / FixedSizeList / ListView / LargeListView | ARRAY | Text | | Interval | INTERVAL | Text | | Struct | STRUCT | Text | ## Limitations[​](#limitations "Direct link to Limitations") ### LargeUtf8 Data Type Is Not Supported[​](#largeutf8-data-type-is-not-supported "Direct link to LargeUtf8 Data Type Is Not Supported") To work around this limitation, use [views](https://spiceai.org/docs/components/views) to manually convert `LargeUtf8` columns to `Utf8` by casting them with `::TEXT`. **Example:** ``` views: - name: taxi_zone_lookup sql: | SELECT LocationID as LocationID, Borough::TEXT as Borough, Zone::TEXT as Zone, service_zone::TEXT as service_zone FROM taxi_zone_lookup_temp; ``` ### Date Time Arithmetic Operations Are Not Supported[​](#date-time-arithmetic-operations-are-not-supported "Direct link to Date Time Arithmetic Operations Are Not Supported") Due to lack of support for the `timestampdiff` function in the [DataFusion query engine](https://datafusion.apache.org/user-guide/sql/scalar_functions.html), date and time arithmetic operations—such as subtracting or adding timestamps and intervals—are not supported and will result in an error similar to `Invalid function 'timestampdiff'.\nDid you mean 'to_timestamp'? (Internal; ExecuteQuery)`. For example: ``` (parameter) => let Sorted = Table.Sort(parameter[taxi_table], {"RecordID"}), T2 = Table.SelectColumns(Sorted, {"PULocationID","lpep_pickup_datetime"}), T3 = Table.Sort(T2, {"PULocationID"}), T4 = Table.AddColumn(T3, "Diff1", each [lpep_pickup_datetime] - #datetime(1999,1,5,0,0,0)) TA = Table.FirstN(T6, 4) in TA ``` ``` ADBC: InternalError [] [FlightSQL] [FlightSQL] Error during planning: Invalid function 'timestampdiff'.\nDid you mean 'to_timestamp'? (Internal; ExecuteQuery) ``` Please [report an issue](https://github.com/spiceai/powerbi-connector/issues) if support for date or time arithmetic operations is required. --- # Apache Superset Use [Apache Superset](https://superset.apache.org/) to query and visualize datasets loaded in Spice. > Apache Superset is a modern, enterprise-ready business intelligence web application. It is fast, lightweight, intuitive, and loaded with options that make it easy for users of all skill sets to explore and visualize their data, from simple pie charts to highly detailed deck.gl geospatial charts. > > – [Apache Superset documentation](https://superset.apache.org/docs/intro/) ## Start Apache Superset with Flight SQL & DataFusion SQL Dialect support[​](#start-apache-superset-with-flight-sql--datafusion-sql-dialect-support "Direct link to Start Apache Superset with Flight SQL & DataFusion SQL Dialect support") Superset requires a Python [DB API 2](https://peps.python.org/pep-0249/) database driver and a [SQLAlchemy](https://www.sqlalchemy.org/) dialect to be installed for each connected datastore. Spice implements a Flight SQL server that understands the DataFusion SQL Dialect. The [`flightsql-dbapi`](https://pypi.org/project/flightsql-dbapi/) library for Python provides the required DB API 2 driver and SQLAlchemy dialect. Select the appropriate tab based on whether you are experimenting with this feature or integrating it into an existing Superset instance. * Experimenting * Integrating with Existing Superset The easiest way to connect Apache Superset and Spice is to follow the [`Sales BI` Cookbook Recipe](https://github.com/spiceai/cookbook/tree/trunk/sales-bi). This recipe builds a local Docker image based on Apache Superset that is pre-configured with the `flightsql-dbapi` library needed to connect to Spice. Clone the Spice cookbook repository and navigate to the `sales-bi` directory: ``` git clone https://github.com/spiceai/cookbook.git cd cookbook/sales-bi ``` Start Apache Superset along with the Spice runtime in Docker Compose: ``` make start ``` Log into Apache Superset at with the username and password `admin/admin`. Follow the below steps to configure a database connection to Spice manually, or run `make import-dashboards` to automatically configure the connection and create a sample dashboard. ## Generic / Virtual Machine[​](#generic--virtual-machine "Direct link to Generic / Virtual Machine") Install the `flightsql-dbapi` library in your existing Apache Superset environment: ``` pip install flightsql-dbapi ``` ## Docker Container[​](#docker-container "Direct link to Docker Container") Install the library in the Dockerfile: ``` FROM apache/superset # Switching to root to install the required packages USER root # https://github.com/influxdata/flightsql-dbapi RUN pip install flightsql-dbapi # Switching back to using the `superset` user USER superset ``` Re-deploy Apache Superset with the updated Docker image. ## Temporary Docker Container Modification[​](#temporary-docker-container-modification "Direct link to Temporary Docker Container Modification") It's possible to modify a running Docker container to install the library, but the change will be lost on container restart. ``` docker exec -u root -it superset /bin/bash pip install flightsql-dbapi ``` *** ## Configure a Spice Connection[​](#configure-a-spice-connection "Direct link to Configure a Spice Connection") Once Apache Superset is up and running, and you are logged in, you can configure a connection to Spice. Hover over the `Settings` menu and select `Database Connections`. ![](/img/superset/superset-docs-connection-settings.png) Click the `+ Database` button to configure the connection. ![](/img/superset/superset-docs-new-db.png) Under `Supported Databases` select `Other`. Set the Display Name to `Spice` and the SQL Alchemy URI to `datafusion+flightsql://spiceai_host:[spiceai_port]`. Specify `?insecure=true` to skip connecting over TLS. Example: `datafusion+flightsql://spiceai-sales-bi-demo:50051?insecure=true`. Click `Test Connection` to verify the connection. ![](/img/superset/superset-docs-test-conn.png) Click `Connect` to save the connection. Start exploring the datasets loaded in Spice by creating a new dataset in Apache Superset to match one of the existing tables. --- # Tableau Use instructions below to install the **Spice.ai Tableau Connector** that enables [Tableau](https://www.tableau.com/) users to easily connect to and visualize data loaded in Spice. > Tableau is the world's leading analytics platform. Tableau is the broadest and deepest end-to-end data and analytics platform. Ensure the responsible use of data and drive better business outcomes with fully integrated data management and governance, visual analytics and data storytelling, and collaboration – all with Salesforce’s industry-leading Einstein built right in. > > – [The Tableau platform](https://www.tableau.com/) ## Step 1. Install the Arrow Flight SQL JDBC Driver[​](#step-1-install-the-arrow-flight-sql-jdbc-driver "Direct link to Step 1. Install the Arrow Flight SQL JDBC Driver") [JDBC](https://docs.oracle.com/javase/tutorial/jdbc/basics/index.html) (Java Database Connectivity) is a standard interface for connecting to and interacting with databases. The Flight SQL driver is a JDBC driver implementation based on the [Arrow Flight SQL](https://arrow.apache.org/docs/format/FlightSql.html) protocol. As Spice supports the Flight SQL protocol, the driver helps establish a connection between Tableau and Spice, enabling Tableau to execute queries and retrieve data from Spice efficiently. Download the [flight-sql-jdbc-driver.jar](https://repo1.maven.org/maven2/org/apache/arrow/flight-sql-jdbc-driver/) file to the Tableau drivers folder: * Windows * macOS * Linux **PowerShell Install Script** ``` Invoke-WebRequest -Uri "https://repo1.maven.org/maven2/org/apache/arrow/flight-sql-jdbc-driver/18.2.0/flight-sql-jdbc-driver-18.2.0.jar" -OutFile "C:\Program Files\Tableau\Drivers\flight-sql-jdbc-driver-18.2.0.jar" ``` **Install Script** ``` curl -L https://repo1.maven.org/maven2/org/apache/arrow/flight-sql-jdbc-driver/18.2.0/flight-sql-jdbc-driver-18.2.0.jar -o ~/Library/Tableau/Drivers/flight-sql-jdbc-driver-18.2.0.jar ``` **Install Script** ``` curl -L https://repo1.maven.org/maven2/org/apache/arrow/flight-sql-jdbc-driver/18.2.0/flight-sql-jdbc-driver-18.2.0.jar -o /opt/tableau/tableau_driver/jdbc/flight-sql-jdbc-driver-18.2.0.jar ``` ## Step 2. Install Spice.ai Tableau Connector[​](#step-2-install-spiceai-tableau-connector "Direct link to Step 2. Install Spice.ai Tableau Connector") ### Tableau Server[​](#tableau-server "Direct link to Tableau Server") 1. Download the latest `spiceai.taco` file from [Releases](https://github.com/spicehq/tableau-connector/releases) 2. Copy to the Tableau connectors directory * Windows * Linux **PowerShell Install Script** ``` Invoke-WebRequest -Uri "https://github.com/spicehq/tableau-connector/releases/latest/download/spiceai.taco" -OutFile "C:\Program Files\Tableau\Connectors\spiceai.taco" ``` **Install Script** ``` curl -L https://github.com/spicehq/tableau-connector/releases/latest/download/spiceai.taco -o /opt/tableau/connectors/spiceai.taco ``` 3. Restart server: `tsm restart` ### Tableau Desktop[​](#tableau-desktop "Direct link to Tableau Desktop") 1. Download the latest `spiceai.taco` file from [Releases](https://github.com/spiceai/tableau-connector/releases) 2. Copy to the Tableau connectors directory * Windows * macOS * Linux **PowerShell Install Script** ``` Invoke-WebRequest -Uri "https://github.com/spicehq/tableau-connector/releases/latest/download/spiceai.taco" -OutFile "C:\Users\[USERNAME]\Documents\My Tableau Repository\Connectors\spiceai.taco" ``` **Install Script** ``` curl -L https://github.com/spicehq/tableau-connector/releases/latest/download/spiceai.taco -o ~/Documents/My\ Tableau\ Repository/Connectors/spiceai.taco ``` **Install Script** ``` curl -L https://github.com/spicehq/tableau-connector/releases/latest/download/spiceai.taco -o /opt/tableau/connectors/spiceai.taco ``` ## Configure a Spice connection[​](#configure-a-spice-connection "Direct link to Configure a Spice connection") 1. Open **Tableau** 2. In the **Connect** column, under **To a Server**, select **Spice.ai by Spice AI, Inc**. 3. Configure a Spice connection to **Spice.ai OSS Self-Hosted** instance or to **Spice Cloud Platform**. ![Spice Tableau Connection Dialog](/img/tableau/tableau-spice-dialog.png) 4. Click **Sign In** ## Working with Spice datasets[​](#working-with-spice-datasets "Direct link to Working with Spice datasets") After establishing a connection, Spice datasets appear under their respective schemas, with the default schema being `spice.public`. When writing queries, use the `PostgreSQL` dialect, as Spice is built on this standard. ![Spice Tableau Example](/img/tableau/tableau-spice-example.png) --- # Runtime Components ## [🗃Data Connectors](/docs/v1.10/components/data-connectors) [32 items](/docs/v1.10/components/data-connectors) ## [🗃Data Accelerators](/docs/v1.10/components/data-accelerators) [5 items](/docs/v1.10/components/data-accelerators) ## [🗃Secret Stores](/docs/v1.10/components/secret-stores) [4 items](/docs/v1.10/components/secret-stores) ## [🗃Catalog Connectors](/docs/v1.10/components/catalogs) [5 items](/docs/v1.10/components/catalogs) ## [🗃Embeddings](/docs/v1.10/components/embeddings) [7 items](/docs/v1.10/components/embeddings) ## [🗃Vector Engines](/docs/v1.10/components/vectors) [1 item](/docs/v1.10/components/vectors) ## [📄️Views](/docs/v1.10/components/views) [Documentation for defining Views in Spice](/docs/v1.10/components/views) ## [📄️Workers Overview](/docs/v1.10/components/workers) [Detailed documentation for workers in the Spice runtime.](/docs/v1.10/components/workers) ## [🗃Model Providers](/docs/v1.10/components/models) [10 items](/docs/v1.10/components/models) ## [🗃LLM Tools](/docs/v1.10/components/tools) [2 items](/docs/v1.10/components/tools) --- # Catalog Connectors In Spice, datasets are organized hierarchically with catalogs, schemas, and tables. A catalog, at the top level, contains multiple schemas. Each schema, in turn, contains multiple tables where the actual data is stored. By default a catalog named `spice` is created with all of the datasets defined in the `datasets` section of the Spicepod. ![](/img/catalog-schema-table.png) Creating schemas and tables within the `spice` catalog is configured by the `name` field in the dataset configuration. A name with a period (`.`) will create schema, i.e. a dataset defined with `name: foo.bar` would have a full path of `spice.foo.bar`. If the name does not contain a period, the dataset will be created in the `public` schema of the `spice` catalog. For example, a dataset defined with `name: foo` would have a full path of `spice.public.foo`. Attempting to create a dataset with a name that contains a catalog name will result in an error. Adding catalogs to Spice is done via Catalog Connectors. Catalog Connectors connect to external catalog providers and make their tables available for federated SQL query in Spice. Configuring accelerations for tables in external catalogs is not supported. The schema hierarchy of the external catalog is preserved in Spice. Supported Catalog Connectors include: | Name | Description | Status | Protocol/Format | | --------------- | ----------------------- | ------ | ---------------------------- | | `unity_catalog` | Unity Catalog | Stable | Delta Lake | | `databricks` | Databricks | Beta | Spark Connect, S3/Delta Lake | | `iceberg` | Apache Iceberg | Beta | Parquet | | `spice.ai` | Spice.ai Cloud Platform | Beta | Arrow Flight | | `glue` | AWS Glue | Alpha | Parquet, Iceberg | ## Catalog Connector Docs[​](#catalog-connector-docs "Direct link to Catalog Connector Docs") Catalog are configured using a Catalog Connector in the `catalogs` section of the Spicepod. See the specific Catalog Connector documentation for configuration details. ### `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. Example: ``` catalogs: - from: spice.ai name: spiceai include: - 'tpch.*' # Include only the "tpch" tables. ``` ## [📄️Databricks](/docs/v1.10/components/catalogs/databricks) [Connect to a Databricks Unity Catalog provider.](/docs/v1.10/components/catalogs/databricks) ## [📄️Unity Catalog](/docs/v1.10/components/catalogs/unity-catalog) [Connect to a Unity Catalog provider.](/docs/v1.10/components/catalogs/unity-catalog) ## [📄️Spice.ai](/docs/v1.10/components/catalogs/spiceai) [Connect to the Spice.ai built-in catalog.](/docs/v1.10/components/catalogs/spiceai) ## [📄️Iceberg](/docs/v1.10/components/catalogs/iceberg) [Connect to an Iceberg catalog provider.](/docs/v1.10/components/catalogs/iceberg) ## [📄️Glue](/docs/v1.10/components/catalogs/glue) [Connect to an AWS Glue Data Catalog.](/docs/v1.10/components/catalogs/glue) --- # Databricks Catalog Connector Connect to a [Databricks Unity Catalog](https://www.databricks.com/product/unity-catalog) as a catalog provider for federated SQL query using [Spark Connect](https://www.databricks.com/blog/2022/07/07/introducing-spark-connect-the-power-of-apache-spark-everywhere.html), directly from [Delta Lake](https://delta.io/) tables, or using the [SQL Statement Execution API](https://docs.databricks.com/aws/en/dev-tools/sql-execution-tutorial). ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog # tables from this catalog will be available in the "uc_catalog" catalog in Spice include: - '*.my_table_name' # include only the "my_table_name" tables params: mode: delta_lake # or spark_connect or sql_warehouse databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com dataset_params: # delta_lake S3 parameters databricks_aws_region: us-west-2 databricks_aws_access_key_id: ${secrets:aws_access_key_id} databricks_aws_secret_access_key: ${secrets:aws_secret_access_key} databricks_aws_endpoint: s3.us-west-2.amazonaws.com # spark_connect parameters databricks_cluster_id: 1234-567890-abcde123 # sql_warehouse parameters databricks_sql_warehouse_id: 2b4e24cff378fb24 ``` ## `from`[​](#from "Direct link to from") The `from` field is used to specify the catalog provider. For Databricks, use `databricks:`. The `catalog_name` is the name of the catalog in the Databricks Unity Catalog you want to connect to. ## `name`[​](#name "Direct link to name") The `name` field is used to specify the name of the catalog in Spice. Tables from the Databricks catalog will be available in the schema with this name in Spice. The schema hierarchy of the external catalog is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. ## `params`[​](#params "Direct link to params") The following parameters are supported for configuring the connection to the Databricks Unity Catalog: | Parameter Name | Definition | | --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mode` | The execution mode for querying against Databricks. `spark_connect` uses Spark Connect to query against Databricks requires a Spark cluster to be available. `delta_lake` queries directly from Delta Tables and requires the object store credentials to be provided. Default is `spark_connect`. | | `databricks_endpoint` | The Databricks workspace endpoint, e.g. `dbc-a12cd3e4-56f7.cloud.databricks.com` | | `databricks_token` | The Databricks API token to authenticate with the Unity Catalog API. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to reference a secret, e.g. `${secrets:my_databricks_token}`. | | `databricks_use_ssl` | If true, use a TLS connection to connect to the Databricks endpoint. Default is `true`. | To locate the Databricks endpoint, do the following: 1. Log in to your Databricks workspace. 2. In the sidebar, click Compute. 3. In the list of available clusters, click the target cluster's name. 4. On the Configuration tab, expand Advanced options. 5. Click the JDBC/ODBC tab. 6. The endpoint is the Server Hostname. ## Authentication[​](#authentication "Direct link to Authentication") ### Personal access token[​](#personal-access-token "Direct link to Personal access token") To learn more about how to set up personal access tokens, see [Databricks PAT docs](https://docs.databricks.com/aws/en/dev-tools/auth/pat). ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog include: - '*.my_table_name' params: databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com databricks_token: ${secrets:DATABRICKS_TOKEN} # PAT ``` ### Databricks service principal[​](#databricks-service-principal "Direct link to Databricks service principal") Spice supports the Machine-to-Machine (M2M) OAuth flow with service principal credentials by utilizing the `databricks_client_id` and `databricks_client_secret` parameters. The runtime will automatically refresh the token. Ensure that you grant your service principal the "Data Reader" privilege preset for the catalog and "Can Attach" cluster permissions when using Spark Connect mode. To learn more about how to set up the service principal, see [Databricks M2M OAuth docs](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m). ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog include: - '*.my_table_name' params: databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} # service principal client id databricks_client_secret: ${secrets:DATABRICKS_CLIENT_SECRET} # service principal client secret ``` ## `dataset_params`[​](#dataset_params "Direct link to dataset_params") The `dataset_params` field is used to configure the dataset-specific parameters for the catalog. The following parameters are supported: ### Spark Connect parameters[​](#spark-connect-parameters "Direct link to Spark Connect parameters") | Dataset Parameter Name | Definition | | ----------------------- | ---------------------------------------------------------------------------------------------- | | `databricks_cluster_id` | The ID of the compute cluster in Databricks to use for the query. e.g. `1234-567890-abcde123`. | To locate the cluster ID, do the following: 1. Log in to your Databricks workspace. 2. In the sidebar, click Compute. 3. In the list of available clusters, click the target cluster's name. 4. On the Configuration tab, expand Advanced options. 5. Click the JDBC/ODBC tab. 6. The cluster ID is the prefix of the Server Hostname. ### Delta Lake object store parameters[​](#delta-lake-object-store-parameters "Direct link to Delta Lake object store parameters") Configure the connection to the object store when using `mode: delta_lake`. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to reference a secret, e.g. `${secrets:aws_access_key_id}`. ### SQL Warehouse parameters[​](#sql-warehouse-parameters "Direct link to SQL Warehouse parameters") * `databricks_sql_warehouse_id`: The ID of the SQL Warehouse in Databricks to use for the query. e.g. `2b4e24cff378fb24`. To locate your SQL Warehouse ID, do the following: 1. Log in to your Databricks workspace. 2. In the sidebar, click SQL -> SQL Warehouses. 3. In the list of available warehouses, click the target warehouse's name. 4. Next to the **Name** field, the ID follows the name in parentheses. For example: `My Serverless Warehouse (ID: 2b4e24cff378fb24)` #### AWS S3[​](#aws-s3 "Direct link to AWS S3") | Dataset Parameter Name | Definition | | ---------------------------------- | ------------------------------------------------------------------------ | | `databricks_aws_region` | The AWS region for the S3 object store. E.g. `us-west-2`. | | `databricks_aws_access_key_id` | The access key ID for the S3 object store. | | `databricks_aws_secret_access_key` | The secret access key for the S3 object store. | | `databricks_aws_endpoint` | The endpoint for the S3 object store. E.g. `s3.us-west-2.amazonaws.com`. | Example: ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog include: - '*.my_table_name' params: mode: delta_lake databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com dataset_params: databricks_aws_region: us-west-2 databricks_aws_access_key_id: ${secrets:aws_access_key_id} databricks_aws_secret_access_key: ${secrets:aws_secret_access_key} databricks_aws_endpoint: s3.us-west-2.amazonaws.com ``` #### Azure Blob[​](#azure-blob "Direct link to Azure Blob") Note One of the following auth values must be provided for Azure Blob: * `databricks_azure_storage_account_key`, * `databricks_azure_storage_client_id` and `databricks_azure_storage_client_secret`, or * `databricks_azure_storage_sas_key`. | Dataset Parameter Name | Definition | | ---------------------------------------- | ---------------------------------------------------------------------- | | `databricks_azure_storage_account_name` | The Azure Storage account name. | | `databricks_azure_storage_account_key` | The Azure Storage master key for accessing the storage account. | | `databricks_azure_storage_client_id` | The service principal client id for accessing the storage account. | | `databricks_azure_storage_client_secret` | The service principal client secret for accessing the storage account. | | `databricks_azure_storage_sas_key` | The shared access signature key for accessing the storage account. | | `databricks_azure_storage_endpoint` | The endpoint for the Azure Blob storage account. | Example: ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog include: - '*.my_table_name' params: mode: delta_lake databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com dataset_params: databricks_azure_storage_account_name: myaccount databricks_azure_storage_account_key: ${secrets:azure_storage_account_key} databricks_azure_storage_endpoint: myaccount.blob.core.windows.net ``` #### Google Storage (GCS)[​](#google-storage-gcs "Direct link to Google Storage (GCS)") | Dataset Parameter Name | Definition | | ----------------------------------- | ------------------------------------------------------------ | | `databricks_google_service_account` | Filesystem path to the Google service account JSON key file. | Example: ``` catalogs: - from: databricks:my_uc_catalog name: uc_catalog include: - '*.my_table_name' params: mode: delta_lake databricks_endpoint: dbc-a12cd3e4-56f7.cloud.databricks.com dataset_params: databricks_google_service_account: /path/to/service-account.json ``` ## Limitations[​](#limitations "Direct link to Limitations") * Databricks catalog connector (mode: delta\_lake) does not support reading Delta tables with the `V2Checkpoint` feature enabled. To use the Databricks catalog connector (mode: delta\_lake) with such tables, drop the `V2Checkpoint` feature by executing the following command: ``` ALTER TABLE DROP FEATURE v2Checkpoint [TRUNCATE HISTORY]; ``` For more details on dropping Delta table features, refer to the official documentation: [Drop Delta table features](https://docs.databricks.com/en/delta/drop-feature.html#:~:text=Databricks%20provides%20limited%20support%20for,data%20files%20backing%20the%20table.) * The Databricks Catalog Connector (`mode: spark_connect`) does not yet support streaming query results from Spark. Memory Considerations When using the Databricks (mode: delta\_lake) Catalog connector without acceleration, data is loaded into memory during query execution. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. --- # Glue Catalog Connector Connect to an [AWS Glue Data Catalog](https://docs.aws.amazon.com/glue/latest/dg/start-data-catalog.html) as a catalog provider for federated SQL query. ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: glue name: my_glue_catalog # tables from this catalog will be available in the "my_glue_catalog" catalog in Spice include: - '*.my_table_name' # include only the "my_table_name" tables params: glue_region: us-east-1 # Region of the AWS Glue Data Catalog. glue_key: ${secrets:aws_access_key_id} # Optional. Access key ID for the AWS Glue Data Catalog. glue_secret: ${secrets:aws_secret_access_key} # Optional. Secret access key for the AWS Glue Data Catalog. ``` ### `from`[​](#from "Direct link to from") The `from` field is used to specify the catalog provider. For Glue, you need only specify `glue`. The catalog is unique for each AWS account and AWS region. ### `name`[​](#name "Direct link to name") The `name` field is used to specify the name of the catalog in Spice. Tables from the AWS Glue Data Catalog will be available in the schema with this name in Spice. The schema hierarchy of the external catalog is preserved in Spice. ### `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. ### `params`[​](#params "Direct link to params") The following parameters are supported for configuring the connection to the Glue Data Catalog: | Parameter Name | Definition | | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | `glue_region` | The AWS region for the Glue Data Catalog. E.g. `us-west-2`. | | `glue_key` | Access key (e.g. AWS\_ACCESS\_KEY\_ID for AWS). If not provided, credentials will be loaded from environment variables or IAM roles. | | `glue_secret` | Secret key (e.g. AWS\_SECRET\_ACCESS\_KEY for AWS). If not provided, credentials will be loaded from environment variables or IAM roles. | | `glue_session_token` | Session token (e.g. AWS\_SESSION\_TOKEN for AWS) for temporary credentials | ## Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. These credentials will be used to connect to the S3 bucket as well as the Glue catalog. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3/Glue permissions (e.g., `s3:ListBucket`, `glue:GetTable`) to access the tables. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ### Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access Iceberg tables in S3/Glue: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::company-bucketname-datasets" }, { "Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::company-bucketname-datasets/*" }, { "Effect": "Allow", "Action": [ "glue:GetCatalog", "glue:GetDatabases", "glue:GetDatabase", "glue:GetTable", "glue:GetTables" ], "Resource": "*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------- | -------------------------------------------------------------- | | `s3:ListBucket` | Required. Allows scanning all objects from the bucket | | `s3:GetObject` | Required. Allows fetching objects | | `glue:GetCatalog` | Required. Retrieve metadata about the specified catalog. | | `glue:GetDatabases` | Required. List the databases available in the current catalog. | | `glue:GetDatabase` | Required. Retrieve metadata about the specified database. | | `glue:GetTable` | Required. Retrieve metadata about the specified table. | | `glue:GetTables` | Required. List the tables available in the current database. | ## Limitations[​](#limitations "Direct link to Limitations") warning * This catalog connector is limited to tables that use the S3 data source. Kinesis and Kafka data sources are not currently supported. * This catalog connector is currently limited to Iceberg tables, tables with parquet or CSV data format only. ## Cookbook[​](#cookbook "Direct link to Cookbook") There is a [cookbook recipe](https://github.com/spiceai/cookbook/tree/trunk/catalogs/glue) to configure an AWS Glue Data Connector in Spice. --- # Iceberg Catalog Connector The Iceberg Catalog Connector helps connect Spice to an [Apache Iceberg](https://iceberg.apache.org/) catalog, making Iceberg tables and schemas available for federated SQL queries. Every Iceberg table must be registered in a catalog, which manages table metadata and access. Using a catalog connector is the recommended approach for working with multiple Iceberg datasets, as it helps organize tables and schemas efficiently and mirrors the structure of the source catalog provider. For connecting to a single Iceberg table, see the [Iceberg Data Connector documentation](/docs/v1.10/components/data-connectors/iceberg). For AWS Glue-based catalogs, see the [AWS Glue Catalog Connector documentation](/docs/v1.10/components/catalogs/glue). Iceberg catalogs can be of several types: * **Iceberg REST Catalog**: The most common and recommended approach. REST Catalogs expose Iceberg table metadata over HTTP(S) endpoints and are compatible with most managed Iceberg services and cloud providers. * **AWS Glue Catalog**: Integrates with AWS Glue as a catalog provider, supporting Iceberg tables stored in S3. This is the preferred method for AWS environments. * **Hadoop-style Catalogs**: Use file-based storage (e.g., `file://`, `s3://`, `s3a://`) to manage table metadata. This approach is typically used for local development or legacy deployments. Hadoop-style Catalogs For production and cloud environments, REST and AWS Glue catalogs are recommended. Hadoop-style catalogs are supported but less common and not recommended for most new deployments. ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_catalog name: ice # tables from this catalog will be available in the "ice" catalog in Spice include: - '*.my_table_name' # include only the "my_table_name" tables params: iceberg_token: ${secrets:iceberg_token} # Optional. Bearer token value to use for Authorization header. iceberg_oauth2_credential: ${secrets:client_id}:${secrets:client_secret} # Optional. Credential to use for OAuth2 client credential flow when initializing the catalog. Separated by a colon as :. iceberg_oauth2_scope: catalog # Optional. Scope to use for OAuth2 client credential flow when initializing the catalog (default: catalog). iceberg_oauth2_server_url: https://iceberg-catalog-host.com/oauth2/token # Optional. URL of the OAuth2 server tokens endpoint for the client credential flow. iceberg_s3_endpoint: http://localhost:9000 # Optional. S3-compatible endpoint where the Iceberg tables are stored. iceberg_s3_region: us-west-2 # Optional. Region of the S3-compatible endpoint. iceberg_s3_access_key_id: ${secrets:aws_access_key_id} # Optional. Access key ID for the S3-compatible endpoint. iceberg_s3_secret_access_key: ${secrets:aws_secret_access_key} # Optional. Secret access key for the S3-compatible endpoint. iceberg_s3_session_token: ${secrets:aws_session_token} # Optional. Session token for the S3-compatible endpoint. iceberg_s3_role_arn: arn:aws:iam::123456789012:role/my-role # Optional. ARN of the IAM role to assume when accessing the S3-compatible endpoint. iceberg_s3_role_session_name: my-session # Optional. Session name to use when assuming the IAM role. iceberg_s3_connect_timeout: 60 # Optional. Connection timeout for the S3-compatible endpoint (default: 60). # AWS Glue Catalog (see also the [AWS Glue Catalog Connector documentation](./glue)) - from: iceberg:https://glue.us-east-1.amazonaws.com/iceberg/v1/catalogs/123456789012/namespaces name: glue params: iceberg_sigv4_enabled: true ``` ## `from`[​](#from "Direct link to from") The `from` field specifies the catalog provider. For Iceberg, use `iceberg:`, where `namespace_path` is the URL to the Iceberg namespace in the catalog provider. The format is `http[s]:///v1/{prefix}/namespaces/`. For AWS Glue catalogs, the URL format is `https://glue..amazonaws.com/iceberg/v1/catalogs//namespaces`, where `` is the AWS account ID. While possible to connect to Iceberg tables hosted by Glue using this generic connector, it is recommended to instead use the [AWS Glue Catalog Connector](/docs/v1.10/components/catalogs/glue) for connecting to Iceberg tables managed by Glue for a better experience. The selected namespace must have sub-namespaces where the tables are stored. Example: With this Iceberg catalog structure: ``` . ├── blockchain │ └── eth │ ├── blocks │ └── transactions ├── spice │ ├── tpch │ │ ├── orders │ │ └── customers │ ├── info │ └── extra │ └── tpch_orders_metadata └── unity └── very └── nested └── namespace └── foobar ``` A valid `from` value would be `iceberg:https://iceberg-catalog-host.com/v1/namespaces/spice`, and would load the following tables: * `.tpch.orders` * `.tpch.customers` * `.extra.tpch_orders_metadata` For loading a multi-part namespace, separate the namespace parts with the `%1F` character. For example, `/v1/namespaces/unity%1Fvery%1Fnested` would load the `foobar` table from the `unity/very/nested/namespace` namespace as `.namespace.foobar`. To connect to a single Iceberg table directly, see the [Iceberg Data Connector documentation](/docs/v1.10/components/data-connectors/iceberg). ## `name`[​](#name "Direct link to name") The `name` field is used to specify the name of the catalog in Spice. Tables from the Iceberg catalog will be available in the schema with this name in Spice. The schema hierarchy of the external catalog is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. ## `params`[​](#params "Direct link to params") The following parameters are supported for configuring the connection to the Iceberg catalog, file, or S3 storage: | Parameter Name | Description | | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `iceberg_token` | Bearer token value to use for Authorization header. | | `iceberg_oauth2_credential` | Credential to use for OAuth2 client credential flow when initializing the catalog. Separated by a colon as `:`. | | `iceberg_oauth2_scope` | The scope to use for OAuth2 token endpoint (default: `catalog`). | | `iceberg_oauth2_server_url` | URL of the OAuth2 server tokens endpoint. | | `iceberg_sigv4_enabled` | Enable SigV4 authentication for the catalog (for connecting to AWS Glue). | | `iceberg_signing_region` | The region to use when signing the request for SigV4. Defaults to the region in the catalog URL if available. | | `iceberg_signing_name` | The name to use when signing the request for SigV4. Default: `glue`. | | `iceberg_s3_endpoint` | Configure an alternative endpoint for the S3 service. This can be any S3-compatible object storage service (e.g., Minio, R2). | | `iceberg_s3_access_key_id` | The AWS access key ID to use for S3 storage. If not provided, credentials will be loaded from environment variables or IAM roles. | | `iceberg_s3_secret_access_key` | The AWS secret access key to use for S3 storage. If not provided, credentials will be loaded from environment variables or IAM roles. | | `iceberg_s3_session_token` | Configure the static session token used for S3 storage. | | `iceberg_s3_region` | The AWS S3 region to use. | | `iceberg_s3_role_session_name` | An optional identifier for the assumed role session for auditing purposes. | | `iceberg_s3_role_arn` | The Amazon Resource Name (ARN) of the role to assume. If provided instead of iceberg\_s3\_access\_key\_id and iceberg\_s3\_secret\_access\_key, temporary credentials will be fetched by assuming this role. | | `iceberg_s3_connect_timeout` | Configure socket connection timeout, in seconds (default: `60`). | The Iceberg Catalog Connector supports both REST and Hadoop-style Catalogs. In both cases, the warehouse path (for example, s3://bucket/warehouse/) specifies the object store location where tables are physically stored. With a Hadoop-style Catalog, the metadata is resolved directly from the filesystem using the Hadoop convention (by reading a `version-hint.txt`), rather than through a catalog service. The warehouse path itself does not change between using a REST Catalog and a Hadoop-style Catalog — only how the metadata is discovered and managed differs. The warehouse path is discovered automatically from the catalog service, but must be explicitly specified when using Hadoop-style Iceberg tables. Hadoop-style catalogs are most commonly used for local development or legacy deployments. Example using Hadoop Catalog with a local warehouse: ``` catalogs: - from: iceberg:file:///tmp/hadoop_warehouse/ name: local_hadoop ``` Example using Hadoop Catalog with S3: ``` catalogs: - from: iceberg:s3a://my-bucket/hadoop_warehouse/ name: s3_hadoop ``` ### AWS Authentication[​](#aws-authentication "Direct link to AWS Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. These credentials will be used to connect to the S3 bucket as well as the Glue catalog (if configured). 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3/Glue permissions (e.g., `s3:ListBucket`, `s3:GetObject`) to access the tables. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access Iceberg tables in S3/Glue: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::company-bucketname-datasets" }, { "Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::company-bucketname-datasets/*" }, { "Effect": "Allow", "Action": [ "glue:GetCatalog", "glue:GetDatabases", "glue:GetDatabase", "glue:GetTable", "glue:GetTables" ], Resource: "*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------- | -------------------------------------------------------------- | | `s3:ListBucket` | Required. Allows scanning all objects from the bucket | | `s3:GetObject` | Required. Allows fetching objects | | `glue:GetCatalog` | Required. Retrieve metadata about the specified catalog. | | `glue:GetDatabases` | Required. List the databases available in the current catalog. | | `glue:GetDatabase` | Required. Retrieve metadata about the specified database. | | `glue:GetTable` | Required. Retrieve metadata about the specified table. | | `glue:GetTables` | Required. List the tables available in the current database. | ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Iceberg as a catalog connector in Spice. [Iceberg Catalog Connector](https://github.com/spiceai/cookbook/tree/trunk/catalogs/iceberg#readme) --- # Spice.ai Catalog Connector Query datasets hosted in the [Spice.ai Cloud Platform](https://spice.ai). Discover available public datasets on [Spicerack](https://spicerack.org). ## Configuration[​](#configuration "Direct link to Configuration") Create a [Spice.ai Cloud Platform](https://spice.ai) account and login with the CLI using `spice login`. Example: ``` catalogs: - from: spice.ai:demo-org/tpch # Load tables from the `demo-org` organization's `tpch` app name: marketplace # Tables will be available in the "marketplace" catalog params: spiceai_api_key: ${secrets:SPICEAI_API_KEY} # Spice.ai API key to login to the organization and app include: - "tpch.part*" # include only the tables from the "tpch" schema and that start with "part" - "tpch.supplier" # also include the "supplier" table ``` This configuration would load the following tables: ``` sql> show tables; +---------------+--------------+----------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------+----------------+------------+ | marketplace | tpch | part | BASE TABLE | | marketplace | tpch | partsupp | BASE TABLE | | marketplace | tpch | supplier | BASE TABLE | +---------------+--------------+----------------+------------+ ``` ## `from`[​](#from "Direct link to from") The `from` field specifies which organization and application to load tables from. The format is: ``` spice.ai/[organization]/[application][/catalog_name] ``` * `organization`: The Spice.ai organization that owns the application * `application`: The specific application to load tables from * `catalog_name`: (optional): A specific catalog within the application. If not specified, the application's default catalog is used For example: 1. Load the default catalog from the "tpch" application in the "demo-org" organization ``` catalogs: - from: spice.ai/demo-org/tpch name: marketplace ``` This will create tables like: ``` SHOW TABLES; +---------------+--------------+----------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------+----------------+------------+ | marketplace | tpch | region | BASE TABLE | | marketplace | tpch | orders | BASE TABLE | | marketplace | tpch | part | BASE TABLE | | marketplace | tpch | supplier | BASE TABLE | | marketplace | tpch | customer | BASE TABLE | | marketplace | tpch | partsupp | BASE TABLE | | marketplace | tpch | lineitem | BASE TABLE | | marketplace | tpch | nation | BASE TABLE | +---------------+--------------+----------------+------------+ ``` 2. Load a specific catalog named "custom\_catalog" from an application ``` catalogs: - from: spice.ai/demo-org/tpch/custom_catalog name: marketplace ``` If the remote application has tables in `custom_catalog` like: ``` SHOW TABLES; +------------------+--------------+----------------+------------+ | table_catalog | table_schema | table_name | table_type | +------------------+--------------+----------------+------------+ | custom_catalog | ice_schema1 | table1 | BASE TABLE | | custom_catalog | ice_schema1 | table2 | BASE TABLE | | custom_catalog | ice_schema2 | table1 | BASE TABLE | +------------------+--------------+----------------+------------+ ``` They will be available locally as: ``` SHOW TABLES; +------------------+--------------+----------------+------------+ | table_catalog | table_schema | table_name | table_type | +------------------+--------------+----------------+------------+ | marketplace | ice_schema1 | table1 | BASE TABLE | | marketplace | ice_schema1 | table2 | BASE TABLE | | marketplace | ice_schema2 | table1 | BASE TABLE | +------------------+--------------+----------------+------------+ ``` ## `name`[​](#name "Direct link to name") The name field defines what catalog name the tables will be available under in the local Spice instance. For example, with the following configuration: ``` from: spice.ai/demo-org/tpch name: marketplace ``` Then tables that exist in the remote application as: ``` spice # Default catalog |- schema1 |- table1 |- table2 ``` Will be available locally as: ``` marketplace |- schema1 |- table1 |- table2 ``` Queries are run against the `marketplace` catalog, like `SELECT * FROM marketplace.schema1.table1`. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables: * `schema_name.*` - Include all tables from a specific schema * `*.table_name` - Include all tables with a specific name from any schema * `schema_name.table_name` - Include a specific table from a schema * `schema_name.table_prefix*` - Include all tables in a schema that start with a prefix Multiple include patterns can be specified and are OR'ed together. For example: ``` include: - "tpch.part*" # Include all tables from the "tpch" schema that start with "part" - "tpch.supplier" # Include the "supplier" table from the "tpch" schema ``` ## `params`[​](#params "Direct link to params") The following parameters are supported for configuring the connection to the Spice Cloud catalog/tables: | Parameter Name | Definition | | ----------------- | ------------------------------------------------------------------------------------------------------------ | | `spiceai_api_key` | Authorization API key from the Spice.ai Cloud Platform, used to login to the specified organization and app. | ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Spice Cloud as a catalog connector in Spice. [Spice Cloud Catalog Connector](https://github.com/spiceai/cookbook/tree/trunk/catalogs/spiceai#readme) --- # Unity Catalog Catalog Connector Connect to a [Unity Catalog](https://www.unitycatalog.io/) as a catalog provider for federated SQL query against [Delta Lake](https://delta.io/) tables. ## Configuration[​](#configuration "Direct link to Configuration") ``` catalogs: - from: unity_catalog:https://my_unity_catalog_host.com/api/2.1/unity-catalog/catalogs/my_catalog name: uc include: - "*.my_table" dataset_params: # delta_lake S3 parameters unity_catalog_aws_region: us-west-2 unity_catalog_aws_access_key_id: ${secrets:aws_access_key_id} unity_catalog_aws_secret_access_key: ${secrets:aws_secret_access_key} unity_catalog_aws_endpoint: s3.us-west-2.amazonaws.com ``` ## `from`[​](#from "Direct link to from") The `from` field is used to specify the catalog provider. For Unity Catalog, use `unity_catalog:`. The `catalog_path` is the URL to the [`getCatalog`](https://github.com/unitycatalog/unitycatalog/blob/main/api/Apis/CatalogsApi) endpoint of the Unity Catalog API. It should be formatted as `https:///api/2.1/unity-catalog/catalogs/`. ## `name`[​](#name "Direct link to name") The `name` field is used to specify the name of the catalog in Spice. The schema hierarchy of the external catalog is preserved in Spice. ## `include`[​](#include "Direct link to include") Use the `include` field to specify which tables to include from the catalog. The `include` field supports glob patterns to match multiple tables. For example, `*.my_table_name` would include all tables with the name `my_table_name` in the catalog from any schema. Multiple `include` patterns are OR'ed together and can be specified to include multiple tables. ## `params`[​](#params "Direct link to params") The `params` field is used to configure the connection to the Unity Catalog. The following parameters are supported: * `unity_catalog_token`: The [personal access token](https://docs.unitycatalog.io/server/auth/#use-admin-token-to-verify-admin-user-is-in-local-database) used to authenticate against the Unity Catalog API. ## `dataset_params`[​](#dataset_params "Direct link to dataset_params") The `dataset_params` field is used to configure the dataset-specific parameters for the catalog. ### Unity catalog object store parameters[​](#unity-catalog-object-store-parameters "Direct link to Unity catalog object store parameters") #### AWS S3[​](#aws-s3 "Direct link to AWS S3") * `unity_catalog_aws_region`: The AWS region for the S3 object store. E.g. `us-west-2`. * `unity_catalog_aws_access_key_id`: The access key ID for the S3 object store. * `unity_catalog_aws_secret_access_key`: The secret access key for the S3 object store. * `unity_catalog_aws_endpoint`: The endpoint for the S3 object store. E.g. `s3.us-west-2.amazonaws.com`. #### Azure Blob[​](#azure-blob "Direct link to Azure Blob") Note One of the following auth values must be provided for Azure Blob: * `unity_catalog_azure_storage_account_key`, * `unity_catalog_azure_storage_client_id` and `unity_catalog_azure_storage_client_secret`, or * `unity_catalog_azure_storage_sas_key`. - `unity_catalog_azure_storage_account_name`: The Azure Storage account name. - `unity_catalog_azure_storage_account_key`: The Azure Storage master key for accessing the storage account. - `unity_catalog_azure_storage_client_id`: The service principal client id for accessing the storage account. - `unity_catalog_azure_storage_client_secret`: The service principal client secret for accessing the storage account. - `unity_catalog_azure_storage_sas_key`: The shared access signature key for accessing the storage account. - `unity_catalog_azure_storage_endpoint`: The endpoint for the Azure Blob storage account. #### Google Storage (GCS)[​](#google-storage-gcs "Direct link to Google Storage (GCS)") * `unity_catalog_google_service_account`: Filesystem path to the Google service account JSON key file. ## Limitations[​](#limitations "Direct link to Limitations") * Unity Catalog does not support reading Delta tables with the `V2Checkpoint` feature enabled. To use the Unity Catalog connector with such tables, drop the `V2Checkpoint` feature by executing the following command: ``` ALTER TABLE DROP FEATURE v2Checkpoint [TRUNCATE HISTORY]; ``` For more details on dropping Delta table features, refer to the official documentation: [Drop Delta table features](https://docs.delta.io/latest/delta-drop-feature.html) --- # Data Accelerators Data sourced by Data Connectors can be locally materialized and accelerated using a Data Accelerator. A Data Accelerator will query/fetch data from a connected data source and store/update it locally in an embedded acceleration engine, such as DuckDB or SQLite. To set data refresh behavior, such as refreshing data on an interval see [Data Refresh](/docs/features/data-acceleration/data-refresh). Dataset acceleration is enabled by setting the acceleration configuration. E.g. ``` datasets: - name: accelerated_dataset acceleration: enabled: true ``` For the complete reference specification, see [datasets](/docs/reference/spicepod/datasets). By default, datasets will be locally materialized using in-memory Arrow records. A choice of DuckDB, SQLite, or PostgreSQL engines can be used to materialize data, in-memory, on disk, or in attached databases. Supported Data Accelerators include: | Name | Description | Status | Engine Modes | | ---------- | ------------------------------------------------------------------ | -------------------- | ------------------------------------ | | `arrow` | In-Memory Arrow Records | Stable | `memory` | | `cayenne` | [Cayenne](/docs/components/data-accelerators/cayenne) | Alpha (v1.9.0-rc.1+) | `file`, `file_create`, `file_update` | | `duckdb` | Embedded [DuckDB](/docs/components/data-accelerators/duckdb) | Stable | `memory`, `file`, `file_create` | | `postgres` | Attached [PostgreSQL](/docs/components/data-accelerators/postgres) | Release Candidate | N/A | | `sqlite` | Embedded [SQLite](/docs/components/data-accelerators/sqlite) | Release Candidate | `memory`, `file`, `file_create` | ## Choosing an Accelerator[​](#choosing-an-accelerator "Direct link to Choosing an Accelerator") Select the appropriate accelerator based on dataset size, query patterns, and resource constraints: | Use Case | Recommended Accelerator | Rationale | | --------------------------------------------------- | ----------------------- | ------------------------------------------------------- | | Small datasets (under 1 GB), maximum speed | `arrow` | In-memory storage provides lowest latency | | Medium datasets (1-100 GB), complex SQL | `duckdb` | Mature SQL support with memory management | | Large datasets (100 GB - 1+ TB), scalable analytics | `cayenne` | Vortex columnar format scales beyond single-file limits | | Point lookups on large datasets | `cayenne` | Vortex provides 100x faster random access vs Parquet | | Simple queries, low resource usage | `sqlite` | Lightweight, minimal overhead | | External database integration | `postgres` | Leverage existing PostgreSQL infrastructure | ### Spice Cayenne vs DuckDB[​](#spice-cayenne-vs-duckdb "Direct link to Spice Cayenne vs DuckDB") Both [Spice Cayenne](/docs/components/data-accelerators/cayenne) and [DuckDB](/docs/components/data-accelerators/duckdb) support file-based acceleration, but differ in architecture and performance characteristics: **Choose Spice Cayenne when:** * Datasets exceed \~1 TB * Multi-file data ingestion is required (e.g., partitioned S3 data) * Lower memory overhead is preferred * Workloads benefit from Vortex's [10-20x faster scans](https://bench.vortex.dev) * Point lookups and random access patterns are common ([100x faster than Parquet](https://bench.vortex.dev)) **Choose DuckDB when:** * Datasets are under \~1 TB * Complex SQL features are required (window functions, CTEs) * Existing DuckDB tooling integration is beneficial * Explicit index control is required ## Data Types[​](#data-types "Direct link to Data Types") Data Accelerators may not support all possible Apache Arrow data types. For complete compatibility, see [specifications](/docs/reference/datatypes/accelerators). Memory Considerations When accelerating a dataset using `mode: memory` (the default), some or all of the dataset is loaded into memory. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. In-memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/components/data-accelerators/duckdb) and [`sqlite`](/docs/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. ## Data Accelerator Docs[​](#data-accelerator-docs "Direct link to Data Accelerator Docs") ## [📄️Cayenne Data Accelerator](/docs/v1.10/components/data-accelerators/cayenne) [Cayenne Data Accelerator (Vortex) Documentation](/docs/v1.10/components/data-accelerators/cayenne) ## [📄️In-Memory Arrow Data Accelerator](/docs/v1.10/components/data-accelerators/arrow) [In-Memory Arrow Data Accelerator Documentation](/docs/v1.10/components/data-accelerators/arrow) ## [📄️DuckDB Data Accelerator](/docs/v1.10/components/data-accelerators/duckdb) [DuckDB Data Accelerator Documentation](/docs/v1.10/components/data-accelerators/duckdb) ## [📄️SQLite Data Accelerator](/docs/v1.10/components/data-accelerators/sqlite) [SQLite Data Accelerator Documentation](/docs/v1.10/components/data-accelerators/sqlite) ## [📄️PostgreSQL Data Accelerator](/docs/v1.10/components/data-accelerators/postgres) [PostgreSQL Data Accelerator Documentation](/docs/v1.10/components/data-accelerators/postgres) ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") * [Managing Memory Usage](/docs/reference/memory) - Memory configuration reference * [Data Refresh](/docs/features/data-acceleration/data-refresh) - Refresh mode configuration * [Indexes](/docs/features/data-acceleration/indexes) - Index configuration for DuckDB and SQLite --- # In-Memory Arrow Data Accelerator The In-Memory Arrow Data Accelerator is the default data accelerator in Spice. It uses Apache Arrow to store data in-memory for fast access and query performance. ## Configuration[​](#configuration "Direct link to Configuration") To use the In-Memory Arrow Data Accelerator, no additional configuration is required beyond enabling acceleration. Example: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: enabled: true ``` However Arrow can be specified explicitly using `arrow` as the `engine` for acceleration. ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: enabled: true engine: arrow ``` ## Hash Index[​](#hash-index "Direct link to Hash Index") Experimental Hash index is an experimental feature available in Spice v1.11.0-rc.2 and later. The In-Memory Arrow Data Accelerator supports an optional hash index for O(1) point lookups on primary key columns. To enable, set `hash_index: enabled` in the dataset params: ``` datasets: - from: s3://bucket/orders.parquet name: orders acceleration: engine: arrow primary_key: order_id params: hash_index: enabled ``` See the configuration example above for usage details. ## Limitations[​](#limitations "Direct link to Limitations") * The In-Memory Arrow Data Accelerator does not support persistent storage. Data is stored in-memory and will be lost when the Spice runtime is stopped. * The In-Memory Arrow Data Accelerator does not support `Decimal256` (76 digits), as it exceeds Arrow's maximum Decimal width of 38 digits. * The In-Memory Arrow Data Accelerator does not support traditional [indexes](/docs/features/data-acceleration/indexes), but does support hash indexes (experimental) for point lookups. * The In-Memory Arrow Data Accelerator only supports primary-key [constraints](/docs/features/data-acceleration/constraints), not `unique` constraints. * With Arrow acceleration, mathematical operations like `value1 / value2` are treated as integer division if the values are integers. For example, `1 / 2` will result in 0 instead of the expected 0.5. Use casting to FLOAT to ensure conversion to a floating-point value: `CAST(1 AS FLOAT) / CAST(2 AS FLOAT)` (or `CAST(1 AS FLOAT) / 2`). Memory Considerations When accelerating a dataset using the In-Memory Arrow Data Accelerator, some or all of the dataset is loaded into memory. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. In-memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/v1.10/components/data-accelerators/duckdb) and [`sqlite`](/docs/v1.10/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure In-Memory Arrow as data accelerator in Spice. [In-Memory Arrow Data Accelerator](https://github.com/spiceai/cookbook/tree/trunk/arrow#readme) --- # Cayenne Data Accelerator Alpha The Cayenne Data Accelerator is in Alpha. Features and configuration may change. Available in Spice v1.9.0-rc.1 and later. Cayenne is a Spice data acceleration engine designed for high-performance, scalable query on large-scale datasets. Built on [Vortex](https://github.com/vortex-data/vortex), a next-generation columnar file format, Cayenne combines columnar storage with in-process metadata management to provide fast query performance to scale to datasets beyond 1TB. ## Why Vortex?[​](#why-vortex "Direct link to Why Vortex?") Cayenne uses Vortex as its storage format, providing significant performance advantages: * **100x faster random access reads** compared to modern Apache Parquet * **10-20x faster scans** for analytical queries * **5x faster writes** with similar compression ratios * **Zero-copy compatibility** with Apache Arrow for efficient data processing * **Extensible architecture** with pluggable encoding, compression, and layout strategies Vortex is a Linux Foundation (LF AI & Data) project under Apache-2.0 license with neutral governance. While [DuckDB](/docs/v1.10/components/data-accelerators/duckdb) excels for datasets up to approximately 1TB, Spice Cayenne with Vortex is designed to scale beyond these limits. For detailed Vortex performance benchmarks, visit [bench.vortex.dev](https://bench.vortex.dev). ## Configuration[​](#configuration "Direct link to Configuration") To use Cayenne as the data accelerator, specify `cayenne` as the `engine` for acceleration. Cayenne supports `mode: file`, `mode: file_create`, and `mode: file_update` and stores data on disk. ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: cayenne mode: file ``` ## Features[​](#features "Direct link to Features") ### High-Performance Columnar Storage[​](#high-performance-columnar-storage "Direct link to High-Performance Columnar Storage") Cayenne uses Vortex's advanced columnar format, which provides: * **Efficient Compression**: Cascading compression with nested encoding schemes including RLE, dictionary encoding, FastLanes, FSST, and ALP * **Rich Statistics**: Lazy-loaded summary statistics for query optimization * **Extensible Encodings**: Pluggable physical layouts optimized for different data patterns * **Wide Table Support**: Efficient handling of tables with many columns through zero-copy metadata access ## Limitations[​](#limitations "Direct link to Limitations") Consider the following limitations when using Cayenne acceleration: * **Alpha Status**: Cayenne is in active development. Configuration options may change between releases. * **File Mode Only**: Cayenne only supports `mode: file`, `mode: file_create`, and `mode: file_update` and does not support in-memory (`mode: memory`) acceleration. * **No `on_conflict` Support**: Cayenne does not yet support the [`on_conflict`](/docs/v1.10/reference/spicepod/datasets#accelerationon_conflict) configuration for handling duplicate keys during data refresh. * **Data Cleanup Requires `retention_sql`**: Data deletion and cleanup operations require configuring [`retention_sql`](/docs/v1.10/reference/spicepod/datasets#accelerationretention_sql) to define retention policies. Manual `DELETE` statements can also be executed directly. * **No Snapshot Support**: Cayenne does not yet support [acceleration snapshots](/docs/v1.10/features/data-acceleration/snapshots) for bootstrapping from object storage. * **Data Types**: Some advanced data types may have limited support. Test your specific schema requirements. * **Index Support**: Index capabilities are still being developed. Check release notes for the latest supported features. Alpha Software As an Alpha feature, Cayenne should be thoroughly tested in development environments before production deployment. Monitor release notes for updates, breaking changes, and new capabilities. ## Resource Considerations[​](#resource-considerations "Direct link to Resource Considerations") Resource requirements for Cayenne depend on dataset size, query patterns, and metastore configuration. ### Memory[​](#memory "Direct link to Memory") Cayenne manages memory efficiently through columnar storage and selective caching. Allocate sufficient memory based on: * Dataset size and schema complexity * Query concurrency requirements * Caching configuration ### Storage[​](#storage "Direct link to Storage") Cayenne stores data in a columnar format optimized for analytical queries. Ensure adequate disk space for: * **Acceleration data**: Compressed Vortex files (typically 30-50% of raw data size with btrblocks) * **Metadata**: SQLite database for catalog and statistics (\~10 MB per 1000 files) * **Temporary files**: Query spill files during complex operations ### CPU[​](#cpu "Direct link to CPU") Query performance scales with available CPU cores. Vortex's columnar format supports parallel decompression and scanning across multiple threads. Allocate sufficient CPU for: * Query execution parallelism * Data refresh and compression operations * Concurrent query workloads ## Limitations[​](#limitations-1 "Direct link to Limitations") Consider the following limitations when using Spice Cayenne acceleration: * **Alpha Status**: Spice Cayenne is in active development. Configuration options may change between releases. * **File Mode Only**: Spice Cayenne only supports `mode: file` and does not support in-memory (`mode: memory`) acceleration. * **No Snapshot Support**: Spice Cayenne does not yet support [acceleration snapshots](/docs/v1.10/features/data-acceleration/snapshots) for bootstrapping from object storage. * **S3 Express Only**: Standard S3 buckets are not supported for remote storage. Only S3 Express One Zone directory buckets are supported. * **Unsupported Data Types**: `Interval`, `Duration`, and `FixedSizeBinary` types require `unsupported_type_action` configuration. * **No Traditional Indexes**: Spice Cayenne does not support explicit index creation via the `indexes` configuration. Vortex's segment statistics and fast random access encodings provide equivalent or better performance for most point lookup workloads. * **No MVCC**: Multi-version concurrency control is not yet implemented. Snapshots and time-travel queries are planned for future releases. * **No File Compaction**: Automatic file compaction to reclaim space from deleted rows is not yet available. ALPHA SOFTWARE As an Alpha feature, Spice Cayenne should be thoroughly tested in development environments before production deployment. Monitor release notes for updates, breaking changes, and new capabilities. ## Example Spicepod[​](#example-spicepod "Direct link to Example Spicepod") Complete example configuration using Cayenne: ``` version: v1 kind: Spicepod name: cayenne-example datasets: - from: s3://my-bucket/data/ name: analytics_data params: file_format: parquet acceleration: engine: cayenne enabled: true refresh_mode: full refresh_check_interval: 1h ``` ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") **Spice Documentation:** * [Managing Memory Usage](/docs/v1.10/reference/memory) - Memory configuration reference * [Data Acceleration](/docs/v1.10/features/data-acceleration) - Data acceleration overview **External References:** * [Apache DataFusion](https://datafusion.apache.org/) - Query execution engine * [DataFusion Configuration](https://datafusion.apache.org/user-guide/configs.html) - DataFusion settings and tuning * [Vortex Project](https://github.com/vortex-data/vortex) - Columnar file format * [Vortex Benchmarks](https://bench.vortex.dev/) - Performance benchmarks * [FSST Paper](https://www.vldb.org/pvldb/vol13/p2649-boncz.pdf) - Fast Static Symbol Table compression * [FastLanes Paper](https://www.vldb.org/pvldb/vol16/p2132-afroozeh.pdf) - High-performance integer encoding * [ALP Paper](https://ir.cwi.nl/pub/33334/33334.pdf) - Adaptive floating-point compression * [BtrBlocks Paper](https://www.cs.cit.tum.de/fileadmin/w00cfj/dis/papers/btrblocks.pdf) - Compression algorithm * [AWS S3 Express One Zone](https://aws.amazon.com/s3/storage-classes/express-one-zone/) - Low-latency object storage --- # DuckDB Data Accelerator The DuckDB Data Accelerator helps improve query performance by using [DuckDB](https://duckdb.org/), an embedded analytical database engine optimized for efficient data processing. It supports in-memory and file-based operation modes, enabling workloads that exceed available memory and optionally providing persistent storage for datasets. To enable DuckDB acceleration, set the dataset's `acceleration.engine` to `duckdb`: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: duckdb mode: file ``` ## Modes[​](#modes "Direct link to Modes") ### Memory Mode[​](#memory-mode "Direct link to Memory Mode") By default, DuckDB acceleration uses `mode: memory`, loading datasets into memory. ### File Mode[​](#file-mode "Direct link to File Mode") When using `mode: file`, datasets are stored by default in a DuckDB file on disk in the `.spice/data` directory relative to the spicepod.yaml. Specify the `duckdb_file` parameter to store the DuckDB file in a different location. For datasets intended to be joined, set the same `duckdb_file` path for all related datasets. ## Configuration Parameters[​](#configuration-parameters "Direct link to Configuration Parameters") DuckDB acceleration supports the following optional parameters under `acceleration.params`: * `duckdb_file` (string, default:`.spice/data/accelerated_duckdb.db`): Path to the DuckDB database file. Applies if `mode` is set to `file`. If the file does not exist, Spice creates it automatically. * `duckdb_data_dir` (string, default:`.spice/data/`): Path to the directory the DuckDB database file(s) will be placed in. This is useful when using the `partition_by` acceleration parameter. If both `duckdb_data_dir` and `duckdb_file` are specified, `duckdb_file` will be used and `duckdb_data_dir` will be ignored. * `duckdb_memory_limit` (string, default: none): Limits DuckDB's memory usage for instance. Acceptable units are KB, MB, GB, TB (decimal: 1000^i) or KiB, MiB, GiB, TiB (binary: 1024^i). See [DuckDB memory limit documentation](https://duckdb.org/docs/stable/configuration/overview). * `duckdb_preserve_insertion_order` (boolean, default: `true`): Controls whether DuckDB preserves the insertion order of rows in tables. When set to `true`, rows are returned in the order they were inserted. See [DuckDB preserve insertion order documentation](https://duckdb.org/docs/stable/guides/performance/how_to_tune_workloads#the-preserve_insertion_order-option) and [order preservation documentation](https://duckdb.org/docs/stable/sql/dialect/order_preservation). * `connection_pool_size` (integer, default: `10` or the number of datasets sharing the same DuckDB file, whichever is larger): Controls the maximum number of connections to keep open in the connection pool for concurrent query execution. * `on_refresh_recompute_statistics` (string, default: `enabled`, `disabled` when `refresh_mode` is `changes`): Triggers automatic `ANALYZE` execution after data refreshes. This keeps DuckDB optimizer statistics up-to-date for efficient query plans and performance. Set to `disabled` to turn automatic statistics recomputation off. See [DuckDB ANALYZE statement documentation](https://duckdb.org/docs/stable/sql/statements/analyze). * `partition_mode` (string, default: `files`): Controls how partitioned data is stored. Can only be used with `partition_by`. Set to `tables` to store partitions as separate tables within a single DuckDB database, improving resource usage through single shared connection pool for all partitions. Default `files` mode creates separate database files per partition with individual connection pools and generally faster query performance. * `duckdb_partitioned_write_flush_threshold_rows` (integer, default: `122880`): The number of rows buffered per partition before flushing data to acceleration storage. Only applicable when using `partition_mode: tables`. Using a larger value can improve write performance but requires more memory. * `optimizer_duckdb_aggregate_pushdown` (string, default: `disabled`): Enables aggregate pushdown optimization to execute supported aggregate queries directly in DuckDB. Set to `enabled` to push down aggregations for improved query performance on supported functions like `count`, `sum`, `avg`, `min`, and `max`. Requires `query_federation` to be `disabled`. Refer to the [datasets configuration reference](/docs/v1.10/reference/spicepod/datasets#acceleration) for additional supported fields. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: duckdb mode: file params: duckdb_file: /my/chosen/location/duckdb.db duckdb_memory_limit: '2GB' ``` ## Limitations[​](#limitations "Direct link to Limitations") Consider the following limitations when using DuckDB acceleration: * DuckDB does not support [enum and dictionary field types](https://duckdb.org/docs/sql/data_types/overview). * DuckDB's maximum decimal precision is 38 digits. `Decimal256` (76 digits) is unsupported. * Queries using `on_zero_results: use_source` cannot filter binary columns directly (e.g., `WHERE col_blob <> ''`). Instead, cast binary columns to another type (e.g., `WHERE CAST(col_blob AS TEXT) <> ''`). * DuckDB indexes currently do not support spilling to disk. * Hot-reloading dataset configurations while the Spice Runtime is active disables DuckDB query federation until the runtime restarts. ## Resource Considerations[​](#resource-considerations "Direct link to Resource Considerations") Resource requirements depend on workload, dataset size, query complexity, and refresh modes. ### Memory[​](#memory "Direct link to Memory") DuckDB manages memory through streaming execution, intermediate spilling, and buffer management. By default, each DuckDB instance (one per DuckDB file) uses up to 80% of available system memory. To control memory usage, set the `duckdb_memory_limit` parameter: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: duckdb mode: file params: duckdb_file: '/data/shared_duckdb_instance.db' duckdb_memory_limit: '4GB' ``` Note that `duckdb_memory_limit` only limits the DuckDB instance it is set on, not the entire runtime process. Additionally, it does not cover all DuckDB operations, such as some insert operations. Index creation and scans are limited by the `duckdb_memory_limit` so ensure adequate memory is provisioned. Allocate at least 30% more container/machine memory for the runtime process. ### Indexes and Memory[​](#indexes-and-memory "Direct link to Indexes and Memory") DuckDB indexes currently do not support spilling to disk. While index memory usage is registered through the buffer manager, index buffers are not managed by the buffer eviction mechanism. As a result, indexes may consume significant memory, impacting memory-intensive query performance. Indexes are serialized to disk and loaded lazily upon database reopening, ensuring they do not affect database opening performance. Also consider index serialization when allocating disk storage. For more details, see DuckDB's [Indexes and Memory documentation](https://duckdb.org/docs/stable/guides/performance/indexing.html#indexes-and-memory). ### CPU[​](#cpu "Direct link to CPU") Query performance, data load, and refresh operations scale with available CPU resources. Allocate sufficient CPU cores based on query complexity and concurrency. ### Storage[​](#storage "Direct link to Storage") Ensure adequate disk space for temporary files, swap files, WAL files, and intermediate spilling. Monitor disk usage regularly and adjust storage capacity based on dataset growth and query patterns. ## Temporary Directory[​](#temporary-directory "Direct link to Temporary Directory") The Spice runtime supports configuring a temporary directory for query and acceleration operations that spill to disk. By default, this is the directory of the `duckdb_file`. Set the `runtime.query.temp_directory` parameter to specify a custom temporary directory. This can help distribute I/O operations across multiple volumes for improved throughput. For example, setting `runtime.temp_directory` to a high-IOPS volume separate from the DuckDB data file can improve performance for workloads exceeding available memory. Example configuration: ``` runtime: query: temp_directory: /tmp/spice ``` Use this parameter when: * Handling workloads that frequently spill to disk. * Distributing swap and data I/O operations across multiple storage volumes. For more details, refer to the [runtime parameters documentation](/docs/v1.10/reference/spicepod/runtime#runtimequerytemp_directory). For detailed DuckDB limits, see the [DuckDB Memory Management Guide](https://duckdb.org/docs/operations_manual/limits.html). ## Cookbook[​](#cookbook "Direct link to Cookbook") For practical examples, see the [DuckDB Data Accelerator Cookbook Recipe](https://github.com/spiceai/cookbook/tree/trunk/duckdb/accelerator#readme). ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") * [Managing Memory Usage](/docs/v1.10/reference/memory) - Memory configuration reference * [Data Refresh](/docs/v1.10/features/data-acceleration/data-refresh) - Refresh mode configuration --- # PostgreSQL Data Accelerator To use PostgreSQL as Data Accelerator, specify `postgres` as the `engine` for acceleration. ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: postgres ``` ## Configuration[​](#configuration "Direct link to Configuration") The connection to PostgreSQL can be configured by providing the following `params`: * `pg_host`: The hostname of the PostgreSQL server. * `pg_port`: The port of the PostgreSQL server. * `pg_db`: The name of the database to connect to. * `pg_user`: The username to connect with. * `pg_pass`: The password to connect with. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to load the password from a secret store, e.g. `${secrets:my_pg_pass}`. * `pg_sslmode`: Optional. Specifies the SSL/TLS behavior for the connection, supported values: * `verify-full`: (default) This mode requires an SSL connection, a valid root certificate, and the server host name to match the one specified in the certificate. * `verify-ca`: This mode requires a TLS connection and a valid root certificate. * `require`: This mode requires a TLS connection. * `prefer`: This mode will try to establish a secure TLS connection if possible, but will connect insecurely if the server does not support TLS. * `disable`: This mode will not attempt to use a TLS connection, even if the server supports it. * `pg_sslrootcert`: Optional parameter specifying the path to a custom PEM certificate that the connector will trust. * `pg_connection_pool_min`: Optional. The minimum number of connections to keep open in the pool, lazily created when requested. Default is `5`. * `connection_pool_size`: Optional. The maximum number of connections created in the connection pool. Default is `10`. Configuration `params` are provided either in the `acceleration` section of a dataset. ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: postgres params: pg_host: localhost pg_port: 5432 pg_db: my_database pg_user: my_user pg_pass: ${secrets:my_pg_pass} pg_sslmode: require ``` Specify different secrets for a PostgreSQL source and acceleration: ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset params: pg_host: localhost pg_port: 5432 pg_db: data_store pg_user: my_user pg_pass: ${secrets:pg1_pass} acceleration: engine: postgres params: pg_host: localhost pg_port: 5433 pg_db: acceleration pg_user: two_user_two_furious pg_pass: ${secrets:pg2_pass} ``` Limitations * The Postgres accelerator does not support `Map` types. * The Postgres federated queries may result in unexpected result types due to the difference in DataFusion and Postgres size increase rules. Please explicitly specify the expected output type of aggregation functions when writing query involving Postgres table in Spice. For example, rewrite `SUM(int_col)` into `CAST (SUM(int_col) as BIGINT`. ## Arrow to PostgreSQL Type Mapping[​](#arrow-to-postgresql-type-mapping "Direct link to Arrow to PostgreSQL Type Mapping") The table below lists the supported [Apache Arrow data types](https://arrow.apache.org/rust/arrow/datatypes/enum.DataType.html) and their mappings to [PostgreSQL types](https://www.postgresql.org/docs/current/datatype.html) when stored | Arrow Type | sea\_query ColumnType | PostgreSQL Type | | -------------------------------------- | ----------------------- | ----------------------------- | | `Int8` | `TinyInteger` | `smallint` | | `Int16` | `SmallInteger` | `smallint` | | `Int32` | `Integer` | `integer` | | `Int64` | `BigInteger` | `bigint` | | `UInt8` | `TinyUnsigned` | `smallint` | | `UInt16` | `SmallUnsigned` | `smallint` | | `UInt32` | `Unsigned` | `bigint` | | `UInt64` | `BigUnsigned` | `numeric` | | `Decimal128` / `Decimal256` | `Decimal` | `decimal` | | `Float32` | `Float` | `real` | | `Float64` | `Double` | `double precision` | | `Utf8 / LargeUtf8` | `Text` | `text` | | `Boolean` | `Boolean` | `bool` | | `Binary / LargeBinary` | `VarBinary` | `bytea` | | `FixedSizeBinary` | `Binary` | `bytea` | | `Timestamp` (no Timezone) | `Timestamp` | `timestamp` without time zone | | `Timestamp` (with Timezone) | `TimestampWithTimeZone` | `timestamp` with time zone | | `Date32` / `Date64` | `Date` | `date` | | `Time32` / `Time64` | `Time` | `time` | | `Interval` | `Interval` | `interval` | | `Duration` | `BigInteger` | `bigint` | | `List` / `LargeList` / `FixedSizeList` | `Array` | `array` | | `Struct` | `N/A` | `Composite` (Custom type) | ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure PostgreSQL as a data accelerator in Spice. [PostgreSQL Data Accelerator](https://github.com/spiceai/cookbook/tree/trunk/postgres/accelerator#readme) --- # SQLite Data Accelerator To use SQLite as Data Accelerator, specify `sqlite` as the `engine` for acceleration. ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: sqlite ``` ## Configuration[​](#configuration "Direct link to Configuration") The connection to SQLite can be configured by providing the following `params`: * `sqlite_file`: The filename for the file to back the SQLite database. Only applies if `mode` is `file`. * `busy_timeout`: Optional. Specifies the duration for the SQLite [busy timeout](https://www.sqlite.org/c3ref/busy_timeout.html) when connecting to the database file. Default: 5000 ms. Configuration `params` are provided in the `acceleration` section of a dataset. Other common `acceleration` fields can be configured for sqlite, see see [datasets](/docs/v1.10/reference/spicepod/datasets). ``` datasets: - from: spice.ai:path.to.my_dataset name: my_dataset acceleration: engine: sqlite mode: file params: sqlite_file: /my/chosen/location/sqlite.db ``` Limitations * The SQLite accelerator doesn't support arrow `Interval` types, as [SQLite](https://www.sqlite.org/lang_datefunc.html) doesn't have a native interval type. * The SQLite accelerator only supports arrow `List` types of primitive data types; lists with structs are not supported. * The SQLite accelerator doesn't support `Dictionary` or `Map` types. * SQLite may not be suitable for high row count use cases with complex join queries. Use [DuckDB](/docs/v1.10/components/data-accelerators/duckdb) instead. * The SQLite accelerator doesn't support advanced grouping features such as `ROLLUP` and `GROUPING`. * In SQLite, `CAST(value AS DECIMAL)` doesn't convert an integer to a floating-point value if the casted value is an integer. Operations like `CAST(1 AS DECIMAL) / CAST(2 AS DECIMAL)` will be treated as integer division, resulting in 0 instead of the expected 0.5. Use `FLOAT` to ensure conversion to a floating-point value: `CAST(1 AS FLOAT) / CAST(2 AS FLOAT)`. * Updating a dataset with SQLite acceleration while the Spice Runtime is running (hot-reload) will cause SQLite accelerator query federation to disable until the Runtime is restarted. Memory Considerations When accelerating a dataset using `mode: memory` (the default), some or all of the dataset is loaded into memory. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. In-memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/v1.10/components/data-accelerators/duckdb) and [`sqlite`](/docs/v1.10/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure SQLite as a data accelerator in Spice. [SQLite Data Accelerator](https://github.com/spiceai/cookbook/tree/trunk/sqlite/accelerator#readme) --- # Data Connectors Data Connectors provide connections to databases, data warehouses, and data lakes for federated SQL queries and data replication. Supported Data Connectors include: | Name | Description | Status | Protocol/Format | | ---------------------------------- | ---------------------------------------------------------------------------------------------- | ----------------- | --------------------------------------------------------------------------------- | | `postgres` | PostgreSQL, Amazon Redshift | Stable | PostgreSQL-line | | `mysql` | MySQL | Stable | | | `s3` | [S3](https://github.com/spiceai/cookbook/tree/trunk/s3#readme) | Stable | Parquet, CSV, JSON | | `file` | File | Stable | Parquet, CSV, JSON | | `duckdb` | DuckDB | Stable | Embedded | | `dremio` | [Dremio](https://github.com/spiceai/cookbook/tree/trunk/dremio#readme) | Stable | Arrow Flight | | `spice.ai` | [Spice.ai OSS & Cloud](https://github.com/spiceai/cookbook/tree/trunk/spiceai#readme) | Stable | Arrow Flight | | `databricks (mode: delta_lake)` | [Databricks](https://github.com/spiceai/cookbook/tree/trunk/databricks#readme) | Stable | S3/Delta Lake | | `delta_lake` | Delta Lake | Stable | Delta Lake | | `github` | GitHub | Stable | GitHub API | | `graphql` | GraphQL | Release Candidate | JSON | | `databricks (mode: spark_connect)` | [Databricks](https://github.com/spiceai/cookbook/tree/trunk/databricks#readme) | Beta | [Spark Connect](https://spark.apache.org/docs/latest/spark-connect-overview.html) | | `flightsql` | FlightSQL | Beta | Arrow Flight SQL | | `mssql` | Microsoft SQL Server | Beta | Tabular Data Stream (TDS) | | `odbc` | ODBC | Beta | ODBC | | `snowflake` | Snowflake | Beta | Arrow | | `spark` | Spark | Beta | [Spark Connect](https://spark.apache.org/docs/latest/spark-connect-overview.html) | | `iceberg` | [Apache Iceberg](https://github.com/spiceai/cookbook/tree/trunk/catalogs/iceberg#readme) | Beta | Parquet | | `abfs` | Azure BlobFS | Alpha | Parquet, CSV, JSON | | `ftp`, `sftp` | FTP/SFTP | Alpha | Parquet, CSV, JSON | | `glue` | [Glue](https://github.com/spiceai/cookbook/tree/trunk/glue/README.md) | Alpha | Iceberg, Parquet, CSV | | `http`, `https` | HTTP(s) | Alpha | Parquet, CSV, JSON | | `imap` | IMAP | Alpha | IMAP Emails | | `localpod` | [Local dataset replication](https://github.com/spiceai/cookbook/blob/trunk/localpod/README.md) | Alpha | | | `oracle` | Oracle | Alpha | [Oracle ODPI-C](https://oracle.github.io/odpi/) | | `sharepoint` | Microsoft SharePoint | Alpha | Unstructured UTF-8 documents | | `clickhouse` | Clickhouse | Alpha | | | `debezium` | Debezium CDC | Alpha | Kafka + JSON | | `kafka` | Kafka | Alpha | Kafka + JSON | | `dynamodb` | DynamoDB | Release Candidate | | | `mongodb` | MongoDB | Alpha | | | `elasticsearch` | ElasticSearch | Roadmap | | ## Object Store File Formats[​](#object-store-file-formats "Direct link to Object Store File Formats") For data connectors that are object store compatible, if a folder is provided, the file format must be specified with `params.file_format`. If a file is provided, the file format will be inferred, and `params.file_format` is unnecessary. File formats currently supported are: ``` datasets: - from: s3://bucket/data/sales/ name: sales params: file_format: parquet ``` When connecting to a **specific file**, the format is inferred from the file extension: ``` datasets: - from: sftp://files.example.com/reports/quarterly.parquet name: quarterly_report ``` ### Supported Formats[​](#supported-formats "Direct link to Supported Formats") | Name | Parameter | Status | Description | | --------------------------------------------- | ---------------------- | ------- | --------------------------------------------- | | [Apache Parquet](https://parquet.apache.org/) | `file_format: parquet` | Stable | Columnar format optimized for analytics | | [CSV](/docs/reference/file_format#csv) | `file_format: csv` | Stable | Comma-separated values | | JSON | `file_format: json` | Roadmap | JavaScript Object Notation | | [Apache Iceberg](https://iceberg.apache.org/) | `file_format: iceberg` | Roadmap | Open table format for large analytic datasets | | Microsoft Excel | `file_format: xlsx` | Roadmap | Excel spreadsheet format | | Markdown | `file_format: md` | Stable | Plain text with formatting (document format) | | Text | `file_format: txt` | Stable | Plain text files (document format) | | PDF | `file_format: pdf` | Alpha | Portable Document Format (document format) | | Microsoft Word | `file_format: docx` | Alpha | Word document format (document format) | ### Format-Specific Parameters[​](#format-specific-parameters "Direct link to Format-Specific Parameters") File formats support additional parameters for fine-grained control. Common examples include: | Parameter | Applies To | Description | | ---------------- | ---------- | ------------------------------------------------ | | `csv_has_header` | CSV | Whether the first row contains column headers | | `csv_delimiter` | CSV | Field delimiter character (default: `,`) | | `csv_quote` | CSV | Quote character for fields containing delimiters | For complete format options, see [File Formats Reference](/docs/reference/file_format). ### Applicable Connectors[​](#object-store-file-formats "Direct link to Applicable Connectors") The following data connectors support file format configuration: | Connector Type | Connectors | | ---------------------------- | -------------------------------------- | | **Object Stores** | S3, Azure Blob (ABFS), GCS, HTTP/HTTPS | | **Network-Attached Storage** | FTP, SFTP, SMB, NFS | | **Local Storage** | File | ### Hive Partitioning[​](#hive-partitioning "Direct link to Hive Partitioning") File-based connectors support Hive-style partitioning, which extracts partition columns from folder names. Enable with `hive_partitioning_enabled: true`. Given a folder structure: ``` /data/ year=2024/ month=01/ data.parquet month=02/ data.parquet ``` Configure the dataset: ``` datasets: - from: s3://bucket/data/ name: partitioned_data params: file_format: parquet hive_partitioning_enabled: true ``` Query with partition filters: ``` SELECT * FROM partitioned_data WHERE year = '2024' AND month = '01'; ``` Partition pruning improves query performance by reading only the relevant files. | Name | Parameter | Supported | Is Document Format | | --------------------------------------------- | ---------------------- | --------- | ------------------ | | [Apache Parquet](https://parquet.apache.org/) | `file_format: parquet` | ✅ | ❌ | | [CSV](/docs/reference/file_format#csv) | `file_format: csv` | ✅ | ❌ | | [Apache Iceberg](https://iceberg.apache.org/) | `file_format: iceberg` | Roadmap | ❌ | | JSON | `file_format: json` | Roadmap | ❌ | | Microsoft Excel | `file_format: xlsx` | Roadmap | ❌ | | Markdown | `file_format: md` | ✅ | ✅ | | Text | `file_format: txt` | ✅ | ✅ | | PDF | `file_format: pdf` | Alpha | ✅ | | Microsoft Word | `file_format: docx` | Alpha | ✅ | File formats support additional parameters in the `params` (like `csv_has_header`) described in [File Formats](/docs/reference/file_format) If a format is a document format, each file will be treated as a document, as per [document support](#document-support) below. Note Document formats in Alpha (e.g. pdf, docx) may not parse all structure or text from the underlying documents correctly. ### Document Support[​](#document-support "Direct link to Document Support") If a Data Connector supports documents, when the appropriate file format is specified (see [above](#object-store-file-formats)), each file will be treated as a row in the table, with the contents of the file within the `content` column. Additional columns will exist, dependent on the data connector. #### Example[​](#example "Direct link to Example") Consider a local filesystem ``` >>> ls -la total 232 drwxr-sr-x@ 22 jeadie staff 704 30 Jul 13:12 . drwxr-sr-x@ 18 jeadie staff 576 30 Jul 13:12 .. -rw-r--r--@ 1 jeadie staff 1329 15 Jan 2024 DR-000-Template.md -rw-r--r--@ 1 jeadie staff 4966 11 Aug 2023 DR-001-Dremio-Architecture.md -rw-r--r--@ 1 jeadie staff 2307 28 Jul 2023 DR-002-Data-Completeness.md ``` And the spicepod ``` datasets: - name: my_documents from: file:docs/decisions/ params: file_format: md ``` A Document table will be created. ``` >>> SELECT * FROM my_documents LIMIT 3 +----------------------------------------------------+--------------------------------------------------+ | location | content | +----------------------------------------------------+--------------------------------------------------+ | Users/docs/decisions/DR-000-Template.md | # DR-000: DR Template | | | **Date:** <> | | | **Decision Makers:** | | | - @<> | | | - @<> | | | ... | | Users/docs/decisions/DR-001-Dremio-Architecture.md | # DR-001: Add "Cached" Dremio Dataset | | | | | | ## Context | | | | | | We use [Dremio](https://www.dremio.com/) to p... | | Users/docs/decisions/DR-002-Data-Completeness.md | # DR-002: Append-Only Data Completeness | | | | | | ## Context | | | | | | Our Ethereum append-only dataset is incomple... | +----------------------------------------------------+--------------------------------------------------+ ``` ## Data Connector Docs[​](#data-connector-docs "Direct link to Data Connector Docs") ## [📄️Redshift Data Connector](/docs/v1.10/components/data-connectors/redshift) [Connect to Amazon Redshift using the PostgreSQL connector in Spice.](/docs/v1.10/components/data-connectors/redshift) ## [📄️Azure BlobFS Data Connector](/docs/v1.10/components/data-connectors/abfs) [Azure BlobFS Data Connector Documentation](/docs/v1.10/components/data-connectors/abfs) ## [📄️ClickHouse Data Connector](/docs/v1.10/components/data-connectors/clickhouse) [ClickHouse Data Connector Documentation](/docs/v1.10/components/data-connectors/clickhouse) ## [📄️Databricks Data Connector](/docs/v1.10/components/data-connectors/databricks) [Databricks Data Connector Documentation](/docs/v1.10/components/data-connectors/databricks) ## [📄️Debezium Data Connector](/docs/v1.10/components/data-connectors/debezium) [Debezium Data Connector Documentation](/docs/v1.10/components/data-connectors/debezium) ## [📄️Delta Lake Data Connector](/docs/v1.10/components/data-connectors/delta-lake) [Delta Lake Data Connector Documentation](/docs/v1.10/components/data-connectors/delta-lake) ## [📄️Dremio Data Connector](/docs/v1.10/components/data-connectors/dremio) [Dremio Data Connector Documentation](/docs/v1.10/components/data-connectors/dremio) ## [📄️DuckDB Data Connector](/docs/v1.10/components/data-connectors/duckdb) [DuckDB Data Connector Documentation](/docs/v1.10/components/data-connectors/duckdb) ## [📄️DynamoDB Data Connector](/docs/v1.10/components/data-connectors/dynamodb) [DynamoDB Data Connector Documentation](/docs/v1.10/components/data-connectors/dynamodb) ## [📄️File Data Connector](/docs/v1.10/components/data-connectors/file) [File Data Connector Documentation](/docs/v1.10/components/data-connectors/file) ## [📄️Flight SQL Data Connector](/docs/v1.10/components/data-connectors/flightsql) [Flight SQL Data Connector Documentation](/docs/v1.10/components/data-connectors/flightsql) ## [📄️FTP/SFTP Data Connector](/docs/v1.10/components/data-connectors/ftp) [FTP/SFTP Data Connector Documentation](/docs/v1.10/components/data-connectors/ftp) ## [📄️GitHub Data Connector](/docs/v1.10/components/data-connectors/github) [GitHub Data Connector Documentation](/docs/v1.10/components/data-connectors/github) ## [📄️Glue Data Connector](/docs/v1.10/components/data-connectors/glue) [Glue Data Connector Documentation](/docs/v1.10/components/data-connectors/glue) ## [📄️GraphQL Data Connector](/docs/v1.10/components/data-connectors/graphql) [GraphQL Data Connector Documentation](/docs/v1.10/components/data-connectors/graphql) ## [📄️HTTP(s) Data Connector](/docs/v1.10/components/data-connectors/https) [HTTP(s) Data Connector Documentation](/docs/v1.10/components/data-connectors/https) ## [📄️Iceberg Data Connector](/docs/v1.10/components/data-connectors/iceberg) [Connect to and query Apache Iceberg tables](/docs/v1.10/components/data-connectors/iceberg) ## [📄️IMAP Data Connector](/docs/v1.10/components/data-connectors/imap) [IMAP Data Connector Documentation](/docs/v1.10/components/data-connectors/imap) ## [📄️Kafka Data Connector](/docs/v1.10/components/data-connectors/kafka) [Kafka Data Connector Documentation](/docs/v1.10/components/data-connectors/kafka) ## [📄️Localpod Data Connector](/docs/v1.10/components/data-connectors/localpod) [Localpod Data Connector Documentation](/docs/v1.10/components/data-connectors/localpod) ## [📄️Memory Data Connector](/docs/v1.10/components/data-connectors/memory) [Memory Data Connector Documentation](/docs/v1.10/components/data-connectors/memory) ## [📄️MongoDB Data Connector](/docs/v1.10/components/data-connectors/mongodb) [MongoDB Data Connector Documentation](/docs/v1.10/components/data-connectors/mongodb) ## [📄️Microsoft SQL Server](/docs/v1.10/components/data-connectors/mssql) [Microsoft SQL Server Data Connector](/docs/v1.10/components/data-connectors/mssql) ## [📄️MySQL Data Connector](/docs/v1.10/components/data-connectors/mysql) [MySQL Data Connector Documentation](/docs/v1.10/components/data-connectors/mysql) ## [📄️ODBC Data Connector](/docs/v1.10/components/data-connectors/odbc) [ODBC Data Connector Documentation](/docs/v1.10/components/data-connectors/odbc) ## [📄️Oracle Data Connector](/docs/v1.10/components/data-connectors/oracle) [Oracle Data Connector Documentation](/docs/v1.10/components/data-connectors/oracle) ## [📄️PostgreSQL Data Connector](/docs/v1.10/components/data-connectors/postgres) [PostgreSQL Data Connector Documentation](/docs/v1.10/components/data-connectors/postgres) ## [📄️S3 Data Connector](/docs/v1.10/components/data-connectors/s3) [S3 Data Connector Documentation](/docs/v1.10/components/data-connectors/s3) ## [📄️SharePoint Data Connector](/docs/v1.10/components/data-connectors/sharepoint) [SharePoint Data Connector Documentation](/docs/v1.10/components/data-connectors/sharepoint) ## [📄️Snowflake Data Connector](/docs/v1.10/components/data-connectors/snowflake) [Snowflake Data Connector Documentation](/docs/v1.10/components/data-connectors/snowflake) ## [📄️Apache Spark Connector](/docs/v1.10/components/data-connectors/spark) [Apache Spark Connector Documentation](/docs/v1.10/components/data-connectors/spark) ## [📄️Spice.ai Data Connector](/docs/v1.10/components/data-connectors/spiceai) [Spice.ai Data Connector Documentation](/docs/v1.10/components/data-connectors/spiceai) --- # Azure BlobFS Data Connector The Azure BlobFS (ABFS) Data Connector enables federated SQL queries on files stored in Azure Blob-compatible endpoints. This includes Azure Data Lake Storage Gen2 endpoints accessed via the `abfs://` and `abfss://` schemes. When a folder path is provided, all the contained files will be loaded. File formats are specified using the `file_format` parameter, as described in [Object Store File Formats](/docs/v1.10/components/data-connectors/#object-store-file-formats). ``` datasets: - from: abfs://foocontainer/taxi_sample.csv name: azure_test params: abfs_account: spiceadls abfs_access_key: ${ secrets:access_key } file_format: csv ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Defines the ABFS-compatible URI to a folder or object: * `from: abfs:///` with the account name configured using `abfs_account` parameter, or * `from: abfs://@.dfs.core.windows.net/` ### `name`[​](#name "Direct link to name") Defines the dataset name, which is used as the table name within Spice. Example: ``` datasets: - from: abfs://foocontainer/taxi_sample.csv name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") #### Basic parameters[​](#basic-parameters "Direct link to Basic parameters") | Parameter name | Description | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `file_format` | Specifies the data format. Required if not inferrable from `from`. Options: `parquet`, `csv`. Refer to [Object Store File Formats](/docs/v1.10/components/data-connectors/#object-store-file-formats) for details. | | `abfs_account` | Azure storage account name | | `abfs_sas_string` | SAS (Shared Access Signature) Token to use for authorization | | `abfs_endpoint` | Storage endpoint, default: `https://{account}.blob.core.windows.net` | | `abfs_use_emulator` | Use `true` or `false` to connect to a local emulator | | `abfs_authority_host` | Alternative authority host, default: `https://login.microsoftonline.com` | | `abfs_proxy_url` | Proxy URL | | `abfs_proxy_ca_certificate` | CA certificate for the proxy | | `abfs_proxy_excludes` | A list of hosts to exclude from proxy connections | | `abfs_disable_tagging` | Disable tagging objects. Use this if your backing store doesn't support tags | | `allow_http` | Allow insecure HTTP connections | | `hive_partitioning_enabled` | Enable partitioning using hive-style partitioning from the folder structure. Defaults to `false` | | `schema_source_path` | Specifies the URL used to infer the dataset schema. Default to the most recently modified file | #### Authentication parameters[​](#authentication-parameters "Direct link to Authentication parameters") The following authentication methods are mutually exclusive — only one can be used at a time: * `abfs_access_key` * `abfs_bearer_token` * `abfs_sas_string` * Client credentials (`abfs_client_id` + `abfs_client_secret` + `abfs_tenant_id`) * `abfs_use_cli` * `abfs_msi_endpoint` * `abfs_federated_token_file` * `abfs_skip_signature` If none of these are set the connector will default to using a [managed identity](https://learn.microsoft.com/en-us/entra/identity/managed-identities-azure-resources/overview) | Parameter name | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `abfs_access_key` | Secret access key | | `abfs_bearer_token` | `BEARER` access token for user authentication. The token can be obtained from the OAuth2 flow (see [access token authentication](#access-token-authentication)). | | `abfs_client_id` | Client ID for client authentication flow | | `abfs_client_secret` | Client Secret to use for client authentication flow | | `abfs_tenant_id` | Tenant ID to use for client authentication flow | | `abfs_skip_signature` | Skip credentials and request signing for public containers | | `abfs_msi_endpoint` | Endpoint for managed identity tokens | | `abfs_federated_token_file` | File path for federated identity token in Kubernetes | | `abfs_use_cli` | Set to `true` to use the Azure CLI to acquire access tokens | #### Retry parameters[​](#retry-parameters "Direct link to Retry parameters") | Parameter name | Description | | ------------------------------- | -------------------------------------------- | | `abfs_max_retries` | Maximum retries | | `abfs_retry_timeout` | Total timeout for retries (e.g., `5s`, `1m`) | | `abfs_backoff_initial_duration` | Initial retry delay (e.g., `5s`) | | `abfs_backoff_max_duration` | Maximum retry delay (e.g., `1m`) | | `abfs_backoff_base` | Exponential backoff base (e.g., `0.1`) | ## Authentication[​](#authentication "Direct link to Authentication") ABFS connector supports three types of authentication, as detailed in the [authentication parameters](#authentication-parameters) ### Service principal authentication[​](#service-principal-authentication "Direct link to Service principal authentication") Configure service principal authentication by setting the `abfs_client_secret` parameter. 1. Create a new Azure AD application in the [Azure portal](https://portal.azure.com/#view/Microsoft_AAD_IAM/ActiveDirectoryMenuBlade/~/Overview) and generate a `client secret` under `Certificates & secrets`. 2. Grant the Azure AD application read access to the storage account under `Access Control (IAM)`, this can typically be done using the `Storage Blob Data Reader` built-in role. ### Access key authentication[​](#access-key-authentication "Direct link to Access key authentication") Configure service principal authentication by setting the `abfs_access_key` parameter to [Azure Storage Account Access Key](https://learn.microsoft.com/en-us/azure/storage/common/storage-account-keys-manage?tabs=azure-portal) ### Access token authentication[​](#access-token-authentication "Direct link to Access token authentication") Configure access token authentication by setting the `abfs_bearer_token` parameter, typically obtained through the following the OAuth2 flow with `spice login abfs`. 1. Create a new Azure AD application in the [Azure portal](https://portal.azure.com/#view/Microsoft_AAD_IAM/ActiveDirectoryMenuBlade/~/Overview). 2. Under the application's `API permissions`, add the permission: `Azure Storage - user_impersonation`. 3. Under the applications's `Authentication`, add `http://localhost` as Mobile and desktop applications redirect URI. 4. Grant the user read access to the storage account under `Access Control (IAM)`, this can typically be done using the `Storage Blob Data Reader` built-in role. 5. Obtain the `abfs_bearer_token` using the following command. The `abfs_bearer_token`, `abfs_client_id`, `abfs_tenant_id` will be automatically filled in environment secret after login. Refere to \[`spice login`(../../cli/reference/login) documentation for more details. ``` spice login abfs --tenant-id $TENANT_ID --client-id $CLIENT_ID ``` ## Supported file formats[​](#supported-file-formats "Direct link to Supported file formats") Specify the file format using `file_format` parameter. More details in [Object Store File Formats](/docs/v1.10/components/data-connectors/#object-store-file-formats). ## Examples[​](#examples "Direct link to Examples") ### Reading a CSV file with an Access Key[​](#reading-a-csv-file-with-an-access-key "Direct link to Reading a CSV file with an Access Key") ``` datasets: - from: abfs://foocontainer/taxi_sample.csv name: azure_test params: abfs_account: spiceadls abfs_access_key: ${ secrets:ACCESS_KEY } file_format: csv ``` ### Using Public Containers[​](#using-public-containers "Direct link to Using Public Containers") ``` datasets: - from: abfs://pubcontainer/taxi_sample.csv name: pub_data params: abfs_account: spiceadls abfs_skip_signature: true file_format: csv ``` ### Connecting to the Storage Emulator[​](#connecting-to-the-storage-emulator "Direct link to Connecting to the Storage Emulator") ``` datasets: - from: abfs://test_container/test_csv.csv name: test_data params: abfs_use_emulator: true file_format: csv ``` ### Using secrets for Account name[​](#using-secrets-for-account-name "Direct link to Using secrets for Account name") ``` datasets: - from: abfs://my_container/my_csv.csv name: prod_data params: abfs_account: ${ secrets:PROD_ACCOUNT } file_format: csv ``` ### Authenticating using Client Authentication[​](#authenticating-using-client-authentication "Direct link to Authenticating using Client Authentication") ``` datasets: - from: abfs://my_data/input.parquet name: my_data params: abfs_tenant_id: ${ secrets:MY_TENANT_ID } abfs_client_id: ${ secrets:MY_CLIENT_ID } abfs_client_secret: ${ secrets:MY_CLIENT_SECRET } ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the \[secret stores documentation(../secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the \[using referenced secrets guide(../secret-stores#using-secrets). --- # ClickHouse Data Connector ClickHouse is a fast, open-source columnar database management system designed for online analytical processing (OLAP) and real-time analytics. This connector enables federated SQL queries from a ClickHouse server. ``` datasets: - from: clickhouse:my.dataset name: my_dataset ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field for the ClickHouse connector takes the form of `from:db.dataset` where `db.dataset` is the path to the Dataset within ClickHouse. In the example above it would be `my.dataset`. If `db` is not specified in either the `from` field or the `clickhouse_db` parameter, it will default to the `default` database. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. ``` datasets: - from: clickhouse:my.dataset name: cool_dataset ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords) or any of the following keywords that are reserved by ClickHouse: * `PREWHERE` * `SETTINGS` * `FORMAT` ### `params`[​](#params "Direct link to params") The ClickHouse data connector can be configured by providing the following `params`: | Parameter Name | Definition | | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `clickhouse_connection_string` | The connection string to use to connect to the ClickHouse server. This can be used instead of providing individual connection parameters. | | `clickhouse_host` | The hostname of the ClickHouse server. | | `clickhouse_tcp_port` | The port of the ClickHouse server. | | `clickhouse_db` | The name of the database to connect to. | | `clickhouse_user` | The username to connect with. | | `clickhouse_pass` | The password to connect with. | | `clickhouse_secure` | Optional. Specifies the SSL/TLS behavior for the connection, supported values:
- `true`: (default) This mode requires an SSL connection. If a secure connection cannot be established, server will not connect.
- `false`: This mode will not attempt to use an SSL connection, even if the server supports it. | | `connection_timeout` | Optional. Specifies the connection timeout in milliseconds. Default is `10000` (10 seconds). | ## Examples[​](#examples "Direct link to Examples") ### Connecting to localhost[​](#connecting-to-localhost "Direct link to Connecting to localhost") ``` datasets: - from: clickhouse:my.dataset name: my_dataset params: clickhouse_host: localhost clickhouse_tcp_port: 9000 clickhouse_db: my_database clickhouse_user: my_user clickhouse_pass: ${secrets:my_clickhouse_pass} connection_timeout: 10000 clickhouse_secure: false ``` ### Specifying a connection timeout[​](#specifying-a-connection-timeout "Direct link to Specifying a connection timeout") ``` datasets: - from: clickhouse:my.dataset name: my_dataset params: clickhouse_connection_string: tcp://my_user:${secrets:my_clickhouse_pass}@localhost:9000/my_database connection_timeout: 10000 clickhouse_secure: true ``` ### Using a connection string[​](#using-a-connection-string "Direct link to Using a connection string") ``` datasets: - from: clickhouse:my.dataset name: my_dataset params: clickhouse_connection_string: tcp://my_user:${secrets:my_clickhouse_pass}@localhost:9000/my_database?connection_timeout=10000&secure=true ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure ClickHouse as data connector in Spice. [Clickhouse Data Connector](https://github.com/spiceai/cookbook/tree/trunk/clickhouse#readme) --- # Databricks Data Connector Databricks as a connector for federated SQL query against Databricks using [Spark Connect](https://www.databricks.com/blog/2022/07/07/introducing-spark-connect-the-power-of-apache-spark-everywhere.html), directly from [Delta Lake](https://delta.io/) tables, or using the [SQL Statement Execution API](https://docs.databricks.com/aws/en/dev-tools/sql-execution-tutorial). ``` datasets: - from: databricks:spiceai.datasets.my_awesome_table # A reference to a table in the Databricks unity catalog name: my_delta_lake_table params: mode: delta_lake databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_token: ${secrets:my_token} databricks_aws_access_key_id: ${secrets:aws_access_key_id} databricks_aws_secret_access_key: ${secrets:aws_secret_access_key} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field for the Databricks connector takes the form `databricks:catalog.schema.table` where `catalog.schema.table` is the fully-qualified path to the table to read from. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: databricks:spiceai.datasets.my_awesome_table name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to reference a secret, e.g. `${secrets:my_token}`. | Parameter Name | Description | | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mode` | The execution mode for querying against Databricks. The default is `spark_connect`. Possible values:
- `spark_connect`: Use Spark Connect to query against Databricks. Requires a Spark cluster to be available.
- `delta_lake`: Query directly from Delta Tables. Requires the object store credentials to be provided.
- `sql_warehouse`: Query through a Databricks SQL Warehouse. Requires `databricks_sql_warehouse_id`. | | `databricks_endpoint` | The endpoint of the Databricks instance. Required for all modes. | | `databricks_sql_warehouse_id` | The ID of the SQL Warehouse in Databricks to use for the query. Only valid when `mode` is `sql_warehouse`. | | `databricks_cluster_id` | The ID of the compute cluster in Databricks to use for the query. Only valid when `mode` is `spark_connect`. | | `databricks_use_ssl` | If true, use a TLS connection to connect to the Databricks endpoint. Default is `true`. | | `client_timeout` | Optional. Specifies timeout for operations. In `delta_lake` mode, applies to object store operations. In `sql_warehouse` mode, applies per-HTTP-call. Default value is `30s`. E.g. `client_timeout: 60s` | | `databricks_token` | The Databricks API token to authenticate with the Unity Catalog API. Can't be used with `databricks_client_id` and `databricks_client_secret`. | | `databricks_client_id` | The Databricks Service Principal Client ID. Can't be used with `databricks_token`. | | `databricks_client_secret` | The Databricks Service Principal Client Secret. Can't be used with `databricks_token`. | ## Authentication[​](#authentication "Direct link to Authentication") ### Personal access token[​](#personal-access-token "Direct link to Personal access token") To learn more about how to set up personal access tokens, see [Databricks PAT docs](https://docs.databricks.com/aws/en/dev-tools/auth/pat). ``` datasets: - from: databricks:spiceai.datasets.my_awesome_table name: my_awesome_table params: databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_cluster_id: 1234-567890-abcde123 databricks_token: ${secrets:DATABRICKS_TOKEN} # PAT ``` ### Databricks service principal[​](#databricks-service-principal "Direct link to Databricks service principal") Spice supports the Machine-to-Machine (M2M) OAuth flow with service principal credentials by utilizing the `databricks_client_id` and `databricks_client_secret` parameters. The runtime will automatically refresh the token. Ensure that you grant your service principal the "Data Reader" privilege preset for the catalog and "Can Attach" cluster permissions when using Spark Connect mode. To Learn more about how to set up the service principal, see [Databricks M2M OAuth docs](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m). ``` datasets: - from: databricks:spiceai.datasets.my_awesome_table name: my_awesome_table params: databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_cluster_id: 1234-567890-abcde123 databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} # service principal client id databricks_client_secret: ${secrets:DATABRICKS_CLIENT_SECRET} # service principal client secret ``` ## Delta Lake object store parameters[​](#delta-lake-object-store-parameters "Direct link to Delta Lake object store parameters") Configure the connection to the object store when using `mode: delta_lake`. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to reference a secret, e.g. `${secrets:aws_access_key_id}`. ### AWS S3[​](#aws-s3 "Direct link to AWS S3") | Parameter Name | Description | | ---------------------------------- | ---------------------------------------------------------------------------------------------- | | `databricks_aws_region` | Optional. The AWS region for the S3 object store. E.g. `us-west-2`. | | `databricks_aws_access_key_id` | The access key ID for the S3 object store. | | `databricks_aws_secret_access_key` | The secret access key for the S3 object store. | | `databricks_aws_endpoint` | Optional. The endpoint for the S3 object store. E.g. `s3.us-west-2.amazonaws.com`. | | `databricks_aws_allow_http` | Optional. Enables insecure HTTP connections to `databricks_aws_endpoint`. Defaults to `false`. | ### Azure Blob[​](#azure-blob "Direct link to Azure Blob") Note **One** of the following auth values must be provided for Azure Blob: * `databricks_azure_storage_account_key`, * `databricks_azure_storage_client_id` and `databricks_azure_storage_client_secret`, or * `databricks_azure_storage_sas_key`. | Parameter Name | Description | | ---------------------------------------- | ---------------------------------------------------------------------- | | `databricks_azure_storage_account_name` | The Azure Storage account name. | | `databricks_azure_storage_account_key` | The Azure Storage key for accessing the storage account. | | `databricks_azure_storage_client_id` | The Service Principal client ID for accessing the storage account. | | `databricks_azure_storage_client_secret` | The Service Principal client secret for accessing the storage account. | | `databricks_azure_storage_sas_key` | The shared access signature key for accessing the storage account. | | `databricks_azure_storage_endpoint` | Optional. The endpoint for the Azure Blob storage account. | ### Google Storage (GCS)[​](#google-storage-gcs "Direct link to Google Storage (GCS)") | Parameter Name | Description | | ----------------------------------- | ------------------------------------------------------------ | | `databricks_google_service_account` | Filesystem path to the Google service account JSON key file. | ## Examples[​](#examples "Direct link to Examples") ### Spark Connect[​](#spark-connect "Direct link to Spark Connect") ``` - from: databricks:spiceai.datasets.my_spark_table # A reference to a table in the Databricks unity catalog name: my_delta_lake_table params: mode: spark_connect databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_cluster_id: 1234-567890-abcde123 databricks_token: ${secrets:my_token} ``` ### SQL Warehouse[​](#sql-warehouse "Direct link to SQL Warehouse") ``` - from: databricks:spiceai.datasets.my_table # A reference to a table in the Databricks unity catalog name: my_table params: mode: sql_warehouse databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_sql_warehouse_id: 2b4e24cff378fb24 databricks_token: ${secrets:my_token} ``` ### Delta Lake (S3)[​](#delta-lake-s3 "Direct link to Delta Lake (S3)") ``` - from: databricks:spiceai.datasets.my_delta_table # A reference to a table in the Databricks unity catalog name: my_delta_lake_table params: mode: delta_lake databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_token: ${secrets:my_token} databricks_aws_region: us-west-2 # Optional databricks_aws_access_key_id: ${secrets:aws_access_key_id} databricks_aws_secret_access_key: ${secrets:aws_secret_access_key} databricks_aws_endpoint: s3.us-west-2.amazonaws.com # Optional ``` ### Delta Lake (Azure Blobs)[​](#delta-lake-azure-blobs "Direct link to Delta Lake (Azure Blobs)") ``` - from: databricks:spiceai.datasets.my_adls_table # A reference to a table in the Databricks unity catalog name: my_delta_lake_table params: mode: delta_lake databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_token: ${secrets:my_token} # Account Name + Key databricks_azure_storage_account_name: my_account databricks_azure_storage_account_key: ${secrets:my_key} # OR Service Principal + Secret databricks_azure_storage_client_id: my_client_id databricks_azure_storage_client_secret: ${secrets:my_secret} # OR SAS Key databricks_azure_storage_sas_key: my_sas_key ``` ### Delta Lake (GCP)[​](#delta-lake-gcp "Direct link to Delta Lake (GCP)") ``` - from: databricks:spiceai.datasets.my_gcp_table # A reference to a table in the Databricks unity catalog name: my_delta_lake_table params: mode: delta_lake databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_token: ${secrets:my_token} databricks_google_service_account: /path/to/service-account.json ``` ## Types[​](#types "Direct link to Types") ### mode: delta\_lake[​](#mode-delta_lake "Direct link to mode: delta_lake") The table below shows the Databricks (mode: delta\_lake) data types supported, along with the type mapping to Apache Arrow types in Spice. | Databricks SQL Type | Arrow Type | | ------------------- | ------------------------------------- | | `STRING` | `Utf8` | | `BIGINT` | `Int64` | | `INT` | `Int32` | | `SMALLINT` | `Int16` | | `TINYINT` | `Int8` | | `FLOAT` | `Float32` | | `DOUBLE` | `Float64` | | `BOOLEAN` | `Boolean` | | `BINARY` | `Binary` | | `DATE` | `Date32` | | `TIMESTAMP` | `Timestamp(Microsecond, Some("UTC"))` | | `TIMESTAMP_NTZ` | `Timestamp(Microsecond, None)` | | `DECIMAL` | `Decimal128` | | `ARRAY` | `List` | | `STRUCT` | `Struct` | | `MAP` | `Map` | ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## Limitations[​](#limitations "Direct link to Limitations") * Databricks connector (mode: delta\_lake) does not support reading Delta tables with the `V2Checkpoint` feature enabled. To use the Databricks connector (mode: delta\_lake) with such tables, drop the `V2Checkpoint` feature by executing the following command: ``` ALTER TABLE DROP FEATURE v2Checkpoint [TRUNCATE HISTORY]; ``` For more details on dropping Delta table features, refer to the official documentation: [Drop Delta table features](https://docs.databricks.com/en/delta/drop-feature.html#:~:text=Databricks%20provides%20limited%20support%20for,data%20files%20backing%20the%20table.) * When using `mode: spark_connect`, correlated scalar subqueries can only be used in filters, aggregations, projections, and UPDATE/MERGE/DELETE commands. [Spark Docs](https://spark.apache.org/docs/latest/sql-error-conditions-unsupported-subquery-expression-category-error-class.html#unsupported_correlated_scalar_subquery) Memory Considerations When using the Databricks (mode: delta\_lake) Data connector without acceleration, data is loaded into memory during query execution. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. Memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/v1.10/components/data-accelerators/duckdb) and [`sqlite`](/docs/v1.10/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. * The Databricks Connector (`mode: spark_connect`) does not yet support streaming query results from Spark. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Databricks as a data connector in Spice. [Spice on Databricks](https://github.com/spiceai/cookbook/tree/trunk/databricks) --- # Debezium Data Connector [Debezium](https://debezium.io/) is an open-source platform that enables \[Change Data Capture (CDC)(../../features/cdc) for efficient real-time updates of locally accelerated datasets. Spice supports connecting to a Kafka topic managed by Debezium to keep datasets up-to-date with the source data. ``` datasets: - from: debezium:my_kafka_topic_with_debezium_changes name: my_dataset params: debezium_transport: kafka # Optional. Only `kafka` is currently supported. debezium_message_format: json # Optional. Only `json` is currently supported. kafka_bootstrap_servers: broker1:9092,broker2:9092,broker3:9092 # Required. A comma separated list of Kafka broker servers. kafka_security_protocol: sasl_ssl # Default is `sasl_ssl`. Valid values are `plaintext`, `ssl`, `sasl_plaintext`, `sasl_ssl`. kafka_sasl_mechanism: SCRAM-SHA-512 # Default is `SCRAM-SHA-512`. Valid values are `PLAIN`, `SCRAM-SHA-256`, `SCRAM-SHA-512`. kafka_sasl_username: kafka # Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. kafka_sasl_password: ${secrets:kafka_sasl_password} # Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. kafka_ssl_ca_location: ./certs/kafka_ca_cert.pem # Optional. Used to verify the SSL/TLS certificate of the Kafka broker. kafka_enable_ssl_certificate_verification: true # Default is `true`. Set to `false` to disable SSL/TLS certificate verification. kafka_ssl_endpoint_identification_algorithm: https # Default is `https`. Valid values are `none` and `https`. acceleration: enabled: true # Acceleration is required for the debezium connector. engine: duckdb # `duckdb`, `sqlite` and `postgres` are supported acceleration engines for Debezium. refresh_mode: changes # Optional. If specified, this is required to be set to `changes` - any other value is an error. mode: file # Persistence is recommended to not have to rebuild the table each time Spice starts. ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form of `debezium:kafka_topic` where `kafka_topic` is the name of the Kafka topic where Debezium is notifying consumers about any upstream changes. In the example above it would listen to the `my_kafka_topic_with_debezium_changes` topic. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. ``` datasets: - from: debezium:my_kafka_topic_with_debezium_changes name: cool_dataset ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a \[reserved keyword(../../reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `debezium_transport` | Optional. The message broker transport to use. The default is `kafka`. Possible values:- `kafka`: Use Kafka as the message broker transport. Spice may support additional transports in the future. | | `debezium_message_format` | Optional. The message format to use. The default is `json`. Possible values: - `json`: Use JSON as the message format. Spice is expected to support additional message formats in the future, like `avro`. | | `kafka_bootstrap_servers` | **Required**. A list of host/port pairs for establishing the initial Kafka cluster connection. The client will use all servers, regardless of the bootstrapping servers specified here. This list only affects the initial hosts used to discover the full server set and should be formatted as `host1:port1,host2:port2,...`. | | `kafka_security_protocol` | Security protocol for Kafka connections. Default: `sasl_ssl`. Options: - `plaintext`
- `ssl`
- `sasl_plaintext`
- `sasl_ssl` | | `kafka_sasl_mechanism` | SASL (Simple Authentication and Security Layer) authentication mechanism. Default: `SCRAM-SHA-512`. Options: - `PLAIN`
- `SCRAM-SHA-256`
- `SCRAM-SHA-512` | | `kafka_sasl_username` | SASL username. | | `kafka_sasl_password` | SASL password. | | `kafka_ssl_ca_location` | Path to the SSL/TLS CA certificate file for server verification. | | `kafka_enable_ssl_certificate_verification` | Enable SSL/TLS certificate verification. Default: `true`. | | `kafka_ssl_endpoint_identification_algorithm` | SSL/TLS endpoint identification algorithm. Default: `https`. Options: - `none`
- `https` | | `kafka_consumer_group_id` | Kafka consumer group id to use. If not set, a unique id will be generated. | ### `metrics`[​](#metrics "Direct link to metrics") The connector supports the following optional \[component metrics(../../features/observability/component\_metrics): | Metric Name | Type | Description | | ------------------------ | ------- | ------------------------------------------------------------------------------------ | | `bytes_consumed_total` | Counter | Total number of bytes consumed from the Kafka topic | | `records_consumed_total` | Counter | Total number of records (messages) consumed from Kafka topics | | `records_lag` | Gauge | Total consumer lag across all topic partitions (number of messages not yet consumed) | These metrics are not enabled by default, enable them by setting the `metrics` parameter: ``` datasets: - from: debezium:my_kafka_topic_with_debezium_changes name: cool_dataset metrics: - name: records_lag - name: records_consumed_total - name: bytes_consumed_total params: ... ``` ### Acceleration Settings[​](#acceleration-settings "Direct link to Acceleration Settings") warning Using the Debezium connector **requires** \[acceleration(../data-accelerators) to be enabled. The following settings are required: | Parameter Name | Description | | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `enabled` | Required. Must be set to `true` to enable acceleration. | | `engine` | Required. The acceleration engine to use. Possible valid values: - `duckdb`: Use \[DuckDB(../data-accelerators/duckdb) as the acceleration engine.
- `sqlite`: Use \[SQLite(../data-accelerators/sqlite) as the acceleration engine.
- `postgres`: Use \[PostgreSQL(../data-accelerators/postgres) as the acceleration engine. | | `refresh_mode` | Optional. The refresh mode to use. If specified, this must be set to `changes`. Any other value is an error. | | `mode` | Optional. The persistence mode to use. When using the `duckdb` and `sqlite` engines, it is recommended to set this to `file` to persist the data across restarts. Spice also persists metadata about the dataset, so it can resume from the last known state of the dataset instead of re-fetching the entire dataset. | ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the \[secret stores documentation(../secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the \[using referenced secrets guide(../secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * See an example of configuring a dataset to use CDC with Debezium by following the sample [Streaming changes in real-time with Debezium CDC](https://github.com/spiceai/cookbook/tree/trunk/cdc-debezium#readme). * An example of [Streaming changes in real-time with Debezium CDC and SASL/SCRAM authentication](https://github.com/spiceai/cookbook/tree/trunk/cdc-debezium/sasl-scram#readme) is available as well. --- # Delta Lake Data Connector Delta Lake data connector enables SQL queries from [Delta Lake](https://delta.io/) tables. ``` datasets: - from: delta_lake:s3://my_bucket/path/to/s3/delta/table/ name: my_delta_lake_table params: delta_lake_aws_access_key_id: ${secrets:aws_access_key_id} delta_lake_aws_secret_access_key: ${secrets:aws_secret_access_key} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field for the Delta Lake connector takes the form of `delta_lake:path` where `path` is any supported path, either local or to a cloud storage location. See the [examples](#examples) section below. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: delta_lake:s3://my_bucket/path/to/s3/delta/table/ name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") Use the [secret replacement syntax](/docs/components/secret-stores) to reference a secret, e.g. `${secrets:aws_access_key_id}`. | Parameter Name | Description | | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | | `client_timeout` | Optional. Specifies timeout for object store operations. No default; uses the underlying HTTP client timeout when not set. E.g. `client_timeout: 60s` | ## Delta Lake object store parameters[​](#delta-lake-object-store-parameters "Direct link to Delta Lake object store parameters") ### AWS S3[​](#aws-s3 "Direct link to AWS S3") | Parameter Name | Description | | ---------------------------------- | ---------------------------------------------------------------------------------------------- | | `delta_lake_aws_region` | Optional. The AWS region for the S3 object store. E.g. `us-west-2`. | | `delta_lake_aws_access_key_id` | The access key ID for the S3 object store. | | `delta_lake_aws_secret_access_key` | The secret access key for the S3 object store. | | `delta_lake_aws_session_token` | Optional. The AWS session token for S3 object store. | | `delta_lake_aws_endpoint` | Optional. The endpoint for the S3 object store. E.g. `s3.us-west-2.amazonaws.com`. | | `delta_lake_aws_allow_http` | Optional. Enables insecure HTTP connections to `delta_lake_aws_endpoint`. Defaults to `false`. | ### Azure Blob[​](#azure-blob "Direct link to Azure Blob") Note **One** of the following auth values must be provided for Azure Blob: * `delta_lake_azure_storage_account_key`, * `delta_lake_azure_storage_client_id` and `delta_lake_azure_storage_client_secret`, or * `delta_lake_azure_storage_sas_key`. | Parameter Name | Description | | ---------------------------------------- | ---------------------------------------------------------------------- | | `delta_lake_azure_storage_account_name` | The Azure Storage account name. | | `delta_lake_azure_storage_account_key` | The Azure Storage master key for accessing the storage account. | | `delta_lake_azure_storage_client_id` | The service principal client id for accessing the storage account. | | `delta_lake_azure_storage_client_secret` | The service principal client secret for accessing the storage account. | | `delta_lake_azure_storage_sas_key` | The shared access signature key for accessing the storage account. | | `delta_lake_azure_storage_endpoint` | Optional. The endpoint for the Azure Blob storage account. | ### Google Storage (GCS)[​](#google-storage-gcs "Direct link to Google Storage (GCS)") | Parameter Name | Description | | ----------------------------------- | ------------------------------------------------------------ | | `delta_lake_google_service_account` | Filesystem path to the Google service account JSON key file. | ## Examples[​](#examples "Direct link to Examples") ### Delta Lake + Local[​](#delta-lake--local "Direct link to Delta Lake + Local") ``` - from: delta_lake:/path/to/local/delta/table # A local filesystem path to a Delta Lake table name: my_delta_lake_table ``` ### Delta Lake + S3[​](#delta-lake--s3 "Direct link to Delta Lake + S3") ``` - from: delta_lake:s3://my_bucket/path/to/s3/delta/table/ # A reference to a table in S3 name: my_delta_lake_table params: delta_lake_aws_region: us-west-2 # Optional delta_lake_aws_access_key_id: ${secrets:aws_access_key_id} delta_lake_aws_secret_access_key: ${secrets:aws_secret_access_key} delta_lake_aws_endpoint: s3.us-west-2.amazonaws.com # Optional ``` ### Delta Lake + MinIO[​](#delta-lake--minio "Direct link to Delta Lake + MinIO") ``` - from: delta_lake:s3://my_bucket/path/to/s3/delta/table/ # A reference to a table in MinIO name: my_delta_lake_table params: delta_lake_aws_region: us-east-1 # Best practice for MinIO delta_lake_aws_access_key_id: ${secrets:aws_access_key_id} delta_lake_aws_secret_access_key: ${secrets:aws_secret_access_key} delta_lake_aws_endpoint: http://localhost:9000 # MinIO Endpoint delta_lake_aws_allow_http: true ``` ### Delta Lake + Azure Blob[​](#delta-lake--azure-blob "Direct link to Delta Lake + Azure Blob") ``` - from: delta_lake:abfss://my_container@my_account.dfs.core.windows.net/path/to/azure/delta/table/ # A reference to a table in Azure Blob name: my_delta_lake_table params: # Account Name + Key delta_lake_azure_storage_account_name: my_account delta_lake_azure_storage_account_key: ${secrets:my_key} # OR Service Principal + Secret delta_lake_azure_storage_client_id: my_client_id delta_lake_azure_storage_client_secret: ${secrets:my_secret} # OR SAS Key delta_lake_azure_storage_sas_key: my_sas_key ``` ### Delta Lake + Google Storage[​](#delta-lake--google-storage "Direct link to Delta Lake + Google Storage") ``` params: delta_lake_google_service_account: /path/to/service-account.json ``` ## Types[​](#types "Direct link to Types") The table below shows the Delta Lake data types supported, along with the type mapping to Apache Arrow types in Spice. | Delta Lake Type | Arrow Type | | --------------- | ------------------------------------- | | `String` | `Utf8` | | `Long` | `Int64` | | `Integer` | `Int32` | | `Short` | `Int16` | | `Byte` | `Int8` | | `Float` | `Float32` | | `Double` | `Float64` | | `Boolean` | `Boolean` | | `Binary` | `Binary` | | `Date` | `Date32` | | `Timestamp` | `Timestamp(Microsecond, Some("UTC"))` | | `TimestampNtz` | `Timestamp(Microsecond, None)` | | `Decimal` | `Decimal128` | | `Array` | `List` | | `Struct` | `Struct` | | `Variant` | `Struct` | | `Map` | `Map` | ## Limitations[​](#limitations "Direct link to Limitations") * Delta Lake connector does not support reading Delta tables with the `V2Checkpoint` feature enabled. To use the Delta Lake connector with such tables, drop the `V2Checkpoint` feature by executing the following command: ``` ALTER TABLE DROP FEATURE v2Checkpoint [TRUNCATE HISTORY]; ``` For more details on dropping Delta table features, refer to the official documentation: [Drop Delta table features](https://docs.delta.io/latest/delta-drop-feature.html) ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/components/secret-stores#using-secrets). --- # Dremio Data Connector [Dremio](https://www.dremio.com/) is a data lake engine that enables high-performance SQL queries directly on data lake storage. It provides a unified interface for querying and analyzing data from various sources without the need for complex data movement or transformation. This connector enables using Dremio as a data source for federated SQL queries. ``` - from: dremio:datasets.dremio_dataset name: dremio_dataset params: dremio_endpoint: grpc://127.0.0.1:32010 dremio_username: demo dremio_password: ${secrets:my_dremio_pass} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `dremio:dataset` where `dataset` is the fully qualified name of the dataset to read from. \[Limitations] Currently, only up to three levels of nesting are supported for dataset names (e.g., a.b.c). Additional levels are not supported at this time. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: dremio:datasets.dremio_dataset name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `dremio_endpoint` | The endpoint used to connect to the Dremio server. | | `dremio_username` | The username used to connect to the Dremio endpoint. | | `dremio_password` | The password used to connect to the Dremio endpoint. Use the [secret replacement syntax](#secrets) to load the password from a secret store, e.g. `${secrets:my_dremio_pass}`. | ## Examples[​](#examples "Direct link to Examples") ### Connecting to a GRPC endpoint[​](#connecting-to-a-grpc-endpoint "Direct link to Connecting to a GRPC endpoint") ``` - from: dremio:datasets.dremio_dataset name: dremio_dataset params: dremio_endpoint: grpc://127.0.0.1:32010 dremio_username: demo dremio_password: ${secrets:my_dremio_pass} ``` ## Types[​](#types "Direct link to Types") The table below shows the Dremio data types supported, along with the type mapping to Apache Arrow types in Spice. | Dremio Type | Arrow Type | | ----------- | ------------------------------ | | `INT` | `Int32` | | `BIGINT` | `Int64` | | `FLOAT` | `Float32` | | `DOUBLE` | `Float64` | | `DECIMAL` | `Decimal128` | | `VARCHAR` | `Utf8` | | `VARBINARY` | `Binary` | | `BOOL` | `Boolean` | | `DATE` | `Date64` | | `TIME` | `Time32` | | `TIMESTAMP` | `Timestamp(Millisecond, None)` | | `INTERVAL` | `Interval` | | `LIST` | `List` | | `STRUCT` | `Struct` | | `MAP` | `Map` | ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## Limitations[​](#limitations "Direct link to Limitations") * Dremio connector does not support queries with the EXCEPT and INTERSECT keywords in Spice REPL. Use DISTINCT and IN/NOT IN instead. See the example below. ``` # fail SELECT ws_item_sk FROM web_sales INTERSECT SELECT ss_item_sk FROM store_sales; # success SELECT DISTINCT ws_item_sk FROM web_sales WHERE ws_item_sk IN ( SELECT DISTINCT ss_item_sk FROM store_sales ); # fail SELECT ws_item_sk FROM web_sales EXCEPT SELECT ss_item_sk FROM store_sales; # success SELECT DISTINCT ws_item_sk FROM web_sales WHERE ws_item_sk NOT IN ( SELECT DISTINCT ss_item_sk FROM store_sales ); ``` ::: ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Dremio as data connector in Spice. [Dremio Data Connector](https://github.com/spiceai/cookbook/tree/trunk/dremio#readme) --- # DuckDB Data Connector DuckDB is an in-process SQL OLAP (Online Analytical Processing) database management system designed for analytical query workloads. It is optimized for fast execution and can be embedded directly into applications, providing efficient data processing without the need for a separate database server. This connector supports DuckDB [persistent databases](https://duckdb.org/docs/connect/overview#persistent-database) as a data source for federated SQL queries. ``` datasets: - from: duckdb:database.schema.table name: my_dataset params: duckdb_open: path/to/duckdb_file.duckdb ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field supports one of two forms: | `from` | Description | | ------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `duckdb:database.schema.table` | Read data from a table named `database.schema.table` in the DuckDB file | | `duckdb:*` | Read data using any DuckDB function that produces a table. For example one of the [data import](https://duckdb.org/docs/data/overview) functions such as `read_json`, `read_parquet` or `read_csv`. | ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: duckdb:database.schema.table name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The DuckDB data connector can be configured by providing the following `params`: | Parameter Name | Description | | -------------- | ---------------------------------------- | | `duckdb_open` | The name of the DuckDB database to open. | Configuration `params` are provided either in the top level `dataset` for a dataset source, or in the `acceleration` section for a data store. ## Examples[​](#examples "Direct link to Examples") ### Reading from a relative path[​](#reading-from-a-relative-path "Direct link to Reading from a relative path") A generic example of DuckDB data connector configuration. ``` datasets: - from: duckdb:database.schema.table name: my_dataset params: duckdb_open: path/to/duckdb_file.duckdb ``` ### Reading from an absolute path[​](#reading-from-an-absolute-path "Direct link to Reading from an absolute path") ``` datasets: - from: duckdb:sample_data.nyc.rideshare name: nyc_rideshare params: duckdb_open: /my/path/my_database.db ``` ### DuckDB Functions[​](#duckdb-functions "Direct link to DuckDB Functions") Common [data import](https://duckdb.org/docs/data/overview) DuckDB functions can also define datasets. Instead of a fixed table reference (e.g. `database.schema.table`), a DuckDB function is provided in the `from:` key. For example ``` datasets: - from: duckdb:database.schema.table name: my_dataset params: duckdb_open: path/to/duckdb_file.duckdb - from: duckdb:read_csv('test.csv', header = false) name: from_function ``` Datasets created from DuckDB functions are similar to a standard `SELECT` query. For example: ``` datasets: - from: duckdb:read_csv('test.csv', header = false) ``` is equivalent to: ``` -- from_function SELECT * FROM read_csv('test.csv', header = false); ``` Many DuckDB data imports can be rewritten as DuckDB functions, making them usable as Spice datasets. For example: ``` SELECT * FROM 'todos.json'; -- As a DuckDB function SELECT * FROM read_json('todos.json'); ``` Limitations * The DuckDB connector does not support enum, dictionary, or map [field types](https://duckdb.org/docs/sql/data_types/overview). For example: * Unsupported: * `SELECT MAP(['key1', 'key2', 'key3'], [10, 20, 30])` * The DuckDB connector does not support `Decimal256` (76 digits), as it exceeds DuckDB's maximum Decimal width of 38 digits. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure DuckDB as a data connector in Spice. [DuckDB Data Connector](https://github.com/spiceai/cookbook/tree/trunk/duckdb/connector#readme) --- # DynamoDB Data Connector Amazon DynamoDB is a fully managed NoSQL database service that provides fast and predictable performance with seamless scalability. This connector enables using DynamoDB tables as data sources for federated SQL queries in Spice. ``` datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 dynamodb_aws_access_key_id: ${secrets:aws_access_key_id} # Optional dynamodb_aws_secret_access_key: ${secrets:aws_secret_access_key} # Optional dynamodb_aws_session_token: ${secrets:aws_session_token} # Optional ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field should specify the DynamoDB table name: | `from` | Description | | ---------------- | --------------------------------------------- | | `dynamodb:table` | Read data from a DynamoDB table named `table` | note If an expected table is not found, verify the `dynamodb_aws_region` parameter. DynamoDB tables are region-specific. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: dynamodb:users name: my_users params: ... ``` ``` SELECT COUNT(*) FROM my_users; ``` The dataset name cannot be a \[reserved keyword(../../reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The DynamoDB data connector supports the following configuration parameters: | Parameter Name | Description | | -------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | | `dynamodb_aws_region` | Required. The AWS region containing the DynamoDB table | | `dynamodb_aws_access_key_id` | Optional. AWS access key ID for authentication. If not provided, credentials will be loaded from environment variables or IAM roles | | `dynamodb_aws_secret_access_key` | Optional. AWS secret access key for authentication. If not provided, credentials will be loaded from environment variables or IAM roles | | `dynamodb_aws_session_token` | Optional. AWS session token for authentication | | `unnest_depth` | Optional. Maximum nesting depth for unnesting embedded documents into a flattened structure. Higher values expand deeper nested fields. | | `schema_infer_max_records` | Optional. The number of documents to use to infer the schema. Defaults to 10 | | `scan_segments` | Optional. Number of segments for `Scan` request. 'auto' by default, which will calculate number of segments based on number of the records in a table | | `scan_interval` | Optional. Interval between polling for new records in a DynamoDB stream. Default: `0s`. See [Streams](#streams). | | `ready_lag` | Optional. When using Streams, once the table reaches this lag the dataset will be reported as Ready. Default: `2s`. See [Streams](#streams). | | `endpoint_url` | Optional. Custom endpoint URL for DynamoDB-compatible services (e.g., DynamoDB Local, ScyllaDB Alternator). | ### Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3 permissions (e.g., `s3:ListBucket`, `s3:GetObject`) to access the files. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access DynamoDB tables: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "dynamodb:Scan", "dynamodb:Query", "dynamodb:DescribeTable" ], "Resource": [ "arn:aws:dynamodb:*:*:table/YOUR_TABLE_NAME" ] } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------------ | ----------------------------------------------------------------- | | `dynamodb:Scan` | Required. Allows reading all items from the table | | `dynamodb:Query` | Required. Allows reading items from the table using partition key | | `dynamodb:DescribeTable` | Required. Allows fetching table metadata and schema information | ### Example IAM Policies[​](#example-iam-policies "Direct link to Example IAM Policies") #### Minimal Policy (Read-only access to specific table)[​](#minimal-policy-read-only-access-to-specific-table "Direct link to Minimal Policy (Read-only access to specific table)") ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "dynamodb:Scan", "dynamodb:Query", "dynamodb:DescribeTable" ], "Resource": "arn:aws:dynamodb:us-west-2:123456789012:table/users" } ] } ``` #### Access to Multiple Tables[​](#access-to-multiple-tables "Direct link to Access to Multiple Tables") ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "dynamodb:Scan", "dynamodb:Query", "dynamodb:DescribeTable" ], "Resource": [ "arn:aws:dynamodb:us-west-2:123456789012:table/users", "arn:aws:dynamodb:us-west-2:123456789012:table/orders" ] } ] } ``` #### Access to All Tables in a Region[​](#access-to-all-tables-in-a-region "Direct link to Access to All Tables in a Region") ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "dynamodb:Scan", "dynamodb:Query", "dynamodb:DescribeTable" ], "Resource": "arn:aws:dynamodb:us-west-2:123456789012:table/*" } ] } ``` Security Considerations * Avoid using `dynamodb:*` permissions as it grants more access than necessary. * Consider using more restrictive policies in production environments. * When using IAM roles with EKS, ensure the [service account is properly configured with IRSA](https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html). ## Data Types[​](#data-types "Direct link to Data Types") The table below shows the DynamoDB data types supported, along with the type mapping to Apache Arrow types in Spice. | DynamoDB Type | Description | Arrow Type | Notes | | ------------- | ----------- | ------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------- | | `Bool` | Boolean | `Boolean` | | | `S` | String | `Utf8` | | | `S` | String | `Timestamp(Millisecond)` | Naive timestamp if it matches `time_format` without timezone | | `S` | String | `Timestamp(Millisecond, )` | Timezone-aware timestamp if it matches `time_format` with timezone | | `S` | String | `Date32` | Date-only string in `YYYY-MM-DD` format (when it does not match `time_format`) | | `Ss` | String Set | `List` | | | `N` | Number | `Int64` \| `Float64` | | | `Ns` | Number Set | `List` | | | `B` | Binary | `Binary` | | | `Bs` | Binary Set | `List` | | | `L` | List | `List` | DynamoDB arrays can be heterogeneous e.g. `[1, "foo", true]`, Arrow arrays must be homogeneous - use strings to preserve all data | | `M` | Map | `Utf8` or Unflattened | Depending on `unnest_depth` value | ## Time format[​](#time-format "Direct link to Time format") Since DynamoDB stores timestamps as strings, Spice supports parsing timestamps using a customizable format. By default, Spice will try to parse timestamps using ISO8601 format, but you can provide a custom format using the `time_format` parameter. Once Spice is able to parse a timestamp, it will convert it to a `Timestamp(Millisecond)` Arrow type, and will use the same format to serialize it back to DynamoDB for filter pushdown. This parameter uses Go-style time formatting, which uses a reference time of `Mon Jan 2 15:04:05 MST 2006`. | Format Pattern | Example Value | Description | | ------------------------------- | ------------------------------- | ---------------------------------------------------------- | | `2006-01-02T15:04:05.000Z07:00` | `2024-03-15T14:30:00.000Z` | ISO8601 / RFC3339 with milliseconds and timezone (default) | | `2006-01-02T15:04:05.999Z07:00` | `2024-03-15T14:30:00.123-07:00` | ISO8601 with milliseconds and timezone | | `2006-01-02T15:04:05` | `2024-03-15T14:30:00` | ISO8601 without timezone (naive timestamp) | | `2006-01-02 15:04:05` | `2024-03-15 14:30:00` | Date and time with space separator | | `01/02/2006 15:04:05` | `03/15/2024 14:30:00` | US-style date with time | | `02/01/2006 15:04:05` | `15/03/2024 14:30:00` | European-style date with time | | `Jan 2, 2006 3:04:05 PM` | `Mar 15, 2024 2:30:00 PM` | Human-readable with 12-hour clock | | `20060102150405` | `20240315143000` | Compact format (no separators) | Go's format uses specific reference values that must appear exactly as shown: | Component | Reference Value | Alternatives | | ------------ | --------------- | ------------------------------------- | | Year | `2006` | `06` (2-digit) | | Month | `01` | `1`, `Jan`, `January` | | Day | `02` | `2` | | Hour (24h) | `15` | — | | Hour (12h) | `03` | `3` | | Minute | `04` | `4` | | Second | `05` | `5` | | AM/PM | `PM` | `pm` | | Timezone | `Z07:00` | `-0700`, `MST` | | Milliseconds | `.000` | `.999` (trailing zeros trimmed) | | Microseconds | `.000000` | `.999999` (trailing zeros trimmed) | | Nanoseconds | `.000000000` | `.999999999` (trailing zeros trimmed) | ::: ## Unnesting[​](#unnesting "Direct link to Unnesting") Consider the following document: ``` { "a": 1, "b": { "x": 2, "y": { "z": 3 } } } ``` Using `unnest_depth` you can control the unnesting behavior. Here are the examples: ### unnest\_depth: 0[​](#unnest_depth-0 "Direct link to unnest_depth: 0") ``` sql> select * from test_table; +-----------+---------------------+ | a (Int32) | b (Utf8) | +-----------+---------------------+ | 1 | {"x":2,"y":{"z":3}} | +---+-----------------------------+ ``` ### unnest\_depth: 1[​](#unnest_depth-1 "Direct link to unnest_depth: 1") ``` sql> select * from test_table; +-----------+-------------+------------+ | a (Int32) | b.x (Int32) | b.y (Utf8) | +-----------+-------------+------------+ | 1 | 2 | {"z":3} | +-----------+-------------+------------+ ``` ### unnest\_depth: 2[​](#unnest_depth-2 "Direct link to unnest_depth: 2") ``` sql> select * from test_table; +-----------+-------------+---------------+ | a (Int32) | b.x (Int32) | b.y.z (Int32) | +-----------+-------------+---------------+ | 1 | 2 | 3 | +-----------+-------------+---------------+ ``` ## Examples[​](#examples "Direct link to Examples") ### Basic Configuration with Environment Credentials[​](#basic-configuration-with-environment-credentials "Direct link to Basic Configuration with Environment Credentials") ``` version: v1 kind: Spicepod name: dynamodb datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 acceleration: enabled: true ``` ### Configuration with Explicit Credentials[​](#configuration-with-explicit-credentials "Direct link to Configuration with Explicit Credentials") ``` version: v1 kind: Spicepod name: dynamodb datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 dynamodb_aws_access_key_id: ${secrets:aws_access_key_id} dynamodb_aws_secret_access_key: ${secrets:aws_secret_access_key} acceleration: enabled: true ``` ### Configuration with time\_format[​](#configuration-with-time_format "Direct link to Configuration with time_format") ``` version: v1 kind: Spicepod name: dynamodb datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 time_format: 2006-01-02 15:04:05 acceleration: enabled: true ``` ### Querying Nested Structures[​](#querying-nested-structures "Direct link to Querying Nested Structures") DynamoDB supports complex nested JSON structures. These fields can be queried using SQL: ``` -- Query nested structs SELECT metadata.registration_ip, metadata.user_agent FROM users LIMIT 5; -- Query nested structs in arrays SELECT address.city FROM ( SELECT unnest(addresses) AS address FROM users ) WHERE address.city = 'San Francisco'; ``` Limitations * The DynamoDB connector will scan the first 10 items to determine the schema of the table. This may miss columns that are not present in the first 10 items. * The DynamoDB connector does not support Decimal type. Example schema from a users table: ``` describe users; ``` ``` +----------------+------------------+-------------+ | column_name | data_type | is_nullable | +----------------+------------------+-------------+ | email | Utf8 | YES | | id | Int64 | YES | | metadata | Struct | YES | | addresses | List(Struct) | YES | | preferences | Struct | YES | | created_at | Utf8 | YES | ... +----------------+------------------+-------------+ ``` ## Streams[​](#streams "Direct link to Streams") The DynamoDB Data Connector integrates with [DynamoDB Streams](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.html) to enable real-time streaming of table changes. This feature supports both initial table bootstrapping and continuous change data capture (CDC), allowing Spice to automatically detect and stream inserts, updates, and deletes from DynamoDB tables. warning Using DynamoDB Streams **requires** \[acceleration(../data-accelerators/index) with `refresh_mode: changes`. ### Basic Configuration[​](#basic-configuration "Direct link to Basic Configuration") To enable streaming from DynamoDB, enable acceleration and set the `refresh_mode` to `changes` in your dataset configuration. You also need to configure the `on_conflict` parameter to specify how the connector should handle updates to existing records. The keys defined in `on_conflict` must match your DynamoDB table's partition key and range key (if your table has one) ``` datasets: - from: dynamodb:my_table name: orders_stream acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes on_conflict: (id, version): upsert ``` ### Configuration Parameters[​](#configuration-parameters "Direct link to Configuration Parameters") #### Dataset Parameters[​](#dataset-parameters "Direct link to Dataset Parameters") * **`ready_lag`** - Defines the maximum lag threshold before the dataset is reported as "Ready". Once the stream lag falls below this value, queries can be executed against the dataset. Default behavior reports ready immediately after bootstrap completes. * **`scan_interval`** - Controls the polling frequency for checking new records in the DynamoDB stream. Lower values provide more real-time updates but increase API calls. Higher values reduce API usage but may introduce additional latency. #### Acceleration Parameters[​](#acceleration-parameters "Direct link to Acceleration Parameters") * **`on_conflict`** - Specifies the conflict resolution strategy when streaming changes that match existing records. The keys in the tuple should correspond to your DynamoDB table's partition key and range key (if applicable). The `upsert` action will insert new records or update existing ones based on these key columns. **Examples:** * Single partition key: `id: upsert` * Partition key + range key: `(partition_key, sort_key): upsert` * **`snapshots_trigger_threshold`** - Determines how frequently snapshots are created during streaming. A value of `5` means a snapshot is created every 5 batch updates. Snapshots enable faster recovery and better query performance but consume additional storage. ### Metrics[​](#metrics "Direct link to Metrics") The following [Component Metrics](/docs/v1.10/features/observability/component_metrics) are provided for monitoring streaming performance and health: | Metric | Type | Description | | ------------------------ | ------- | -------------------------------------------------------------------------- | | `shards_active` | Gauge | Current number of active shards in the stream | | `records_consumed_total` | Counter | Total number of records consumed from the stream | | `lag_ms` | Gauge | Current lag in milliseconds between stream watermark and the current time | | `errors_transient_total` | Counter | Total number of transient errors encountered while polling from the stream | These metrics are not enabled by default, enable them by setting the metrics parameter: ``` datasets: - from: kafka:user_events name: events metrics: - name: shards_active - name: lag_ms ``` You can find an example dashboard for DynamoDB Streams in [monitoring/grafana-dashboard.json](https://github.com/spiceai/spiceai/blob/trunk/monitoring/grafana-dashboard.json). ## Advanced Configuration[​](#advanced-configuration "Direct link to Advanced Configuration") For production workloads requiring fine-tuned control over streaming behavior and performance characteristics: ``` datasets: - from: dynamodb:my_table name: orders_stream params: ready_lag: 1s # Dataset reports as Ready when lag is below 1 second scan_interval: 100ms # Poll for new stream records every 100 milliseconds acceleration: enabled: true engine: duckdb mode: file refresh_mode: changes on_conflict: (id, version): upsert params: snapshots_trigger_threshold: 5 # Create snapshot every 5 batch updates metrics: - name: shards_active enabled: true - name: records_consumed_total enabled: true - name: lag_ms enabled: true - name: errors_transient_total enabled: true ``` ## Cookbooks[​](#cookbooks "Direct link to Cookbooks") * A cookbook recipe to configure DynamoDB as a data connector in Spice. [DynamoDB Data Connector](https://github.com/spiceai/cookbook/tree/trunk/dynamodb#readme) * A cookbook recipe to configure DynamoDB Streams as a data connector in Spice. [DynamoDB Streams Data Connector](https://github.com/spiceai/cookbook/tree/trunk/dynamodb/streams#readme) --- # File Data Connector The File Data Connector enables federated SQL queries on files stored by locally accessible filesystems. It supports querying individual files or entire directories, where all child files within the directory will be loaded and queried. File formats are specified using the `file_format` parameter, as described in [Object Store File Formats](/docs/v1.10/components/data-connectors/#object-store-file-formats). Example `spicepod.yml` ``` datasets: - from: file://path/to/customer.parquet name: customer params: file_format: parquet ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field for the File connector takes the form `file://path` where `path` is the path to the file to read from. See the [examples](#examples) below for examples of relative and absolute paths ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: file://path/to/customer.parquet name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter name | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `file_format` | Specifies the data file format. Required if the format cannot be inferred from the `from` path. Refer to [Object Store File Formats](/docs/v1.10/components/data-connectors/#object-store-file-formats) for details. | | `hive_partitioning_enabled` | Enable partitioning using hive-style partitioning from the folder structure. Defaults to `false` | | `schema_source_path` | Specifies the path used to infer the dataset schema. Default to the most recently modified file | For additional CSV, JSON, and Parquet specific parameters, see [File Formats](/docs/v1.10/reference/file_format). ## Trigger data refresh on file change[​](#trigger-data-refresh-on-file-change "Direct link to Trigger data refresh on file change") In addition to standard [Data Refresh](/docs/v1.10/features/data-acceleration/data-refresh), a data refresh can also be triggered when the source file is modified. The File Data Connector uses a file system watcher to be notified the file has changed. The file watcher is disabled by default and can be enabled by setting the `file_watcher` parameter to `enabled` in the acceleration parameters. ``` datasets: - from: file://path/to/my_file.csv name: my_file acceleration: enabled: true refresh_mode: full params: file_watcher: enabled ``` When the file is modified, the acceleration will be refreshed and will include the latest data. ## Types[​](#types "Direct link to Types") Refer to [Object Store Data Types](/docs/v1.10/reference/datatypes/object_store) for data type mapping from object store files to arrow data type. ## Examples[​](#examples "Direct link to Examples") ### Absolute path[​](#absolute-path "Direct link to Absolute path") In this example, `path` is an absolute path to the file on the filesystem. ``` datasets: - from: file:///path/to/customer.parquet name: customer params: file_format: parquet ``` ### Relative path[​](#relative-path "Direct link to Relative path") In this example, the path is relative to the directory where the `spicepod.yaml` is located. ``` ├── foo │   └── yellow_tripdata_2024-01.parquet └── spicepod.yaml ``` ``` datasets: - from: file://foo/yellow_tripdata_2024-01.parquet name: trip_data params: file_format: parquet ``` Performance Considerations When using the File Data connector without acceleration, data is loaded into memory during query execution. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. Memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/v1.10/components/data-accelerators/duckdb) and [`sqlite`](/docs/v1.10/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. ## Cookbook[​](#cookbook "Direct link to Cookbook") Refer to the [File cookbook recipe](https://github.com/spiceai/cookbook/tree/trunk/file) to see an example of the File connector in use. --- # Flight SQL Data Connector Connect to any Flight SQL compatible server (e.g. Influx 3.0, CnosDB, other Spice runtimes!) as a connector for federated SQL queries. ``` - from: flightsql:my_catalog.good_schemas.cool_dataset name: cool_dataset params: flightsql_endpoint: http://127.0.0.1:50051 flightsql_username: spicy flightsql_password: ${secrets:my_flightsql_pass} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `flightsql:dataset` where `dataset` is the fully qualified name of the dataset to read from. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter name | Description | | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `flightsql_endpoint` | Required. The Apache Flight endpoint used to connect to the Flight SQL server. | | `flightsql_username` | Optional. The username to use in the underlying Apache flight Handshake Request to authenticate to the server (see [reference](https://arrow.apache.org/docs/format/Flight.html#authentication)). | | `flightsql_password` | Optional. The password to use in the underlying Apache flight Handshake Request to authenticate to the server. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to load the password from a secret store, e.g. `${secrets:my_flightsql_pass}`. | ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). --- # FTP/SFTP Data Connector FTP (File Transfer Protocol) and SFTP (SSH File Transfer Protocol) are network protocols used for transferring files between a client and server, with FTP being less secure and SFTP providing encrypted file transfer over SSH. The FTP/SFTP Data Connector enables federated SQL query across [supported file formats](/docs/v1.10/components/data-connectors/#supported-formats) stored on FTP/SFTP servers. ## Quickstart[​](#quickstart "Direct link to Quickstart") Connect to an SFTP server and query CSV files: ``` datasets: - from: sftp://remote-sftp-server.com/path/to/folder/ name: my_dataset params: file_format: csv sftp_port: 22 sftp_user: my-sftp-user sftp_pass: ${secrets:my_sftp_password} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes one of two forms: `ftp:///` or `sftp:///` where `` is the host to connect to and `` is the path to the file or directory to read from. If a folder is provided, all child files will be loaded. ### `name`[​](#name "Direct link to name") The dataset name used as the table name in SQL queries. Cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") #### FTP[​](#ftp "Direct link to FTP") | Parameter Name | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------- | | `file_format` | Required when connecting to a directory. See [File Formats](/docs/v1.10/components/data-connectors/#supported-formats). | | `ftp_user` | Required. Username for FTP authentication. | | `ftp_pass` | Required. Password for FTP authentication. Use [secrets](/docs/v1.10/components/secret-stores) syntax: `${secrets:my_ftp_pass}`. | | `ftp_port` | FTP server port. Default: `21`. | | `client_timeout` | Connection timeout duration. E.g. `30s`, `1m`. No timeout when unset. | | `hive_partitioning_enabled` | Enable Hive-style partitioning from folder structure. Default: `false`. | #### SFTP[​](#sftp "Direct link to SFTP") | Parameter Name | Description | | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- | | `file_format` | Required when connecting to a directory. See [File Formats](/docs/v1.10/components/data-connectors/#supported-formats). | | `sftp_user` | Required. Username for SFTP authentication. | | `sftp_pass` | Required. Password for SFTP authentication. Use [secrets](/docs/v1.10/components/secret-stores) syntax: `${secrets:my_sftp_pass}`. | | `sftp_port` | SFTP server port. Default: `22`. | | `client_timeout` | Connection timeout duration. E.g. `30s`, `1m`. No timeout when unset. | | `hive_partitioning_enabled` | Enable Hive-style partitioning from folder structure. Default: `false`. | ## Examples[​](#examples "Direct link to Examples") ### Connecting to FTP[​](#connecting-to-ftp "Direct link to Connecting to FTP") ``` - from: ftp://remote-ftp-server.com/path/to/folder/ name: my_dataset params: file_format: csv ftp_user: my-ftp-user ftp_pass: ${secrets:my_ftp_password} hive_partitioning_enabled: false ``` ### Connecting to SFTP[​](#connecting-to-sftp "Direct link to Connecting to SFTP") ``` - from: sftp://remote-sftp-server.com/path/to/folder/ name: my_dataset params: file_format: csv sftp_port: 22 sftp_user: my-sftp-user sftp_pass: ${secrets:my_sftp_password} hive_partitioning_enabled: false ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores for secure credential management. Store FTP/SFTP credentials in a secret store and reference them using the `${secrets:key}` syntax. ``` datasets: - from: sftp://files.example.com/data/ name: secure_data params: file_format: parquet sftp_user: ${secrets:sftp_username} sftp_pass: ${secrets:sftp_password} ``` For detailed information, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") ### Connection Timeouts[​](#connection-timeouts "Direct link to Connection Timeouts") If connections frequently timeout, increase the `client_timeout` value: ``` params: client_timeout: 120s ``` ### Authentication Failures[​](#authentication-failures "Direct link to Authentication Failures") Verify credentials are correctly stored in your secret store and that the user has read access to the specified path on the server. ### File Format Errors[​](#file-format-errors "Direct link to File Format Errors") When connecting to a directory, ensure `file_format` is specified and matches the actual file types in the directory. Spice expects all files in a directory to have the same format. ## Cookbook[​](#cookbook "Direct link to Cookbook") Refer to the [FTP cookbook recipe](https://github.com/spiceai/cookbook/tree/trunk/ftp) to see an example of the FTP connector in use. --- # GitHub Data Connector The GitHub Data Connector enables federated SQL queries on various GitHub resources such as files, issues, pull requests, and commits by specifying `github` as the selector in the `from` value for the dataset. ## Common Configuration[​](#common-configuration "Direct link to Common Configuration") ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field specifies the GitHub resource to query. It supports the following formats: | Format | Description | | ---------------------------------------------- | --------------------------------------------------------- | | `github:github.com///files/` | Query files from a repository at a specific branch or tag | | `github:github.com///issues` | Query issues from a repository | | `github:github.com///pulls` | Query pull requests from a repository | | `github:github.com///commits` | Query commits from a repository | | `github:github.com///stargazers` | Query stargazers from a repository | | `github:github.com//members` | Query members from an organization | ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. The dataset name cannot be a \[reserved keyword(../../reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") #### Personal Access Token[​](#personal-access-token "Direct link to Personal Access Token") | Parameter Name | Description | | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `github_token` | Required. GitHub personal access token to use to connect to the GitHub API. [Learn more](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens). | #### GitHub App Installation[​](#github-app-installation "Direct link to GitHub App Installation") GitHub Apps provide a secure and scalable way to integrate with GitHub's API, and works well when interacting with one or more GitHub organizations. [Learn more](https://docs.github.com/en/apps). | Parameter Name | Description | | ------------------------ | ------------------------------------------------------------------------------ | | `github_client_id` | Required. Specifies the client ID for GitHub App Installation auth mode. | | `github_private_key` | Required. Specifies the private key for GitHub App Installation auth mode. | | `github_installation_id` | Required. Specifies the installation ID for GitHub App Installation auth mode. | The client ID and private key are generated when creating the GitHub app. **Getting the Installation ID** If the app is installed on a GitHub organization: * Visit the settings page for the organization (`https://github.com/organizations//settings/installations`) * Click "Configure" on the app * The URL of the page will be of the form `https://github.com/organizations//settings/installations/` If the app is installed on a GitHub user: * Visit [the settings page](https://github.com/settings/installations) * Click "Configure" on the app * The URL of the page will be of the form `https://github.com/settings/installations/` Limitations With GitHub App Installation authentication, the connector's functionality depends on the permissions and scope of the GitHub App. Ensure that the app is installed on the repositories and configured with content, commits, issues and pull permissions to allow the corresponding datasets to work. #### Common Parameters[​](#common-parameters "Direct link to Common Parameters") | Parameter Name | Description | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `github_query_mode` | Optional. Specifies whether the connector should use the GitHub [search API](https://docs.github.com/en/graphql/reference/queries#search) for improved filter performance. Defaults to `auto`, possible values of `auto` or `search`. | | `owner` | Required. Specifies the owner of the GitHub repository. | | `repo` | Required. Specifies the name of the GitHub repository. | ## Advanced Configuration[​](#advanced-configuration "Direct link to Advanced Configuration") When using multiple GitHub datasets sharing the same GitHub token or GitHub app credentials, it is possible to exceed GitHub's primary and secondary rate limits. To mitigate this, use the `github_max_concurrent_connections` runtime parameter. This connections limit applies per GitHub token and per GitHub app installation, following GitHub's rate limit policy. Example Configuration: ``` # ... other configuration ... runtime: params: github_max_concurrent_connections: 5 # Defaults to 10 datasets: - from: github:github.com/spiceai/spiceai/files/v0.17.2-beta name: spiceai.files params: github_token: ${secrets:GITHUB_TOKEN} include: '**/*.txt' acceleration: enabled: true - from: github:github.com///issues name: spiceai.issues params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true # ... other configuration ... ``` ## Filter Push Down[​](#filter-push-down "Direct link to Filter Push Down") GitHub queries support a `github_query_mode` parameter, which can be set to either `auto` or `search` for the following types: * **Issues**: Defaults to `auto`. Query filters are only pushed down to the GitHub API in `search` mode. * **Pull Requests**: Defaults to `auto`. Query filters are only pushed down to the GitHub API in `search` mode. Commits only supports `auto` mode. Query with filter push down is only enabled for the `committed_date` column. `committed_date` supports exact matches, or greater/less than matches for dates provided in [ISO8601](https://www.iso.org/iso-8601-date-and-time-format.html) format, like `WHERE committed_date > '2024-09-24'`. When set to `search`, Issues and Pull Requests will use the GitHub [Search API](https://docs.github.com/en/search-github/searching-on-github/searching-issues-and-pull-requests) for improved filter performance when querying against the columns: * `author` and `state`; supports exact matches, or NOT matches. For example, `WHERE author = 'peasee'` or `WHERE author <> 'peasee'`. * `body` and `title`; supports exact matches, or LIKE matches. For example, `WHERE body LIKE '%duckdb%'`. * `updated_at`, `created_at`, `merged_at` and `closed_at`; supports exact matches, or greater/less than matches with dates provided in [ISO8601](https://www.iso.org/iso-8601-date-and-time-format.html) format. For example, `WHERE created_at > '2024-09-24'`. All other filters are supported when `github_query_mode` is set to `search`, but cannot be pushed down to the GitHub API for improved performance. Limitations * GitHub has a limitation in the Search API where it may return more stale data than the standard API used in the default query mode. * GitHub has a limitation in the Search API where it only returns a maximum of 1000 results for a query. Use [append mode acceleration](/docs/v1.10/features/data-acceleration/data-refresh) to retrieve more results over time. See the [append example](#append-example) for pull requests. ## Examples[​](#examples "Direct link to Examples") ### Querying GitHub Files[​](#querying-github-files "Direct link to Querying GitHub Files") Limitations * `content` column is fetched only when acceleration is enabled. * Querying GitHub files does not support filter push down, which may result in long query times when acceleration is disabled. * Setting `github_query_mode` to `search` is not supported. - `ref` - Required. Specifies the GitHub branch or tag to fetch files from. - `include` - Optional. Specifies a pattern to include specific files. Supports glob patterns. If not specified, all files are included by default. ``` datasets: - from: github:github.com///files/ name: spiceai.files params: github_token: ${secrets:GITHUB_TOKEN} include: '**/*.json; **/*.yaml' acceleration: enabled: true ``` #### Schema[​](#schema "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ------------- | --------- | ----------- | | name | Utf8 | YES | | path | Utf8 | YES | | size | Int64 | YES | | sha | Utf8 | YES | | mode | Utf8 | YES | | url | Utf8 | YES | | download\_url | Utf8 | YES | | content | Utf8 | YES | #### Example[​](#example "Direct link to Example") ``` datasets: - from: github:github.com/spiceai/spiceai/files/v0.17.2-beta name: spiceai.files params: github_token: ${secrets:GITHUB_TOKEN} include: '**/*.txt' # include txt files only acceleration: enabled: true ``` ``` sql> select * from spiceai.files +-------------+-------------+------+------------------------------------------+--------+-------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------+-------------+ | name | path | size | sha | mode | url | download_url | content | +-------------+-------------+------+------------------------------------------+--------+-------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------+-------------+ | version.txt | version.txt | 12 | ee80f747038c30e776eecb2c2ae155dec9a68187 | 100644 | https://api.github.com/repos/spiceai/spiceai/git/blobs/ee80f747038c30e776eecb2c2ae155dec9a68187 | https://raw.githubusercontent.com/spiceai/spiceai/v0.17.2-beta/version.txt | 0.17.2-beta | | | | | | | | | | +-------------+-------------+------+------------------------------------------+--------+-------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------+-------------+ Time: 0.005067 seconds. 1 rows. ``` ### Querying GitHub Issues[​](#querying-github-issues "Direct link to Querying GitHub Issues") Limitations * Querying with filters using date columns requires the use of [ISO8601 formatted dates](https://www.iso.org/iso-8601-date-and-time-format.html). For example, `WHERE created_at > '2024-09-24'`. ``` datasets: - from: github:github.com///issues name: spiceai.issues params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true ``` #### Schema[​](#schema-1 "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ---------------- | ------------ | ----------- | | assignees | List(Utf8) | YES | | author | Utf8 | YES | | body | Utf8 | YES | | closed\_at | Timestamp | YES | | comments | List(Struct) | YES | | created\_at | Timestamp | YES | | id | Utf8 | YES | | labels | List(Utf8) | YES | | milestone\_id | Utf8 | YES | | milestone\_title | Utf8 | YES | | comments\_count | Int64 | YES | | number | Int64 | YES | | state | Utf8 | YES | | title | Utf8 | YES | | updated\_at | Timestamp | YES | | url | Utf8 | YES | #### Example[​](#example-1 "Direct link to Example") ``` datasets: - from: github:github.com/spiceai/spiceai/issues name: spiceai.issues params: github_token: ${secrets:GITHUB_TOKEN} ``` ``` sql> select title, state, labels from spiceai.issues where title like '%duckdb%' +-----------------------------------------------------------------------------------------------------------+--------+----------------------+ | title | state | labels | +-----------------------------------------------------------------------------------------------------------+--------+----------------------+ | Limitation documentation duckdb accelerator about nested struct and decimal256 | CLOSED | [kind/documentation] | | Inconsistent duckdb connector params: `params.open` and `params.duckdb_file` | CLOSED | [kind/bug] | | federation across multiple duckdb acceleration tables. | CLOSED | [] | | Integration tests to cover "On Conflict" behaviors for duckdb accelerator | CLOSED | [kind/task] | | Permission denied issue while using duckdb data connector with spice using HELM for Kubernetes deployment | CLOSED | [kind/bug] | +-----------------------------------------------------------------------------------------------------------+--------+----------------------+ Time: 0.011877542 seconds. 5 rows. ``` ### Querying GitHub Pull Requests[​](#querying-github-pull-requests "Direct link to Querying GitHub Pull Requests") Limitations * Querying with filters using date columns requires the use of [ISO8601 formatted dates](https://www.iso.org/iso-8601-date-and-time-format.html). For example, `WHERE created_at > '2024-09-24'`. ``` datasets: - from: github:github.com///pulls name: spiceai.pulls params: github_token: ${secrets:GITHUB_TOKEN} # Specifies the types of comments to fetch: 'all', 'review', 'discussion', or 'none'. Defaults to 'none'. github_include_comments: none # Number of comments to fetch per discussion or review thread. # Defaults to 100, and is capped at 100 github_max_comments_fetched: 100 ``` #### Schema[​](#schema-2 "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ---------------- | -------------------------------------------------------------- | ----------- | | additions | Int64 | YES | | assignees | List(Utf8) | YES | | author | Utf8 | YES | | body | Utf8 | YES | | changed\_files | Int64 | YES | | closed\_at | Timestamp | YES | | comments\_count | Int64 | YES | | commits\_count | Int64 | YES | | created\_at | Timestamp | YES | | deletions | Int64 | YES | | discussion | List(Struct(body: Utf8, author: Utf8, created\_at: Timestamp)) | YES | | hashes | List(Utf8) | YES | | id | Utf8 | YES | | labels | List(Utf8) | YES | | merged\_at | Timestamp | YES | | number | Int64 | YES | | review\_comments | List(Struct(body: Utf8, author: Utf8, created\_at: Timestamp)) | YES | | reviews\_count | Int64 | YES | | state | Utf8 | YES | | title | Utf8 | YES | | url | Utf8 | YES | **Note**: The `discussion` and `review_comments` columns are only included in the schema when the `github_include_comments` parameter is set accordingly. #### Example[​](#example-2 "Direct link to Example") ``` datasets: - from: github:github.com/spiceai/spiceai/pulls name: spiceai.pulls params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true ``` ``` sql> select title, url, state from spiceai.pulls where title like '%GitHub connector%' +---------------------------------------------------------------------+----------------------------------------------+--------+ | title | url | state | +---------------------------------------------------------------------+----------------------------------------------+--------+ | GitHub connector: convert `labels` and `hashes` to primitive arrays | https://github.com/spiceai/spiceai/pull/2452 | MERGED | +---------------------------------------------------------------------+----------------------------------------------+--------+ Time: 0.034996667 seconds. 1 rows. ``` #### Append Example[​](#append-example "Direct link to Append Example") ``` datasets: - from: github:github.com/spiceai/spiceai/pulls name: spiceai.pulls params: github_token: ${secrets:GITHUB_TOKEN} github_query_mode: search time_column: created_at acceleration: enabled: true refresh_mode: append refresh_check_interval: 6h # check for new results every 6 hours refresh_data_window: 90d # at initial load, load the last 90 days of pulls ``` #### Comments Example[​](#comments-example "Direct link to Comments Example") ``` datasets: - from: github:github.com/spiceai/spiceai/pulls name: spiceai.pulls params: github_token: ${secrets:GITHUB_TOKEN} github_include_comments: all github_max_comments_fetched: 100 acceleration: enabled: true ``` ``` sql> select unnest(unnest(review_comments)) from spiceai.pulls where number = 6 limit 1; +------------------------------------------------------------------+------------------------------------------------------------------------+--------------------------------------------------------------------+ | __unnest_placeholder(UNNEST(spiceai.pulls.review_comments)).body | __unnest_placeholder(UNNEST(spiceai.pulls.review_comments)).created_at | __unnest_placeholder(UNNEST(spiceai.pulls.review_comments)).author | +------------------------------------------------------------------+------------------------------------------------------------------------+--------------------------------------------------------------------+ | Nitpick - extra space. | 2021-08-11T17:36:23 | haardvark | +------------------------------------------------------------------+------------------------------------------------------------------------+--------------------------------------------------------------------+ Time: 0.034283334 seconds. 1 rows. ``` ``` sql> select unnest(unnest(discussion)) from spiceai.pulls where number = 148 limit 1; +-------------------------------------------------------------+-------------------------------------------------------------------+---------------------------------------------------------------+ | __unnest_placeholder(UNNEST(spiceai.pulls.discussion)).body | __unnest_placeholder(UNNEST(spiceai.pulls.discussion)).created_at | __unnest_placeholder(UNNEST(spiceai.pulls.discussion)).author | +-------------------------------------------------------------+-------------------------------------------------------------------+---------------------------------------------------------------+ | Do not merge until after repo goes public. | 2021-09-06T08:00:45 | lukekim | +-------------------------------------------------------------+-------------------------------------------------------------------+---------------------------------------------------------------+ Time: 0.036530584 seconds. 1 rows. ``` ### Querying GitHub Commits[​](#querying-github-commits "Direct link to Querying GitHub Commits") Limitations * Querying with filters using date columns requires the use of [ISO8601 formatted dates](https://www.iso.org/iso-8601-date-and-time-format.html). For example, `WHERE committed_date > '2024-09-24'`. * Setting `github_query_mode` to `search` is not supported. ``` datasets: - from: github:github.com///commits name: spiceai.commits params: github_token: ${secrets:GITHUB_TOKEN} ``` #### Schema[​](#schema-3 "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ------------------- | --------- | ----------- | | additions | Int64 | YES | | author\_email | Utf8 | YES | | author\_name | Utf8 | YES | | committed\_date | Timestamp | YES | | deletions | Int64 | YES | | id | Utf8 | YES | | message | Utf8 | YES | | message\_body | Utf8 | YES | | message\_head\_line | Utf8 | YES | | sha | Utf8 | YES | #### Example[​](#example-3 "Direct link to Example") ``` datasets: - from: github:github.com/spiceai/spiceai/commits name: spiceai.commits params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true ``` ``` sql> select sha, message_head_line from spiceai.commits limit 10 +------------------------------------------+------------------------------------------------------------------------+ | sha | message_head_line | +------------------------------------------+------------------------------------------------------------------------+ | 2a9fab7905737e1af182e17f40aecc5c4b5dd236 | wait 2 seconds for the status to turn ready in refreshing status tes… | | b9c210a818abeaf14d2493fde5227781f47faed8 | Update README.md - Remove bigquery from tablet of connectors (#1434) | | d61e1af61ebf826f83703b8dd939f19e8b2ba426 | Add databricks_use_ssl parameter (#1406) | | f1ec55c5986e3e5d57eff94197182ffebbae1045 | wording and logs change reflected on readme (#1435) | | bfc74185584d1e048ef66c72ce3572a0b652bfd9 | Update acknowledgements (#1433) | | 0d870f1791d456e7924b4ecbbda5f3b762db1e32 | Update helm version and use v0.13.0-alpha (#1436) | | 12f930cbad69833077bd97ea43599a75cff985fc | Enable push-down federation by default (#1429) | | 6e4521090aaf39664bd61d245581d34398ce77db | Add functional tests for federation push-down (#1428) | | fa3279b7d9fcaa5e8baaa2425f69b556bb30e309 | Add LRU cache support for http-based sql queries (#1410) | | a3f93dde9d1312bfbf14f7ae3b75bdc468289212 | Add guides and examples about error handling (#1427) | +------------------------------------------+------------------------------------------------------------------------+ Time: 0.0065395 seconds. 10 rows. ``` ### Querying GitHub stars (Stargazers)[​](#querying-github-stars-stargazers "Direct link to Querying GitHub stars (Stargazers)") Limitations * Querying with filters using date columns requires the use of [ISO8601 formatted dates](https://www.iso.org/iso-8601-date-and-time-format.html). For example, `WHERE starred_at > '2024-09-24'`. * Setting `github_query_mode` to `search` is not supported. ``` datasets: - from: github:github.com///stargazers name: spiceai.stargazers params: github_token: ${secrets:GITHUB_TOKEN} ``` #### Schema[​](#schema-4 "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ----------- | --------- | ----------- | | starred\_at | Timestamp | YES | | login | Utf8 | YES | | email | Utf8 | YES | | name | Utf8 | YES | | company | Utf8 | YES | | x\_username | Utf8 | YES | | location | Utf8 | YES | | avatar\_url | Utf8 | YES | | bio | Utf8 | YES | #### Example[​](#example-4 "Direct link to Example") ``` datasets: - from: github:github.com/spiceai/spiceai/stargazers name: spiceai.stargazers params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true ``` ``` sql> select starred_at, login from spiceai.stargazers order by starred_at DESC limit 10 +----------------------+----------------------+ | starred_at | login | +----------------------+----------------------+ | 2024-09-15T13:22:09Z | cisen | | 2024-09-14T18:04:22Z | tyan-boot | | 2024-09-13T10:38:01Z | yofriadi | | 2024-09-13T10:01:33Z | FourSpaces | | 2024-09-13T04:02:11Z | d4x1 | | 2024-09-11T18:10:28Z | stephenakearns-insta | | 2024-09-09T22:17:42Z | Lrs121 | | 2024-09-09T19:56:26Z | jonathanfinley | | 2024-09-09T07:02:10Z | leookun | | 2024-09-09T03:04:27Z | royswale | +----------------------+----------------------+ Time: 0.0088075 seconds. 10 rows. ``` ### Querying Members of a GitHub Organization[​](#querying-members-of-a-github-organization "Direct link to Querying Members of a GitHub Organization") Limitations * Querying with filters using date columns requires the use of [ISO8601 formatted dates](https://www.iso.org/iso-8601-date-and-time-format.html). For example, `WHERE created_at > '2024-09-24'`. * Setting `github_query_mode` to `search` is not supported. ``` datasets: - from: github:github.com//members name: members params: github_token: ${secrets:GITHUB_TOKEN} ``` #### Schema[​](#schema-5 "Direct link to Schema") | Column Name | Data Type | Is Nullable | | ----------- | --------- | ----------- | | username | Utf8 | YES | | name | Utf8 | YES | | avatar\_url | Utf8 | YES | | url | Utf8 | YES | | email | Utf8 | YES | | location | Utf8 | YES | | company | Utf8 | YES | | created\_at | Timestamp | YES | | bio | Utf8 | YES | #### Example[​](#example-5 "Direct link to Example") ``` datasets: - from: github:github.com/apache/members name: apache.members params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true ``` ``` sql> select created_at, username from apache.members order by created_at desc limit 10; +---------------------+-------------------+ | created_at | username | +---------------------+-------------------+ | 2023-10-09T13:14:13 | heliang666s | | 2023-04-14T11:26:44 | cortlepp | | 2023-02-16T08:28:58 | ChengJie1053 | | 2023-02-11T03:51:52 | FinalT | | 2022-11-20T12:12:56 | Yanshuming1 | | 2022-10-10T23:29:29 | bernardodemarco | | 2022-10-07T05:06:37 | coldgust | | 2022-09-06T14:38:44 | No-SilverBullet | | 2022-08-18T13:31:44 | harshithasudhakar | | 2022-07-05T10:44:08 | bearslyricattack | +---------------------+-------------------+ Time: 0.054390375 seconds. 10 rows. ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Github as a data connector in Spice. [GitHub Data Connector](https://github.com/spiceai/cookbook/tree/trunk/github#readme) --- # Glue Data Connector The Glue Data Connector enables federated SQL querying on tables in an AWS Glue Data Catalog. ``` datasets: - from: glue:tpch.lineitem name: lineitem params: glue_region: us-east-1 glue_key: ${env:SPICE_AWS_KEY} # Optional. glue_secret: ${env:SPICE_AWS_SECRET} # Optional. ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Specify a table using the format, `glue:.
` by replacing `` with the name of the Glue database and `
`with the name of the table inside of the ``. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` SELECT COUNT(*) FROM lineitem; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a \[reserved keyword(../../reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The following parameters are supported for configuring the connection to the Glue Data Catalog: | Parameter Name | Definition | | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `glue_region` | The AWS region for the Glue Data Catalog. E.g. `us-west-2`. | | `glue_catalog_id` | The Glue catalog ID. For Amazon S3 Tables, use the format `:s3tablescatalog/`. If not provided, the default catalog for the account is used. | | `glue_key` | Access key (e.g. AWS\_ACCESS\_KEY\_ID for AWS). If not provided, credentials will be loaded from environment variables or IAM roles. | | `glue_secret` | Secret key (e.g. AWS\_SECRET\_ACCESS\_KEY for AWS). If not provided, credentials will be loaded from environment variables or IAM roles. | | `glue_session_token` | Session token (e.g. AWS\_SESSION\_TOKEN for AWS) for temporary credentials | ## Examples[​](#examples "Direct link to Examples") ### Basic Glue Table[​](#basic-glue-table "Direct link to Basic Glue Table") ``` datasets: - from: glue:tpch.lineitem name: lineitem params: glue_region: us-east-1 glue_key: ${env:AWS_ACCESS_KEY_ID} glue_secret: ${env:AWS_SECRET_ACCESS_KEY} ``` ### Amazon S3 Tables[​](#amazon-s3-tables "Direct link to Amazon S3 Tables") Connect to tables in [Amazon S3 Tables](https://aws.amazon.com/s3/features/tables/) using the `glue_catalog_id` parameter with the S3 Tables catalog format: ``` datasets: - from: glue:my_namespace.orders name: orders params: glue_catalog_id: 123635965758:s3tablescatalog/my-table-bucket glue_region: us-east-2 glue_key: ${env:AWS_ACCESS_KEY_ID} glue_secret: ${env:AWS_SECRET_ACCESS_KEY} ``` ## Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. These credentials will be used to connect to the S3 bucket as well as the Glue catalog. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3/Glue permissions (e.g., `s3:ListBucket`, `glue:GetTable`) to access the tables. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ### Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access Iceberg tables in S3/Glue: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::company-bucketname-datasets" }, { "Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::company-bucketname-datasets/*" }, { "Effect": "Allow", "Action": [ "glue:GetCatalog", "glue:GetDatabases", "glue:GetDatabase", "glue:GetTable", "glue:GetTables" ], "Resource": "*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------- | -------------------------------------------------------------- | | `s3:ListBucket` | Required. Allows scanning all objects from the bucket | | `s3:GetObject` | Required. Allows fetching objects | | `glue:GetCatalog` | Required. Retrieve metadata about the specified catalog. | | `glue:GetDatabases` | Required. List the databases available in the current catalog. | | `glue:GetDatabase` | Required. Retrieve metadata about the specified database. | | `glue:GetTable` | Required. Retrieve metadata about the specified table. | | `glue:GetTables` | Required. List the tables available in the current database. | ## Limitations[​](#limitations "Direct link to Limitations") Data Source/Data Format Restrictions This catalog connector is limited to tables that use the S3 data source. Kinesis and Kafka data sources are not currently supported. Additionally, this catalog connector is currently limited to Iceberg tables, tables with parquet or CSV data format only. Performance Considerations When using the Glue Data connector without acceleration, data is loaded into memory during query execution. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. Memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/v1.10/components/data-accelerators/duckdb) and [`sqlite`](/docs/v1.10/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. Each query retrieves data from the S3 source, which might result in significant network requests and bandwidth consumption. This can affect network performance and incur costs related to data transfer from S3. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Glue as a data connector in Spice. [Glue Data Connector](https://github.com/spiceai/cookbook/tree/trunk/glue#readme) --- # GraphQL Data Connector The [GraphQL](https://graphql.org/) Data Connector enables federated SQL queries on any GraphQL endpoint by specifying `graphql` as the selector in the `from` value for the dataset. ``` datasets: - from: graphql:your-graphql-endpoint name: my_dataset params: json_pointer: /data/some/nodes graphql_query: | { some { nodes { field1 field2 } } } ``` Limitations * The GraphQL data connector does not support variables in the query. * Filter pushdown, with the exclusion of `LIMIT`, is not currently supported. Using a `LIMIT` will reduce the amount of data requested from the GraphQL server. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form of `graphql:your-graphql-endpoint`. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The GraphQL data connector can be configured by providing the following `params`. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to load the password from a secret store, e.g. `${secrets:my_graphql_auth_token}`. | Parameter Name | Description | | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `unnest_depth` | Depth level to automatically unnest objects to. By default, disabled if unspecified or `0`. | | `graphql_auth_token` | The authentication token to use to connect to the GraphQL server. Uses bearer authentication. | | `graphql_auth_user` | The username to use for basic auth. E.g. `graphql_auth_user: my_user` | | `graphql_auth_pass` | The password to use for basic auth. E.g. `graphql_auth_pass: ${secrets:my_graphql_auth_pass}` | | `graphql_query` | The GraphQL query to execute. See [examples](#examples) for a sample GraphQL query | | `json_pointer` | The [JSON pointer](https://datatracker.ietf.org/doc/html/rfc6901) into the response body. When `graphql_query` is [paginated](#pagination), the `json_pointer` can be inferred. | #### GraphQL Query Example[​](#graphql-query-example "Direct link to GraphQL Query Example") ``` graphql_query: | { some { nodes { field1 field2 } } } ``` ### Examples[​](#examples "Direct link to Examples") Example using the GitHub GraphQL API and Bearer Auth. The following will use `json_pointer` to retrieve all of the nodes in starredRepositories: ``` from: graphql:https://api.github.com/graphql name: stars params: graphql_auth_token: ${env:GITHUB_TOKEN} graphql_auth_user: ${env:GRAPHQL_USER} ... graphql_auth_pass: ${env:GRAPHQL_PASS} json_pointer: /data/viewer/starredRepositories/nodes graphql_query: | { viewer { starredRepositories { nodes { name stargazerCount languages (first: 10) { nodes { name } } } } } } ``` ## Pagination[​](#pagination "Direct link to Pagination") The GraphQL Data Connector supports automatic pagination of the response for queries using [cursor pagination](https://graphql.org/learn/pagination/). The `graphql_query` must include the `pageInfo` field as per [spec](https://relay.dev/graphql/connections.htm#sec-undefined.PageInfo). The connector will parse the `graphql_query`, and when `pageInfo` is present, will retrieve data until pagination completes. The query must have the correct pagination arguments in the associated paginated field. ### Example[​](#example "Direct link to Example") **Forward Pagination:** ``` { something_paginated(first: 100) { nodes { foo bar } pageInfo { endCursor hasNextPage } } } ``` **Backward Pagination:** ``` { something_paginated(last: 100) { nodes { foo bar } pageInfo { startCursor hasPreviousPage } } } ``` ## Working with JSON Data[​](#working-with-json-data "Direct link to Working with JSON Data") Tips for working with JSON data. For more information see [Datafusion Docs](https://datafusion.apache.org/user-guide/sql/scalar_functions.html#array-functions). ### Accessing objects fields[​](#accessing-objects-fields "Direct link to Accessing objects fields") You can access the fields of the object using the square bracket notation. Arrays are indexed from 1. Example for the stargazers query from [pagination section](#pagination): ``` sql> select node['login'] as login, node['name'] as name from stargazers limit 5; +--------------+----------------------+ | login | name | +--------------+----------------------+ | simsieg | Simon Siegert | | davidmathers | David Mathers | | ahmedtadde | Ahmed Tadde | | lordhamlet | Shih-Fen Cheng | | thinmy | Thinmy Patrick Alves | +--------------+----------------------+ ``` ### Piping array into rows[​](#piping-array-into-rows "Direct link to Piping array into rows") You can use Datafusion `unnest` function to pipe values from array into rows. We'll be using [countries GraphQL api](https://countries.trevorblades.com) as an example. ``` from: graphql:https://countries.trevorblades.com name: countries params: json_pointer: /data/continents graphql_query: | { continents { name countries { name capital } } } description: countries acceleration: enabled: true refresh_mode: full refresh_check_interval: 30m ``` Example query: ``` sql> select continent, country['name'] as country, country['capital'] as capital from (select name as continent, unnest(countries) as country from countries) where continent = 'North America' limit 5; +---------------+---------------------+--------------+ | continent | country | capital | +---------------+---------------------+--------------+ | North America | Antigua and Barbuda | Saint John's | | North America | Anguilla | The Valley | | North America | Aruba | Oranjestad | | North America | Barbados | Bridgetown | | North America | Saint Barthélemy | Gustavia | +---------------+---------------------+--------------+ ``` ### Unnesting object properties[​](#unnesting-object-properties "Direct link to Unnesting object properties") You can also use the `unnest_depth` parameter to control automatic unnesting of objects from GraphQL responses. This examples uses the GitHub stargazers endpoint: ``` from: graphql:https://api.github.com/graphql name: stargazers params: graphql_auth_token: ${env:GITHUB_TOKEN} unnest_depth: 2 json_pointer: /data/repository/stargazers/edges graphql_query: | { repository(name: "spiceai", owner: "spiceai") { id name stargazers(first: 100) { edges { node { id name login } } pageInfo { hasNextPage endCursor } } } } ``` If `unnest_depth` is set to 0, or unspecified, object unnesting is disabled. When enabled, unnesting automatically moves nested fields to the parent level. Without unnesting, stargazers data looks like this in a query: ``` sql> select node from stargazers limit 1; +------------------------------------------------------------+ | node | +------------------------------------------------------------+ | {id: MDQ6VXNlcjcwNzIw, login: ashtom, name: Thomas Dohmke} | +------------------------------------------------------------+ ``` With unnesting, these properties are automatically placed into their own columns: ``` sql> select node from stargazers limit 1; +------------------+--------+---------------+ | id | login | name | +------------------+--------+---------------+ | MDQ6VXNlcjcwNzIw | ashtom | Thomas Dohmke | +------------------+--------+---------------+ ``` #### Unnesting Duplicate Columns[​](#unnesting-duplicate-columns "Direct link to Unnesting Duplicate Columns") By default, the Spice Runtime will error when a duplicate column is detected during unnesting. For example, this example `spicepod.yml` query would fail due to `name` fields: ``` from: graphql:https://localhost name: stargazers params: unnest_depth: 2 json_pointer: /data/users graphql_query: | query { users { name emergency_contact { name } } } ``` This example would fail with a runtime error: ``` WARN runtime: Invalid object access. Column 'name' already exists in the object. ``` Avoid this error by [using aliases in the query](https://www.apollographql.com/docs/kotlin/advanced/using-aliases/) where possible. In the example above, a duplicate error was introduced from `emergency_contact { name }`. The example below uses a GraphQL alias to rename `emergency_contact.name` as `emergencyContactName`. ``` from: graphql:https://localhost name: stargazers params: unnest_depth: 2 json_pointer: /data/people graphql_query: | query { users { name emergency_contact { emergencyContactName: name } } } ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure GraphQL as a data connector in Spice. [GraphQL Data Connector](https://github.com/spiceai/cookbook/tree/trunk/graphql#readme) --- # HTTP(s) Data Connector The HTTP(s) Data Connector enables federated SQL query across [supported file formats](/docs/v1.10/components/data-connectors/#object-store-file-formats) stored at an HTTP(s) endpoint. The connector supports dynamic query and data refresh through SQL-based filtering. ``` datasets: - from: http://static_username@localhost:3001/report.csv name: local_report params: http_password: ${env:MY_HTTP_PASS} ``` ## Examples[​](#examples "Direct link to Examples") ### Basic Example[​](#basic-example "Direct link to Basic Example") ``` datasets: - from: https://github.com/LAION-AI/audio-dataset/raw/7fd6ae3cfd7cde619f6bed817da7aa2202a5bc28/metadata/freesound/parquet/freesound_parquet.parquet name: laion_freesound ``` ### Using Basic Authentication[​](#using-basic-authentication "Direct link to Using Basic Authentication") ``` datasets: - from: http://static_username@localhost:3001/report.csv name: local_report params: http_password: ${env:MY_HTTP_PASS} ``` ### Using Custom Headers[​](#using-custom-headers "Direct link to Using Custom Headers") Custom HTTP headers can be specified for authentication, API keys, or other requirements. Headers are treated as sensitive data and will not be logged. `http_headers` applies to **dynamic JSON API endpoints only**. A structured HTTP file dataset — `csv`, `tsv`, `parquet`, `arrow`, or `avro` — is served by the object-store listing path, which cannot carry request headers, so the headers are ignored. To authenticate a structured file download, use [Basic authentication](#using-basic-authentication) (`http_username` / `http_password`, or user info in the URL). ``` datasets: - from: https://api.example.com name: api_data params: http_headers: 'Authorization:Bearer ${secrets:api_token},Accept:application/json' ``` Headers can also be separated by semicolons: ``` datasets: - from: https://api.example.com name: api_data params: http_headers: 'Authorization: Bearer ${secrets:api_token}; X-API-Key: ${secrets:api_key}' ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field specifies the HTTP(s) endpoint and can be configured in two ways: 1. **Direct URL to a file**: A complete URL pointing to a specific [supported file](/docs/v1.10/components/data-connectors/#object-store-file-formats). ``` from: https://example.com/data/report.csv ``` 2. **Base domain/path**: A base URL that will be combined with special metadata fields to construct the complete request. ``` from: https://api.example.com/v1 ``` The connector supports templated URLs with query parameters that can be dynamically populated using `refresh_sql` filters and special metadata fields. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: http://static_username@localhost:3001/report.csv name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The connector supports authentication, timeout, connection pooling, and retry configuration via `params`. | Parameter Name | Description | | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `http_port` | Optional. Port to create HTTP(s) connection over. Default: 80 and 443 for HTTP and HTTPS respectively. | | `http_username` | Optional. Username for HTTP basic authentication. Default: None. | | `http_password` | Optional. Password for HTTP basic authentication. Default: None. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to load the password from a secret store, e.g. `${secrets:my_http_pass}`. | | `http_headers` | Optional. Custom HTTP headers as a comma-separated list of `key:value` pairs. Example: `Content-Type:application/json,Accept:application/json`. Applies to dynamic JSON API endpoints only; structured HTTP file datasets ignore these headers. Default: None. | | `allowed_request_paths` | **Required** for using `request_path` filters. Comma-separated list of allowed paths. Example: `/api/users,/api/posts`. Paths must start with `/` and cannot contain `..` segments. | | `request_query_filters` | Optional. Set to `enabled` to enable `request_query` filters. Default: `disabled`. When disabled, query parameter filters will be rejected. | | `request_body_filters` | Optional. Set to `enabled` to enable `request_body` filters for POST requests. Default: `disabled`. When disabled, request body filters will be rejected. | | `client_timeout` | Optional. Maximum time to wait for a response from the HTTP server (in seconds). Default: `30`. Supports duration formats like `30s`, `1m`, `500ms`, `2m30s`. Applied to the entire request-response cycle. | | `connect_timeout` | Optional. Timeout for establishing HTTP(s) connections (in seconds). Default: `10`. | | `pool_max_idle_per_host` | Optional. Maximum number of idle connections to keep alive per host. Default: `10`. | | `pool_idle_timeout` | Optional. Timeout for idle connections in the pool (in seconds). Default: `90`. | | `max_retries` | Optional. Maximum number of retries for failed HTTP requests. Default: `3`. | | `retry_backoff_method` | Optional. Retry backoff strategy: `fibonacci` (default), `linear`, or `exponential`. | | `retry_max_duration` | Optional. Maximum total duration for all retries (e.g., `30s`, `5m`). If not set, retries continue up to `max_retries`. | | `retry_jitter` | Optional. Randomization factor for retry delays (0.0 to 1.0). Default: `0.3` (30% randomization). Set to `0` for no jitter. | | `max_request_query_length` | Optional. Maximum length in characters for `request_query` filter values. Default: `1024`. Maximum: `4096`. | | `max_request_body_bytes` | Optional. Maximum size in bytes for `request_body` filter values. Default: `16384` (16 KiB). Maximum: `65536` (64 KiB). | | `health_probe` | Optional. Custom health probe path for endpoint validation during initialization (e.g., `/health`, `/api/status`). The endpoint must return a 2xx status code to pass validation. If not set, a random path is used and any status (including 404) is accepted. Must start with `/`. | ## HTTP Response Headers[​](#http-response-headers "Direct link to HTTP Response Headers") When querying HTTP(s) datasets, Spice respects standard HTTP caching headers in responses. The connector supports the following cache-related response headers: ### `Cache-Control`[​](#cache-control "Direct link to cache-control") The `Cache-Control` response header from the HTTP(s) endpoint is passed through to clients querying Spice. When the HTTP(s) server returns a `Cache-Control` header with the `stale-while-revalidate` directive, clients can use this value to determine appropriate caching behavior. For example, if the HTTP(s) endpoint returns: ``` Cache-Control: max-age=10, stale-while-revalidate=10 ``` Clients querying Spice will receive this header and can: 1. Serve fresh data for 10 seconds after fetching. 2. Between 10-20 seconds, serve stale data while fetching fresh data in the background. 3. After 20 seconds, fetch fresh data before serving the next request. The stale-while-revalidate behavior in Spice is controlled by the `stale_while_revalidate_ttl` parameter in the [caching configuration](/docs/v1.10/features/caching#stale-while-revalidate). When `stale_while_revalidate_ttl` is set to `0` (default), stale data will not be served. When set to a non-zero value, Spice serves stale cache entries while revalidating in the background. ## Advanced Features[​](#advanced-features "Direct link to Advanced Features") The HTTP connector provides advanced capabilities for working with dynamic APIs and RESTful services through special metadata fields. ### Special Metadata Fields[​](#special-metadata-fields "Direct link to Special Metadata Fields") The HTTP connector supports special metadata fields that provide fine-grained control over HTTP requests. These fields can be included in your dataset schema to dynamically construct request URLs and payloads. Security Requirements For security, these metadata fields require explicit configuration to prevent unauthorized access: * `request_path` requires `allowed_request_paths` to be configured with glob patterns * `request_query` requires `request_query_filters: enabled` * `request_body` requires `request_body_filters: enabled` | Field Name | Type | Description | | --------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `request_path` | String | Specifies the URL path to append to the base URL from the `from` field. When using a base domain/path in `from`, `request_path` constructs the complete endpoint. Example: If `from: https://api.example.com` and `request_path: /users/123`, the request will be made to `https://api.example.com/users/123`. **Requires `allowed_request_paths` parameter.** | | `request_query` | String | Defines query parameters to append to the request URL. Formatted as a query string (e.g., `key1=value1&key2=value2`). These parameters are appended to the URL after any path specified in `request_path`. **Requires `request_query_filters: enabled`.** Maximum length: configurable via `max_request_query_length` (default: 1024 characters). | | `request_body` | String | Contains the request body for POST/PUT requests. Typically used with REST APIs that require a JSON or form-encoded payload. The content type should be specified using `http_headers`. **Requires `request_body_filters: enabled`.** Maximum size: configurable via `max_request_body_bytes` (default: 16 KiB). | These metadata fields work in combination: * If `from` specifies a complete file URL, these fields are ignored * If `from` specifies a base URL, these fields construct the full request dynamically * `request_path` is appended to the base URL * `request_query` is appended as query parameters * `request_body` is sent as the request payload (requires appropriate HTTP method configuration) ### Endpoint Validation[​](#endpoint-validation "Direct link to Endpoint Validation") The HTTP connector validates the configured endpoint during initialization to detect issues such as DNS errors, connection problems, or invalid URLs early in the startup process. #### Default Validation Behavior[​](#default-validation-behavior "Direct link to Default Validation Behavior") By default, the connector performs a health check by requesting a randomly generated path (e.g., `/__spice_health_check_abc123def456`) that is expected to return a 404 status. Any HTTP response, including 404 Not Found, indicates that the endpoint is reachable and the dataset will initialize successfully. This default behavior works for most HTTP endpoints but may not be suitable for APIs that: * Return error responses for unknown paths without proper HTTP status codes * Have strict path validation that rejects requests to non-existent endpoints * Require authentication for all paths, including health check endpoints #### Custom Health Probe[​](#custom-health-probe "Direct link to Custom Health Probe") For endpoints that require a specific health check path, configure the `health_probe` parameter: ``` datasets: - from: https://api.example.com/v1 name: api_data params: health_probe: /health ``` When a custom health probe is configured: * The connector validates the endpoint by requesting the specified path * The health probe endpoint must return a 2xx status code (200-299) for validation to succeed * If the health probe returns a non-2xx status code, the dataset will fail to initialize with an error message This provides more reliable validation for APIs with dedicated health check endpoints. ##### Example with Authentication[​](#example-with-authentication "Direct link to Example with Authentication") ``` datasets: - from: https://api.example.com name: authenticated_api params: http_headers: 'Authorization:Bearer ${secrets:api_token}' health_probe: /api/status ``` In this configuration, the health probe request to `/api/status` will include the authentication header, ensuring that the validation succeeds for APIs that require authentication on all endpoints. ##### Health Probe Requirements[​](#health-probe-requirements "Direct link to Health Probe Requirements") The `health_probe` parameter has the following requirements: * Must start with `/` * Cannot exceed 2048 characters in length * The target endpoint must return a 2xx HTTP status code for validation to succeed ## Advanced Usage[​](#advanced-usage "Direct link to Advanced Usage") ### Using Special Metadata Fields with Base URL[​](#using-special-metadata-fields-with-base-url "Direct link to Using Special Metadata Fields with Base URL") When using a base URL with special metadata fields, you can dynamically construct different API endpoints: ``` datasets: - from: https://api.example.com/v1 name: api_requests params: http_headers: 'Content-Type:application/json' allowed_request_paths: '/users,/data/upload,/api/**' request_query_filters: enabled request_body_filters: enabled ``` With the above configuration, you can query different endpoints by providing values for the special metadata fields: ``` -- Query a specific user endpoint SELECT * FROM api_requests WHERE request_path = '/users/123' AND request_query = 'include=profile,settings'; -- Make a POST request with a body SELECT * FROM api_requests WHERE request_path = '/data/upload' AND request_body = '{"name":"example","value":42}'; ``` The connector will construct requests like: * `https://api.example.com/v1/users/123?include=profile,settings` * `https://api.example.com/v1/data/upload` with the JSON body #### Securing Paths with Glob Patterns[​](#securing-paths-with-glob-patterns "Direct link to Securing Paths with Glob Patterns") The `allowed_request_paths` parameter supports glob patterns to flexibly and securely match request paths. This provides a powerful way to configure path filtering without listing every possible endpoint. **Pattern Types:** * **Single wildcard (`*`)**: Matches any characters within a single path segment * Example: `/shows/*` matches `/shows/123` and `/shows/breaking-bad` * Does not match across path separators: `/shows/*` does not match `/shows/123/episodes` * \*\*Recursive wildcard (`**`)\*\*: Matches any number of path segments * Example: `/api/**` matches `/api/users`, `/api/v1/users`, and `/api/v2/posts/123` * Use for flexible API version matching or deep hierarchies * **Character classes (`[...]`)**: Matches one character from a set * Example: `/api/v[0-9]/*` matches `/api/v1/users` and `/api/v2/posts` * Example: `/api/v[1-3]/*` matches `/api/v1/users`, `/api/v2/posts`, and `/api/v3/data` **Examples:** ``` datasets: - from: https://api.tvmaze.com name: tv_api params: # Match any show ID allowed_request_paths: '/shows/*' ``` ``` -- Matches because /shows/82 matches the pattern /shows/* SELECT * FROM tv_api WHERE request_path = '/shows/82'; ``` ``` datasets: - from: https://api.example.com name: versioned_api params: # Match all endpoints under any API version allowed_request_paths: '/api/**' ``` ``` -- All of these match the pattern /api/** SELECT * FROM versioned_api WHERE request_path = '/api/users'; SELECT * FROM versioned_api WHERE request_path = '/api/v1/users'; SELECT * FROM versioned_api WHERE request_path = '/api/v2/products/electronics'; ``` ``` datasets: - from: https://api.example.com name: specific_versions params: # Match only API versions 1-9 allowed_request_paths: '/api/v[0-9]/*' ``` ``` -- Matches because /api/v1/users matches /api/v[0-9]/* SELECT * FROM specific_versions WHERE request_path = '/api/v1/users'; -- Does NOT match because v10 has two digits SELECT * FROM specific_versions WHERE request_path = '/api/v10/users'; ``` ### Dynamic Filters with Metadata Fields[​](#dynamic-filters-with-metadata-fields "Direct link to Dynamic Filters with Metadata Fields") The special metadata fields can be combined with dynamic filters to create sophisticated data refresh patterns. #### Dynamic API Queries with SQL[​](#dynamic-api-queries-with-sql "Direct link to Dynamic API Queries with SQL") ``` datasets: - from: https://api.tvmaze.com name: tv_shows params: http_headers: 'Accept:application/json' allowed_request_paths: '/search/shows,/shows/*,/shows/*/episodes' request_query_filters: enabled ``` Query specific API endpoints dynamically: ``` -- Search for shows by name SELECT * FROM tv_shows WHERE request_path = '/search/shows' AND request_query = 'q=game+of+thrones'; -- Get a specific show by ID (matches /shows/* pattern) SELECT * FROM tv_shows WHERE request_path = '/shows/82'; -- Get episodes for a show with filters (matches /shows/*/episodes pattern) SELECT * FROM tv_shows WHERE request_path = '/shows/82/episodes' AND request_query = 'season=1'; ``` #### Incremental Loading with Metadata Fields[​](#incremental-loading-with-metadata-fields "Direct link to Incremental Loading with Metadata Fields") ``` datasets: - from: https://api.example.com name: events params: allowed_request_paths: '/events,/events/*' request_query_filters: enabled acceleration: enabled: true refresh_mode: append refresh_sql: | SELECT * FROM events WHERE request_path = '/events' AND request_query = CONCAT('since=', (SELECT MAX(created_at) FROM events)) ``` This configuration: * Uses `request_path` to specify the `/events` endpoint * Dynamically constructs the `request_query` parameter using the latest timestamp from existing data * On each refresh, only fetches events created after the last refresh #### Paginated Data Loading[​](#paginated-data-loading "Direct link to Paginated Data Loading") ``` datasets: - from: https://api.example.com/v2 name: paginated_data params: http_headers: 'Content-Type:application/json' allowed_request_paths: '/data' request_query_filters: enabled acceleration: enabled: true refresh_mode: append refresh_sql: | SELECT * FROM paginated_data WHERE request_path = '/data' AND request_query = CONCAT('page=', COALESCE((SELECT MAX(page_number) FROM paginated_data) + 1, 1), '&limit=100') ``` This incrementally loads pages of data by: * Tracking the last loaded page number * Constructing the next page query parameter * Fetching 100 records per page #### POST Request with Dynamic Body[​](#post-request-with-dynamic-body "Direct link to POST Request with Dynamic Body") ``` datasets: - from: https://api.example.com name: search_results params: http_headers: 'Content-Type:application/json' allowed_request_paths: '/search' request_body_filters: enabled acceleration: enabled: true refresh_mode: full refresh_sql: | SELECT * FROM search_results WHERE request_path = '/search' AND request_body = '{"query": {"match": {"status": "active"}}, "from": 0, "size": 1000}' ``` This example demonstrates: * Using `_body` to send a JSON payload for a POST request * Executing complex search queries against REST APIs * Fetching results based on structured query syntax ### Processing JSON Responses[​](#processing-json-responses "Direct link to Processing JSON Responses") APIs often return JSON data that requires parsing to extract specific fields. Spice provides [JSON functions](/docs/v1.10/reference/sql/json) to process and transform JSON responses directly in SQL queries. #### Extracting Fields from JSON[​](#extracting-fields-from-json "Direct link to Extracting Fields from JSON") ``` datasets: - from: https://api.tvmaze.com name: tvmaze params: file_format: json allowed_request_paths: '/shows/*' ``` Extract specific fields from JSON responses: ``` -- Extract the show name from a JSON response SELECT json_get_str(content, 'name') as name FROM tvmaze WHERE request_path = '/shows/169'; ``` #### Working with Nested JSON[​](#working-with-nested-json "Direct link to Working with Nested JSON") APIs often return deeply nested JSON structures that require parsing to extract specific fields. Use chained JSON functions to navigate nested objects: ``` -- Extract nested fields from a show's network information SELECT json_get_str(content, 'name') as show_name, json_get_str(json_get(content, 'network'), 'name') as network_name, json_get_str(json_get(json_get(content, 'network'), 'country'), 'name') as country, json_get_str(json_get(json_get(content, 'network'), 'country'), 'code') as country_code FROM tvmaze WHERE request_path = '/shows/82'; ``` This demonstrates extracting nested objects step by step: * `json_get(content, 'network')` extracts the network object * `json_get_str(json_get(content, 'network'), 'name')` gets the network name from the nested object * Multiple `json_get` calls can be chained to navigate deeper levels #### Extracting Multiple Fields[​](#extracting-multiple-fields "Direct link to Extracting Multiple Fields") ``` -- Parse multiple fields from a TV show API response SELECT json_get_str(content, 'name') as show_name, json_get_str(content, 'type') as show_type, json_get_str(content, 'language') as language, json_get_int(content, 'runtime') as runtime_minutes, json_get_str(content, 'premiered') as premiere_date, json_get_str(content, 'status') as status FROM tvmaze WHERE request_path = '/shows/169'; ``` #### Processing JSON Arrays[​](#processing-json-arrays "Direct link to Processing JSON Arrays") ``` -- Extract genres from a JSON array SELECT json_get_str(content, 'name') as show_name, json_get_array(content, 'genres') as genres_array FROM tvmaze WHERE request_path = '/shows/82'; ``` For more details on available JSON functions including `json_get`, `json_get_str`, `json_get_int`, `json_get_bool`, and others, refer to the [JSON functions reference](/docs/v1.10/reference/sql/json). ### Refresh SQL with Dynamic Filters[​](#refresh-sql-with-dynamic-filters "Direct link to Refresh SQL with Dynamic Filters") The HTTP connector supports dynamic URL construction through `refresh_sql` with templated query parameters. This enables incremental data loading by appending filter conditions from the SQL query to the HTTP request URL. #### How It Works[​](#how-it-works "Direct link to How It Works") When `refresh_sql` is specified with filters, the connector extracts filter conditions and appends them as query parameters to the URL. This is particularly useful for APIs that support filtering via query parameters. #### Time-Based Incremental Loading[​](#time-based-incremental-loading "Direct link to Time-Based Incremental Loading") ``` datasets: - from: https://api.example.com/data.csv?start_time={start_time}&end_time={end_time} name: incremental_data acceleration: enabled: true refresh_mode: append refresh_sql: | SELECT * FROM incremental_data WHERE timestamp > (SELECT MAX(timestamp) FROM incremental_data) ``` In this example: * The `{start_time}` and `{end_time}` placeholders in the URL are replaced with values extracted from the `WHERE` clause in `refresh_sql` * Each refresh appends only new data since the last refresh * The connector automatically maps SQL filter conditions to URL query parameters #### Supported Filter Operations[​](#supported-filter-operations "Direct link to Supported Filter Operations") The dynamic filter feature supports the following SQL operations: * Equality comparisons (`=`) * Greater than (`>`) * Less than (`<`) * Greater than or equal (`>=`) * Less than or equal (`<=`) * Range queries with `BETWEEN` #### Notes[​](#notes "Direct link to Notes") * URL parameters must match filter column names in the `refresh_sql` * Only filters that can be pushed down to the HTTP source will be applied to the URL * Complex filters may not be supported for URL templating ## Limitations[​](#limitations "Direct link to Limitations") ### Security Constraints[​](#security-constraints "Direct link to Security Constraints") For security and to prevent unauthorized access, the HTTP connector enforces the following constraints on special metadata fields: #### Request Path Limitations[​](#request-path-limitations "Direct link to Request Path Limitations") * **Explicit Allow-List Required**: The `request_path` field cannot be used without configuring `allowed_request_paths` * **Path Pattern Format**: All patterns in `allowed_request_paths` must: * Start with `/` * Not contain `..` path traversal segments * Not exceed 2048 characters in length * **Glob Pattern Matching**: Query filters are matched against glob patterns in the `allowed_request_paths` list using: * `*` matches a single path segment (e.g., `/shows/*` matches `/shows/123` but not `/shows/123/episodes`) * `**` matches multiple path segments recursively (e.g., `/api/**` matches `/api/v1/users` and `/api/v2/posts/123`) * `[...]` character classes (e.g., `/api/v[0-9]/*` matches `/api/v1/users` but not `/api/v10/users`) * **Empty Paths**: Empty `request_path` filters are rejected Example error when `allowed_request_paths` is not configured: ``` request_path filters are disabled for this dataset. Configure allowed_request_paths to enable them. ``` #### Request Query Limitations[​](#request-query-limitations "Direct link to Request Query Limitations") * **Explicit Enable Required**: The `request_query` field requires `request_query_filters: enabled` * **Length Limit**: Query strings are limited to 1024 characters by default (configurable up to 4096 via `max_request_query_length`) * **Control Characters**: Query strings cannot contain control characters * **Leading Question Mark**: The connector automatically strips leading `?` if present Example error when query filters are not enabled: ``` request_query filters are disabled for this dataset. Enable request_query_filters to use them. ``` #### Request Body Limitations[​](#request-body-limitations "Direct link to Request Body Limitations") * **Explicit Enable Required**: The `request_body` field requires `request_body_filters: enabled` * **Size Limit**: Request bodies are limited to 16 KiB (16,384 bytes) by default (configurable up to 64 KiB via `max_request_body_bytes`) * **POST Method**: When a `request_body` filter is present, the HTTP method automatically changes to POST Example error when body filters are not enabled: ``` request_body filters are disabled for this dataset. Enable request_body_filters to use them. ``` ### Configuration Requirements[​](#configuration-requirements "Direct link to Configuration Requirements") To use the special metadata fields (`request_path`, `request_query`, `request_body`), you must: 1. **For `request_path`**: Configure `allowed_request_paths` with a comma-separated list of allowed path patterns (supports glob patterns) 2. **For `request_query`**: Set `request_query_filters: enabled` in params 3. **For `request_body`**: Set `request_body_filters: enabled` in params Example minimal configuration for all three fields: ``` datasets: - from: https://api.example.com name: my_api params: allowed_request_paths: '/users,/posts,/comments,/api/**' request_query_filters: enabled request_body_filters: enabled ``` ### Performance Considerations[​](#performance-considerations "Direct link to Performance Considerations") * **Connection Pooling**: The connector maintains up to 10 idle connections per host by default * **Retry Overhead**: With the default 3 retries and Fibonacci backoff, failed requests may take several seconds before returning an error * **Cache Behavior**: HTTP responses are cached based on the combination of path, query, and body parameters ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). --- # Iceberg Data Connector The Iceberg Data Connector helps query [Apache Iceberg](https://iceberg.apache.org/) tables using federated SQL. Every Iceberg dataset requires an Iceberg catalog to provide table metadata and manage access. When working with multiple datasets, it is recommended to use a catalog connector (instead of a data connector), such as the \[Iceberg Catalog Connector(../catalogs/iceberg) or \[AWS Glue Catalog Connector(../catalogs/glue) instead of configuring individual datasets. Iceberg catalogs can be of several types: * **Iceberg REST Catalog**: The most common and recommended approach. REST Catalogs expose Iceberg tables over HTTP(S) endpoints and are compatible with most managed Iceberg services and cloud providers. * **AWS Glue Catalog**: Integrates with AWS Glue as a catalog provider, supporting Iceberg tables stored in S3. This is the preferred method for AWS environments. * **Hadoop-style Catalogs**: Use file-based storage (e.g., `file://`, `s3://`, `s3a://`) to manage table metadata. This approach is typically used for local development or legacy deployments. Hadoop-style Catalogs For production and cloud environments, REST and AWS Glue catalogs are recommended. Hadoop-style catalogs are supported but less common and not recommended for most new deployments. ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: my_table ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field specifies the Iceberg table to connect to, in the format `iceberg:`. The `table_path` is the URL to the Iceberg table in the catalog provider. For REST Catalogs, use the format `http[s]:///v1/{prefix}/namespaces//tables/`. For AWS Glue catalogs, the URL format is `https://glue..amazonaws.com/iceberg/v1/catalogs//namespaces`, where `` is the AWS account ID. While possible to connect to Iceberg tables hosted by Glue using this generic connector, it is recommended to instead use the \[AWS Glue Data Connector(./glue) for connecting to Iceberg tables managed by Glue for a better experience. Example (REST Catalog): ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: my_table ``` Example (AWS Glue Catalog): ``` datasets: - from: iceberg:https://glue.us-east-1.amazonaws.com/iceberg/v1/catalogs/123456789012/namespaces/my_namespace/tables/my_table name: glue_table ``` Hadoop-style catalogs use file-based paths such as `file://`, `s3://`, or `s3a://`. For these, specify the warehouse path as the table location. This is typically only used for local development or legacy setups. Example (Hadoop Catalog, local): ``` datasets: - from: iceberg:file:///tmp/hadoop_warehouse/test/my_table_1 name: local_hadoop ``` Example (Hadoop Catalog, S3): ``` datasets: - from: iceberg:s3a://my-bucket/hadoop_warehouse/test/my_table_2 name: s3_hadoop ``` ### `name`[​](#name "Direct link to name") The `name` field sets the table name within Spice. This name is used to reference the dataset in SQL queries. The name cannot be a \[reserved keyword(../../reference/spicepod/keywords). Example: ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: transactions params: iceberg_token: ${secrets:iceberg_token} ``` ``` SELECT COUNT(*) FROM transactions; ``` ``` +----------+ | count(*) | +----------+ | 1234567 | +----------+ ``` ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `iceberg_token` | Bearer token value to use for Authorization header. | | `iceberg_oauth2_credential` | Credential to use for OAuth2 client credential flow when connecting to the table. Format: `:` | | `iceberg_oauth2_scope` | Scope to use for OAuth2 client credential flow when connecting to the table. Default: `catalog` | | `iceberg_oauth2_server_url` | URL of the OAuth2 server tokens endpoint for the client credential flow. | | `iceberg_s3_endpoint` | S3-compatible endpoint where the Iceberg table data is stored. | | `iceberg_s3_region` | Region of the S3-compatible endpoint. | | `iceberg_s3_access_key_id` | The AWS access key ID to use for S3 storage. If not provided, credentials will be loaded from environment variables or IAM roles. | | `iceberg_s3_secret_access_key` | The AWS secret access key to use for S3 storage. If not provided, credentials will be loaded from environment variables or IAM roles. | | `iceberg_s3_session_token` | Session token for the S3-compatible endpoint. | | `iceberg_s3_role_arn` | ARN of the IAM role to assume when accessing the S3-compatible endpoint. | | `iceberg_s3_role_session_name` | Session name to use when assuming the IAM role. | | `iceberg_s3_connect_timeout` | Connection timeout in seconds for the S3-compatible endpoint. Default: `60` | | `iceberg_sigv4_enabled` | Enable SigV4 (AWS Signature Version 4) authentication when connecting to the catalog. Automatically enabled if the URL in `from` is an AWS Glue catalog. Default: `false` | | `iceberg_signing_region` | Region to use for SigV4 authentication. Extracted from the URL in `from` if not specified. | | `iceberg_signing_name` | Service name to use for SigV4 authentication. Default: `glue`. | | `metadata_path` | The path including scheme to the metadata file for the Hadoop table. Must specify a path to a `.json` file. For example, `s3a://my-bucket/warehouse/namespace/table/metadata/v1.metadata.json` | ## Authentication[​](#authentication "Direct link to Authentication") Authentication to the Iceberg catalog. Supported methods include: * **Bearer Token**: Use `iceberg_token` for Authorization header. * **OAuth2 Client Credentials**: Use `iceberg_oauth2_credential`, `iceberg_oauth2_scope`, and `iceberg_oauth2_server_url`. * **AWS SigV4**: For AWS Glue, set `iceberg_sigv4_enabled: true` (or use a Glue URL). * **S3 Authentication**: Use `iceberg_s3_*` parameters for S3 data access. ### AWS Authentication[​](#aws-authentication "Direct link to AWS Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. These credentials will be used to connect to the S3 bucket as well as the Glue catalog (if configured). 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3/Glue permissions (e.g., `s3:ListBucket`, `s3:GetObject`) to access the tables. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ### Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access Iceberg tables in S3/Glue: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::company-bucketname-datasets" }, { "Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::company-bucketname-datasets/*" }, { "Effect": "Allow", "Action": [ "glue:GetCatalog", "glue:GetDatabases", "glue:GetDatabase", "glue:GetTable", "glue:GetTables" ], Resource: "*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------- | -------------------------------------------------------------- | | `s3:ListBucket` | Required. Allows scanning all objects from the bucket | | `s3:GetObject` | Required. Allows fetching objects | | `glue:GetCatalog` | Required. Retrieve metadata about the specified catalog. | | `glue:GetDatabases` | Required. List the databases available in the current catalog. | | `glue:GetDatabase` | Required. Retrieve metadata about the specified database. | | `glue:GetTable` | Required. Retrieve metadata about the specified table. | | `glue:GetTables` | Required. List the tables available in the current database. | ## Examples[​](#examples "Direct link to Examples") ### Basic Example (REST Catalog)[​](#basic-example-rest-catalog "Direct link to Basic Example (REST Catalog)") Connect to an Iceberg table with token authentication: ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: my_table params: iceberg_token: ${secrets:iceberg_token} ``` ### AWS Glue Catalog Example[​](#aws-glue-catalog-example "Direct link to AWS Glue Catalog Example") Connect to an Iceberg table in AWS Glue catalog: ``` datasets: - from: iceberg:https://glue.us-east-1.amazonaws.com/iceberg/v1/catalogs/123456789012/namespaces/my_namespace/tables/my_table name: glue_table params: iceberg_sigv4_enabled: true ``` ### OAuth2 Authentication Example[​](#oauth2-authentication-example "Direct link to OAuth2 Authentication Example") Connect to an Iceberg table using OAuth2 authentication: ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: oauth_table params: iceberg_oauth2_credential: ${secrets:client_id}:${secrets:client_secret} iceberg_oauth2_scope: catalog iceberg_oauth2_server_url: https://iceberg-catalog-host.com/oauth2/token ``` ### S3 Storage Example[​](#s3-storage-example "Direct link to S3 Storage Example") Connect to an Iceberg table with custom S3 storage configuration: ``` datasets: - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table name: s3_table params: iceberg_token: ${secrets:iceberg_token} iceberg_s3_endpoint: http://localhost:9000 iceberg_s3_region: us-west-2 iceberg_s3_access_key_id: ${secrets:aws_access_key_id} iceberg_s3_secret_access_key: ${secrets:aws_secret_access_key} ``` ### Hadoop Catalog Example[​](#hadoop-catalog-example "Direct link to Hadoop Catalog Example") Connect to an Iceberg table using Hadoop Catalog with a local warehouse: ``` datasets: - from: iceberg:file:///tmp/hadoop_warehouse/test/my_table_1 name: local_hadoop params: metadata_path: file:///tmp/hadoop_warehouse/test/my_table_1/metadata/v1.metadata.json ``` Connect to an Iceberg table using Hadoop Catalog with S3: ``` datasets: - from: iceberg:s3a://my-bucket/hadoop_warehouse/test/my_table_2 name: s3_hadoop params: metadata_path: s3a://my-bucket/hadoop_warehouse/test/my_table_2/metadata/v1.metadata.json ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the \[secret stores documentation(../secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the \[using referenced secrets guide(../secret-stores#using-secrets). ## Limitations[​](#limitations "Direct link to Limitations") Performance Considerations When querying Iceberg tables, performance depends on the size of the table, the complexity of the query, and the underlying storage system. For large tables, consider using appropriate filtering to limit the amount of data scanned. The connector needs to access both the Iceberg catalog metadata and the underlying data files (typically stored in S3 or a compatible object store). Ensure proper network connectivity and authentication for both systems. --- # IMAP Data Connector The IMAP Data Connector enables federated SQL query across emails stored in an IMAP email server. ``` datasets: - from: imap:myawesomeemail@example.com name: emails params: imap_password: ${secrets:IMAP_PASSWORD} ``` ## Schema[​](#schema "Direct link to Schema") | Field Name | Data Type | Nullable | Description | | ------------- | ------------ | -------- | -------------------------------------------------------------------------------- | | `date` | `Date64` | No | The date and time when the email was sent. | | `subject` | `Utf8` | Yes | The subject line of the email. | | `from` | `List` | Yes | The sender(s) of the email. | | `to` | `List` | Yes | The primary recipient(s) of the email. | | `cc` | `List` | Yes | The carbon copy recipient(s) of the email. | | `bcc` | `List` | Yes | The blind carbon copy recipient(s) of the email. | | `reply_to` | `List` | Yes | The email address(es) to which replies should be sent. | | `message_id` | `Utf8` | Yes | A unique identifier for the email message. | | `in_reply_to` | `Utf8` | Yes | The `message_id` of the email this message is replying to, if applicable. | | `content` | `Utf8` | Yes | The raw email body of this message. Not retrieved when acceleration is disabled. | If a MIME-encoded value is retrieved for a field, it is not decoded and the MIME-encoded value is returned in SQL queries. Most fields are optional, and depend on the implementation of the specific IMAP server being connected to. For example, the IMAP RFC specifies the `message_id` field SHOULD be supplied but the field is an optional field. For more information, refer to the [IMAP RFC 2822 - Section 3.6](https://www.rfc-editor.org/rfc/rfc2822#section-3.6). ## Retrieving email body contents[​](#retrieving-email-body-contents "Direct link to Retrieving email body contents") When the IMAP Data Connector is used without acceleration, the email body will not be retrieved - only header/subject values. To load the email body contents, specify an acceleration: ``` datasets: - from: imap:myawesomeemail@example.com name: emails params: imap_password: ${secrets:IMAP_PASSWORD} acceleration: enabled: true ``` With an acceleration enabled, the `content` field will be populated with the complete email body including headers, without any decoding applied. This field could be used for post-processing the email, like retrieving custom header values or decoding MIME-encoded content. Limitations * Email attachments are currently not parsed from the email body into separate dataset fields. To read email attachments, parse the multipart encodings from the `content` field. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field must contain the email address for the mailbox to connect to. For example, `me@outlook.com`, or `jsmith@example.com`. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: imap:jsmith@example.com name: emails params: ... ``` ``` SELECT COUNT(*) FROM emails; ``` ``` +----------+ | count(*) | +----------+ | 1234 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The IMAP connector supports the following connection and authentication parameters: | Parameter Name | Description | | --------------- | ---------------------------------------------------------------------------------------------------------------------------- | | `imap_username` | Optional. The username to use for the IMAP connection. Defaults to the value of the `from:` mailbox field. | | `imap_password` | Required. The password to use for the IMAP connection, in plaintext authentication mode. | | `imap_host` | Optional. The host or IP address of the IMAP server to connect to. Not required for known connections like Outlook or Gmail. | | `imap_port` | Optional. The port of the IMAP server to connect to. Defaults to `993`. | | `imap_mailbox` | Optional. The mailbox to read mail from. Defaults to `INBOX`, the standard email inbox. | | `imap_ssl_mode` | Optional. The IMAP SSL mode to use. Defaults to `auto`, permitted values of `tls`, `starttls`, `disabled` or `auto`. | ## Examples[​](#examples "Direct link to Examples") ### Basic example[​](#basic-example "Direct link to Basic example") ``` datasets: - from: imap:jsmith@example.com name: emails params: imap_host: mail.example.com imap_password: ${ secrets:IMAP_PASSWORD } ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure IMAP as a data connector in Spice. [IMAP Data Connector](https://github.com/spiceai/cookbook/tree/trunk/imap/#readme) --- # Kafka Data Connector The Kafka Data Connector enables direct acceleration of data from [Apache Kafka](https://kafka.apache.org/) topics using `refresh_mode: append` \[acceleration(../data-accelerators). This allows seamless integration with existing Kafka-based event streaming infrastructure for real-time data acceleration and analytics. ``` datasets: - from: kafka:my_kafka_topic name: my_dataset params: kafka_bootstrap_servers: broker1:9092,broker2:9092,broker3:9092 # Required. A comma separated list of Kafka broker servers. kafka_security_protocol: sasl_ssl # Default is `sasl_ssl`. Valid values are `plaintext`, `ssl`, `sasl_plaintext`, `sasl_ssl`. kafka_sasl_mechanism: SCRAM-SHA-512 # Default is `SCRAM-SHA-512`. Valid values are `PLAIN`, `SCRAM-SHA-256`, `SCRAM-SHA-512`. kafka_sasl_username: kafka # Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. kafka_sasl_password: ${secrets:kafka_sasl_password} # Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. kafka_ssl_ca_location: ./certs/kafka_ca_cert.pem # Optional. Used to verify the SSL/TLS certificate of the Kafka broker. kafka_enable_ssl_certificate_verification: true # Default is `true`. Set to `false` to disable SSL/TLS certificate verification. kafka_ssl_endpoint_identification_algorithm: https # Default is `https`. Valid values are `none` and `https`. acceleration: enabled: true # Acceleration is required for the kafka connector. engine: duckdb # `duckdb`, `sqlite` and `postgres` are supported acceleration engines for Kafka. refresh_mode: append # Required. Must be set to `append` for the Kafka connector. mode: file # Persistence is recommended to not have to fully rebuild the table each time Spice starts. ``` ## Overview[​](#overview "Direct link to Overview") Upon startup, Spice fetches all messages for the specified topic using a uniquely generated consumer group. If a persistent acceleration engine is used (with `mode: file`), data is fetched starting from the last processed record, allowing Spice to resume without reprocessing all historical data. Schema is automatically inferred from the first available topic message in JSON format. The connector creates the appropriate table schema for acceleration based on the detected data structure. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form of `kafka:kafka_topic` where `kafka_topic` is the name of the Kafka topic to consume from. ``` datasets: - from: kafka:user_events name: events ... ``` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. ``` datasets: - from: kafka:orders_events name: orders ... ``` ``` SELECT COUNT(*) FROM orders; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a \[reserved keyword(../../reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `kafka_bootstrap_servers` | **Required**. A list of host/port pairs for establishing the initial Kafka cluster connection. The client will use all servers, regardless of the bootstrapping servers specified here. This list only affects the initial hosts used to discover the full server set and should be formatted as `host1:port1,host2:port2,...`. | | `kafka_security_protocol` | Security protocol for Kafka connections. Default: `sasl_ssl`. Options: - `plaintext`
- `ssl`
- `sasl_plaintext`
- `sasl_ssl` | | `kafka_sasl_mechanism` | SASL (Simple Authentication and Security Layer) authentication mechanism. Default: `SCRAM-SHA-512`. Options: - `PLAIN`
- `SCRAM-SHA-256`
- `SCRAM-SHA-512` | | `kafka_sasl_username` | SASL username. Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. | | `kafka_sasl_password` | SASL password. Required if `kafka_security_protocol` is `sasl_plaintext` or `sasl_ssl`. | | `kafka_ssl_ca_location` | Path to the SSL/TLS CA certificate file for server verification. | | `kafka_enable_ssl_certificate_verification` | Enable SSL/TLS certificate verification. Default: `true`. | | `kafka_ssl_endpoint_identification_algorithm` | SSL/TLS endpoint identification algorithm. Default: `https`. Options: - `none`
- `https` | | `kafka_consumer_group_id` | Kafka consumer group id to use. If not set, a unique id will be generated. | | `schema_infer_max_records` | Number of Kafka messages to sample for schema inference. Default: `1`. Increase if your data has optional fields or varying structure. | | `flatten_json` | Set `true` to flatten nested structs in JSON as separate columns. | ### `metrics`[​](#metrics "Direct link to metrics") The connector supports the following optional \[component metrics(../../features/observability/component\_metrics): | Metric Name | Type | Description | | ------------------------ | ------- | ------------------------------------------------------------------------------------ | | `bytes_consumed_total` | Counter | Total number of bytes consumed from the Kafka topic | | `records_consumed_total` | Counter | Total number of records (messages) consumed from Kafka topics | | `records_lag` | Gauge | Total consumer lag across all topic partitions (number of messages not yet consumed) | These metrics are not enabled by default, enable them by setting the `metrics` parameter: ``` datasets: - from: kafka:user_events name: events metrics: - name: records_lag - name: records_consumed_total - name: bytes_consumed_total params: ... ``` ### Acceleration Settings[​](#acceleration-settings "Direct link to Acceleration Settings") warning Using the Kafka connector **requires** \[acceleration(../data-accelerators) with `refresh_mode: append` enabled. The following settings are required: | Parameter Name | Description | | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `enabled` | Required. Must be set to `true` to enable acceleration. | | `engine` | Required. The acceleration engine to use. Possible valid values: - `duckdb`: Use \[DuckDB(../data-accelerators/duckdb) as the acceleration engine.
- `sqlite`: Use \[SQLite(../data-accelerators/sqlite) as the acceleration engine.
- `postgres`: Use \[PostgreSQL(../data-accelerators/postgres) as the acceleration engine. | | `refresh_mode` | Required. The refresh mode to use. Must be set to `append` for the Kafka connector. | | `mode` | Optional. The persistence mode to use. When using the `duckdb` and `sqlite` engines, it is recommended to set this to `file` to persist the data across restarts. Spice persists metadata about the dataset, allowing it to resume from the last known state instead of re-processing all messages. | ## Data Format Support[​](#data-format-support "Direct link to Data Format Support") The Kafka connector currently supports JSON-formatted messages. Schema is automatically inferred from the first available message in the topic, and all subsequent messages are expected to follow a compatible structure. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the \[secret stores documentation(../secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the \[using referenced secrets guide(../secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * See how to query Kafka real-time data with other datasets using federated queries in [Live Orders Analytics example](https://github.com/spiceai/cookbook/blob/trunk/kafka/README.md). --- # Localpod Data Connector The Localpod Data Connector enables setting up a parent/child relationship between datasets in the current Spicepod. This can be used for configuring multiple/tiered accelerations for a single dataset, and ensuring that the data is only downloaded once from the remote source. For example, you can use the `localpod` connector to create a child dataset that is accelerated in-memory, while the parent dataset is accelerated to a file. The dataset created by the `localpod` connector will logically have the same data as the parent dataset. ## Synchronized Refreshes[​](#synchronized-refreshes "Direct link to Synchronized Refreshes") The `localpod` connector supports synchronized refreshes, which ensures that the child dataset is refreshed from the same data as the parent dataset. Synchronized refreshes require that both the parent and child datasets are accelerated with `refresh_mode: full` (which is the default). When synchronization is enabled, the following logs will be emitted: ``` 2024-10-28T15:45:24.220665Z INFO runtime::datafusion: Localpod dataset test_local synchronizing refreshes with parent table test ``` ### Examples[​](#examples "Direct link to Examples") ``` datasets: - from: postgres:cleaned_sales_data name: test params: ... acceleration: enabled: true # This dataset will be accelerated into a DuckDB file engine: duckdb mode: file refresh_check_interval: 10s - from: localpod:test name: test_local acceleration: enabled: true # This dataset accelerates the parent `test` dataset into in-memory Arrow records and is synchronized with the parent ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Localpod as a data connector in Spice. [Local dataset replication (Localpod)](https://github.com/spiceai/cookbook/tree/trunk/localpod#readme) --- # Memory Data Connector The Memory Data Connector enables configuring an in-memory dataset for tables used, or produced by the Spice runtime. Only certain tables, with predefined schemas, can be defined by the connector. These are: * `store`: Defines a table that LLMs, with \[memory tooling(../../features/large-language-models/memory), can store data in. Requires `access: read_write`. ### Examples[​](#examples "Direct link to Examples") ``` datasets: - from: memory:store name: llm_memory access: read_write columns: - name: value embeddings: # Easily make your LLM learnings searchable. - from: all-MiniLM-L6-v2 embeddings: - name: all-MiniLM-L6-v2 from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to provide persistent memory capabilities for language models in Spice. [LLM Memory](https://github.com/spiceai/cookbook/tree/trunk/llm-memory#readme) --- # MongoDB Data Connector MongoDB is an open-source NoSQL database that stores data in flexible, JSON-like documents, allowing for dynamic schemas and easy scalability. The MongoDB Data Connector enables federated/accelerated SQL queries on data stored in MongoDB databases. ``` datasets: - from: mongodb:mytable name: my_dataset params: mongodb_host: localhost mongodb_port: 27017 mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} mongodb_pool_min: 1 mongodb_pool_max: 5 ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `mongodb:{table_name}` where `table_name` is the table identifer in the MongoDB server to read from. ``` datasets: - from: mongodb:mytable name: my_dataset params: mongodb_db: my_database ... ``` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: mongodb:my_dataset name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords) ### `params`[​](#params "Direct link to params") The MongoDB data connector can be configured by providing the following `params`. Use the [secret replacement syntax](/docs/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:my_mongodb_conn_string}`. | Parameter Name | Description | | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mongodb_connection_string` | The connection string to use to connect to the MongoDB server. This can be used instead of providing individual connection parameters. | | `mongodb_user` | The MongoDB username. | | `mongodb_pass` | The password to connect with. | | `mongodb_host` | The hostname of the MongoDB server. Defaults to `localhost`. | | `mongodb_port` | The port of the MongoDB server. Defaults to `27017`. | | `mongodb_db` | The name of the database to connect to. Defaults to `default`. | | `mongodb_sslmode` | Optional. Specifies the SSL/TLS behavior for the connection, supported values:
- `required`: (default) This mode requires an SSL connection. If a secure connection cannot be established, server will not connect.
- `preferred`: Establishes an encrypted TLS/SSL connection but does not validate the server certificate or hostname (accepts invalid or self-signed certificates). If the server does not support TLS, the connection fails; it is never downgraded to plaintext.
- `disabled`: This mode will not attempt to use an SSL connection, even if the server supports it. | | `mongodb_sslrootcert` | Optional parameter specifying the path to a custom PEM certificate that the connector will trust. | | `mongodb_time_zone` | Optional. Specifies connection time zone. Default is `UTC`. Accepts:
- Fixed offsets (e.g., `+02:00`).
- IANA time zone names (e.g., `America/Los_Angeles`) | | `mongodb_auth_source` | Optional. Authentication source database. Overrides the default auth source in the connection string. | | `mongodb_unnest_depth` | Optional. Maximum nesting depth for unnesting embedded documents into a flattened structure. Higher values expand deeper nested fields. Default: `0` | | `mongodb_num_docs_to_infer_schema` | Optional. Number of documents to use to infer the schema. Defaults to 400. | | `mongodb_pool_min` | The minimum number of connections to keep open in the pool, lazily created when requested. Default: `1` | | `mongodb_pool_max` | The maximum number of connections to allow in the pool. Default: `5` | ## Types[​](#types "Direct link to Types") The table below shows the MongoDB data types supported, along with the type mapping to Apache Arrow types in Spice. | MongoDB Type | Arrow Type | | ------------------------- | ------------------------------------ | | `String` | `Utf8` | | `Boolean` | `Boolean` | | `Int32` | `Int32` | | `Int64` | `Int64` | | `Double` | `Float64` | | `Decimal128` | `Decimal128` | | `Binary` | `Binary` | | `Datetime` without time | `Date32` | | `Datetime` with time | `Timestamp(Millisecond, )` | | `Timestamp` | `Timestamp(Millisecond, None)` | | `Array` | `List` | | `Null` | `Null` | | `Undefined` | `Null` | | `RegularExpression` | `Utf8` | | `JavaScriptCode` | `Utf8` | | `JavaScriptCodeWithScope` | `Utf8` | | `Symbol` | `Utf8` | | `MaxKey` | `Utf8` | | `MinKey` | `Utf8` | | `DbPointer` | `Utf8` | | `ObjectId` | `Utf8` | | `Document` | See unnesting section | note * The MongoDB `Datetime` value is [retrieved as a UTC time value](https://www.mongodb.com/docs/manual/reference/method/Date/) by default. Use the `mongodb_time_zone` configuration parameter to specify the desired time zone for interpreting `TIMESTAMP` values during data retrieval. ## Unnesting[​](#unnesting "Direct link to Unnesting") Consider the following document: ``` { "a": 1, "b": { "x": 2, "y": { "z": 3 } } } ``` Using `mongodb_unnest_depth` you can control the unnesting behavior. Here are the examples: ### mongodb\_unnest\_depth: 0[​](#mongodb_unnest_depth-0 "Direct link to mongodb_unnest_depth: 0") ``` sql> select * from test_table; +-----------+---------------------+ | a (Int32) | b (Utf8) | +-----------+---------------------+ | 1 | {"x":2,"y":{"z":3}} | +---+-----------------------------+ ``` ### mongodb\_unnest\_depth: 1[​](#mongodb_unnest_depth-1 "Direct link to mongodb_unnest_depth: 1") ``` sql> select * from test_table; +-----------+-------------+------------+ | a (Int32) | b.x (Int32) | b.y (Utf8) | +-----------+-------------+------------+ | 1 | 2 | {"z":3} | +-----------+-------------+------------+ ``` ### mongodb\_unnest\_depth: 2[​](#mongodb_unnest_depth-2 "Direct link to mongodb_unnest_depth: 2") ``` sql> select * from test_table; +-----------+-------------+---------------+ | a (Int32) | b.x (Int32) | b.y.z (Int32) | +-----------+-------------+---------------+ | 1 | 2 | 3 | +-----------+-------------+---------------+ ``` ## Examples[​](#examples "Direct link to Examples") ### Connecting using username and password and custom auth table[​](#connecting-using-username-and-password-and-custom-auth-table "Direct link to Connecting using username and password and custom auth table") ``` datasets: - from: mongodb:my_dataset name: my_dataset params: mongodb_host: localhost mongodb_port: 27017 mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} mongodb_auth_source: admin ``` ### Connecting using SSL[​](#connecting-using-ssl "Direct link to Connecting using SSL") ``` datasets: - from: mongodb:my_dataset name: my_dataset params: mongodb_host: localhost mongodb_port: 27017 mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} mongodb_sslmode: preferred mongodb_sslrootcert: ./custom_cert.pem ``` ### Connecting using a Connection String[​](#connecting-using-a-connection-string "Direct link to Connecting using a Connection String") ``` datasets: - from: mongodb:my_dataset name: my_dataset params: mongodb_connection_string: mongodb://${secrets:my_user}:${secrets:my_password}@localhost:27017/my_db?authSource=admin ``` ### With custom connection pool settings[​](#with-custom-connection-pool-settings "Direct link to With custom connection pool settings") ``` datasets: - from: mongodb:my_dataset name: my_dataset params: mongodb_host: localhost mongodb_port: 27017 mongodb_db: my_database mongodb_user: my_user mongodb_pass: ${secrets:mongodb_pass} mongodb_pool_min: 5 mongodb_pool_max: 5 ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure MongoDB as a data connector in Spice. [MongoDB Data Connector](https://github.com/spiceai/cookbook/tree/trunk/mongodb/connector#readme) --- # Microsoft SQL Server Data Connector [Microsoft SQL Server](https://www.microsoft.com/en-us/sql-server) is a relational database management system developed by Microsoft. The Microsoft SQL Server Data Connector enables federated/accelerated SQL queries on data stored in MSSQL databases. Limitations 1. The connector supports SQL Server authentication (SQL Login and Password) only. 2. Spatial types (`geography`) are not supported, and columns with these types will be ignored. 3. `DATETIME2` and `DATETIMEOFFSET` columns are mapped to Arrow `Timestamp(Nanosecond)`. Timestamps outside the nanosecond range (approximately years 1677–2262) will silently return `1970-01-01 UTC`. This is an inherent limitation of Arrow's nanosecond timestamp representation. ``` datasets: - from: mssql:path.to.my_dataset name: my_dataset params: mssql_connection_string: ${secrets:mssql_connection_string} ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `mssql:database.schema.table` where `database.schema.table` is the fully-qualified table name in the SQL server. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: mssql:path.to.my_dataset name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords) or any of the following keywords that are reserved by Microsoft SQL Server: * `OUTER` * `SET` * `QUALIFY` * `WINDOW` * `END` * `FOR` ### `params`[​](#params "Direct link to params") The data connector supports the following `params`. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:my_mssql_conn_string}`. | Parameter Name | Description | | -------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `mssql_connection_string` | The ADO connection string to use to connect to the server. This can be used instead of providing individual connection parameters. | | `mssql_host` | The hostname or IP address of the Microsoft SQL Server instance. | | `mssql_port` | (Optional) The port of the Microsoft SQL Server instance. Default value is 1433. | | `mssql_database` | (Optional) The name of the database to connect to. The default database (`master`) will be used if not specified. | | `mssql_username` | The username for the SQL Server authentication. | | `mssql_password` | The password for the SQL Server authentication. | | `mssql_encrypt` | (Optional) Specifies whether encryption is required for the connection.
- `true` or `require`: (default) This mode requires an SSL connection. If a secure connection cannot be established, server will not connect.
- `false` or `disable`: This mode will not attempt to use an SSL connection, even if the server supports it. Only the login procedure is encrypted. | | `mssql_trust_server_certificate` | (Optional) Specifies whether the server certificate should be trusted without validation when encryption is enabled.
- `true`: The server certificate will not be validated and it is accepted as-is.
- `false`: (default) Server certificate will be validated against system's certificate storage. | ### Example[​](#example "Direct link to Example") ``` datasets: - from: mssql:SalesLT.Customer name: customer params: mssql_host: mssql-host.database.windows.net mssql_database: my_catalog mssql_username: my_user mssql_password: ${secrets:mssql_pass} mssql_encrypt: true mssql_trust_server_certificate: true ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Microsoft SQL Server as a data connector in Spice. [MSSQL (Microsoft SQL Server) Connector](https://github.com/spiceai/cookbook/tree/trunk/mssql#readme) --- # MySQL Data Connector MySQL is an open-source relational database management system that uses structured query language (SQL) for managing and manipulating databases. The MySQL Data Connector enables federated/accelerated SQL queries on data stored in MySQL databases. ``` datasets: - from: mysql:mytable name: my_dataset params: mysql_host: localhost mysql_tcp_port: 3306 mysql_db: my_database mysql_user: my_user mysql_pass: ${secrets:mysql_pass} mysql_pool_min: 1 mysql_pool_max: 5 ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `mysql:database_name.table_name` where `database_name` is the fully-qualified table name in the SQL server. If the `database_name` is omitted in the `from` field, the connector will use the database specified in the `mysql_db` parameter. If the `mysql_db` parameter is not provided, it will default to the user's default database. These two examples are identical: ``` datasets: - from: mysql:mytable name: my_dataset params: mysql_db: my_database ... ``` ``` datasets: - from: mysql:my_database.mytable name: my_dataset params: ... ``` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: mysql:path.to.my_dataset name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords) or any of the following keywords that are reserved by MySQL: * `PARTITION` ### `params`[​](#params "Direct link to params") The MySQL data connector can be configured by providing the following `params`. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:my_mysql_conn_string}`. | Parameter Name | Description | | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `mysql_connection_string` | The connection string to use to connect to the MySQL server. This can be used instead of providing individual connection parameters. | | `mysql_host` | The hostname of the MySQL server. | | `mysql_tcp_port` | The port of the MySQL server. | | `mysql_db` | The name of the database to connect to. | | `mysql_user` | The MySQL username. | | `mysql_pass` | The password to connect with. | | `mysql_sslmode` | Optional. Specifies the SSL/TLS behavior for the connection, supported values:
- `required`: (default) Require TLS and verify the server certificate against system root CAs and domain name (equivalent to `verify_identity`). Set `mysql_sslrootcert` to verify against a specific CA bundle instead of the system trust store.
- `preferred`: Attempt a TLS connection but skip certificate and hostname verification (invalid or self-signed certificates are accepted). The connection is still encrypted and does **not** fall back to plaintext — if the server does not support TLS, the connection fails. Not recommended for production, as it does not protect against man-in-the-middle attacks.
- `disabled`: Do not attempt to use an SSL connection, even if the server supports it. | | `mysql_sslrootcert` | Optional parameter specifying the path to a custom PEM certificate that the connector will trust. | | `mysql_time_zone` | Optional. Specifies connection time zone. Default is `+00:00` (UTC). Accepts:
- Fixed offsets (e.g., `+02:00`).
- IANA time zone names (e.g., `America/Los_Angeles`), if supported by the MySQL server.
- `system`: The MySQL server host’s OS time zone.
- `local_system`: The local runtime OS time zone. | | `mysql_pool_min` | The minimum number of connections to keep open in the pool, lazily created when requested. Default: `1` | | `mysql_pool_max` | The maximum number of connections to allow in the pool. Default: `5` | ### `metrics`[​](#metrics "Direct link to metrics") The MySQL data connector supports the following optional [component metrics](/docs/v1.10/features/observability/component_metrics): | Metric Name | Type | Description | | ------------------------------------ | ------- | ------------------------------------------------------------------------------------------------------------------ | | `connection_count` | Gauge | Gauge of active connections to the database server | | `connections_in_pool` | Gauge | Gauge of active connections that are idling in the pool | | `active_wait_requests` | Gauge | Gauge of requests that are waiting for a connection to be returned to the pool | | `create_failed` | Counter | Counter of connections that failed to be created | | `discarded_superfluous_connection` | Counter | Counter of connections that were closed because there were already enough idle connections in the pool | | `discarded_unestablished_connection` | Counter | Counter of connections that were closed because they could not be established | | `dirty_connection_return` | Counter | Counter of connections that were returned to the pool but were dirty (ie. open transactions, pending queries, etc) | | `discarded_expired_connection` | Counter | Counter of connections that were discarded because they were expired by the pool constraints (i.e. TTL expired) | | `resetting_connection` | Counter | Counter of connections that were reset | | `discarded_error_during_cleanup` | Counter | Counter of connections that were discarded because they returned an error during cleanup | | `connection_returned_to_pool` | Counter | Counter of connections that were returned to the pool | These metrics are not enabled by default, enable them by setting the `metrics` parameter: ``` datasets: - from: mysql:mytable name: my_dataset metrics: - name: connection_count - name: connections_in_pool - name: active_wait_requests - name: create_failed - name: discarded_superfluous_connection - name: discarded_unestablished_connection - name: dirty_connection_return - name: discarded_expired_connection - name: resetting_connection - name: discarded_error_during_cleanup - name: connection_returned_to_pool params: ¶ms mysql_host: localhost mysql_tcp_port: 3306 mysql_user: my_user mysql_pass: ${secrets:mysql_pass} ``` ## Types[​](#types "Direct link to Types") The table below shows the MySQL data types supported, along with the type mapping to Apache Arrow types in Spice. | MySQL Type | Arrow Type | | ------------ | ------------------------------ | | `TINYINT` | `Int8` | | `SMALLINT` | `Int16` | | `INT` | `Int32` | | `MEDIUMINT` | `Int32` | | `BIGINT` | `Int64` | | `DECIMAL` | `Decimal128` / `Decimal256` | | `FLOAT` | `Float32` | | `DOUBLE` | `Float64` | | `DATETIME` | `Timestamp(Microsecond, None)` | | `TIMESTAMP` | `Timestamp(Microsecond, None)` | | `YEAR` | `Int16` | | `TIME` | `Time64(Nanosecond)` | | `DATE` | `Date32` | | `CHAR` | `Utf8` | | `BINARY` | `Binary` | | `VARCHAR` | `Utf8` | | `VARBINARY` | `Binary` | | `TINYBLOB` | `Binary` | | `TINYTEXT` | `Utf8` | | `BLOB` | `Binary` | | `TEXT` | `Utf8` | | `MEDIUMBLOB` | `Binary` | | `MEDIUMTEXT` | `Utf8` | | `LONGBLOB` | `LargeBinary` | | `LONGTEXT` | `LargeUtf8` | | `JSON` | `LargeUtf8` | | `SET` | `Utf8` | | `ENUM` | `Dictionary(UInt16, Utf8)` | | `BIT` | `UInt64` | note * The MySQL `TIMESTAMP` value is [retrieved as a UTC time value](https://dev.mysql.com/doc/refman/8.4/en/datetime.html) by default. Use the `mysql_time_zone` configuration parameter to specify the desired time zone for interpreting `TIMESTAMP` values during data retrieval. ## Examples[​](#examples "Direct link to Examples") ### Connecting using username and password[​](#connecting-using-username-and-password "Direct link to Connecting using username and password") ``` datasets: - from: mysql:path.to.my_dataset name: my_dataset params: mysql_host: localhost mysql_tcp_port: 3306 mysql_db: my_database mysql_user: my_user mysql_pass: ${secrets:mysql_pass} ``` ### Connecting using SSL[​](#connecting-using-ssl "Direct link to Connecting using SSL") ``` datasets: - from: mysql:path.to.my_dataset name: my_dataset params: mysql_host: localhost mysql_tcp_port: 3306 mysql_db: my_database mysql_user: my_user mysql_pass: ${secrets:mysql_pass} mysql_sslmode: preferred mysql_sslrootcert: ./custom_cert.pem ``` ### Connecting using a Connection String[​](#connecting-using-a-connection-string "Direct link to Connecting using a Connection String") ``` datasets: - from: mysql:path.to.my_dataset name: my_dataset params: mysql_connection_string: mysql://${secrets:my_user}:${secrets:my_password}@localhost:3306/my_db ``` ### Connecting to the default database[​](#connecting-to-the-default-database "Direct link to Connecting to the default database") ``` datasets: - from: mysql:mytable name: my_dataset params: mysql_host: localhost mysql_tcp_port: 3306 mysql_user: my_user mysql_pass: ${secrets:mysql_pass} ``` ### With custom connection pool settings[​](#with-custom-connection-pool-settings "Direct link to With custom connection pool settings") ``` datasets: - from: mysql:path.to.my_dataset name: my_dataset params: mysql_host: localhost mysql_tcp_port: 3306 mysql_db: my_database mysql_user: my_user mysql_pass: ${secrets:mysql_pass} mysql_pool_min: 5 mysql_pool_max: 10 ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure MySQL as a data connector in Spice. [MySQL Data Connector](https://github.com/spiceai/cookbook/tree/trunk/mysql/connector#readme) * A cookbook recipe to configure AWS RDS Aurora (MySQL Compatible) as a data connector in Spice. [AWS RDS Aurora (MySQL Data Connector)](https://github.com/spiceai/cookbook/tree/trunk/mysql/rds-aurora#readme) * A cookbook recipe to configure Planetscale as a data connector in Spice. [Planetscale (MySQL Data Connector)](https://github.com/spiceai/cookbook/tree/trunk/mysql/planetscale#readme) --- # ODBC Data Connector ODBC (Open Database Connectivity) is a standard API that allows applications to connect to and interact with various database management systems using a common interface. To connect to any ODBC database for federated/accelerated SQL queries, specify `odbc` as the selector in the `from` value for the dataset. The `odbc_connection_string` parameter is required. warning Spice must be [built with the `odbc` feature](#building-spice-with-odbc), and the host/container must have a [valid ODBC configuration](https://www.unixodbc.org/odbcinst.html). The published `spiceai/spiceai` Docker images do **not** include ODBC support — they are built without the `odbc` feature and do not ship an ODBC Driver Manager. To run ODBC in a container, [bake your own image](#baking-an-image-with-odbc-support) from an ODBC-enabled build. ``` datasets: - from: odbc:path.to.my_dataset name: my_dataset params: odbc_connection_string: Driver={Foo Driver};Host=db.foo.net;Param=Value ``` An ODBC connection requires a compatible ODBC driver and valid driver configuration. ODBC drivers are available from their respective vendors. Here are a few examples: * [PostgreSQL](https://odbc.postgresql.org/) * [MySQL](https://dev.mysql.com/downloads/connector/odbc/) * [Databricks](https://www.databricks.com/spark/odbc-drivers-download) * [AWS Athena](https://docs.aws.amazon.com/athena/latest/ug/connect-with-odbc.html) Non-Windows systems additionally require the installation of an ODBC Driver Manager like `unixodbc`. * Ubuntu: `sudo apt-get install unixodbc` * MacOS: `brew install unixodbc` info For the best `JOIN` performance, ensure all ODBC datasets from the same database are configured with the exact same `odbc_connection_string` in Spice. ## ODBC Connection String[​](#odbc-connection-string "Direct link to ODBC Connection String") The ODBC connection string requires the use of an installed and registered driver based on your system type: * Unix systems; ODBC driver installations can be managed using [unixODBC](https://www.unixodbc.org/), or directly edited through `/etc/odbc.ini` or `/etc/odbcinst.ini`. For example, in the [Databricks DSN Connection Setup Guide](https://docs.databricks.com/en/integrations/odbc/dsn.html#linux) for Linux. * Windows systems; ODBC driver installations are managed using the [ODBC Data Source Administrator](https://support.microsoft.com/en-au/office/administer-odbc-data-sources-b19f856b-5b9b-48c9-8b93-07484bfab5a7). For an example Unix system with an installed PostgreSQL driver where the contents of `/etc/odbcinst.ini` is: ``` [PostgreSQL Unicode] Description=PostgreSQL ODBC driver (Unicode version) Driver=psqlodbcw.so Setup=libodbcpsqlS.so Debug=0 CommLog=1 UsageCount=1 ``` The Spice Runtime can use this driver installation where `Driver={PostgreSQL Unicode}` is used in the connection string, like: ``` datasets: - from: odbc:my_table name: my_dataset params: odbc_connection_string: Driver={PostgreSQL Unicode};Server=localhost;Port=5432;Database=postgres;Uid=myuser;Pwd=mypass ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `odbc:path.to.my.dataset` where `path.to.my.dataset` is the table name in the ODBC-supporting server to read from. ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: odbc:my.cool.table name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter | Type | Description | | ----------------------------- | -------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `sql_dialect` | string | Override what SQL dialect is used for the ODBC connection. Supports `postgresql`, `mysql`, `sqlite`, `athena` or `databricks` values. Default is unset (auto-detected). | | `odbc_max_bytes_per_batch` | number (bytes) | Maximum number of bytes transferred in each query record batch. A lower value may improve performance on low-memory systems. Default is `512_000_000`. | | `odbc_max_num_rows_per_batch` | number (rows) | Maximum number of rows transferred in each query record batch. A higher value may speed up query results, but requires more memory in conjunction with `odbc_max_bytes_per_batch`. Default is `4000`. | | `odbc_max_text_size` | number (bytes) | A limit for the maximum size of text columns transmitted between the ODBC driver and the Runtime. Default is unset (allocates driver-reported max column size). | | `odbc_max_binary_size` | number (bytes) | A limit for the maximum size of binary columns transmitted between the ODBC driver and the Runtime. Default is unset (allocates driver-reported max column size). | | `odbc_connection_string` | string | Connection string to use to connect to the ODBC server | ``` datasets: - from: odbc:path.to.my_dataset name: my_dataset params: odbc_connection_string: Driver={Foo Driver};Host=db.foo.net;Param=Value ``` ## Selecting SQL Dialect[​](#selecting-sql-dialect "Direct link to Selecting SQL Dialect") The default SQL dialect may not be supported by every ODBC connection. The `sql_dialect` parameter supports overriding the selected SQL dialect for a specified connection. The runtime will attempt to detect the dialect to use for a connection based on the contents of `Driver=` in the `odbc_connection_string`. The runtime will detect the correct SQL dialect for the following connection types, when setup with a standard driver configuration: * PostgreSQL * MySQL * SQLite * Databricks * AWS Athena These connection types are also the supported values for overriding dialect in `sql_dialect`, in lowercase format: `postgresql`, `mysql`, `sqlite`, `databricks`, `athena`. For example, overriding the dialect for your connection to a `postgresql` style dialect: ``` datasets: - from: odbc:path.to.my_dataset name: my_dataset params: sql_dialect: postgresql odbc_connection_string: Driver={Foo Driver};Host=db.foo.net;Param=Value ``` ## Building Spice with ODBC[​](#building-spice-with-odbc "Direct link to Building Spice with ODBC") ODBC support is not included in the released binaries or in the published `spiceai/spiceai` Docker images. To use ODBC with Spice, you need to [checkout and compile the code](https://github.com/spiceai/spiceai/blob/trunk/CONTRIBUTING.md#building) with the `--features odbc` flag (`cargo build --release --features odbc`). To build a container image, pass the feature through to the repository `Dockerfile`, which also installs the `unixodbc` Driver Manager when the feature is enabled: ``` docker build --build-arg CARGO_FEATURES=release,models,odbc -t spiceai-odbc:local . ``` ## Baking an image with ODBC Support[​](#baking-an-image-with-odbc-support "Direct link to Baking an image with ODBC Support") There are many dozens of ODBC adapters; this recipe covers making a custom image and configuring it to work with Spice. The base image must already contain an ODBC-enabled `spiced` build — the published `spiceai/spiceai` images do not (see [Building Spice with ODBC](#building-spice-with-odbc)). ``` FROM spiceai-odbc:local RUN apt update \ && apt install --yes libsqliteodbc --no-install-recommends \ && rm -rf /var/lib/{apt,dpkg,cache,log} ``` Build the container: ``` docker build -t spice-libsqliteodbc . ``` Validate that the ODBC configuration was updated to reference the newly installed driver: Note Since `libsqliteodbc` is vendored by Debian, the package install hooks append the driver configuration to `/etc/odbcinst.ini`. When using a custom driver (e.g. [Databricks Simba](https://www.databricks.com/spark/odbc-drivers-download)), it is your responsibility to update `/etc/odbcinst.ini` to point at the location of the newly installed driver. ``` $ docker run --entrypoint /bin/bash -it spice-libsqliteodbc root@f8ceccc94d6a:/# odbcinst -j unixODBC 2.3.11 DRIVERS............: /etc/odbcinst.ini SYSTEM DATA SOURCES: /etc/odbc.ini FILE DATA SOURCES..: /etc/ODBCDataSources USER DATA SOURCES..: /root/.odbc.ini SQLULEN Size.......: 8 SQLLEN Size........: 8 SQLSETPOSIROW Size.: 8 root@f8ceccc94d6a:/# cat /etc/odbcinst.ini [SQLite] Description=SQLite ODBC Driver Driver=libsqliteodbc.so Setup=libsqliteodbc.so UsageCount=1 [SQLite3] Description=SQLite3 ODBC Driver Driver=libsqlite3odbc.so Setup=libsqlite3odbc.so UsageCount=1 ``` ### `test.db`[​](#testdb "Direct link to testdb") To fully test the image, make an example SQLite database (`test.db`) and spicepod on your host: ``` $ sqlite3 test.db SQLite version 3.43.2 2023-10-10 13:08:14 Enter ".help" for usage hints. sqlite> create table spice_test (name text not null); sqlite> insert into spice_test values ("Lala"); sqlite> insert into spice_test values ("Hopper"); sqlite> insert into spice_test values ("Linus"); ``` ### `spicepod.yaml`[​](#spicepodyaml "Direct link to spicepodyaml") Make sure that the `DRIVER` parameter matches the name of the driver section in `odbcinst.ini`. ``` version: v1 kind: Spicepod name: sqlite datasets: - from: odbc:spice_test name: spice_test mode: read acceleration: enabled: false params: odbc_connection_string: DRIVER={SQLite3};SERVER=localhost;DATABASE=test.db;Trusted_connection=yes ``` All together now: ``` $ docker run -p8090:8090 -p50051:50051 -v $(pwd)/spicepod.yaml:/spicepod.yaml -v $(pwd)/test.db:/test.db -it spice-libsqliteodbc --http=0.0.0.0:8090 --flight=0.0.0.0:50051 $ spice sql Welcome to the interactive Spice.ai SQL Query Utility! Type 'help' for help. show tables; -- list available tables sql> show tables; +------------+ | table_name | +------------+ | spice_test | +------------+ Query took: 0.059305583 seconds. 1/1 rows displayed. sql> select * from spice_test; +--------+ | name | +--------+ | Hopper | | Lala | | Linus | +--------+ Query took: 1.8504053329999999 seconds. 3/3 rows displayed. ``` ## Examples[​](#examples "Direct link to Examples") ### Connecting to an SQLite database[​](#connecting-to-an-sqlite-database "Direct link to Connecting to an SQLite database") ``` version: v1 kind: Spicepod name: sqlite datasets: - from: odbc:spice_test name: spice_test mode: read acceleration: enabled: false params: odbc_connection_string: DRIVER={SQLite3};SERVER=localhost;DATABASE=test.db;Trusted_connection=yes ``` ### Connecting to Postgres[​](#connecting-to-postgres "Direct link to Connecting to Postgres") Ensure that the Postgres ODBC driver is installed. On Unix systems, this will create an entry in `/etc/odbcinst.ini` similar to: ``` [PostgreSQL Unicode] Description=PostgreSQL ODBC driver (Unicode version) Driver=psqlodbcw.so Setup=libodbcpsqlS.so Debug=0 CommLog=1 UsageCount=1 ``` Then, in your `spicepod.yaml` the `odbc_connection_string` parameter can be used for the ODBC connection string: ``` version: v1 kind: Spicepod name: odbc-demo datasets: - from: odbc:taxi_trips name: taxi_trips params: odbc_connection_string: Driver={PostgreSQL Unicode};Server=localhost;Port=5432;Database=spice_demo;Uid=postgres ``` See the [ODBC Cookbook](https://github.com/spiceai/cookbook/blob/trunk/odbc/README.md) for more help on getting started with ODBC and Postgres. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure ODBC as a data connector in Spice. [ODBC Data Connector](https://github.com/spiceai/cookbook/tree/trunk/odbc#readme) --- # Oracle Data Connector The Oracle Data Connector enables SQL queries on data stored in Oracle databases, including on-premises instances, Oracle Cloud User-Managed Databases, and Oracle Cloud Autonomous Databases (ADB). ``` datasets: - from: oracle:"SH"."PRODUCTS" name: my_dataset params: oracle_host: localhost oracle_port: 1521 oracle_username: scott oracle_password: ${secrets:oracle_password} oracle_service_name: XEPDB1 ``` Limitations 1. Only basic filter predicates are currently pushed down to the Oracle database. Full query federation is not currently supported. Joins, subqueries, and complex query constructs are not pushed down to the Oracle database; these operations are performed in-memory after data retrieval. **Enable [Data Acceleration](/docs/v1.10/features/data-acceleration) for full federation support**. 2. The Oracle connector does not support filter push-down optimization for datetime columns. Filtering on these columns is performed in-memory after data retrieval. 3. The following Oracle data types are not currently supported; columns with these types will be ignored: `INTERVAL YEAR TO MONTH` (Code 182), `INTERVAL DAY TO SECOND` (Code 183), `UROWID` (Code 208), `BFILE` (Code 114), `JSON` (Code 119). ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `oracle:"schema_name"."table_name"` where both schema and table names should be quoted to handle case sensitivity properly. Example: ``` datasets: - from: oracle:"SH"."PRODUCTS" name: products params: oracle_host: localhost oracle_username: scott oracle_password: ${secrets:ORACLE_PASSWORD} ``` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: oracle:"SH"."PRODUCTS" name: products params: ... ``` ``` SELECT COUNT(*) FROM products; ``` ``` +----------+ | count(*) | +----------+ | 10500 | +----------+ ``` ### `params`[​](#params "Direct link to params") The Oracle data connector can be configured by providing the following `params`. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:MY_ORACLE_PASSWORD}`. | Parameter Name | Description | | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `oracle_connection_string` | The connection string to use to connect to the Oracle server. This can be a TNS alias from `tnsnames.ora` for local mTLS/Wallet connections or an [Easy Connect](https://download.oracle.com/ocomdocs/global/Oracle-Net-Easy-Connect-Plus.pdf) string. | | `oracle_host` | The hostname or IP address of the Oracle Database instance. Required when not using `oracle_connection_string`. | | `oracle_port` | Optional. The port of the Oracle Database server. Default: `1521` | | `oracle_username` | The Oracle username. Required. | | `oracle_password` | The password to connect with. Required. | | `oracle_service_name` | The Oracle Database service name to connect to. Default: `XEPDB1` | | `oracle_wallet_sso_cert` | The base64-encoded `cwallet.sso` (wallet auto-login certificate) to use for mTLS authentication with Oracle Cloud. | | `oracle_wallet` | Specifies the Oracle wallet directory for mTLS connections — either an existing/pre-downloaded wallet, or the destination to save the decoded `oracle_wallet_sso_cert`. The `.oracle` default applies only as the save destination when `oracle_wallet_sso_cert` is set; otherwise no wallet directory is used unless this is specified. | ## Types[​](#types "Direct link to Types") The table below shows the Oracle data types supported, along with the type mapping to Apache Arrow types in Spice. | Oracle Type | Arrow Type | | -------------------------------- | -------------------------------------------------------------------------------- | | `ROWID` | `Utf8` | | `CHAR` | `Utf8` | | `NCHAR` | `Utf8` | | `VARCHAR2` | `Utf8` | | `NVARCHAR2` | `Utf8` | | `LONG` | `Utf8` | | `CLOB` | `LargeUtf8` | | `NCLOB` | `LargeUtf8` | | `NUMBER` | `Int64` for integer types (scale=0, precision≤18), otherwise `Decimal128` | | `FLOAT` | `Float32` for precision≤24, otherwise `Float64` | | `BINARY_FLOAT` | `Float32` | | `BINARY_DOUBLE` | `Float64` | | `BOOLEAN` | `Boolean` | | `DATE` | `Date32` | | `TIMESTAMP` | `Timestamp(Second)` for precision=0, otherwise `Timestamp(Nanosecond)` | | `TIMESTAMP WITH TIME ZONE` | `Timestamp(Second, UTC)` for precision=0, otherwise `Timestamp(Nanosecond, UTC)` | | `TIMESTAMP WITH LOCAL TIME ZONE` | `Timestamp(Second, UTC)` for precision=0, otherwise `Timestamp(Nanosecond, UTC)` | | `RAW` | `Binary` | | `LONG RAW` | `Binary` | | `BLOB` | `LargeBinary` | note * The Oracle `TIMESTAMP WITH LOCAL TIME ZONE` value is retrieved as a UTC time value. * `TIMESTAMP`, `TIMESTAMP WITH TIME ZONE`, and `TIMESTAMP WITH LOCAL TIME ZONE` columns with non-zero precision are mapped to `Timestamp(Nanosecond)`. Timestamps outside the nanosecond range (approximately years 1677–2262) will silently return `1970-01-01 UTC`. This is an inherent limitation of Arrow's nanosecond timestamp representation. ## Examples[​](#examples "Direct link to Examples") ### Connecting to On-Premises Oracle Database[​](#connecting-to-on-premises-oracle-database "Direct link to Connecting to On-Premises Oracle Database") ``` datasets: - from: oracle:"SH"."PRODUCTS" name: products params: oracle_host: localhost oracle_port: 1521 oracle_username: scott oracle_password: ${secrets:ORACLE_PASSWORD} oracle_service_name: XEPDB1 ``` ### Connecting to Oracle Cloud Autonomous Database with mTLS (Wallet-based)[​](#connecting-to-oracle-cloud-autonomous-database-with-mtls-wallet-based "Direct link to Connecting to Oracle Cloud Autonomous Database with mTLS (Wallet-based)") ### Wallet Folder Exists Locally[​](#wallet-folder-exists-locally "Direct link to Wallet Folder Exists Locally") If your Oracle Cloud Autonomous Database wallet folder is available locally, specify its path using the `oracle_wallet` parameter. Set the `oracle_connection_string` to the TNS alias defined in your wallet's `tnsnames.ora` file. Example: ``` datasets: - from: oracle:"SALES" name: sales params: oracle_username: admin oracle_password: ${secrets:ORACLE_PASSWORD} oracle_connection_string: 'fgp1tqs1e_low' # TNS alias from tnsnames.ora oracle_wallet: '/path/to/wallet_folder' ``` ### Wallet Auto-Login (SSO) Certificate Provided via Application Secret[​](#wallet-auto-login-sso-certificate-provided-via-application-secret "Direct link to Wallet Auto-Login (SSO) Certificate Provided via Application Secret") If your Oracle Cloud Autonomous Database wallet folder is not available locally, provide the base64-encoded wallet auto-login (SSO) certificate (`cwallet.sso`) using the `oracle_wallet_sso_cert` parameter. Set the `oracle_connection_string` to the Easy Connect string from the *Database connection* section. ``` datasets: - from: oracle:"SALES" name: sales params: oracle_username: admin oracle_password: ${secrets:ORACLE_PASSWORD} oracle_wallet_sso_cert: ${secrets:oracle_wallet_sso_cert} oracle_connection_string: 'tcps://adb.us-sanjose-1.oraclecloud.com:1522/g81f1d1d5c853_fgc1e_low.adb.oraclecloud.com?ssl_server_dn_match=yes' ``` To generate a base64-encoded wallet certificate for use as a secret: ``` base64 -i cwallet.sso > cwallet.b64.txt ``` ### Connecting with Easy Connect string (TLS-only, no wallet required)[​](#connecting-with-easy-connect-string-tls-only-no-wallet-required "Direct link to Connecting with Easy Connect string (TLS-only, no wallet required)") ``` datasets: - from: oracle:"SALES" name: sales params: oracle_username: admin oracle_password: ${secrets:ORACLE_PASSWORD} oracle_connection_string: 'tcps://adb.us-sanjose-1.oraclecloud.com:1522/g81f1d1d5c853_fgc1e_low.adb.oraclecloud.com?ssl_server_dn_match=yes' ``` ## Installation Requirements[​](#installation-requirements "Direct link to Installation Requirements") The Oracle data connector requires the Oracle Instant Client or Oracle Database Client libraries to be installed on the system where Spice is running. Follow the [Oracle installation guide](https://oracle.github.io/odpi/) for your platform. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to connect to and accelerate data from an Oracle database in Spice. [Oracle Data Connector](https://github.com/spiceai/cookbook/blob/trunk/oracle/README.md) --- # PostgreSQL Data Connector PostgreSQL is an advanced open-source relational database management system known for its robustness, extensibility, and support for SQL compliance. The PostgreSQL Server Data Connector enables federated/accelerated SQL queries on data stored in PostgreSQL databases. ``` datasets: - from: postgres:my_table name: my_dataset params: ... ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `postgres:my_table` where `my_table` is the table identifer in the PostgreSQL server to read from. The fully-qualified table name (`database.schema.table`) can also be used in the `from` field. ``` datasets: - from: postgres:my_database.my_schema.my_table name: my_dataset params: ... ``` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: postgres:my_database.my_schema.my_table name: cool_dataset params: ... ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` ### `params`[​](#params "Direct link to params") The connection to PostgreSQL can be configured by providing the following `params`: | Parameter Name | Description | | ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `pg_host` | The hostname of the PostgreSQL server. | | `pg_port` | The port of the PostgreSQL server. | | `pg_db` | The name of the database to connect to. | | `pg_user` | The username to connect with. | | `pg_pass` | The password to connect with. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to load the password from a secret store, e.g. `${secrets:my_pg_pass}`. | | `pg_sslmode` | Optional. Specifies the SSL/TLS behavior for the connection, supported values:
- `verify-full`: (default) This mode requires an SSL connection, a valid root certificate, and the server host name to match the one specified in the certificate.
- `verify-ca`: This mode requires a TLS connection and a valid root certificate.
- `require`: This mode requires a TLS connection.
- `prefer`: This mode will try to establish a secure TLS connection if possible, but will connect insecurely if the server does not support TLS.
- `disable`: This mode will not attempt to use a TLS connection, even if the server supports it. | | `pg_sslrootcert` | Optional parameter specifying the path to a custom PEM certificate that the connector will trust. | | `pg_connection_pool_min_idle` | Optional. The minimum number of idle connections to keep open in the pool. Default is `1`. | | `connection_pool_size` | Optional. The maximum number of connections to keep open in the connection pool. Default is `5`. | ## Types[​](#types "Direct link to Types") The table below shows the PostgreSQL data types supported, along with the type mapping to Apache Arrow types in Spice. | PostgreSQL Type | Arrow Type | | --------------- | ----------------------------- | | `int2` | `Int16` | | `int4` | `Int32` | | `int8` | `Int64` | | `money` | `Int64` | | `float4` | `Float32` | | `float8` | `Float64` | | `numeric` | `Decimal128` | | `text` | `Utf8` | | `varchar` | `Utf8` | | `bpchar` | `Utf8` | | `uuid` | `Utf8` | | `bytea` | `Binary` | | `bool` | `Boolean` | | `json` | `Utf8` | | `timestamp` | `Timestamp(Nanosecond, None)` | | `timestamptz` | `Timestamp(Nanosecond, UTC)` | | `date` | `Date32` | | `time` | `Time64(Nanosecond)` | | `interval` | `Interval(MonthDayNano)` | | `point` | `FixedSizeList(Float64[2])` | | `int2[]` | `List(Int16)` | | `int4[]` | `List(Int32)` | | `int8[]` | `List(Int64)` | | `float4[]` | `List(Float32)` | | `float8[]` | `List(Float64)` | | `text[]` | `List(Utf8)` | | `bool[]` | `List(Boolean)` | | `bytea[]` | `List(Binary)` | | `geometry` | `Binary` | | `geography` | `Binary` | | `enum` | `Dictionary(Int8, Utf8)` | | Composite Types | `Struct` | info The Postgres federated queries may result in unexpected result types due to the difference in DataFusion and Postgres size increase rules. Please explicitly specify the expected output type of aggregation functions when writing query involving Postgres table in Spice. For example, rewrite `SUM(int_col)` into `CAST (SUM(int_col) as BIGINT`. ## Examples[​](#examples "Direct link to Examples") ### Connecting using Username/Password[​](#connecting-using-usernamepassword "Direct link to Connecting using Username/Password") ``` datasets: - from: postgres:my_database.my_schema.my_table name: my_dataset params: pg_host: localhost pg_port: 5432 pg_db: my_database pg_user: my_user pg_pass: ${secrets:my_pg_pass} ``` ### Connect using SSL[​](#connect-using-ssl "Direct link to Connect using SSL") ``` datasets: - from: postgres:my_database.my_schema.my_table name: my_dataset params: pg_host: localhost pg_port: 5432 pg_db: my_database pg_user: my_user pg_pass: ${secrets:my_pg_pass} pg_sslmode: verify-ca pg_sslrootcert: ./custom_cert.pem ``` ### Separate dataset/accelerator secrets[​](#separate-datasetaccelerator-secrets "Direct link to Separate dataset/accelerator secrets") Specify different secrets for a PostgreSQL source and acceleration: ``` datasets: - from: postgres:my_schema.my_table name: my_dataset params: pg_host: localhost pg_port: 5432 pg_db: my_database pg_user: my_user pg_pass: ${secrets:pg1_pass} acceleration: engine: postgres params: pg_host: localhost pg_port: 5433 pg_db: acceleration pg_user: two_user_two_furious pg_pass: ${secrets:pg2_pass} ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure PostgreSQL as a data connector in Spice. [PostgreSQL Data Accelerator](https://github.com/spiceai/cookbook/tree/trunk/postgres/accelerator#readme) * A cookbook recipe to configure AWS RDS for PostgreSQL as a data connector in Spice. [AWS RDS for PostgreSQL](https://github.com/spiceai/cookbook/tree/trunk/postgres/rds#readme) * A cookbook recipe to configure Supabase a data connector in Spice. [Supabase (PostgreSQL Data Connector)](https://github.com/spiceai/cookbook/tree/trunk/postgres/supabase#readme) --- # Amazon Redshift Data Connector Amazon Redshift is a columnar OLAP database compatible with PostgreSQL. To connect Redshift to Spice, use the [PostgreSQL data connector](/docs/v1.10/components/data-connectors/postgres) and specify the Redshift cluster connection parameters. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Use the format `postgres:schema.table` to reference a Redshift table. The connector parameters should match your Redshift cluster settings. ### Example Spicepod[​](#example-spicepod "Direct link to Example Spicepod") ``` version: v1beta1 kind: Spicepod name: tpch-read datasets: - from: postgres:public.customer name: customer params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.lineitem name: lineitem params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.nation name: nation params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.orders name: orders params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.part name: part params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.partsupp name: partsupp params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.region name: region params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true - from: postgres:public.supplier name: supplier params: pg_host: ${secrets:PG_HOST} pg_port: 5439 pg_sslmode: prefer pg_db: dev pg_user: ${secrets:PG_USER} pg_pass: ${secrets:PG_PASS} acceleration: enabled: true ``` ### Parameters[​](#parameters "Direct link to Parameters") | Parameter Name | Description | | -------------- | ------------------------------------------------------------------------------------ | | `pg_host` | Hostname or IP address of the Redshift cluster | | `pg_port` | The PostgreSQL TCP port. Redshift uses port `5439` by default — set this explicitly. | | `pg_sslmode` | SSL mode (e.g., `prefer`) | | `pg_db` | Database name | | `pg_user` | Username for authentication | | `pg_pass` | Password for authentication (use secret reference) | ## Supported Types[​](#supported-types "Direct link to Supported Types") Redshift types are mapped to PostgreSQL types. See the [PostgreSQL connector documentation](/docs/v1.10/components/data-connectors/postgres) for details on supported types and configuration. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For details, see the [secret stores documentation](/docs/v1.10/components/secret-stores) and [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## References[​](#references "Direct link to References") * [Amazon Redshift Documentation](https://docs.aws.amazon.com/redshift/latest/mgmt/welcome.html) * [PostgreSQL Connector Documentation](/docs/v1.10/components/data-connectors/postgres) --- # S3 Data Connector The S3 Data Connector enables federated SQL querying on files stored in S3 or S3-compatible systems (e.g., MinIO, Cloudflare R2). If a folder path is specified as the dataset source, all files within the folder will be loaded. File formats are specified using the `file_format` parameter, as described in [Object Store File Formats](/docs/v1.10/components/data-connectors/#object-store-file-formats). ``` datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") S3-compatible URI to a folder or file, in the format `s3:///` Example: `from: s3://my-bucket/path/to/file.parquet` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. Example: ``` datasets: - from: s3://s3-bucket-name/taxi_sample.csv name: cool_dataset params: file_format: csv ``` ``` SELECT COUNT(*) FROM cool_dataset; ``` ``` +----------+ | count(*) | +----------+ | 6001215 | +----------+ ``` The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `file_format` | Specifies the data format. Required if it cannot be inferred from the object URI. Options: `parquet`, `csv`, `json`. Refer to [Object Store File Formats](/docs/v1.10/components/data-connectors/#object-store-file-formats) for details. | | `s3_endpoint` | S3 endpoint URL (e.g., for MinIO). Default is the region endpoint. E.g. `s3_endpoint: https://my.minio.server` | | `s3_region` | S3 bucket region. Default: `us-east-1`. | | `client_timeout` | Timeout for S3 operations. Default: `30s`. | | `hive_partitioning_enabled` | Enable partitioning using hive-style partitioning from the folder structure. Defaults to `false` | | `s3_auth` | Authentication type. Options: `public`, `key` and `iam_role`. Defaults to `public`. If set to `key` the `s3_key` and `s3_secret` parameters must also be set. If set to `iam_role` the credentials will be loaded from environment variables or IAM roles (see [Authentication](#authentication) for details). | | `s3_key` | Access key (e.g. `AWS_ACCESS_KEY_ID` for AWS). Requires `s3_auth` is set to `key`. | | `s3_secret` | Secret key (e.g. `AWS_SECRET_ACCESS_KEY` for AWS). Requires `s3_auth` is set to `key`. | | `s3_session_token` | Session token (e.g. `AWS_SESSION_TOKEN` for AWS) for temporary credentials. Requires `s3_auth` is set to `key`. | | `s3_versioning` | Enables support for S3 buckets with [S3 Versioning](https://docs.aws.amazon.com/AmazonS3/latest/userguide/Versioning.html). Options: `enabled` and `disabled`. Defaults to `enabled`. | | `allow_http` | Enables insecure HTTP connections to `s3_endpoint`. Defaults to `false`. | | `schema_source_path` | Specifies the URL used to infer the dataset schema. Default to the most recently modified file | For additional CSV, JSON, and Parquet specific parameters, see [File Formats](/docs/v1.10/reference/file_format). ## Authentication[​](#authentication "Direct link to Authentication") No authentication is required for public endpoints. For private buckets, set `s3_auth` to `key` or `iam_role`. If `s3_auth` is set to `iam_role`, the connector will automatically load credentials from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials (in .aws/credentials) [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile (in .aws/config) [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts ([IRSA](https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html)) * Relies on the environment variables `AWS_WEB_IDENTITY_TOKEN_FILE` and `AWS_ROLE_ARN` to be present. * These environment variables are automatically injected by EKS when the IAM Role is annotated on the pod Service Account. 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint. * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3 permissions (e.g., `s3:ListBucket`, `s3:GetObject`) to access the files. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. kube2iam [`kube2iam`](https://github.com/jtblin/kube2iam) is a project that provides IAM roles to Kubernetes pods based on annotations. It has been superceded by [IAM Roles for service accounts (IRSA)](https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html), which should be preferred for new deployments. Spice requires `kube2iam >= 0.12` - versions prior to [`0.12`](https://github.com/jtblin/kube2iam/releases/tag/0.12.0) only supported IMDSv1. ### Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") Minimum IAM policy for S3 access: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::company-bucketname-datasets" }, { "Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::company-bucketname-datasets/*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | --------------- | ----------------------------------------------------- | | `s3:ListBucket` | Required. Allows scanning all objects from the bucket | | `s3:GetObject` | Required. Allows fetching objects | ## Types[​](#types "Direct link to Types") Refer to [Object Store Data Types](/docs/v1.10/reference/datatypes/object_store) for data type mapping from object store files to arrow data type. ## Examples[​](#examples "Direct link to Examples") ### Public bucket Example[​](#public-bucket-example "Direct link to Public bucket Example") Create a dataset named `taxi_trips` from a public S3 folder. ``` - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet ``` ### MinIO Example[​](#minio-example "Direct link to MinIO Example") Create a dataset named `cool_dataset` from a Parquet file stored in MinIO. ``` - from: s3://s3-bucket-name/path/to/parquet/cool_dataset.parquet name: cool_dataset params: s3_endpoint: http://my.minio.server s3_region: 'us-east-1' # Best practice for MinIO allow_http: true ``` ### Hive Partitioning Example[​](#hive-partitioning-example "Direct link to Hive Partitioning Example") Hive partitioning is a data organization technique that improves query performance by storing data in a hierarchical directory structure based on partition column values. This allows for efficient data retrieval by skipping unnecessary data scans. For example, a dataset partitioned by year, month, and day might have a directory structure like: ``` s3://bucket/dataset/year=2024/month=03/day=15/data_file.parquet s3://bucket/dataset/year=2024/month=03/day=16/data_file.parquet ``` Spice can automatically infer these partition columns from the directory structure when `hive_partitioning_enabled` is set to `true`. ``` version: v1 kind: Spicepod name: hive_data datasets: - from: s3://spiceai-public-datasets/hive_partitioned_data/ name: hive_data_infer params: file_format: parquet hive_partitioning_enabled: true ``` ### Schema Source Path example[​](#schema-source-path-example "Direct link to Schema Source Path example") Use `schema_source_path` to speed up dataset registration by specifying a URL to use to infer the schema. ``` - from: s3://spiceai-demo-datasets/taxi_trips/ name: taxi_trips params: file_format: parquet schema_source_path: s3://spiceai-demo-datasets/taxi_trips/2014/1/trips_01.parquet # or s3://spiceai-demo-datasets/taxi_trips/2014/1/ ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the \[secret stores documentation(../secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the \[using referenced secrets guide(../secret-stores#using-secrets). ## Limitations[​](#limitations "Direct link to Limitations") Performance Considerations When using the S3 Data connector without acceleration, data is loaded into memory during query execution. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. Memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/v1.10/components/data-accelerators/duckdb) and [`sqlite`](/docs/v1.10/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. Each query retrieves data from the S3 source, which might result in significant network requests and bandwidth consumption. This can affect network performance and incur costs related to data transfer from S3. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure S3 as a data connector in Spice. [S3 Data Connector](https://github.com/spiceai/cookbook/tree/trunk/s3#readme) --- # SharePoint Data Connector The SharePoint Data Connector enables federated SQL queries on documents stored in SharePoint. ``` datasets: - from: sharepoint:drive:Documents/path:/top_secrets/ name: important_documents params: sharepoint_client_id: ${secrets:SPICE_SHAREPOINT_CLIENT_ID} sharepoint_tenant_id: ${secrets:SPICE_SHAREPOINT_TENANT_ID} sharepoint_client_secret: ${secrets:SPICE_SHAREPOINT_CLIENT_SECRET} ``` #### Example[​](#example "Direct link to Example") ``` SELECT * FROM important_documents limit 1 ``` Returns ```` [ { "created_by_id": "cbccd193-f9f1-4603-b01d-ff6f3e6f2108", "created_by_name": "Jack Eadie", "created_at": "2024-09-09T04:57:00", "c_tag": "\"c:{BD4D130F-2C95-4E59-9F93-85BD0A9E1B19},1\"", "e_tag": "\"{BD4D130F-2C95-4E59-9F93-85BD0A9E1B19},1\"", "id": "01YRH3MPAPCNG33FJMLFHJ7E4FXUFJ4GYZ", "last_modified_by_id": "cbccd193-f9f1-4603-b01d-ff6f3e6f2108", "last_modified_by_name": "Jack Eadie", "last_modified_at": "2024-09-09T04:57:00", "name": "ngx_google_perftools_module.md", "size": 959, "web_url": "https://spiceai.sharepoint.com/Shared%20Documents/md/ngx_google_perftools_module.md", "content": "# Module ngx_google_perftools_module\n\nThe `ngx_google_perftools_module` module (0.6.29) enables profiling of nginx worker processes using [Google Performance Tools](https://github.com/gperftools/gperftools). The module is intended for nginx developers.\n\nThis module is not built by default, it should be enabled with the `--with-google_perftools_module` configuration parameter.\n\n> **Note:** This module requires the [gperftools](https://github.com/gperftools/gperftools) library.\n\n## Example Configuration\n\n```nginx\ngoogle_perftools_profiles /path/to/profile;\n```\n\nProfiles will be stored as `/path/to/profile.`.\n\n## Directives\n\n### google_perftools_profiles\n\n- **Syntax:** `google_perftools_profiles file;`\n- **Default:** —\n- **Context:** `main`\n\nSets a file name that keeps profiling information of nginx worker process. The ID of the worker process is always a part of the file name and is appended to the end of the file name, after a dot.\n" } ] ```` Limitations The sharepoint connector does not yet support creating a dataset from a single file (e.g. an Excel spreadsheet). Datasets must be created from a folder of documents (see [Document Support](/docs/v1.10/components/data-connectors/#document-support)). ## Configuration[​](#configuration "Direct link to Configuration") ### Parameters[​](#parameters "Direct link to Parameters") | Name | Required? | Description | | -------------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | | `sharepoint_client_id` | **Yes** | The client ID of the Azure AD (Entra) application | | `sharepoint_tenant_id` | **Yes** | The tenant ID of the Azure AD (Entra) application. | | `sharepoint_client_secret` | Optional | For service principal authentication. The client secret of the Azure AD (Entra) application. | | `sharepoint_bearer_token` | Optional | For user authentication. The bearer access token obtained from the OAuth2 flow (see `spice login sharepoint` [docs](/docs/v1.10/cli/reference/login)). | note Only one of `sharepoint_client_secret` or `sharepoint_bearer_token` is allowed. ### `from` formats[​](#from-formats "Direct link to from-formats") The `from` field in a SharePoint dataset takes the following format: ``` from: 'sharepoint::/:' ``` #### Drives[​](#drives "Direct link to Drives") `drive_type` in a SharePoint Connector `from` field supports the following types: | Drive Type | Description | Example | | ---------- | --------------------------- | ---------------------------------------------------------------- | | `drive` | The SharePoint drive's name | `from: sharepoint:drive:Documents/...` | | `driveId` | The SharePoint drive's ID | `from: sharepoint:driveId:b!Mh8opUGD80ec7zGXgX9r/...` | | `site` | A SharePoint site's name | `from: sharepoint:site:MySite/...` | | `siteId` | A SharePoint site's ID | `from: sharepoint:siteId:b!Mh8opUGD80ec7zGXgX9r/...` | | `group` | A SharePoint group's name | `from: sharepoint:group:MyGroup/...` | | `groupId` | A SharePoint group's ID | `from: sharepoint:groupId:b!Mh8opUGD80ec7zGXgX9r/...` | | `user` | A user's drive by user ID | `from: sharepoint:user:48d31887-5fad-4d73-a9f5-3c356e68a038/...` | | `me` | A user's OneDrive | `from: sharepoint:me/...` | note For the `me` drive type the user is identified based on `sharepoint_bearer_token` and cannot be used with `sharepoint_client_secret` For a name-based `drive_id`, the connector will attempt to resolve the name to an ID at startup. #### Subpaths[​](#subpaths "Direct link to Subpaths") Within a drive, the SharePoint connector can load documents from: | Description | Example | | -------------------------------- | ---------------------------------------------------------------------- | | The root of the drive | `from: sharepoint:me/root` | | A specific path within the drive | `from: sharepoint:drive:Documents/path:/top_secrets` | | A specific folder ID | `from: sharepoint:group:MyGroup/id:01QM2NJSNHBISUGQ52P5AJQ3CBNOXDMVNT` | ## Authentication[​](#authentication "Direct link to Authentication") As outlined in the [connector parameters](#parameters), the SharePoint connector supports two types of authentication: 1. Service principal authentication, by setting the `sharepoint_client_secret` parameter. 2. User authentication, by setting the `sharepoint_bearer_token` parameter. Generally this is obtained by running `spice login sharepoint` and following the OAuth2 flow. ### Creating an Enterprise Application[​](#creating-an-enterprise-application "Direct link to Creating an Enterprise Application") To use the SharePoint connector with service principal authentication, you will need to create an Azure AD application and grant it the necessary permissions. This will also support OAuth2 authentication for users within the tenant (i.e. `sharepoint_bearer_token`). 1. Create a new Azure AD application in the [Azure portal](https://portal.azure.com/#view/Microsoft_AAD_IAM/ActiveDirectoryMenuBlade/~/Overview). 2. Under the application's `API permissions`, add the following permissions: `Sites.Read.All`, `Files.Read.All`, `User.Read`, `GroupMember.Read.All` * For service principal authentication, Application permissions are required. * For user authentication, only delegated permissions are required. 3. (For user authentication): Under the applications's `Authentication`, add `http://localhost` as Mobile and desktop applications redirect URI. 4. Add `sharepoint_client_id` (from the `Application (Client) ID` field) and `sharepoint_tenant_id` to the connector configuration. 5. (For service principal authentication): Under the application's `Certificates & secrets`, create a new client secret. Use this for the `sharepoint_client_secret` parameter. ### Default Spice Application[​](#default-spice-application "Direct link to Default Spice Application") For your convenience, Spice AI maintains a default Entra (Azure AD) application that can be used for authentication against your SharePoint instance. This application requires OAuth2 authentication. To use it: ``` datasets: - from: sharepoint:me/root # Set the drive and subpath as needed. name: my_data params: sharepoint_client_id: f2b3116e-b4c4-464f-80ec-73cd9d9886b4 sharepoint_tenant_id: #${env:TENANT_ID} sharepoint_bearer_token: ${secrets:SPICE_SHAREPOINT_BEARER_TOKEN} ``` And set the `SPICE_SHAREPOINT_BEARER_TOKEN` secret via: ``` spice login sharepoint --tenant-id $TENANT_ID --client-id f2b3116e-b4c4-464f-80ec-73cd9d9886b4 ``` ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the [secret stores documentation](/docs/v1.10/components/secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the [using referenced secrets guide](/docs/v1.10/components/secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Sharepoint as a data connector in Spice. [SharePoint Data Connector](https://github.com/spiceai/cookbook/tree/trunk/sharepoint#readme) --- # Snowflake Data Connector The Snowflake Data Connector enables federated SQL queries across datasets in the [Snowflake Cloud Data Warehouse](https://www.snowflake.com/). ``` datasets: - from: snowflake:DATABASE.SCHEMA.TABLE name: table params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin ``` Hint Unquoted table identifiers should be UPPERCASED in the `from` field. See [Identifier resolution](https://docs.snowflake.com/en/sql-reference/identifiers-syntax#label-identifier-casing). ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") A Snowflake fully qualified table name (database.schema.table). For instance `snowflake:SNOWFLAKE_SAMPLE_DATA.TPCH_SF1.LINEITEM` or `snowflake:TAXI_DATA."2024".TAXI_TRIPS` ### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. The dataset name cannot be a \[reserved keyword(../../reference/spicepod/keywords) or any of the following keywords that are reserved by Snowflake: * `START` * `CONNECT` * `MATCH_RECOGNIZE` * `SAMPLE` * `TABLESAMPLE` * `FROM` ### `params`[​](#params "Direct link to params") | Parameter Name | Description | | ---------------------------------- | ----------------------------------------------------------------------------------------------------------------- | | `snowflake_warehouse` | Optional, specifies the [Snowflake Warehouse](https://docs.snowflake.com/en/user-guide/warehouses-tasks) to use | | `snowflake_role` | Optional, specifies the role to use for accessing Snowflake data | | `snowflake_account` | Required, specifies the Snowflake account-identifier | | `snowflake_username` | Required, specifies the Snowflake username to use for accessing Snowflake data | | `snowflake_password` | Required when `snowflake_auth_type` is `snowflake` (default). Specifies the Snowflake password for authentication | | `snowflake_private_key_path` | Required when `snowflake_auth_type` is `keypair`. Specifies the path to the private key file. | | `snowflake_private_key_passphrase` | Required when the private key is encrypted. Specifies the passphrase to decrypt the private key. | ## Auth[​](#auth "Direct link to Auth") The connector supports password-based and [key-pair](https://docs.snowflake.com/en/user-guide/key-pair-auth) authentication that must be configured using `spice login snowflake` or using \[Secrets Stores(../secret-stores). Login requires the account identifier ('orgname-accountname' format) - use [Finding the organization and account name for an account](https://docs.snowflake.com/en/user-guide/admin-account-identifier#finding-the-organization-and-account-name-for-an-account) instructions. ![](/img/snowflake/ui-snowsight-account-identifier.png) * Env * Kubernetes ``` # Password-based SPICE_SNOWFLAKE_ACCOUNT= \ SPICE_SNOWFLAKE_USERNAME= \ SPICE_SNOWFLAKE_PASSWORD= \ spice run # Key-pair (the `` is an optional parameter and is used for encrypted private key only) SPICE_SNOWFLAKE_ACCOUNT= \ SPICE_SNOWFLAKE_USERNAME= \ SPICE_SNOWFLAKE_SNOWFLAKE_PRIVATE_KEY_PATH= \ SPICE_SNOWFLAKE_SNOWFLAKE_PRIVATE_KEY_PASSPHRASE= \ spice run ``` or using the Spice CLI: ``` # Password-based spice login snowflake -a -u -p # Key-pair (the `` is an optional parameter and is used for encrypted private key only) spice login snowflake -a -u -k -s ``` The CLI will create or update an `.env` file that looks like: ``` SPICE_SNOWFLAKE_ACCOUNT="account" SPICE_SNOWFLAKE_PASSWORD="pass" SPICE_SNOWFLAKE_USERNAME="user" ``` Configure the spicepod to load secrets from the `env` secret store: (Note: This is the default setting) `spicepod.yaml` ``` version: v1 kind: Spicepod name: spice-app secrets: - from: env name: env datasets: - from: snowflake:DATABASE.SCHEMA.TABLE name: table params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin snowflake_username: ${env:SPICE_SNOWFLAKE_USERNAME} snowflake_password: ${env:SPICE_SNOWFLAKE_PASSWORD} snowflake_account: ${env:SPICE_SNOWFLAKE_ACCOUNT} ``` Learn more about \[Env Secret Store(../secret-stores/env). ``` # Password-based kubectl create secret generic snowflake \ --from-literal=account='' \ --from-literal=username='' \ --from-literal=password='' # Key-pair (the `` is an optional parameter and is used for encrypted private key only) kubectl create secret generic snowflake \ --from-literal=account='' \ --from-literal=username='' \ --from-literal=snowflake_private_key_path='' \ --from-literal=snowflake_private_key_passphrase='' ``` `spicepod.yaml` ``` version: v1 kind: Spicepod name: spice-app secrets: - from: kubernetes:snowflake name: snowflake datasets: - from: snowflake:DATABASE.SCHEMA.TABLE name: table params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin snowflake_username: ${snowflake:username} snowflake_password: ${snowflake:password} snowflake_account: ${snowflake:account} ``` `spicepod.yaml` (key-pair with private key content from secret) ```` version: v1 kind: Spicepod name: spice-app secrets: - from: kubernetes:snowflake name: snowflake datasets: - from: snowflake:DATABASE.SCHEMA.TABLE name: table params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin snowflake_auth_type: keypair snowflake_username: ${snowflake:username} snowflake_private_key: ${snowflake:private_key} snowflake_account: ${snowflake:account} ` ``` Learn more about [Kubernetes Secret Store(../secret-stores/kubernetes). Add new keychain entries (macOS) for user and password: ```bash # Password-based security add-generic-password -l "Snowflake Secret" \ -a spiced -s spice_snowflake_password\ -w # Key-pair (the `` is an optional parameter and is used for encrypted private key only) security add-generic-password -l "Snowflake Secret" \ -a spiced -s spice_snowflake_snowflake_private_key_path\ -w $(echo -n '' | base64) ```` `spicepod.yaml` ``` version: v1 kind: Spicepod name: spice-app secrets: - from: keyring name: keyring datasets: - from: snowflake:DATABASE.SCHEMA.TABLE name: table params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin snowflake_username: user_name snowflake_password: ${keyring:spice_snowflake_password} snowflake_account: account_identifier ``` Learn more about \[Keyring Secret Store(../secret-stores/keyring). ## Example[​](#example "Direct link to Example") ``` datasets: - from: snowflake:SNOWFLAKE_SAMPLE_DATA.TPCH_SF1.LINEITEM name: lineitem params: snowflake_warehouse: COMPUTE_WH snowflake_role: accountadmin ``` Limitations 1. Account identifier does not support the [Legacy account locator in a region format](https://docs.snowflake.com/en/user-guide/admin-account-identifier#format-2-legacy-account-locator-in-a-region). Use [Snowflake preferred name in organization format](https://docs.snowflake.com/en/user-guide/admin-account-identifier#format-1-preferred-account-name-in-your-organization). 2. The connector supports password-based and [key-pair](https://docs.snowflake.com/en/user-guide/key-pair-auth) authentication. ## Secrets[​](#secrets "Direct link to Secrets") Spice integrates with multiple secret stores to help manage sensitive data securely. For detailed information on supported secret stores, refer to the \[secret stores documentation(../secret-stores). Additionally, learn how to use referenced secrets in component parameters by visiting the \[using referenced secrets guide(../secret-stores#using-secrets). ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Snowflake as a data connector in Spice. [Snowflake Data Connector](https://github.com/spiceai/cookbook/tree/trunk/snowflake#readme) --- # Apache Spark Connector Apache Spark as a connector for federated SQL query against a Spark Cluster using [Spark Connect](https://spark.apache.org/docs/latest/spark-connect-overview.html) ``` datasets: - from: spark:spiceai.datasets.my_awesome_table name: my_table params: spark_remote: sc://localhost:15002 ``` ## Configuration[​](#configuration "Direct link to Configuration") * `spark_remote`: Required. A [spark remote](https://spark.apache.org/docs/latest/spark-connect-overview.html#set-sparkremote-environment-variable) connection URI. Refer to [spark connect client connection string](https://github.com/apache/spark/blob/master/connector/connect/docs/client-connection-string) for parameters in URI. The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### Auth Examples[​](#auth-examples "Direct link to Auth Examples") Spark clusters configured to accept authenticated requests should not set `spark_remote` as an inline dataset param, as it will contain sensitive data. For this case, use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:my_spark_remote}`. Check [Secrets Stores](/docs/v1.10/components/secret-stores) for more details. * Env * Kubernetes * Keyring ``` SPICE_SPARK_REMOTE= \ spice run # Or using the CLI to configure the secrets into an `.env` file spice login spark --spark_remote ``` `.env` ``` SPICE_SPARK_REMOTE= ``` `spicepod.yaml` ``` version: v1 kind: Spicepod name: spice-app secrets: - from: env name: env datasets: - from: spark:spiceai.datasets.my_awesome_table name: my_table params: spark_remote: ${env:SPICE_SPARK_REMOTE} ``` Learn more about [Env Secret Store](/docs/v1.10/components/secret-stores/env). ``` kubectl create secret generic spark \ --from-literal=spark_remote='' ``` `spicepod.yaml` ``` version: v1 kind: Spicepod name: spice-app secrets: - from: kubernetes:spark name: spark datasets: - from: spark:spiceai.datasets.my_awesome_table name: my_table params: spark_remote: ${spark:spark_remote} ``` Learn more about [Kubernetes Secret Store](/docs/v1.10/components/secret-stores/kubernetes). Add new keychain entry (macOS) with the spark remote: ``` security add-generic-password -l "Spark Remote" \ -a spiced -s spice_spark_remote \ -w ``` `spicepod.yaml` ``` version: v1 kind: Spicepod name: spice-app secrets: - from: keyring name: keyring datasets: - from: spark:spiceai.datasets.my_awesome_table name: my_table params: spark_remote: ${keyring:spice_spark_remote} ``` Learn more about [Keyring Secret Store](/docs/v1.10/components/secret-stores/keyring). ## Limitations[​](#limitations "Direct link to Limitations") * Correlated scalar subqueries are only supported in filters, aggregations, projections, and UPDATE/MERGE/DELETE commands. [Spark Docs](https://spark.apache.org/docs/latest/sql-error-conditions-unsupported-subquery-expression-category-error-class.html#unsupported_correlated_scalar_subquery) * The Spark connector does not yet support streaming query results from Spark. ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Spark as a data connector in Spice. [Apache Spark Data Connector](https://github.com/spiceai/cookbook/tree/trunk/spark#readme) --- # Spice.ai Data Connector The [Spice.ai](https://spice.ai/) Data Connector enables federated SQL query across datasets in the [Spice.ai Cloud Platform](https://docs.spice.ai/building-blocks/datasets). Access to these datasets requires a free [Spice.ai account](https://spice.ai/login). ## Configuration[​](#configuration "Direct link to Configuration") ### Secrets[​](#secrets "Direct link to Secrets") Secrets will be written to a `.env` file by using the `spice login` command and logging in with an active Spice AI account. Learn more about the [Env Secret Store](/docs/v1.10/components/secret-stores/env). * `api_key`: A Spice.ai API key. * `token`: An active personal access token that is configured when logging in to spice via `spice login`. ### Parameters[​](#parameters "Direct link to Parameters") #### `from`[​](#from "Direct link to from") The Spice.ai Cloud Platform dataset URI. To query a dataset in a public Spice.ai App, use the format `spice.ai///datasets/`. #### `name`[​](#name "Direct link to name") The dataset name. This will be used as the table name within Spice. The dataset name cannot be a [reserved keyword](/docs/v1.10/reference/spicepod/keywords). ### `params`[​](#params "Direct link to params") The Spice.ai Cloud Platform data connector can be configured by providing the following `params`. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to load the secret from a secret store, e.g. `${secrets:SPICEAI_API_KEY}`. | Parameter Name | Description | | ----------------- | ---------------------------------------------------- | | `spiceai_api_key` | The Spice.ai Cloud Platform API key to connect with. | ## Example[​](#example "Direct link to Example") ``` - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips ``` ``` - from: spice.ai/spiceai/tpch/datasets/customer name: tpch.customer ``` ## Full Configuration Example[​](#full-configuration-example "Direct link to Full Configuration Example") ``` - from: spice.ai/spiceai/tpch/datasets/customer name: tpch.customer params: spiceai_api_key: ${secrets:spiceai_api_key} acceleration: enabled: true ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure Spice.ai Cloud Platform as a data connector in Spice. [Spice.ai Cloud Platform Data Connector](https://github.com/spiceai/cookbook/tree/trunk/spiceai#readme) ## Limitations[​](#limitations "Direct link to Limitations") * The Spice.ai Data Connector is subject to a maximum limit of 1000 requests per connection, after which the connection is reset by the Spice Cloud Platform. If the error message `Connection is reset by the server. Please retry the request.` is encountered or the `spiceai-retryable` metadata appears in the response, the query should be retried. Memory Considerations When using the Spice.ai Data Connector without acceleration, part of the query execution will be in memory if federating across different Spice Cloud Platform apps. Ensure sufficient memory is available, including overhead for queries and the runtime, especially with concurrent queries. Memory limitations can be mitigated by storing acceleration data on disk, which is supported by [`duckdb`](/docs/v1.10/components/data-accelerators/duckdb) and [`sqlite`](/docs/v1.10/components/data-accelerators/sqlite) accelerators by specifying `mode: file`. --- # Embedding Models Embedding models transform raw text into numerical vectors that machine learning models can use. Spice supports running embedding models locally or via hosted services such as OpenAI, Amazon Bedrock, Databricks MosaicAI, or [la Plateforme](https://console.mistral.ai/). Embeddings enable vector-based and similarity search, such as document retrieval. For chat-based large language models, see [Model Providers](/docs/v1.10/components/models). Spice supports a variety of embedding model sources and formats: | Name | Description | Status | ML Format(s) | LLM Format(s)\* | | -------------------------------------------------------------- | --------------------------------------- | ----------------- | ------------ | ------------------------------- | | [`file`](/docs/v1.10/components/embeddings/local) | Local filesystem | Release Candidate | ONNX | GGUF, GGML, SafeTensor | | [`huggingface`](/docs/v1.10/components/embeddings/huggingface) | Models hosted on HuggingFace | Release Candidate | ONNX | GGUF, GGML, SafeTensor | | [`openai`](/docs/v1.10/components/embeddings/openai) | OpenAI (or compatible) LLM endpoint | Release Candidate | - | OpenAI-compatible HTTP endpoint | | [`azure`](/docs/v1.10/components/embeddings/azure) | Azure OpenAI | Alpha | - | OpenAI-compatible HTTP endpoint | | [`databricks`](/docs/v1.10/components/embeddings/databricks) | Models deployed to Databricks Mosaic AI | Alpha | - | OpenAI-compatible HTTP endpoint | | [`bedrock`](/docs/v1.10/components/embeddings/bedrock) | Models deployed on AWS Bedrock | Alpha | - | OpenAI-compatible HTTP endpoint | | [`model2vec`](/docs/v1.10/components/embeddings/model2vec) | Model2Vec static word embeddings | Alpha | - | Model2Vec format | ## Overview[​](#overview "Direct link to Overview") Spice provides three ways to handle embedding columns in datasets: 1. **[Just-in-Time (JIT) Embeddings](#jit-embeddings):** Embeddings are computed on demand during query execution, with no precomputation. 2. **[Accelerated Embeddings](#accelerated-embeddings):** Embeddings are precomputed and stored, enabling faster queries and searches. 3. **[Passthrough Embeddings](#passthrough-embeddings):** Pre-existing embeddings in the source dataset are used directly, with no additional computation. ## Configuring Embedding Models[​](#configuring-embedding-models "Direct link to Configuring Embedding Models") Define embedding models in the `spicepod.yaml` file as top-level components. Example configuration in `spicepod.yaml`: ``` embeddings: - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: all_minilm_l6_v2 - from: openai:text-embedding-3-large name: xl_embed params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } - name: my_model from: file:model.safetensors files: - path: config.json - path: models/embed/tokenizer.json ``` Embedding models can be used via: * An OpenAI-compatible [endpoint](/docs/v1.10/api/HTTP/post-embeddings) * Augmenting a dataset with column-level [embeddings](/docs/v1.10/reference/spicepod/datasets#embeddings) for vector-based [search functionality](/docs/v1.10/features/search#vector-search) ### Configuring Embedding Columns on Datasets[​](#configuring-embedding-columns-on-datasets "Direct link to Configuring Embedding Columns on Datasets") To create vector embeddings for specific dataset columns, define them under `columns` in the `spicepod.yaml` file, within the `datasets` section. Example configuration in `spicepod.yaml`: ``` embeddings: - from: openai:text-embedding-3-large name: xl_embed params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } datasets: - from: file:sales_data.parquet name: sales columns: - name: address_line1 description: The first line of the address. embeddings: - from: xl_embed row_id: order_number chunking: enabled: true target_chunk_size: 256 overlap_size: 32 ``` See the [embeddings](/docs/v1.10/reference/spicepod/embeddings) and [datasets](/docs/v1.10/reference/spicepod/datasets#embeddings) reference for more details. ## Embedding Methods[​](#embedding-methods "Direct link to Embedding Methods") ### Just-in-Time (JIT) Embeddings[​](#jit-embeddings "Direct link to Just-in-Time (JIT) Embeddings") JIT embeddings are computed at query time. This is useful when precomputing is impractical (e.g., large or rarely queried datasets, or heavy prefiltering). To add a JIT embedding column, specify it in the dataset's column config. ``` datasets: - name: invoices from: sftp://remote-sftp-server.com/invoices/2024/ columns: - name: line_item_details embeddings: - from: my_embedding_model params: file_format: parquet embeddings: # Or any model you like! - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: my_embedding_model ``` ### Accelerated Embeddings[​](#accelerated-embeddings "Direct link to Accelerated Embeddings") To speed up queries, embeddings can be precomputed and stored in a [data accelerator](/docs/v1.10/components/data-accelerators). Enable this by adding: ``` acceleration: enabled: true ``` to the dataset configuration. All other data accelerator configurations are optional, but can be applied as per their respective [documentation](/docs/v1.10/components/data-accelerators). **Full example:** ``` datasets: - name: invoices from: sftp://remote-sftp-server.com/invoices/2024/ acceleration: enabled: true columns: - name: line_item_details embeddings: - from: my_embedding_model params: file_format: parquet ``` ### Passthrough Embeddings[​](#passthrough-embeddings "Direct link to Passthrough Embeddings") If the dataset already contains embedding columns, Spice can use them for vector search and other embedding features. The schema must match that of Spice-generated embeddings (or be adapted with a [view](/docs/v1.10/reference/spicepod#views)). **Example:** A `sales` table with an `address` column and its embedding: ``` sql> describe sales; +-------------------+-----------------------------------------+-------------+ | column_name | data_type | is_nullable | +-------------------+-----------------------------------------+-------------+ | order_number | Int64 | YES | | quantity_ordered | Int64 | YES | | price_each | Float64 | YES | | order_line_number | Int64 | YES | | address | Utf8 | YES | | address_embedding | FixedSizeList( | NO | | | Field { | | | | name: "item", | | | | data_type: Float32, | | | | nullable: false, | | | | dict_id: 0, | | | | dict_is_ordered: false, | | | | metadata: {} | | | | }, | | | | 384 | | +-------------------+-----------------------------------------+-------------+ ``` The same table if it was chunked: ``` sql> describe sales; +-------------------+-----------------------------------------+-------------+ | column_name | data_type | is_nullable | +-------------------+-----------------------------------------+-------------+ | order_number | Int64 | YES | | quantity_ordered | Int64 | YES | | price_each | Float64 | YES | | order_line_number | Int64 | YES | | address | Utf8 | YES | | address_embedding | List(Field { | NO | | | name: "item", | | | | data_type: FixedSizeList( | | | | Field { | | | | name: "item", | | | | data_type: Float32, | | | | }, | | | | 384 | | | | ), | | | | }) | | +-------------------+-----------------------------------------+-------------+ | address_offset | List(Field { | NO | | | name: "item", | | | | data_type: FixedSizeList( | | | | Field { | | | | name: "item", | | | | data_type: Int32, | | | | nullable: false, | | | | dict_id: 0, | | | | dict_is_ordered: false, | | | | metadata: {} | | | | }, | | | | 2 | | | | ), | | | | }) | | +-------------------+-----------------------------------------+-------------+ ``` Passthrough embedding columns must still be defined in the `spicepod.yaml` file. The Spice instance must also have access to the same embedding model used to generate the embeddings. ``` datasets: - from: sftp://remote-sftp-server.com/sales/2024.csv name: sales columns: - name: address embeddings: - from: local_embedding_model embeddings: - name: local_embedding_model # The model originally used for this column ... ``` #### Requirements[​](#requirements "Direct link to Requirements") To ensure compatibility, embedding columns must meet these requirements: 1. **Underlying Column:** * The original column must exist and be of `string` [Arrow data type](/docs/v1.10/reference/datatypes/accelerators). 2. **Naming Convention:** * The embedding column must be named `_embedding` (e.g., `review_embedding` for a `review` column). 3. **Data Type:** * The embedding column must be: * `FixedSizeList[Float32 or Float64, N]` for unchunked data, where `N` is the embedding vector size. * `List[FixedSizeList[Float32 or Float64, N]]` for chunked data. 4. **Offset Column (for chunked data):** * If chunked, an offset column `_offsets` must exist with type `List[FixedSizeList[Int32, 2]]`, where each pair `[start, end]` maps a chunk to its text segment. * Example: `[[0, 100], [101, 200]]` means two chunks covering indices 0–100 and 101–200. Following these guidelines ensures that the dataset's pre-existing embeddings are fully compatible with Spice. ## Advanced Configuration[​](#advanced-configuration "Direct link to Advanced Configuration") ### Chunking[​](#chunking "Direct link to Chunking") Spice supports chunking large text columns before embedding, which is useful for [Document Tables](/docs/v1.10/components/data-connectors#document-support). Chunking helps return only the most relevant text during search. Configure chunking in the embedding config: ``` datasets: - from: github:github.com/spiceai/spiceai/issues name: spiceai.issues acceleration: enabled: true columns: - name: body embeddings: - from: local_embedding_model chunking: enabled: true target_chunk_size: 512 ``` The `body` column will be split into chunks of about 512 tokens, preserving sentence and semantic boundaries. See the [API reference](/docs/v1.10/reference/spicepod/datasets#columns-embeddings-chunking) for details. #### Row Identifiers[​](#row-identifiers "Direct link to Row Identifiers") The `row_id` field specifies which column(s) uniquely identify each row, similar to a primary key. This is important for chunked embeddings, so that operations (e.g., [`v1/search`](/docs/v1.10/api/HTTP/post-search)) can map multiple chunked vectors to a single row. Set `row_id` in `columns[*].embeddings[*].row_id`. ``` datasets: - from: github:github.com/spiceai/spiceai/issues name: spiceai.issues acceleration: enabled: true columns: - name: body embeddings: - from: local_embedding_model chunking: enabled: true target_chunk_size: 512 row_id: id ``` ## [📄️OpenAI](/docs/v1.10/components/embeddings/openai) [To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration.](/docs/v1.10/components/embeddings/openai) ## [📄️Azure OpenAI](/docs/v1.10/components/embeddings/azure) [To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page:](/docs/v1.10/components/embeddings/azure) ## [📄️HuggingFace](/docs/v1.10/components/embeddings/huggingface) [To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice.](/docs/v1.10/components/embeddings/huggingface) ## [📄️Local](/docs/v1.10/components/embeddings/local) [Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services.](/docs/v1.10/components/embeddings/local) ## [📄️Model2Vec](/docs/v1.10/components/embeddings/model2vec) [Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning.](/docs/v1.10/components/embeddings/model2vec) ## [📄️AWS Bedrock](/docs/v1.10/components/embeddings/bedrock) [Instructions for using Amazon Bedrock embedding models](/docs/v1.10/components/embeddings/bedrock) ## [📄️Databricks](/docs/v1.10/components/embeddings/databricks) [Instructions for using Databricks Mosaic AI Models](/docs/v1.10/components/embeddings/databricks) --- # Azure OpenAI Embedding Models To use an embedding model hosted on Azure OpenAI, specify the `azure` path in the `from` field and the following parameters from the [Azure OpenAI Model Deployment](https://ai.azure.com/resource/deployments) page: | Param | Description | Default | | ----------------------- | ----------------------------------------------------------------------------------- | ---------- | | `azure_api_key` | The Azure OpenAI API key from the models deployment page. | - | | `azure_api_version` | The API version used for the Azure OpenAI service. | - | | `azure_deployment_name` | The name of the model deployment. | Model name | | `endpoint` | The Azure OpenAI resource endpoint, e.g., `https://resource-name.openai.azure.com`. | - | | `azure_entra_token` | The Azure Entra token for authentication. | - | Only one of `azure_api_key` or `azure_entra_token` can be provided for model configuration. Example: ``` embeddings: - name: embeddings-model from: azure:text-embedding-3-small params: endpoint: ${ secrets:SPICE_AZURE_AI_ENDPOINT } azure_deployment_name: text-embedding-3-small azure_api_version: 2023-05-15 azure_api_key: ${ secrets:SPICE_AZURE_API_KEY } ``` Refer to the [Azure OpenAI Service models](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models) for more details on available models and configurations. Follow the [Azure OpenAI Models Cookbook](https://github.com/spiceai/cookbook/tree/trunk/azure_openai) to try Azure OpenAI models for vector-based search and chat functionalities with structured (taxi trips) and unstructured (GitHub files) data. --- # Amazon Bedrock Model Provider To use an embedding model deployed to [AWS Bedrock service](https://aws.amazon.com/bedrock/), specify the model endpoint name prefixed with `bedrock:` in the `from` field and include the required parameters in the `params` section. ### Parameters[​](#parameters "Direct link to Parameters") #### AWS Parameters[​](#aws-parameters "Direct link to AWS Parameters") | Parameter | Description | | ---------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | `aws_region` | AWS region. Default: `us-east-1`. | | `aws_profile` | Optional. AWS profile to use when loading credentials. | | `aws_access_key_id` | Optional. AWS access key ID for authentication. If not provided, credentials will be loaded from environment variables or IAM roles | | `aws_secret_access_key` | Optional. AWS secret access key for authentication. If not provided, credentials will be loaded from environment variables or IAM roles | | `aws_session_token` | Optional. AWS session token for authentication | | `max_concurrent_invocations` | Optional. The maximum number of concurrent API invocations. Defaults to `40` | | `requests_per_min_limit` | Optional. The maximum number of requests made per minute. Defaults to `1500` | #### AWS Titan Models[​](#aws-titan-models "Direct link to AWS Titan Models") These parameters are used for [Amazon Titan Text](https://docs.aws.amazon.com/bedrock/latest/userguide/titan-embedding-models.html) embedding model | Parameter | Description | | ------------ | ----------------------------------------------------------------------------------------------------------------------- | | `normalize` | Whether or not to normalize the output embedding. Defaults to `true`. | | `dimensions` | The number of dimensions the output embedding should have. The following values are accepted: 1024 (default), 512, 256. | #### Amazon Nova Models[​](#amazon-nova-models "Direct link to Amazon Nova Models") These parameters are used for [Amazon Nova](https://docs.aws.amazon.com/nova/latest/userguide/embeddings-schema.html) multimodal embedding models | Parameter | Description | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `dimensions` | **Required**. The number of dimensions the output embedding should have. Accepted value: 256, 384, 1024 or 3072. | | `truncation_mode` | Optional. Specifies how the API handles inputs longer than the maximum token length. One of: `START`, `END` or `NONE` (default). | | `embedding_purpose` | Optional. Use the Nova embeddings model optimized for different purposes. Default `GENERIC_INDEX`. See reference [docs](https://docs.aws.amazon.com/nova/latest/userguide/embeddings-schema.html) for all options. | #### Cohere Models[​](#cohere-models "Direct link to Cohere Models") | Parameter | Description | | ------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `truncate` | Specifies how the API handles inputs longer than the maximum token length. One of: `START`, `END` or `NONE` (default). | | `input_type` | Use the Cohere embeddings model optimized for different types of inputs. One of: `search_document` (default), `search_query`, `classification` or `clustering`. | ### Example `spicepod.yaml` configuration, Cohere model[​](#example-spicepodyaml-configuration-cohere-model "Direct link to example-spicepodyaml-configuration-cohere-model") ``` embeddings: - from: bedrock:cohere.embed-english-v3 name: cohere-embeddings params: aws_region: us-east-1 input_type: classification truncate: END aws_access_key_id: ${ secrets:AWS_ACCESS_KEY_ID } aws_secret_access_key: ${ secrets:AWS_SECRET_ACCESS_KEY } ``` ### Example `spicepod.yaml` configuration, Titan model[​](#example-spicepodyaml-configuration-titan-model "Direct link to example-spicepodyaml-configuration-titan-model") ``` - from: bedrock:amazon.titan-embed-text-v2:0 name: titan-embeddings params: dimensions: "256" ``` ### Example `spicepod.yaml` configuration, Amazon Nova model[​](#example-spicepodyaml-configuration-amazon-nova-model "Direct link to example-spicepodyaml-configuration-amazon-nova-model") ``` embeddings: - from: bedrock:amazon.nova-2-multimodal-embeddings-v1:0 name: nova-embeddings params: dimensions: "3072" truncation_mode: START embedding_purpose: GENERIC_RETRIEVAL aws_region: us-east-1 aws_access_key_id: ${ secrets:AWS_ACCESS_KEY_ID } aws_secret_access_key: ${ secrets:AWS_SECRET_ACCESS_KEY } ``` ### Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate bedrock permissions (e.g., `bedrock:InvokeModel`) to access the model. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access DynamoDB tables: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "bedrock:InvokeModel" ], "Resource": [ "arn:aws:bedrock:us-east-1::foundation-model/amazon.titan-*" ] } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | --------------------- | --------------------------------------------- | | `bedrock:InvokeModel` | Required. Used to invoke the embedding model. | ### Additional Information[​](#additional-information "Direct link to Additional Information") Refer to the [Amazon Bedrock documentation](https://docs.aws.amazon.com/bedrock/) for more details on available models and configurations. --- # Databricks Model Provider To use an embedding model deployed to [Databricks Mosaic AI Model Serving](https://docs.databricks.com/aws/en/machine-learning/model-serving/), specify the model endpoint name prefixed with `databricks:` in the `from` field and include the required parameters in the `params` section. ### Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `databricks_endpoint` | The Databricks workspace endpoint, e.g., `dbc-a12cd3e4-56f7.cloud.databricks.com`. | | `databricks_token` | The Databricks API token to authenticate with the Databricks Models API. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to reference a secret, e.g., `${secrets:my_databricks_token}`. | | `databricks_client_id` | The Databricks Service Principal Client ID. Can't be used with `databricks_token`. | | `databricks_client_secret` | The Databricks Service Principal Client Secret. Can't be used with `databricks_token`. | ### Example `spicepod.yaml` configuration, using personal access token[​](#example-spicepodyaml-configuration-using-personal-access-token "Direct link to example-spicepodyaml-configuration-using-personal-access-token") To learn more about how to set up personal access tokens, see [Databricks PAT docs](https://docs.databricks.com/aws/en/dev-tools/auth/pat). ``` embeddings: - from: databricks:databricks-gte-large-en name: gte-large-en params: databricks_endpoint: dbc-46470731-42e5.cloud.databricks.com databricks_token: ${ secrets:SPICE_DATABRICKS_TOKEN } ``` ### Example `spicepod.yaml` configuration, using Databricks service principal[​](#example-spicepodyaml-configuration-using-databricks-service-principal "Direct link to example-spicepodyaml-configuration-using-databricks-service-principal") Spice supports the Machine-to-Machine (M2M) OAuth flow with service principal credentials by utilizing the `databricks_client_id` and `databricks_client_secret` parameters. The runtime will automatically refresh the token. The service principal must be granted the "Can Query" permission for model serving. To learn more about how to set up the service principal, see [Databricks M2M OAuth docs](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m). ``` embeddings: - from: databricks:databricks-gte-large-en name: gte-large-en params: databricks_endpoint: dbc-42424242-4242.cloud.databricks.com databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} databricks_client_secret: ${secrets:DATABRICKS_CLIENT_SECRET} ``` ### Additional Information[​](#additional-information "Direct link to Additional Information") Refer to the [Mosaic AI Model Serving documentation](https://docs.databricks.com/aws/en/machine-learning/model-serving/) for more details on available models and configurations. --- # HuggingFace Text Embedding Models To use an embedding model from HuggingFace with Spice, specify the `huggingface` path in the `from` field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. The following parameters are specific to HuggingFace models: | Parameter | Description | Default | | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- | | `hf_token` | The Huggingface access token. | - | | `pooling` | The [pooling method](https://huggingface.co/docs/text-embeddings-inference/en/cli_arguments) for embedding models. Supported values are `cls`, `mean`, `splade`, `last_token` | - | Here is an example configuration in `spicepod.yaml`: ``` embeddings: - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: all_minilm_l6_v2 ``` Supported models include: * All models tagged as [text-embeddings-inference](https://huggingface.co/models?other=text-embeddings-inference) on Huggingface * Any Huggingface repository with the correct files to be loaded as a [local embedding model](/docs/v1.10/components/embeddings/local). With the same semantics as [language models](/docs/v1.10/components/models/huggingface#access-tokens), `spice` can run private HuggingFace embedding models: ``` embeddings: - from: huggingface:huggingface.co/secret-company/awesome-embedding-model name: top_secret params: hf_token: ${ secrets:HF_TOKEN } ``` --- # Local Filesystem Embedding Models Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") To configure an embedding model using local files, you can specify the details in the `spicepod.yaml` file as shown below: ``` embeddings: - name: all_minilm_l6_v2 from: file:model.safetensors files: - path: /Users/jeadie/Github/spiceai/models/embed/config.json - path: models/embed/tokenizer.json ``` ## Required Files[​](#required-files "Direct link to Required Files") * Model file, one of: `model.safetensors`, `pytorch_model.bin`. * A tokenizer file with the filename `tokenizer.json`. * A config file with the filename `config.json`. --- # Model2Vec Embedding Models [Model2Vec](https://huggingface.co/blog/Pringled/model2vec) is a technique that distills embeddings from sentence transformer models into static word embeddings, providing efficient embedding generation, in parallel, without performing external API calls. This can result in sentence transformer models up to 500x faster and 15x smaller. To use a Model2Vec embedding model with Spice, specify the `model2vec` prefix in the `from` field of your configuration. ## Model Compatibility[​](#model-compatibility "Direct link to Model Compatibility") Find models compatible with `model2vec`: * [Model2Vec base models](https://huggingface.co/collections/minishlab/model2vec-base-models-66fd9dd9b7c3b3c0f25ca90e) (ready to use, pre-distilled) * [Sentence Transformers](https://huggingface.co/collections/sentence-transformers/embedding-model-datasets-6644d7a3673a511914aa7552) (needs [distillation](#distilling-your-own-models)) * [Pre-distilled community models](https://huggingface.co/models?search=model2vec) ## Parameters[​](#parameters "Direct link to Parameters") The following parameters are specific to Model2Vec models: | Parameter | Description | Default | | ------------------------- | ------------------------------------------------------------------------------- | ----------------------- | | `hf_token` | The Hugging Face access token for accessing private models. | - | | `normalize` | Whether to normalize embeddings (defaults to the model's configuration). | Model's default setting | | `subfolder` | Optional subfolder path for models that reside in a subfolder of the repo/path. | - | | `parallelism` | Number of parallel threads to use for embedding computation. | System CPU count | | `embed_max_token_length` | Maximum token length for embeddings. | - | | `embed_custom_batch_size` | Custom batch size override for embedding operations. | - | For more details on Model2Vec parameters and functionality, refer to the [model2vec-rs documentation](https://github.com/MinishLab/model2vec-rs). Example configuration in `spicepod.yaml` for [`minishlab/potion-base-8m`](https://huggingface.co/minishlab/potion-base-8M): ``` embeddings: - from: model2vec:minishlab/potion-base-8M name: potion_base_8m ``` ## Local Models[​](#local-models "Direct link to Local Models") Model2Vec models can also be loaded from the local filesystem by specifying a file path: ``` embeddings: - from: model2vec:/path/to/local/model name: local_model2vec ``` ## Private Models[​](#private-models "Direct link to Private Models") Model2Vec supports private Hugging Face models with authentication: ``` embeddings: - from: model2vec:your-organization/private-model name: private_embeddings params: hf_token: ${ secrets:HF_TOKEN } ``` ## Advanced Configuration[​](#advanced-configuration "Direct link to Advanced Configuration") For performance optimization, configure parallelism and embedding batch sizes: ``` embeddings: - from: model2vec:minishlab/potion-base-8M name: potion_optimized params: parallelism: 8 embed_custom_batch_size: 32 normalize: true ``` ## Distilling Your Own Models[​](#distilling-your-own-models "Direct link to Distilling Your Own Models") Create custom Model2Vec embeddings by distilling existing sentence transformer models. For more detailed instructions, see the [Model2Vec Quickstart guide](https://github.com/MinishLab/model2vec?tab=readme-ov-file#quickstart). Here's how to distill the popular [`sentence-transformers/all-MiniLM-L6-v2`](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) model: 1. **Install the Model2Vec Python library:** ``` pip install model2vec ``` 2. **Create a distillation script:** ``` from model2vec import StaticModel # Load a sentence transformer model and distill it to a static model model = StaticModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2", device="cpu") # Save the static model model.save_pretrained("./all-MiniLM-L6-v2-model2vec") print("Model distilled and saved to ./all-MiniLM-L6-v2-model2vec") ``` 3. **Use the distilled model with Spice:** ``` embeddings: - from: model2vec:./all-MiniLM-L6-v2-model2vec name: distilled_minilm - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: all_minilm_l6_v2 ``` 4. **Race!** Compare the throughput of the distilled embedding model with the full version by declaring both models in the same Spicepod. This uses example [Wikipedia article data from Kaggle](https://www.kaggle.com/datasets/jjinho/wikipedia-20230701): ``` datasets: - from: file://wiki_a.parquet name: wiki_a_distilled acceleration: enabled: true columns: - name: text_trunc embeddings: - from: minilm_distilled - from: file://wiki_a.parquet name: wiki_a_full acceleration: enabled: true refresh_sql: select * from wiki_a_full limit 100; columns: - name: text_trunc embeddings: - from: all_minilm_l6_v2 embeddings: - from: model2vec:./all-MiniLM-L6-v2-model2vec name: distilled_minilm - from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 name: all_minilm_l6_v2 ``` Start Spice with `spice run`: ``` 2025-08-25T15:59:39.033644Z INFO runtime::init::embedding: Embedding Model minilm_distilled ready 2025-08-25T15:59:39.969381Z INFO runtime::init::embedding: Embedding Model all_minilm_l6_v2 ready 2025-08-25T15:59:39.969713Z INFO runtime::init::dataset: Dataset wiki_a_full initializing... 2025-08-25T15:59:39.969713Z INFO runtime::init::dataset: Dataset wiki_a_distilled initializing... 2025-08-25T15:59:39.973287Z INFO runtime::init::dataset: Dataset wiki_a_full registered (file://wiki_a2.parquet), acceleration (arrow), results cache enabled. 2025-08-25T15:59:39.973344Z INFO runtime::init::dataset: Dataset wiki_a_distilled registered (file://wiki_a2.parquet), acceleration (arrow), results cache enabled. 2025-08-25T15:59:39.973637Z INFO runtime::accelerated_table::refresh_task: Loading data for dataset wiki_a_distilled 2025-08-25T15:59:39.973714Z INFO runtime::accelerated_table::refresh_task: Loading data for dataset wiki_a_full 2025-08-25T15:59:50.982854Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 40,960 records 2025-08-25T16:00:02.953192Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 57,344 records 2025-08-25T16:00:14.429262Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 40,960 records 2025-08-25T16:00:26.177483Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 49,152 records 2025-08-25T16:00:36.963177Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 40,960 records 2025-08-25T16:00:49.157552Z INFO runtime::accelerated_table::refresh_task: Dataset wiki_a_distilled received 49,152 records 2025-08-25T16:01:09.789714Z INFO runtime::accelerated_table::refresh_task: Loaded 100 rows (904.41 kiB) for dataset wiki_a_full in 1m 29s 816ms. ``` **Performance Results:** Note: The dramatic results are due to `model2vec` embedding execution being parallelized across all of the host's cores (default configuration). Per core, model2vec achieves a throughput of 300/400 rows/sec with this corpus. This specific test machine has 16 cores. Execution of SBERT models is currently not parallelized. | Model Name | Model Type | Records Processed | Throughput (records/sec) | | ---------------------------------------- | --------------------- | ----------------- | ------------------------ | | `sentence-transformers/all-MiniLM-L6-v2` | Model2Vec (Distilled) | 278,528 | \~4,043 | | `sentence-transformers/all-MiniLM-L6-v2` | SBERT (Full) | 100 | \~1.1 | | **Performance Gain** (model2vec) | - | - | **\~3,675x faster** | --- # OpenAI (or Compatible) Embedding Models To use a hosted OpenAI (or compatible) embedding model, specify the `openai` path in the `from` field of your configuration. For a specific model, include its model ID in the `from` field. If no model ID is specified, it defaults to `"text-embedding-3-small"`. The following parameters are specific to OpenAI models: | Parameter | Description | Default | | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------- | | `openai_api_key` | The API key for accessing OpenAI. | - | | `openai_org_id` | The organization ID for OpenAI. | - | | `openai_project_id` | The project ID for OpenAI. | - | | `openai_usage_tier` | The [OpenAI usage tier](https://platform.openai.com/settings/organization/limits) for the account. This parameter sets the maximum number of concurrent requests based on OpenAI's published limits per tier. Valid values are `free`, `tier1`, `tier2`, `tier3`, `tier4`, or `tier5`. | `tier1` | | `endpoint` | The base endpoint for the OpenAI API. | `https://api.openai.com/v1` | Below is an example configuration in `spicepod.yaml`: ``` embeddings: - from: openai:text-embedding-3-large name: xl_embed params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } - name: mistral from: openai:mistral-embed params: endpoint: https://api.mistral.ai/v1 api_key: ${ secrets:SPICE_MISTRAL_API_KEY } ``` For detailed instructions and examples on running vector searches, refer to the [Vector-Based Search documentation](/docs/v1.10/features/search/vector-search). --- # Model Providers Spice supports various model providers for traditional machine learning (ML) models and large language models (LLMs). | Name | Description | Status | ML Format(s) | LLM Format(s)\* | | ---------------------------------------------------------- | -------------------------------------------- | ----------------- | ------------ | ------------------------------- | | [`openai`](/docs/v1.10/components/models/openai) | OpenAI (or compatible) LLM endpoint | Release Candidate | - | OpenAI-compatible HTTP endpoint | | [`file`](/docs/v1.10/components/models/filesystem) | Local filesystem | Release Candidate | ONNX | GGUF, GGML, SafeTensor | | [`huggingface`](/docs/v1.10/components/models/huggingface) | Models hosted on HuggingFace | Release Candidate | ONNX | GGUF, GGML, SafeTensor | | [`spice.ai`](/docs/v1.10/components/models/spiceai) | Models hosted on the Spice.ai Cloud Platform | Alpha | ONNX | OpenAI-compatible HTTP endpoint | | [`azure`](/docs/v1.10/components/models/azure) | Azure OpenAI | Alpha | - | OpenAI-compatible HTTP endpoint | | [`anthropic`](/docs/v1.10/components/models/anthropic) | Models hosted on Anthropic | Alpha | - | OpenAI-compatible HTTP endpoint | | [`xai`](/docs/v1.10/components/models/xai) | Models hosted on xAI | Alpha | - | OpenAI-compatible HTTP endpoint | | [`databricks`](/docs/v1.10/components/models/databricks) | Models deployed to Databricks Mosaic AI | Alpha | - | OpenAI-compatible HTTP endpoint | Spice also tests and evaluates common models and grades their ability to integrate with Spice. See the [Models Grade Report](/docs/v1.10/reference/models). \*LLM Format(s) may require additional files (e.g., `tokenizer_config.json`). The model type is inferred based on the model source and files. For more detail, refer to the `model` [reference specification](/docs/v1.10/reference/spicepod/models). ## Features[​](#features "Direct link to Features") Spice supports a variety of features for large language models (LLMs): * **Custom Tools**: Provide models with tools to interact with the Spice runtime. See [Tools](/docs/v1.10/features/large-language-models/tools). * **System Prompts**: Declaratively define system prompts and default values for [`v1/chat/completion`](/docs/v1.10/api/HTTP/post-chat-completions) parameters. See [Parameter Overrides](/docs/v1.10/features/large-language-models/parameter_overrides). Use Jinja templating to parameterize system prompts per request. See [Parameterized prompts](/docs/v1.10/features/large-language-models/parameterized_prompts). * **Memory**: Provide LLMs with memory persistence tools to store and retrieve information across conversations. See [Memory](/docs/v1.10/features/large-language-models/memory). * **Vector Search**: Perform advanced vector-based searches using embeddings. See [Vector Search](/docs/v1.10/features/search/vector-search). * **Evals**: Evaluate, track, compare, and improve language model performance for specific tasks. See [Evals](/docs/v1.10/features/large-language-models/evals). * **Local Models**: Load and serve models locally from various sources, including local filesystems and Hugging Face. See [Local Models](/docs/v1.10/features/large-language-models/serving). For more details, refer to the [Large Language Models documentation](/docs/v1.10/features/large-language-models). ## Model Provider Prefix[​](#model-provider-prefix "Direct link to Model Provider Prefix") The model provider prefix identifies the source or provider of a model in Spice configuration files. This prefix is specified before the model identifier in the `from` field of a model definition, and is used in specifying model [default parameter overrides](#example-setting-default-parameter-overrides). It helps the runtime determine how to load and interact with the model. The following provider prefixes are supported: | Prefix | Description | | ------------ | ------------------------------------- | | `openai` | OpenAI or OpenAI-compatible endpoints | | `azure` | Azure OpenAI | | `xai` | xAI | | `anthropic` | Anthropic | | `perplexity` | Perplexity | | `hf` | Hugging Face | | `file` | Local filesystem | | `spiceai` | Spice.ai Cloud Platform | | `databricks` | Databricks Mosaic AI | **Example usage in `spicepod.yaml`:** ``` models: - from: openai:gpt-4o name: openai-model - from: hf:meta-llama/Llama-3-8B-Instruct name: llama3-hf - from: file://absolute/path/to/model.gguf name: local-model ``` ## Model Examples[​](#model-examples "Direct link to Model Examples") The following examples demonstrate how to configure and use various models or model features with Spice. Each example provides a specific use case to help understand the configuration options available. ### Example: Configuring an OpenAI Model[​](#example-configuring-an-openai-model "Direct link to Example: Configuring an OpenAI Model") To use a language model hosted on OpenAI (or compatible), specify the `openai` path and model ID in `from`. For more details, see [OpenAI Model Provider](/docs/components/models/openai). Example `spicepod.yml`: ``` models: - from: openai:gpt-4o-mini name: openai params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } - from: openai:llama3-groq-70b-8192-tool-use-preview name: groq-llama params: endpoint: https://api.groq.com/openai/v1 openai_api_key: ${ secrets:SPICE_GROQ_API_KEY } ``` ### Example: Using an OpenAI Model with Tools[​](#example-using-an-openai-model-with-tools "Direct link to Example: Using an OpenAI Model with Tools") To specify tools for an OpenAI model, include them in the `params.tools` field. For more details, see the [Tools documentation](/docs/features/large-language-models/tools). ``` models: - name: sql-model from: openai:gpt-4o params: tools: list_datasets, sql, table_schema ``` ### Example: Adding Memory to a Model[​](#example-adding-memory-to-a-model "Direct link to Example: Adding Memory to a Model") To enable memory tools for a model, define a `store` memory dataset and specify `memory` in the model's `tools` parameter. For more details, see the [Memory documentation](/docs/features/large-language-models/memory). ``` datasets: - from: memory:store name: llm_memory access: read_write models: - name: memory-enabled-model from: openai:gpt-4o params: tools: memory, sql ``` ### Example: Setting Default Parameter Overrides[​](#example-setting-default-parameter-overrides "Direct link to Example: Setting Default Parameter Overrides") To set default overrides for parameters, use the [model provider prefix](#model-provider-prefix) followed by the parameter name. For more details, see the [Parameter Overrides documentation](/docs/features/large-language-models/parameter_overrides). ``` models: - name: pirate-haikus from: openai:gpt-4o params: openai_temperature: 0.1 openai_response_format: { 'type': 'json_object' } ``` ### Example: Configuring a System Prompt[​](#example-configuring-a-system-prompt "Direct link to Example: Configuring a System Prompt") To configure an additional system prompt, use the `system_prompt` parameter. For more details, see the [Parameter Overrides documentation](/docs/features/large-language-models/parameter_overrides). ``` models: - name: pirate-haikus from: openai:gpt-4o params: system_prompt: | Write everything in Haiku like a pirate ``` ### Example: Serving a Local Model[​](#example-serving-a-local-model "Direct link to Example: Serving a Local Model") To serve a model from the local filesystem, specify the `from` path as `file` and provide the local path. For more details, see [Filesystem Model Provider](/docs/components/models/filesystem). ``` models: - from: file://absolute/path/to/my/model.onnx name: local_fs_model ``` ### Example: Analyzing GitHub Issues with a Chat Model[​](#example-analyzing-github-issues-with-a-chat-model "Direct link to Example: Analyzing GitHub Issues with a Chat Model") This example demonstrates how to pull GitHub issue data from the last 14 days, accelerate the data, create a chat model with memory and tools to access the accelerated data, and use Spice to ask the chat model about the general themes of new issues. #### Step 1: Pull GitHub Issue Data[​](#step-1-pull-github-issue-data "Direct link to Step 1: Pull GitHub Issue Data") First, configure a dataset to pull GitHub issue data from the last 14 days. ``` datasets: - from: github:github.com///issues name: github_issues params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true refresh_mode: append refresh_check_interval: 24h refresh_data_window: 14d ``` #### Step 2: Create a Chat Model with Memory and Tools[​](#step-2-create-a-chat-model-with-memory-and-tools "Direct link to Step 2: Create a Chat Model with Memory and Tools") Next, create a chat model that includes memory and tools to access the accelerated GitHub issue data. ``` datasets: - from: memory:store name: llm_memory access: read_write models: - name: github-issues-analyzer from: openai:gpt-4o params: tools: memory, sql ``` #### Step 3: Query the Chat Model[​](#step-3-query-the-chat-model "Direct link to Step 3: Query the Chat Model") At this step, the `spicepod.yaml` should look like: ``` datasets: - from: github:github.com///issues name: github_issues params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true refresh_mode: append refresh_check_interval: 24h refresh_data_window: 14d - from: memory:store name: llm_memory access: read_write models: - name: github-issues-analyzer from: openai:gpt-4o params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } tools: memory, sql ``` Finally, use Spice to ask the chat model about the general themes of new issues in the last 14 days. The following `curl` command demonstrates how to make this request using the OpenAI-compatible API. ``` curl -X POST http://localhost:8090/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "github-issues-analyzer", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What are the general themes of new issues in the last 14 days?"} ] }' ``` Refer to the [Create Chat Completion API documentation](/docs/api/HTTP/post-chat-completions) for more details on making chat completion requests. ## [📄️OpenAI](/docs/v1.10/components/models/openai) [Instructions for using language models hosted on OpenAI or compatible services with Spice.](/docs/v1.10/components/models/openai) ## [📄️Azure OpenAI](/docs/v1.10/components/models/azure) [Instructions for using Azure OpenAI models](/docs/v1.10/components/models/azure) ## [📄️Anthropic](/docs/v1.10/components/models/anthropic) [Instructions for using language models hosted on Anthropic with Spice.](/docs/v1.10/components/models/anthropic) ## [📄️HuggingFace](/docs/v1.10/components/models/huggingface) [Instructions for using machine learning models hosted on HuggingFace with Spice.](/docs/v1.10/components/models/huggingface) ## [📄️Perplexity](/docs/v1.10/components/models/perplexity) [Instructions for using language models hosted on Perplexity with Spice.](/docs/v1.10/components/models/perplexity) ## [📄️Filesystem](/docs/v1.10/components/models/filesystem) [Instructions for using models hosted on a filesystem with Spice.](/docs/v1.10/components/models/filesystem) ## [📄️Spice Cloud Platform](/docs/v1.10/components/models/spiceai) [Instructions for using models hosted on the Spice Cloud Platform with Spice.](/docs/v1.10/components/models/spiceai) ## [📄️xAI](/docs/v1.10/components/models/xai) [Instructions for using xAI models](/docs/v1.10/components/models/xai) ## [📄️Databricks](/docs/v1.10/components/models/databricks) [Instructions for using Databricks Mosaic AI Models](/docs/v1.10/components/models/databricks) ## [📄️Bedrock](/docs/v1.10/components/models/bedrock) [How to use Amazon Bedrock models with Spice.](/docs/v1.10/components/models/bedrock) --- # Anthropic Models To use a language model hosted on Anthropic, specify `anthropic` in the `from` field. To use a specific model, include its model ID in the `from` field (see example below). If not specified, the default model is `claude-3-5-sonnet-latest`. The following parameters are specific to Anthropic models: | Parameter | Description | Default | | ------------------- | -------------------------------- | ------------------------------ | | `anthropic_api_key` | The Anthropic API key. | - | | `endpoint` | The Anthropic API base endpoint. | `https://api.anthropic.com/v1` | Example `spicepod.yml` configuration: ``` models: - from: anthropic:claude-3-5-sonnet-latest name: claude_3_5_sonnet params: anthropic_api_key: ${ secrets:SPICE_ANTHROPIC_API_KEY } ``` See [Anthropic Model Names](https://docs.anthropic.com/en/docs/about-claude/models#model-names) for a list of supported model names. --- # Azure OpenAI Models To use a language model hosted on Azure OpenAI, specify the `azure` path in the `from` field and the following parameters from the [Azure OpenAI Model Deployment](https://ai.azure.com/resource/deployments) page: | Param | Description | Default | | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------- | | `azure_api_key` | The Azure OpenAI API key from the models deployment page. | - | | `azure_api_version` | The API version used for the Azure OpenAI service. | - | | `azure_deployment_name` | The name of the model deployment. | Model name | | `endpoint` | The Azure OpenAI resource endpoint, e.g., `https://resource-name.openai.azure.com`. | - | | `azure_entra_token` | The Azure Entra token for authentication. | - | | `responses_api` | `enabled` or `disabled`. Whether to enable invoking this model from the `/v1/responses` HTTP endpoint | `disabled` | | `azure_openai_responses_tools` | Comma-separated list of OpenAI-hosted tools exposed via the Responses API for this model. These hosted tools are **not** available from the `/v1/chat/completions` HTTP endpoint. Supported tools: `code_interpreter`, `web_search`. | - | Only one of `azure_api_key` or `azure_entra_token` can be provided for model configuration. Example: ``` models: - from: azure:gpt-4o-mini name: gpt-4o-mini params: endpoint: ${ secrets:SPICE_AZURE_AI_ENDPOINT } azure_api_version: 2024-08-01-preview azure_deployment_name: gpt-4o-mini azure_api_key: ${ secrets:SPICE_AZURE_API_KEY } # Responses API configuration responses_api: enabled azure_openai_responses_tools: web_search ``` Refer to the [Azure OpenAI Service models](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models) for more details on available models and configurations. Follow the [Azure OpenAI Models Cookbook](https://github.com/spiceai/cookbook/tree/trunk/azure_openai) to try Azure OpenAI models for vector-based search and chat functionalities with structured (taxi trips) and unstructured (GitHub files) data. --- # Amazon Bedrock Models Amazon Bedrock provides access to a range of foundation models for generative AI. Spice supports using Bedrock-hosted models by specifying the `bedrock` prefix in the `from` field and configuring the required parameters. ## Supported Model IDs[​](#supported-model-ids "Direct link to Supported Model IDs") The following model IDs are supported: * `amazon.nova-lite-v1:0` * `amazon.nova-micro-v1:0` * `amazon.nova-premier-v1:0` * `amazon.nova-pro-v1:0` Refer to the [Amazon Bedrock documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/model-ids.html) for details on available models and cross-region inference profiles. To request support for a model, file a GitHub Issue or ask us on Slack. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") Specify the Bedrock model ID in the `from` field: ``` models: - from: bedrock:us.amazon.nova-lite-v1:0 name: novash params: aws_region: us-east-1 aws_access_key_id: ${ secrets:AWS_ACCESS_KEY_ID } aws_secret_access_key: ${ secrets:AWS_SECRET_ACCESS_KEY } ``` ### Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | Default | | ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | | `aws_region` | AWS region for Bedrock API requests. | - | | `aws_access_key_id` | AWS access key ID. If not provided, credentials will be loaded from environment variables or IAM roles. | - | | `aws_secret_access_key` | AWS secret access key. If not provided, credentials will be loaded from environment variables or IAM roles. | - | | `aws_session_token` | Session token (e.g. AWS\_SESSION\_TOKEN for AWS) for temporary credentials | - | | `bedrock_guardrail_identifier` | Identifier for the guardrail. See [GuardrailConfiguration](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_GuardrailConfiguration.html). Pattern: `(([a-z0-9]+) \| (arn:aws(-[^:]+)?:bedrock:[a-z0-9-]{1,20}:[0-9]{12}:guardrail/[a-z0-9]+))`. Length: 0-2048. | - | | `bedrock_guardrail_version` | Guardrail version. Pattern: `(([1-9][0-9]{0,7}) \| (DRAFT))` | - | | `bedrock_trace` | Trace behavior for the guardrail. Valid values: `enabled`, `disabled`, `enabled_full`. Default: `disabled`. | disabled | ### OpenAI-Compatible Overrides[​](#openai-compatible-overrides "Direct link to OpenAI-Compatible Overrides") The following OpenAI-compatible parameters are supported and passed in the request payload: * `maxTokens` * `temperature` * `topP` * `topK` * `stopSequences` See [Parameter Overrides](https://spiceai.org/docs/features/large-language-models/parameter_overrides) for details. #### Model Parameters[​](#model-parameters "Direct link to Model Parameters") These parameters control model behavior and are passed in the request payload: | Parameter | Description | | --------------- | --------------------------------------------------------------- | | `maxTokens` | Maximum number of tokens to generate. | | `temperature` | Sampling temperature (0.0 to 1.0). Lower is more deterministic. | | `topP` | Nucleus sampling probability (0.0 to 1.0). | | `topK` | Number of highest probability tokens to consider. | | `stopSequences` | Sequences that stop generation when encountered. | See [Parameter Overrides](/docs/v1.10/features/large-language-models/parameter_overrides) for details on setting default values. ## Examples[​](#examples "Direct link to Examples") ### Basic Configuration[​](#basic-configuration "Direct link to Basic Configuration") ``` models: - from: bedrock:us.amazon.nova-lite-v1:0 name: novash params: aws_region: us-east-1 aws_access_key_id: ${ secrets:AWS_ACCESS_KEY_ID } aws_secret_access_key: ${ secrets:AWS_SECRET_ACCESS_KEY } bedrock_guardrail_identifier: arn:aws:bedrock:abcdefg012927:0123456789876:guardrail/hello bedrock_guardrail_version: DRAFT bedrock_trace: enabled bedrock_temperature: 42 ``` ## Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate bedrock permissions (e.g., `bedrock:InvokeModel`) to access the model. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access DynamoDB tables: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"], "Resource": ["arn:aws:bedrock:us-east-1::foundation-model/amazon.titan-*"] } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | --------------------------------------- | ----------------------------------------------------------------- | | `bedrock:InvokeModel` | Required. Used to invoke the text model. | | `bedrock:InvokeModelWithResponseStream` | Required. Used to invoke the text model with streaming responses. | * [Amazon Bedrock Embeddings](/docs/v1.10/components/embeddings/bedrock) - Use Bedrock for text embeddings * [Parameter Overrides](/docs/v1.10/features/large-language-models/parameter_overrides) - Set default model parameters * [Amazon Bedrock User Guide](https://docs.aws.amazon.com/bedrock/latest/userguide/) - AWS documentation * [Bedrock Model IDs](https://docs.aws.amazon.com/bedrock/latest/userguide/model-ids.html) - Available models and inference profiles --- # Databricks Model Provider To use a language model deployed to [Databricks Mosaic AI Model Serving](https://docs.databricks.com/aws/en/machine-learning/model-serving/), specify the model endpoint name prefixed with `databricks:` in the `from` field and include the required parameters in the `params` section. ### Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `databricks_endpoint` | The Databricks workspace endpoint, e.g., `dbc-a12cd3e4-56f7.cloud.databricks.com`. | | `databricks_token` | The Databricks API token to authenticate with the Databricks Models API. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to reference a secret, e.g., `${secrets:my_databricks_token}`. | | `databricks_client_id` | The Databricks Service Principal Client ID. Can't be used with `databricks_token`. | | `databricks_client_secret` | The Databricks Service Principal Client Secret. Can't be used with `databricks_token`. | ### Example `spicepod.yaml` configuration, using personal access token[​](#example-spicepodyaml-configuration-using-personal-access-token "Direct link to example-spicepodyaml-configuration-using-personal-access-token") To learn more about how to set up personal access tokens, see [Databricks PAT docs](https://docs.databricks.com/aws/en/dev-tools/auth/pat). ``` models: - from: databricks:databricks-llama-4-maverick name: llama-4-maverick params: databricks_endpoint: dbc-46470731-42e5.cloud.databricks.com databricks_token: ${ secrets:SPICE_DATABRICKS_TOKEN } ``` ### Example `spicepod.yaml` configuration, using Databricks service principal[​](#example-spicepodyaml-configuration-using-databricks-service-principal "Direct link to example-spicepodyaml-configuration-using-databricks-service-principal") Spice supports the Machine-to-Machine (M2M) OAuth flow with service principal credentials by utilizing the `databricks_client_id` and `databricks_client_secret` parameters. The runtime will automatically refresh the token. The service principal must be granted the "Can Query" permission for model serving. To learn more about how to set up the service principal, see [Databricks M2M OAuth docs](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m). ``` models: - from: databricks:databricks-llama-4-maverick name: llama-4-maverick params: databricks_endpoint: dbc-46470731-42e5.cloud.databricks.com databricks_client_id: ${secrets:DATABRICKS_CLIENT_ID} databricks_client_secret: ${secrets:DATABRICKS_CLIENT_SECRET} ``` ### Additional Information[​](#additional-information "Direct link to Additional Information") Refer to the [Mosaic AI Model Serving documentation](https://docs.databricks.com/aws/en/machine-learning/model-serving/) for more details on available models and configurations. --- # Filesystem Hosted Models To use a model hosted on a filesystem, specify the path to the model file or folder in the `from` field: ``` models: - from: file://models/llms/llama3.2-1b-instruct/ name: llama3 ``` Supported formats include GGUF, GGML, and SafeTensor for large language models (LLMs) and ONNX for traditional machine learning (ML) models. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") An absolute or relative path to the model file or folder: ``` from: file://absolute/path/models/llms/llama3.2-1b-instruct/ from: file:models/llms/llama3.2-1b-instruct/ ``` ### `params` (optional)[​](#params-optional "Direct link to params-optional") | Param | Description | | --------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `tools` | Which [tools](/docs/v1.10/features/large-language-models/tools) should be made available to the model. Set to `auto` to use all available tools. | | `system_prompt` | An additional system prompt used for all chat completions to this model. | | `chat_template` | Customizes the transformation of OpenAI chat messages into a character stream for the model. See [Overriding the Chat Template](#overriding-the-chat-template). | See [Large Language Models](/docs/v1.10/features/large-language-models) for additional configuration options. * [Tools](/docs/v1.10/features/large-language-models/tools) * [Memory](/docs/v1.10/features/large-language-models/memory) * [Evals](/docs/v1.10/features/large-language-models/evals) * [Parameter overrides](/docs/v1.10/features/large-language-models/parameter_overrides) ### `files` (optional)[​](#files-optional "Direct link to files-optional") The `files` field specifies additional files required by the model, such as tokenizer, configuration, and other files. ``` - name: local-model from: file://models/llms/llama3.2-1b-instruct/model.safetensors files: - path: //models/llms/llama3.2-1b-instruct/tokenizer.json - path: //models/llms/llama3.2-1b-instruct/tokenizer_config.json - path: //models/llms/llama3.2-1b-instruct/config.json ``` ## Examples[​](#examples "Direct link to Examples") ### Loading a GGML Model[​](#loading-a-ggml-model "Direct link to Loading a GGML Model") ``` models: - from: file://absolute/path/to/my/model.ggml name: local_ggml_model files: - path: models/llms/ggml/tokenizer.json - path: models/llms/ggml/tokenizer_config.json - path: models/llms/ggml/config.json ``` ### Example: Loading a SafeTensor Model[​](#example-loading-a-safetensor-model "Direct link to Example: Loading a SafeTensor Model") ``` models: - name: safety from: file:models/llms/llama3.2-1b-instruct/model.safetensors files: - path: models/llms/llama3.2-1b-instruct/tokenizer.json - path: models/llms/llama3.2-1b-instruct/tokenizer_config.json - path: models/llms/llama3.2-1b-instruct/config.json ``` ### Loading LLM from a directory[​](#loading-llm-from-a-directory "Direct link to Loading LLM from a directory") ``` models: - name: llama3 from: file:models/llms/llama3.2-1b-instruct/ ``` Note: The folder provided should contain all the expected files (see examples above). ### Loading an ONNX Model[​](#loading-an-onnx-model "Direct link to Loading an ONNX Model") ``` models: - from: file://absolute/path/to/my/model.onnx name: local_fs_model ``` ### Loading a GGUF Model[​](#loading-a-gguf-model "Direct link to Loading a GGUF Model") ``` models: - from: file://absolute/path/to/my/model.gguf name: local_gguf_model ``` ### Overriding the Chat Template[​](#overriding-the-chat-template "Direct link to Overriding the Chat Template") Chat templates convert the OpenAI compatible chat messages (see [format](https://platform.openai.com/docs/api-reference/chat/create#chat-create-messages)) and other components of a request into a stream of characters for the language model. It follows Jinja3 templating [syntax](https://jinja.palletsprojects.com/en/3.1.x/templates/). Further details on chat templates can be found [here](https://huggingface.co/docs/transformers/main/chat_templating#advanced-how-do-chat-templates-work). ``` models: - name: local_model from: file:path/to/my/model.gguf params: chat_template: | {% set loop_messages = messages %} {% for message in loop_messages %} {% set content = '<|start_header_id|>' + message['role'] + '<|end_header_id|>\n\n'+ message['content'] | trim + '<|eot_id|>' %} {{ content }} {% endfor %} {% if add_generation_prompt %} {{ '<|start_header_id|>assistant<|end_header_id|>\n\n' }} {% endif %} ``` #### Templating Variables[​](#templating-variables "Direct link to Templating Variables") * `messages`: List of chat messages, in the OpenAI [format](https://platform.openai.com/docs/api-reference/chat/create#chat-create-messages). * `add_generation_prompt`: Boolean flag whether to add a [generation prompt](https://huggingface.co/docs/transformers/main/chat_templating#what-are-generation-prompts). * `tools`: List of callable tools, in the OpenAI [format](https://platform.openai.com/docs/api-reference/chat/create#chat-create-tools). Limitations * The throughput, concurrency & latency of a locally hosted model will vary based on the underlying hardware and model size. Spice supports [Apple metal](/docs/v1.10/installation#metal-support) and [CUDA](/docs/v1.10/installation#cuda-support) for accelerated inference. --- # HuggingFace To use a model hosted on HuggingFace, specify the `huggingface.co` path in the `from` field and, when needed, the files to include. ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` key takes the form of `huggingface:model_path`. Below shows 2 common example of `from` key configuration. * `huggingface:username/modelname`: Implies the latest version of `modelname` hosted by `username`. * `huggingface:huggingface.co/username/modelname:revision`: Specifies a particular `revision` of `modelname` by `username`, including the optional domain. The `from` key follows the following regex format. ``` \A(huggingface:)(huggingface\.co\/)?(?[\w\-]+)\/(?[\w\-]+)(:(?[\w\d\-\.]+))?\z ``` The `from` key consists of five components: 1. **Prefix:** The value must start with `huggingface:`. 2. **Domain (Optional):** Optionally includes `huggingface.co/` immediately after the prefix. Currently no other Huggingface compatible services are supported. 3. **Organization/User:** The HuggingFace organization (`org`). 4. **Model Name:** After a `/`, the model name (`model`). 5. **Revision (Optional):** A colon (`:`) followed by the git-like revision identifier (`revision`). ### `name`[​](#name "Direct link to name") The model name. This will be used as the model ID within Spice and Spice's endpoints (i.e. `http://localhost:8090/v1/models`). This can be set to the same value as the model ID in the `from` field. ### `params`[​](#params "Direct link to params") | Param | Description | Default | | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- | | `hf_token` | The Huggingface access token. | - | | `model_type` | The architecture to load the model as. Supported values: `mistral`, `gemma`, `mixtral`, `llama`, `phi2`, `phi3`, `qwen2`, `gemma2`, `starcoder2`, `phi3.5moe`, `deepseekv2`, `deepseekv3` | - | | `tools` | Which \[tools] should be made available to the model. Set to `auto` to use all available tools. | - | | `system_prompt` | An additional system prompt used for all chat completions to this model. | - | ### `files`[​](#files "Direct link to files") The specific file path for Huggingface model. For example, GGUF model formats require a specific file path, other varieties (e.g. `.safetensors`) are inferred. #### Example[​](#example "Direct link to Example") ``` models: - from: huggingface:huggingface.co/lmstudio-community/Qwen2.5-Coder-3B-Instruct-GGUF name: sloth-gguf files: - path: Qwen2.5-Coder-3B-Instruct-Q3_K_L.gguf ``` ## Access Tokens[​](#access-tokens "Direct link to Access Tokens") Access tokens can be provided for Huggingface models in two ways: 1. In the Huggingface token cache (i.e. `~/.cache/huggingface/token`). Default. 2. Via [model params](#params). ``` models: - name: llama_3.2_1B from: huggingface:huggingface.co/meta-llama/Llama-3.2-1B params: hf_token: ${ secrets:HF_TOKEN } ``` ## Examples[​](#examples "Direct link to Examples") ### Load a ML model to predict taxi trips outcomes[​](#load-a-ml-model-to-predict-taxi-trips-outcomes "Direct link to Load a ML model to predict taxi trips outcomes") ``` models: - from: huggingface:huggingface.co/spiceai/darts:latest name: hf_model files: - path: model.onnx datasets: - taxi_trips ``` ### Load a LLM model to generate text[​](#load-a-llm-model-to-generate-text "Direct link to Load a LLM model to generate text") ``` models: - from: huggingface:huggingface.co/microsoft/Phi-3.5-mini-instruct name: phi ``` ### Load a private model[​](#load-a-private-model "Direct link to Load a private model") ``` models: - name: llama_3.2_1B from: huggingface:huggingface.co/meta-llama/Llama-3.2-1B params: hf_token: ${ secrets:HF_TOKEN } ``` For more details on authentication, see [access tokens](#access-tokens). Limitations * The throughput, concurrency & latency of a locally hosted model will vary based on the underlying hardware and model size. Spice supports [Apple metal](/docs/v1.10/installation#metal-support) and [CUDA](/docs/v1.10/installation#cuda-support) for accelerated inference. * ML models currently only support ONNX file format. ## Cookbook[​](#cookbook "Direct link to Cookbook") * Use the Llama family of models locally from HuggingFace using Spice. [Running Llama3 Locally](https://github.com/spiceai/cookbook/blob/trunk/llama/README.md) --- # OpenAI (or Compatible) Language Models To use a language model hosted on OpenAI (or compatible), specify the `openai` path in the `from` field. For a specific model, include it as the model ID in the `from` field (see example below). The default model is `gpt-4o-mini`. ``` models: - from: openai:gpt-4o-mini name: openai_model params: openai_api_key: ${ secrets:OPENAI_API_KEY } # Required for official OpenAI models tools: auto # Optional. Connect the model to datasets via SQL query/vector search tools system_prompt: 'You are a helpful assistant.' # Optional. # Optional parameters endpoint: https://api.openai.com/v1 # Override to use a compatible provider (i.e. NVidia NIM) openai_org_id: ${ secrets:OPENAI_ORG_ID } openai_project_id: ${ secrets:OPENAI_PROJECT_ID } # Override default chat completion request parameters openai_temperature: 0.1 openai_response_format: { 'type': 'json_object' } # OpenAI Responses API configuration responses_api: enabled openai_responses_tools: web_search, code_interpreter ``` ## Configuration[​](#configuration "Direct link to Configuration") ### `from`[​](#from "Direct link to from") The `from` field takes the form `openai:model_id` where `model_id` is the model ID of the OpenAI model, valid model IDs are found in the `{endpoint}/v1/models` API response. Example: ``` curl -H "Authorization: Bearer $OPENAI_API_KEY" https://api.openai.com/v1/models ``` ``` { "object": "list", "data": [ { "id": "gpt-4o-mini", "object": "model", "created": 1727389042, "owned_by": "system" }, ... } ``` ### `name`[​](#name "Direct link to name") The model name. This will be used as the model ID within Spice and Spice's endpoints (i.e. `http://localhost:8090/v1/models`). This can be set to the same value as the model ID in the `from` field. ### `params`[​](#params "Direct link to params") | Param | Description | Default | | ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------- | | `endpoint` | The OpenAI API base endpoint. Can be overridden to use a compatible provider (i.e. Nvidia NIM). | `https://api.openai.com/v1` | | `tools` | Which [tools](/docs/v1.10/features/large-language-models/tools) should be made available to the model. Set to `auto` to use all available tools. | - | | `system_prompt` | An additional system prompt used for all chat completions to this model. | - | | `openai_api_key` | The OpenAI API key. | - | | `openai_org_id` | The OpenAI organization ID. | - | | `openai_project_id` | The OpenAI project ID. | - | | `openai_temperature` | Set the default temperature to use on chat completions. | - | | `openai_response_format` | An object specifying the format that the model must output, see [structured outputs](https://platform.openai.com/docs/guides/structured-outputs). | - | | `openai_reasoning_effort` | For reasoning models, like `o1`, this parameter specifies the reasoning effort used for the model. | - | | `openai_usage_tier` | The [OpenAI usage tier](https://platform.openai.com/settings/organization/limits) for the account. This parameter sets the maximum number of concurrent requests based on OpenAI's published limits per tier. Valid values are `free`, `tier1`, `tier2`, `tier3`, `tier4`, or `tier5`. | `tier1` | | `responses_api` | `enabled` or `disabled`. Whether to enable invoking this model from the `/v1/responses` HTTP endpoint using [OpenAI's Responses API](https://platform.openai.com/docs/api-reference/responses). When using OpenAI-compatible providers, ensure the provider supports OpenAI's Responses API. | `disabled` | | `openai_responses_tools` | Comma-separated list of OpenAI-hosted tools exposed via the Responses API for this model. These hosted tools are **not** available from the `/v1/chat/completions` HTTP endpoint. Supported tools: `code_interpreter`, `web_search`. | - | See [Large Language Models](/docs/v1.10/features/large-language-models) for additional configuration options. * [Tools](/docs/v1.10/features/large-language-models/tools) * [Memory](/docs/v1.10/features/large-language-models/memory) * [Evals](/docs/v1.10/features/large-language-models/evals) * [Parameter overrides](/docs/v1.10/features/large-language-models/parameter_overrides) ## Supported OpenAI Compatible Providers[​](#supported-openai-compatible-providers "Direct link to Supported OpenAI Compatible Providers") Spice supports several OpenAI compatible providers. Specify the appropriate endpoint in the params section. ### Azure OpenAI[​](#azure-openai "Direct link to Azure OpenAI") Follow [Azure AI Models](/docs/v1.10/components/models/azure) instructions. ### Groq[​](#groq "Direct link to Groq") Groq provides OpenAI compatible endpoints. Use the following configuration: ``` models: - from: openai:llama3-groq-70b-8192-tool-use-preview name: groq-llama params: endpoint: https://api.groq.com/openai/v1 openai_api_key: ${ secrets:SPICE_GROQ_API_KEY } ``` ### NVidia NIM[​](#nvidia-nim "Direct link to NVidia NIM") NVidia NIM models are OpenAI compatible endpoints. Use the following configuration: ``` models: - from: openai:my_nim_model_id name: my_nim_model params: endpoint: https://my_nim_host.com/v1 openai_api_key: ${ secrets:SPICE_NIM_API_KEY } ``` View the Spice cookbook for an example of setting up NVidia NIM with Spice [here](https://github.com/spiceai/cookbook/tree/trunk/nvidia-nim/ec2). ### Parasail[​](#parasail "Direct link to Parasail") Parasail also offers OpenAI compatible endpoints. Use the following configuration: ``` models: - from: openai:parasail-model-id name: parasail_model params: endpoint: https://api.parasail.com/v1 openai_api_key: ${ secrets:SPICE_PARASAIL_API_KEY } ``` Refer to the respective provider documentation for more details on available models and configurations. --- # Perplexity Models To use a language model hosted on Perplexity, specify `perplexity` in the `from` field. To use a specific model, include its model ID in the `from` field (see example below). If not specified, the default model is `sonar`. ``` models: - name: webs from: perplexity:sonar params: perplexity_auth_token: ${ secrets:SPICE_PERPLEXITY_AUTH_TOKEN } ``` The following parameters are specific to Perplexity models: | Parameter | Description | Default | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------- | | `perplexity_auth_token` | The Perplexity API authentication token. | - | | `perplexity_*` | Additional, perplexity specific parameters to use on all requests. See [Perplexity API Reference](https://docs.perplexity.ai/api-reference/chat-completions) | - | **Note:** Like other models in Spice, Perplexity can set default overrides for OpenAI parameters. See [Parameter Overrides](/docs/v1.10/features/large-language-models/parameter_overrides). ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` models: - name: webs from: perplexity:sonar params: perplexity_auth_token: ${ secrets:SPICE_PERPLEXITY_AUTH_TOKEN } perplexity_search_domain_filter: - docs.spiceai.org - huggingface.co perplexity_temperature: 0.3141595 ``` --- # Spice Cloud Platform To use a model hosted on the [Spice Cloud Platform](https://docs.spice.ai/building-blocks/spice-models), specify the `spice.ai` path in the `from` field. Example: ``` models: - from: spice.ai/taxi_tech_co/taxi_drives/models/drive_stats name: drive_stats datasets: - drive_stats_inferencing ``` Specific model versions can be referenced using a version label or Training Run ID. ``` models: - from: spice.ai/taxi_tech_co/taxi_drives/models/drive_stats:latest # Label name: drive_stats_a datasets: - drive_stats_inferencing - from: spice.ai/taxi_tech_co/taxi_drives/models/drive_stats:60cb80a2-d59b-45c4-9b68-0946303bdcaf # Training Run ID name: drive_stats_b datasets: - drive_stats_inferencing ``` ## `from` Format[​](#from-format "Direct link to from-format") The from key must conform to the following regex format: ``` \A(?:spice\.ai\/)?(?[\w\-]+)\/(?[\w\-]+)(?:\/models)?\/(?[\w\-]+):(?[\w\d\-\.]+)\z ``` Examples: * `spice.ai/lukekim/smart/models/drive_stats:latest`: Refers to the latest version of the drive\_stats model in the smart application by the user or organization lukekim. * `spice.ai/lukekim/smart/drive_stats:60cb80a2-d59b-45c4-9b68-0946303bdcaf`: Specifies a model with a unique training run ID. ### Specification[​](#specification "Direct link to Specification") 1. **Prefix (Optional):** The value must start with `spice.ai/`. 2. **Organization/User:** The name of the organization or user (`org`) hosting the model. 3. **Application Name**: The name of the application (`app`) which the model belongs to. 4. **Model Name:** The name of the model (`model`). 5. **Version (Optional):** A colon (`:`) followed by the version identifier (`version`), which could be a semantic version, `latest` for the most recent version, or a specific training run ID. --- # xAI Models To use a language model hosted on xAI, specify `xai` path in the `from` field and the associated `xai_api_key` parameter: | Param | Description | Default | | ------------- | ---------------- | ------- | | `xai_api_key` | The xAI API key. | - | Example: ``` models: - from: xai:grok4 name: xai params: xai_api_key: ${secrets:SPICE_GROK_API_KEY} ``` Refer to the [xAI models documentation](https://docs.x.ai/docs/models) for more details on available models and configurations. note Although the xAI [documentation](https://docs.x.ai/docs/guides/structured-outputs) shows that xAI models can return structured outputs, this is not true. --- # Secret Stores A Secret Store is a location where `secrets` are stored and can be used to store sensitive data, like passwords, tokens, and secret keys. Spice supports secret stores: [`env`](/docs/components/secret-stores/env), [`kubernetes`](/docs/components/secret-stores/kubernetes), [`keyring`](/docs/components/secret-stores/keyring) and [`aws_secrets_manager`](/docs/components/secret-stores/aws-secrets-manager). The `env` secret store is loaded by default. ### Default[​](#default "Direct link to Default") The `env` secret store is loaded by default. It reads secrets from environment variables and any `.env.local` or `.env` files in the project directory. ``` secrets: - from: env name: env ``` ## Configured Secret Stores[​](#configured-secret-stores "Direct link to Configured Secret Stores") Secret Stores can be configured using the `secrets` section of the `spicepod.yml` file. The Secret Store type and name are specified using the `from` and `name` fields. The `name` can be referenced by other components, like datasets or models. Some Secret Stores support adding a selector delimited by a colon (`:`), For example, when using the Kubernetes Secret Store, `from: kubernetes:my_secret` selects and enables the `my_secret` secret only to be referenced. Additional parameters may be specified in the `params` field, which are typically specific to the secret store type. Example: ``` secrets: - from: kubernetes:my_secret name: k8s - from: env name: env ``` ## Using referenced secrets in component parameters[​](#using-secrets "Direct link to Using referenced secrets in component parameters") Secrets may be used by components with the syntax `${:}`. For example, to reference a secret stored as an environment variable named `MY_SECRET` in the `env` secret store, use `${env:MY_SECRET}`. Example: ``` datasets: - from: postgres:my_table name: my_table params: pg_host: localhost pg_port: 5432 pg_user: ${env:PG_USER} pg_pass: ${env:FOO_PASSWORD} # The environment variable name may differ from the parameter name. ``` This syntax also works within a larger string, like a connection string: ``` datasets: - from: mysql:my_table name: my_table params: mysql_connection_string: mysql://${env:USER}:${env:PASSWORD}@localhost:3306/mysql_db ``` The `` value in `${:}` is the `name` value defined in the secret store configuration. This can be renamed to any value. Example: ``` secrets: - from: env name: my_env ``` ``` datasets: - from: postgres:my_table name: my_table params: pg_host: localhost pg_port: 5432 pg_user: ${my_env:PG_USER} pg_pass: ${my_env:PG_PASS} ``` ## Load secrets from multiple secret stores[​](#load-secrets-from-multiple-secret-stores "Direct link to Load secrets from multiple secret stores") Spice supports configuring multiple secret stores which are loaded in the order they are defined in the `secrets` section of the `spicepod.yml` configuration file. If a secret is defined in multiple secret stores, the secret store defined last will take precedence. To load a secret from any of the configured secret stores in precedence order, use the `${secrets:}` syntax. Example: ``` secrets: - from: env name: env - from: keyring name: keyring ``` ``` datasets: - from: postgres:my_table name: my_table params: pg_host: localhost pg_port: 5432 pg_user: ${secrets:pg_user} pg_pass: ${secrets:pg_pass} ``` In this example, the runtime would look for `pg_user` and `pg_pass` in the `keyring` secret store first and then in the `env` secret store. The `` value in `${secrets:}` is automatically uppercased for the `env` secret store. To override the `keyring` secret store secrets with environment variables, re-order the secret stores in the configuration file: ``` secrets: - from: keyring name: keyring - from: env name: env ``` ## Secret Stores[​](#secret-stores "Direct link to Secret Stores") ## [📄️Environment Secret Store](/docs/v1.10/components/secret-stores/env) [Environment Variables Secret Store Documentation](/docs/v1.10/components/secret-stores/env) ## [📄️AWS Secrets Manager Secret Store](/docs/v1.10/components/secret-stores/aws-secrets-manager) [AWS Secrets Manager Secret Store Documentation](/docs/v1.10/components/secret-stores/aws-secrets-manager) ## [📄️Kubernetes Secret Store](/docs/v1.10/components/secret-stores/kubernetes) [Kubernetes Secret Store Documentation](/docs/v1.10/components/secret-stores/kubernetes) ## [📄️Keyring Secret Store](/docs/v1.10/components/secret-stores/keyring) [Keyring Secret Store Documentation](/docs/v1.10/components/secret-stores/keyring) --- # AWS Secrets Manager Secret Store The `aws_secrets_manager` store enables Spice to read secrets from [AWS Secrets Manager](https://aws.amazon.com/secrets-manager/) by specifying the secret’s name with a selector. ``` secrets: from: aws_secrets_manager:my_secret_name name: aws ``` The store reads keys from the secret named in the selector. In the above example `my_secret_name` must be defined in [AWS Secrets Manager](https://console.aws.amazon.com/secretsmanager/listsecrets), and any keys referenced using `${aws:my_key}` will look for a key `my_key` within `my_secret_name`. ![](/img/secrets-aws-secrets-manager-1.png) ![](/img/secrets-aws-secrets-manager-2.png) ## Example[​](#example "Direct link to Example") A complete spicepod definition with a dataset that uses a secret from AWS Secrets Manager. ``` version: v1 kind: Spicepod name: taxi_trips secrets: - from: aws_secrets_manager:dremio name: dremio datasets: - from: dremio:datasets.taxi_trips name: taxi_trips description: dremio taxi trips params: dremio_endpoint: grpc://20.163.171.81:32010 dremio_username: ${dremio:username} dremio_password: ${dremio:password} ``` ### Authentication[​](#authentication "Direct link to Authentication") Spice will automatically load credentials to connect to AWS Secrets Manager from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate secretsmanager permissions (e.g., `secretsmanager:GetSecretValue`) to access the secrets. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following permissions to access secrets in Secrets Manager: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "secretsmanager:GetSecretValue" ], "Resource": [ "arn:aws:secretsmanager:us-east-1:123456789012:secret:TestEnv/*" ] } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------------------- | --------------------------------------- | | `secretsmanager:GetSecretValue` | Required. Allows reading secret values. | --- # Environment Secret Store The `env` store type enables Spice to read secrets from environment variables and any `.env.local` or `.env` files in the project directory. This is the default secret store and is loaded automatically as: ``` secrets: - from: env name: env ``` Reference secrets directly in parameters using the syntax `${env:MY_ENV_VAR}`. This will load the value of the environment variable `MY_ENV_VAR` into the parameter. Example: ``` datasets: - from: postgres:my_table name: my_table params: pg_host: localhost pg_port: 5432 pg_user: ${env:MY_PG_USER} pg_pass: ${env:MY_PG_PASSWORD} ``` The `${}` replacement syntax also works within a larger string, like a connection string: ``` datasets: - from: mysql:my_table name: my_table params: connection_string: mysql://${env:MY_USER}:${env:MY_PASSWORD}@localhost:3306/my_db ``` When used with the `${secrets:}` syntax, the `` variable is UPPERCASED to follow the convention of environment variables. Example: ``` datasets: - from: postgres:my_table name: my_table params: pg_host: localhost pg_port: 5432 pg_user: ${secrets:my_pg_user} # same as ${env:MY_PG_USER} pg_pass: ${secrets:my_pg_password} # same as ${env:MY_PG_PASSWORD} ``` ## .env Files[​](#env-files "Direct link to .env Files") The `env` secret store reads secrets from any `.env.local` or `.env` files in the project directory. The `.env.local` file takes precedence over the `.env` file. This enables defining template secrets in the `.env` file which can be checked into source control and overriding them with local secrets in the `.env.local` file. Example `.env` file: ``` MY_PG_USER=postgres MY_PG_PASSWORD=postgres ``` ### Additional Parameters[​](#additional-parameters "Direct link to Additional Parameters") To load environment variables from a specific `.env` file, use the `file_path` parameter. When a `file_path` parameter is specified, environment variables from `.env` or `.env.local` will not be loaded. ``` secrets: - from: env name: env params: file_path: ./custom/path/to/.env ``` To continue loading `.env` or `.env.local`, specify them as additional secret stores: ``` secrets: - from: env name: env - from: env name: env params: file_path: ./custom/path/to/.env ``` --- # Keyring Secret Store The `keyring` store enables Spice to access secrets from the secure/credential store of the host operating system: * Linux: The secret-service and kernel keyutils. * macOS: The keychain. * Windows: The Credential Manager. The Keyring Store will read entries where the entry account or user is set to `spiced`. ## Example[​](#example "Direct link to Example") To set the `spiceai` API Key secret using macOS keychain, create a new keychain entry, and set the value with the API Key. `""` ![](/img/secrets-keychain-example.png) The `keyring` store is configured in the Spicepod manifest: ``` secrets: - from: keyring name: keyring ``` And the secret can be referenced in parameters: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips params: spiceai_api_key: ${keyring:spiceai_api_key} # ${secrets:spiceai_api_key} can also be used ``` --- # Kubernetes Secret Store The `kubernetes` store supports reading specific secrets using a selector with the secret's name. ## Example[​](#example "Direct link to Example") ``` secrets: - from: kubernetes:my_secret name: k8s ``` And the secret can be referenced in parameters: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips params: spiceai_api_key: ${k8s:spiceai_api_key} # ${secrets:spiceai_api_key} can also be used to fallback to other secret stores ``` Load secrets from multiple Kubernetes secrets by defining multiple Kubernetes secret stores with the appropriate selectors for the secrets to read: ``` secrets: - from: kubernetes:my_secret name: k8s - from: kubernetes:my_other_secret name: k8s_other ``` ## Kubernetes Secret Store Configuration[​](#kubernetes-secret-store-configuration "Direct link to Kubernetes Secret Store Configuration") Note: This method requires the Kubernetes service account, which is running the `spiced` pod, to have extended roles for secrets API access. Configure this service account with the necessary permissions to read secrets from the Kubernetes API. Example of Kubernetes role configuration for a custom service account: ``` kind: Role apiVersion: rbac.authorization.k8s.io/v1 metadata: name: spiced-account-role rules: - apiGroups: [''] resources: ['secrets'] verbs: ['get'] ``` --- # LLM Tools (Function Calling) A tool is a function or operation that can be called directly or by a [language model](/docs/v1.10/features/large-language-models) (LLMs). The Spice runtime has several tools available by default, giving LLMs access to various parts of the runtime. Tools can also be added or configured by the user by declaring them in the `tools` section of `spicepod.yaml`. For details about providing LLMs tool access, see [Language Model Tools](/docs/v1.10/features/large-language-models/tools). **Example** ``` tools: - name: arpanet from: websearch description: 'Search the web for information.' params: engine: perplexity perplexity_auth_token: ${ secrets:SPICE_PERPLEXITY_AUTH_TOKEN } ``` For details on tool specifications, see the [Tools Spicepod Reference](/docs/v1.10/reference/spicepod/tools). ### Available Tools[​](#available-tools "Direct link to Available Tools") | Name | Description | Default Group | | ----------------------------------------------------- | ----------------------------------------------------------------- | ------------- | | `list_datasets` | List all available datasets in the runtime. | `auto` | | `sql` | Execute SQL queries on the runtime. | `auto` | | `table_schema` | Get the schema of a specific SQL table. | `auto` | | `search` | Searches a configured dataset based on an input query. | `auto` | | `sample_distinct_columns` | Generate a synthetic sample of data with distinct values. | `auto` | | `random_sample` | Sample random rows from a table. | `auto` | | `top_n_sample` | Sample the top N rows from a table based on a specified ordering. | `auto` | | `memory:load` | Retrieve all stored memories from the last time period. | `memory` | | `memory:store` | Store information from LLM interaction(s) for future reference. | `memory` | | [`websearch`](/docs/v1.10/components/tools/websearch) | Search the web for information. | - | ### Tool Groups[​](#tool-groups "Direct link to Tool Groups") Tool groups are predefined sets of tools that can be provided to LLMs in a single tool name. For example, the `auto` tool group provides all default tools to the LLM (see above table). ``` models: - name: full-runtime from: openai:gpt-4o params: tools: auto # Use all default tools ``` #### Available tool groups[​](#available-tool-groups "Direct link to Available tool groups") | Name | Description | | ----------------------------------------- | -------------------------------------------------------------------------------------------- | | `auto` | All default tools (see above table) | | `memory` | Memory tools for storing and retrieving information across conversations. | | [`MCP`](/docs/v1.10/components/tools/mcp) | Tools provided from an MCP server. Can be run within Spice, or connected to over HTTP(s) SSE | --- # Model Context Protocol Tools Spice integrates with tools and services using the [Model Context Protocol](https://modelcontextprotocol.io/) (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Server-Sent Events (SSE) protocol. ## Overview[​](#overview "Direct link to Overview") MCP helps extend the capabilities of the Spice runtime by enabling integration with external tools and services. This includes: 1. Running stdio-based MCP servers internally. 2. Connecting to external MCP servers over SSE. ## Configuring MCP Tools[​](#configuring-mcp-tools "Direct link to Configuring MCP Tools") To configure MCP tools, define them in the `tools` section of your `spicepod.yaml` file. The `from` field specifies the transport mechanism, such as `mcp:npx` for stdio-based tools or an HTTP URL for SSE-based tools. ### Example: Adding an MCP Tool (Stdio)[​](#example-adding-an-mcp-tool-stdio "Direct link to Example: Adding an MCP Tool (Stdio)") The following example demonstrates how to configure an MCP tool using `npx` to run a Google Maps MCP server: ``` tools: - name: google_maps from: mcp:npx params: mcp_args: -y @modelcontextprotocol/server-google-maps ``` ### Example: Connecting to an External MCP Server (SSE)[​](#example-connecting-to-an-external-mcp-server-sse "Direct link to Example: Connecting to an External MCP Server (SSE)") This example shows how to connect to an external MCP server over SSE: ``` tools: - name: external_mcp_server from: mcp:http://example.com/v1/mcp/sse ``` ## Using MCP Tools with Models[​](#using-mcp-tools-with-models "Direct link to Using MCP Tools with Models") Once configured, MCP tools can be assigned to models via the `tools` parameter. For example: ``` models: - name: model_with_mcp from: openai:gpt-4o params: tools: google_maps ``` ## Spice as an MCP Server[​](#spice-as-an-mcp-server "Direct link to Spice as an MCP Server") Spice can also act as an MCP server, exposing its tools over SSE. This enables other Spice instances or external systems to connect and use the tools. ### Example: Connecting to Another Spice Instance via MCP[​](#example-connecting-to-another-spice-instance-via-mcp "Direct link to Example: Connecting to Another Spice Instance via MCP") ``` tools: - name: spice_instance from: mcp:http://localhost:8090/v1/mcp/sse ``` ## Configuration Options[​](#configuration-options "Direct link to Configuration Options") ### `from`[​](#from "Direct link to from") The `from` field specifies the transport mechanism for the MCP tool: * **SSE**: Use an HTTP URL ending with `/sse` (e.g., `http://localhost:8090/v1/mcp/sse`). * **Stdio**: Use commands like `mcp:npx` or `mcp:docker`. Additional arguments can be passed via `params.mcp_args`. ### `params`[​](#params "Direct link to params") The `params` field provides additional configuration for MCP tools. For stdio-based tools, use `mcp_args` to specify command-line arguments. ``` tools: - name: custom_tool from: mcp:npx params: mcp_args: -y @custom/tool ``` ### `env`[​](#env "Direct link to env") For stdio-based MCP tools, environment variables can be set using the `env` field. ``` tools: - name: tool_with_env from: mcp:docker env: API_KEY: your_api_key ``` ### `description`[​](#description "Direct link to description") The `description` field provides a textual description of the tool. This description is passed to any language model that uses the tool. ``` tools: - name: google_maps from: mcp:npx description: Provides geocoding and mapping capabilities. ``` For more details, see the \[MCP Tools Reference(../../reference/spicepod/tools). --- # Web Search Tool The Web Search Tool enables Spice models to search the web for information. The tool is available through the `websearch` tool, and backed by different search engines. ## Usage[​](#usage "Direct link to Usage") ``` tools: - name: the_internet from: websearch description: "Search the web for information." params: engine: perplexity perplexity_auth_token: ${ secrets:SPICE_PERPLEXITY_AUTH_TOKEN } ``` # Configuration ## `from`[​](#from "Direct link to from") The `from` field is used to specify the tool to use. For the Web Search Tool, use `websearch`. ## `name`[​](#name "Direct link to name") The `name` field is used to specify the name of the tool. This name is used: * To reference the tool in the model's `params.tools` field * To make HTTP requests to the tool via the [API](/docs/v1.10/api/HTTP/post), i.e. `v1/tools/{name}`. * Provided to any language model that uses the tool. ## `description`[​](#description "Direct link to description") The `description` field is used to provide a description of the tool. This description is provided to any language model that uses the tool. * To reference the tool in the model's `params.tools` field * To make HTTP requests to the tool via the \[API(../../api/HTTP/post), i.e. `v1/tools/{name}` * Provided to any language model that uses the tool The `params` field is used to ... The following parameters are supported: * `engine`: The search engine to use. Possible values: * `perplexity`: Use the Perplexity search engine. * `engine_*`: Each search engine has its own set of parameters. See the documentation for the specific search engine for more information. # Search Engines * [Perplexity](https://perplexity.com): Powered by Perplexity [Sonar](https://docs.perplexity.ai/). ## Perplexity[​](#perplexity "Direct link to Perplexity") To define a Perplexity search engine, use the following parameters: * `perplexity_auth_token` (required): The authentication token for the Perplexity API. Use the [secret replacement syntax](/docs/v1.10/components/secret-stores) to reference a secret, e.g. `${secrets:my_perplexity_auth_token}`. To get an authentication token, see Perplexity's [Getting Started](https://docs.perplexity.ai/guides/getting-started). * `perplexity_return_images` (default: false): Determines whether or not a request should return images. * `perplexity_return_related_questions` (default: false): Determines whether or not a request should return related questions. * `perplexity_search_domain_filter`: Given a list of domains, limit the citations used by the online model to URLs from the specified domains. Currently limited to only 3 domains for whitelisting and blacklisting. For blacklisting add a - to the beginning of the domain string. * Example: ``` perplexity_search_domain_filter: - spice.ai - docs.perplexity.ai ``` * `perplexity_search_recency_filter`: Returns search results within the specified time interval - does not apply to images. One of: `month`, `week`, `day`, `hour`. --- # Vector Engines > 🎓 Learn how it works with the [Amazon S3 Vectors with Spice](https://spiceai.org/blog/amazon-s3-vectors-with-spice) engineering blog post. Data sourced by Data Connectors, or views built atop them with vector embedding columns can be indexed and efficiently searched using a vector engine. A vector engine will store all vector embeddings associated with columns in a dataset/view, provide efficient vector search operations and avoid unnecessary recomputation of embeddings. A vector engine is configured by setting the `vectors` configuration. E.g. ``` datasets: - name: dataset_with_embeddings vectors: enabled: true ``` For the complete reference specification see [datasets](/docs/reference/spicepod/datasets). Supported Vector engines: | Name | Description | | --------------------------------------------------- | -------------- | | [`s3_vectors`](/docs/components/vectors/s3_vectors) | AWS S3 vectors | Limitations * A dataset or view must be accelerated (i.e. `datasets[].acceleration.enabled: true`, see [docs](/docs/reference/spicepod/datasets#accelerationenabled)) for a vector engine to be provided the appropriate data to ingest. ## Vector Engine Docs[​](#vector-engine-docs "Direct link to Vector Engine Docs") ## [📄️Amazon S3 Vectors](/docs/v1.10/components/vectors/s3_vectors) [Amazon S3 Vectors Engine Documentation](/docs/v1.10/components/vectors/s3_vectors) --- # Amazon S3 Vectors Engine > 🎓 Learn how it works with the [Amazon S3 Vectors with Spice](https://spiceai.org/blog/amazon-s3-vectors-with-spice) engineering blog post. Amazon S3 Vectors, announced in public preview at AWS Summit New York 2025, is a new S3 bucket type designed for storing and querying vector embeddings at scale. It supports billions of vectors with sub-second similarity queries, reducing costs by up to 90% compared to traditional vector databases by separating storage from compute. Spice AI integrates S3 Vectors as a vector index backend, managing embedding indexing, lifecycle, and queries for hybrid search experiences. To use Amazon S3 Vectors as a Vector Engine, specify `s3_vectors` as the `engine`, and configure the associated location and AWS credentials. ``` datasets: - from: spice.ai:dataset.with.embeddings name: my_dataset vectors: enabled: true engine: s3_vectors params: s3_vectors_bucket: my-s3-vector-bucket columns: - name: 'body' embeddings: - from: bedrock_titan embeddings: - name: bedrock_titan # ... Define an embedding model to use. ``` ## Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | Example Value | | ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- | | `s3_vectors_arn` | The S3 vectors index to use. Incompatible with `s3_vectors_bucket` and `s3_vectors_index`. | `arn:aws:s3vectors:us-east-1:123456789012:bucket/a-bucket/index/index-of-important-embeddings` | | `s3_vectors_aws_access_key_id` | Optional. The access key ID for the S3 vectors index. If not specified, credentials will be loaded from the environment. | - | | `s3_vectors_aws_region` | The AWS region for the S3 vectors index. | `us-east-1` | | `s3_vectors_aws_secret_access_key` | Optional. The secret access key for the S3 vectors index. If not specified, credentials will be loaded from the environment. | - | | `s3_vectors_aws_session_token` | Optional. Session token for the S3 vectors index. | - | | `s3_vectors_batch_write_rows` | Optional. Number of rows each record batch is chunked into when writing vectors, to control memory usage during writes. Default `100000`. | `100000` | | `s3_vectors_bucket` | The S3 vectors bucket to use. If `s3_vectors_index` is not specified, an index will be created based on the underlying embedding column. Incompatible with `s3_vectors_arn` | `a-bucket` | | `s3_vectors_index` | The name of the s3 vectors index to use or create. Incompatible with `s3_vectors_arn`. | `index-of-important-embeddings` | | `s3_vectors_distance_metric` | The distance metric to be used for similarity search. One of: `euclidean`, `cosine`. Default `cosine`. | `euclidean` | | `s3_vectors_index_poll_interval` | The interval to poll for index updates to avoid excessive API calls. Minimum 5 seconds. Default is to poll on every scan. | `5m` | | `client_timeout` | Timeout for S3 operations. Default: `30s`. | `30s`, `9 century`, `1m` | Limitations * `s3_vectors_index` and `s3_vectors_arn` specify a single index for the dataset and therefore should not be used with a dataset containing more than one embedding column. * S3 Vectors uses approximate nearest neighbor (ANN) algorithms for performance, providing probabilistically closest results. ## Overview[​](#overview "Direct link to Overview") Amazon S3 Vectors exposes two core operations: upsert vectors (assign a vector to a key with optional metadata) and vector similarity queries (find closest vectors by metrics like cosine or Euclidean distance). It stores vectors durably in S3 and executes queries on transient compute, avoiding always-on servers. This enables petabyte-scale storage at low cost, ideal for semantic search, recommendations, and RAG in AI applications. ## Configuration[​](#configuration "Direct link to Configuration") Annotate dataset schemas to specify columns for embedding and models. Spice supports local or hosted models like Amazon Titan Embeddings. When data ingests, Spice generates embeddings and stores them in the configured S3 Vectors index. Spice handles index creation, updates, and synchronization with data changes. Example with AWS Titan model: ``` datasets: - from: oracle:"CUSTOMER_REVIEWS" name: reviews vectors: enabled: true engine: s3_vectors params: s3_vectors_bucket: my-s3-vector-bucket columns: - name: body embeddings: from: bedrock_titan embeddings: - from: bedrock:amazon.titan-embed-text-v2:0 name: bedrock_titan params: aws_region: us-east-2 dimensions: '256' ``` ## Querying[​](#querying "Direct link to Querying") Perform vector searches via HTTP API or SQL table-valued function `vector_search(dataset, query)`. HTTP example: ``` curl -X POST http://localhost:8090/v1/search \ -H "Content-Type: application/json" \ -d '{ "datasets": ["reviews"], "text": "issues with same day shipping", "additional_columns": ["rating", "customer_id"], "where": "created_at >= now() - INTERVAL '7 days'", "limit": 2 }' ``` SQL example: ``` SELECT review_id, rating, customer_id, body, score FROM vector_search(reviews, 'issues with same day shipping') WHERE created_at >= to_unixtime(now() - INTERVAL '7 days') ORDER BY score DESC LIMIT 2; ``` Results include matching snippets, additional fields, primary keys, scores, and table names. ## Managing Embeddings[​](#managing-embeddings "Direct link to Managing Embeddings") Spice manages the vector lifecycle: embedding data on ingestion, upserting to S3 Vectors with primary keys as identifiers, and handling updates/deletions. Vectors can be pre-stored in S3 for scalability, avoiding memory limits in accelerators or slow just-in-time computations. ## Query Execution[​](#query-execution "Direct link to Query Execution") Spice pushes similarity computations to S3 Vectors, retrieving top matches as (id, score) pairs. These form a temporary table joinable with the dataset for full records, reducing processing to candidates only. ## Handling Filters[​](#handling-filters "Direct link to Handling Filters") Mark columns as filterable metadata to push filters into S3 Vectors queries: ``` columns: - name: created_at metadata: vectors: filterable ``` This ensures results respect constraints like time ranges during similarity search. ## Optimizations[​](#optimizations "Direct link to Optimizations") Store non-filterable columns as metadata to avoid joins: ``` columns: - name: rating metadata: vectors: non-filterable ``` Queries can then retrieve all needed data directly from the index, improving latency for read-heavy workloads. ## Advanced Features[​](#advanced-features "Direct link to Advanced Features") Spice supports hybrid search (vector + full-text via BM25, merged with RRF), multi-vector queries (weighting columns), and re-ranking (e.g., keyword first-pass then vector on candidates). Compose via SQL CTEs and joins. Hybrid RRF example: ``` WITH vector_results AS ( SELECT review_id, RANK() OVER (ORDER BY score DESC) AS vector_rank FROM vector_search(reviews, 'issues with same day shipping') ), text_results AS ( SELECT review_id, RANK() OVER (ORDER BY score DESC) AS text_rank FROM text_search(reviews, 'issues with same day shipping') ) SELECT COALESCE(v.review_id, t.review_id) AS review_id, (1.0 / (60 + COALESCE(v.vector_rank, 1000)) + 1.0 / (60 + COALESCE(t.text_rank, 1000))) AS fused_score FROM vector_results v FULL OUTER JOIN text_results t ON v.review_id = t.review_id ORDER BY fused_score DESC LIMIT 50; ``` Multi-column example (weighting title higher): ``` WITH body_results AS ( SELECT review_id, score AS body_score FROM vector_search(reviews, 'issues with same day shipping', col => 'body') ), title_results AS ( SELECT review_id, score AS title_score FROM vector_search(reviews, 'issues with same day shipping', col => 'title') ) SELECT COALESCE(body.review_id, title.review_id) AS review_id, COALESCE(body_score, 0) + 2.0 * COALESCE(title_score, 0) AS combined_score FROM body_results FULL OUTER JOIN title_results ON body_results.review_id = title_results.review_id ORDER BY combined_score DESC LIMIT 5; ``` ## Index Partitioning[​](#index-partitioning "Direct link to Index Partitioning") S3 Vectors indexes can be partitioned using an arbitrary logical expression. This allows Spice to compose many actual vector indexes as one logical vector index, enabling elastic scalability for vector storage. To partition your S3 vector indexes: ``` vectors: enabled: true engine: s3_vectors partition_by: - 'bucket(50, PULocationID)' ``` This example uses a `bucket` user-defined function (UDF) to hash the `PULocationID` column and split the associated vectors into one of 50 partitioned indexes. The runtime will use the `s3_vectors_index` parameter as a prefix and generate partition-specific names. Limitations * `partition_by` must have only 1 expression. * Expression must reference exactly one column from the dataset. * Expression must produce a scalar value * Expression cannot contain a subquery ## Authentication[​](#authentication "Direct link to Authentication") If AWS credentials are not explicitly provided in the configuration, the connector will automatically load credentials from the following sources in order. 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate S3 Vectors permissions (e.g., `s3vectors:QueryVectors`, `s3vectors:GetVectors`) to access the vectors. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Required IAM Permissions[​](#required-iam-permissions "Direct link to Required IAM Permissions") The IAM role or user needs the following minimum permissions to access S3 Vectors: ``` { "Version": "2012-10-17", "Statement": [ { "Sid": "AllowApplicationVectorAccess", "Effect": "Allow", "Action": [ "s3vectors:QueryVectors", "s3vectors:GetIndex", "s3vectors:PutVectors", "s3vectors:ListVectors", "s3vectors:GetVectors", "s3vectors:ListIndexes" ], "Resource": [ "arn:aws:s3vectors:aws-region:123456789012:bucket/amzn-s3-demo-vector-bucket/index/*", ] }, { "Sid": "AllowGetVectorBucket", "Effect": "Allow", "Action": "s3vectors:GetVectorBucket", "Resource": "arn:aws:s3vectors:aws-region:123456789012:bucket/*" }, { "Sid": "AllowCreateVectorBucket", "Effect": "Allow", "Action": "s3vectors:CreateVectorBucket", "Resource": "arn:aws:s3vectors:aws-region:123456789012:bucket/*" }, { "Sid": "AllowCreateVectorBucketIndex", "Effect": "Allow", "Action": "s3vectors:CreateIndex", "Resource": "arn:aws:s3vectors:aws-region:123456789012:bucket/amzn-s3-demo-vector-bucket/index/*" } ] } ``` ### Permission Details[​](#permission-details "Direct link to Permission Details") | Permission | Purpose | | ------------------------------ | --------------------------------------------------------------------------------------- | | `s3vectors:GetIndex` | Required. Used to verify if the index already exists or needs to be created. | | `s3vectors:GetVectorBucket` | Required. Used to verify if the vector bucket already exists or needs to be created. | | `s3vectors:GetVectors` | Required. Used to read vector data when the `*_embeddings` column is projected. | | `s3vectors:ListIndexes` | Required when using index partitioning, to enumerate the physical indexes. | | `s3vectors:ListVectors` | Required. Used to populate the `*_embeddings` column on vector tables in Spice. | | `s3vectors:PutVectors` | Required. Used to populate the vector index with Spice-computed embeddings. | | `s3vectors:QueryVectors` | Required. Used to query for vectors using the `vector_search` table function. | | `s3vectors:CreateIndex` | Optional. Spice can automatically create indexes if this permission is given. | | `s3vectors:CreateVectorBucket` | Optional. Spice can automatically create the vector bucket if this permission is given. | ### `metrics`[​](#metrics "Direct link to metrics") Spice supports the following [S3 Vector engine metrics](/docs/v1.10/features/observability/component_metrics): | Metric Name | Type | Description | | ------------------------------------------------- | --------- | ---------------------------------------------------------------------------- | | `s3_vectors_create_index_errors` | counter | Number of errors returned from create\_index operation. | | `s3_vectors_create_index_latency` | histogram | Total duration of create\_index operation, in milliseconds. | | `s3_vectors_create_index_requests` | counter | Number of requests to create\_index operation. | | `s3_vectors_create_vector_bucket_errors` | counter | Number of errors returned from create\_vector\_bucket operation. | | `s3_vectors_create_vector_bucket_latency` | histogram | Total duration of create\_vector\_bucket operation, in milliseconds. | | `s3_vectors_create_vector_bucket_requests` | counter | Number of requests to create\_vector\_bucket operation. | | `s3_vectors_delete_index_errors` | counter | Number of errors returned from delete\_index operation. | | `s3_vectors_delete_index_latency` | histogram | Total duration of delete\_index operation, in milliseconds. | | `s3_vectors_delete_index_requests` | counter | Number of requests to delete\_index operation. | | `s3_vectors_delete_vector_bucket_errors` | counter | Number of errors returned from delete\_vector\_bucket operation. | | `s3_vectors_delete_vector_bucket_latency` | histogram | Total duration of delete\_vector\_bucket operation, in milliseconds. | | `s3_vectors_delete_vector_bucket_policy_errors` | counter | Number of errors returned from delete\_vector\_bucket\_policy operation. | | `s3_vectors_delete_vector_bucket_policy_latency` | histogram | Total duration of delete\_vector\_bucket\_policy operation, in milliseconds. | | `s3_vectors_delete_vector_bucket_policy_requests` | counter | Number of requests to delete\_vector\_bucket\_policy operation. | | `s3_vectors_delete_vector_bucket_requests` | counter | Number of requests to delete\_vector\_bucket operation. | | `s3_vectors_delete_vectors_errors` | counter | Number of errors returned from delete\_vectors operation. | | `s3_vectors_delete_vectors_latency` | histogram | Total duration of delete\_vectors operation, in milliseconds. | | `s3_vectors_delete_vectors_requests` | counter | Number of requests to delete\_vectors operation. | | `s3_vectors_get_index_errors` | counter | Number of errors returned from get\_index operation. | | `s3_vectors_get_index_latency` | histogram | Total duration of get\_index operation, in milliseconds. | | `s3_vectors_get_index_requests` | counter | Number of requests to get\_index operation. | | `s3_vectors_get_vector_bucket_errors` | counter | Number of errors returned from get\_vector\_bucket operation. | | `s3_vectors_get_vector_bucket_latency` | histogram | Total duration of get\_vector\_bucket operation, in milliseconds. | | `s3_vectors_get_vector_bucket_policy_errors` | counter | Number of errors returned from get\_vector\_bucket\_policy operation. | | `s3_vectors_get_vector_bucket_policy_latency` | histogram | Total duration of get\_vector\_bucket\_policy operation, in milliseconds. | | `s3_vectors_get_vector_bucket_policy_requests` | counter | Number of requests to get\_vector\_bucket\_policy operation. | | `s3_vectors_get_vector_bucket_requests` | counter | Number of requests to get\_vector\_bucket operation. | | `s3_vectors_get_vectors_errors` | counter | Number of errors returned from get\_vectors operation. | | `s3_vectors_get_vectors_latency` | histogram | Total duration of get\_vectors operation, in milliseconds. | | `s3_vectors_get_vectors_requests` | counter | Number of requests to get\_vectors operation. | | `s3_vectors_list_indexes_errors` | counter | Number of errors returned from list\_indexes operation. | | `s3_vectors_list_indexes_latency` | histogram | Total duration of list\_indexes operation, in milliseconds. | | `s3_vectors_list_indexes_requests` | counter | Number of requests to list\_indexes operation. | | `s3_vectors_list_vector_buckets_errors` | counter | Number of errors returned from list\_vector\_buckets operation. | | `s3_vectors_list_vector_buckets_latency` | histogram | Total duration of list\_vector\_buckets operation, in milliseconds. | | `s3_vectors_list_vector_buckets_requests` | counter | Number of requests to list\_vector\_buckets operation. | | `s3_vectors_list_vectors_errors` | counter | Number of errors returned from list\_vectors operation. | | `s3_vectors_list_vectors_latency` | histogram | Total duration of list\_vectors operation, in milliseconds. | | `s3_vectors_list_vectors_requests` | counter | Number of requests to list\_vectors operation. | | `s3_vectors_put_vector_bucket_policy_errors` | counter | Number of errors returned from put\_vector\_bucket\_policy operation. | | `s3_vectors_put_vector_bucket_policy_latency` | histogram | Total duration of put\_vector\_bucket\_policy operation, in milliseconds. | | `s3_vectors_put_vector_bucket_policy_requests` | counter | Number of requests to put\_vector\_bucket\_policy operation. | | `s3_vectors_put_vectors_errors` | counter | Number of errors returned from put\_vectors operation. | | `s3_vectors_put_vectors_latency` | histogram | Total duration of put\_vectors operation, in milliseconds. | | `s3_vectors_put_vectors_requests` | counter | Number of requests to put\_vectors operation. | | `s3_vectors_query_vectors_errors` | counter | Number of errors returned from query\_vectors operation. | | `s3_vectors_query_vectors_latency` | histogram | Total duration of query\_vectors operation, in milliseconds. | | `s3_vectors_query_vectors_requests` | counter | Number of requests to query\_vectors operation. | ## Cookbook[​](#cookbook "Direct link to Cookbook") * A cookbook recipe to configure a dataset with an S3 vectors engine in Spice. [S3 Vectors engine](https://github.com/spiceai/cookbook/tree/trunk/vectors/s3#readme) ## References[​](#references "Direct link to References") * [Spice.ai announcement](https://spiceai.org/blog/amazon-s3-vectors-with-spice) * [Amazon S3 Vectors official page](https://aws.amazon.com/s3/vectors/) --- # Views Views in Spice are virtual tables defined by SQL queries. They help simplify complex queries and promote reuse across different applications by encapsulating query logic in a single, reusable entity. ## Defining a View[​](#defining-a-view "Direct link to Defining a View") To define a view in the `spicepod.yaml` configuration file, specify the `views` section. Each view definition must include a `name` and a `sql` field. ### Example[​](#example "Direct link to Example") The following example demonstrates how to define a view named `rankings` that lists the top five products based on the total count of orders: ``` views: - name: rankings sql: | WITH a AS ( SELECT products.id, SUM(count) AS count FROM orders INNER JOIN products ON orders.product_id = products.id GROUP BY products.id ) SELECT name, count FROM products LEFT JOIN a ON products.id = a.id ORDER BY count DESC LIMIT 5 ``` ### Fields[​](#fields "Direct link to Fields") * `name`: The view's identifier, used for referencing in queries. * `sql`: The SQL query defining the view, supporting joins, subqueries, and aggregations. * `acceleration`: Views can be [locally accelerated](/docs/v1.10/features/data-acceleration). ## Limitations and Considerations[​](#limitations-and-considerations "Direct link to Limitations and Considerations") * Views are read-only; insert, update, and delete operations are not supported. * Performance depends on SQL complexity and underlying data. * Ensure queries are optimized to prevent slow execution. --- # Workers Overview Workers in the Spice runtime represent configurable units of compute that help coordinate and manage interactions between models and tools. Each worker is defined as a component in the `spicepod.yaml` file, specifying its behavior and interaction logic. ## Configuration[​](#configuration "Direct link to Configuration") Workers are configured in the `workers` section of the `spicepod.yaml` file. Each worker definition includes a name, description, and a list of models or tools it encapsulates. **Example `spicepod.yaml` configuration:** ``` workers: - name: round-robin description: | Distributes requests between 'llama3_2' and 'gpt4_1' models in a round-robin fashion. load_balance: routing: - from: llama3_2 - from: gpt4_1 - name: fallback description: | Attempts 'gpt4_1' first, then 'llama3_2', then 'anth_haiku' if previous models fail. load_balance: routing: - from: llama3_2 order: 2 - from: gpt4_1 order: 1 - from: anth_haiku order: 3 - name: weighted description: | Routes 80% of traffic to 'llama3_2'. load_balance: routing: - from: llama3_2 weight: 4 - from: gpt4_1 weight: 1 ``` ## Use-Cases[​](#use-cases "Direct link to Use-Cases") Workers currently help implement: * Model fallback and error handling * Load balancing across multiple models ## Usage[​](#usage "Direct link to Usage") Workers can be invoked using the same API endpoints as individual models. For example, to call a worker named `fallback` using the OpenAI-compatible HTTP API: ``` curl http://localhost:8090/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "fallback", "messages": [{ "role": "user", "content": "Tell me a joke"}] }' ``` ## Roadmap[​](#roadmap "Direct link to Roadmap") The vision for workers includes support for dynamic serverless compute, enabling execution of user-defined functions within the Spice runtime. This direction aims to help developers define custom logic and orchestration patterns directly in the worker configuration, supporting more advanced workflows and automation. Further details and implementation timelines will be provided in future updates. For ongoing progress, refer to the project repository and documentation. ## Further Reading[​](#further-reading "Direct link to Further Reading") For a complete specification of worker configuration, routing rules, and available options, refer to the [Spicepod Workers Reference](/docs/reference/spicepod/workers). --- # Deployment Learn how to deploy Spice.ai in your environment. ## Deployment Architectures[​](#deployment-architectures "Direct link to Deployment Architectures") * [Overview](/docs/v1.10/deployment/architectures) * [Sidecar Deployment](/docs/v1.10/deployment/architectures/sidecar) * [Microservice Deployment (Single or Multiple Replicas)](/docs/v1.10/deployment/architectures/microservice) * [Tiered Deployment](/docs/v1.10/deployment/architectures/tiered) * [Cloud-Hosted in the Spice Cloud Platform](/docs/v1.10/deployment/architectures/hosted) * [Sharded Deployment](/docs/v1.10/deployment/architectures/sharded) * [Cluster Deployment (Spice.ai Enterprise)](/docs/v1.10/deployment/architectures/cluster) ## Deployment Guides[​](#deployment-guides "Direct link to Deployment Guides") * [Kubernetes (Helm)](/docs/v1.10/deployment/kubernetes) * [Docker](/docs/v1.10/deployment/docker) * [Spice Cloud](/docs/v1.10/deployment/cloud) * [AWS](/docs/v1.10/deployment/aws) --- # Deployment Architectures ![Spice ai OSS as a data and AI compute engine over disaggregated storage](https://github.com/user-attachments/assets/da3c0e90-4c48-48ca-b4bd-72eda816cfec) * [Sidecar Deployment](/docs/deployment/architectures/sidecar) * [Microservice Deployment (Single or Multiple Replicas)](/docs/deployment/architectures/microservice) * [Tiered Deployment](/docs/deployment/architectures/tiered) * [Cloud-Hosted in the Spice Cloud Platform](/docs/deployment/architectures/hosted) * [Sharded Deployment](/docs/deployment/architectures/sharded) * [Cluster Deployment (Spice.ai Enterprise)](/docs/deployment/architectures/cluster) --- # Cluster-Based Deployment (Spice.ai Enterprise) A full cluster-based deployment leveraging **Spice.ai Enterprise**, which includes advanced services and integrations for Kubernetes. This method is ideal for organizations requiring large-scale or complex deployments, including specialized clustering capabilities. ![cluster](https://github.com/user-attachments/assets/643e0a5c-6745-40c0-8695-0955c795179b) **Benefits** * Provides **enterprise-grade features**: advanced security, monitoring, and support. * Simplifies **managing multiple nodes** for high availability and large workloads. * Offers **direct integration** with Spice Cloud or on-prem Kubernetes clusters. **Considerations** * **Requires a commercial license** or subscription to Spice Enterprise. * More **complex initial setup**, typically involving specialized DevOps expertise. **Use This Approach When** * You operate at **significant scale** or have stringent availability requirements. * You need **enterprise-level support** and advanced monitoring, security, or compliance features. * Your team can manage a **robust Kubernetes environment** or you plan to integrate with the Spice Cloud at scale. **Example Use Case** A large financial services firm requiring a highly available, secure environment. They run Spice.ai across multiple clusters using Spice Enterprise for advanced monitoring, role-based access control, and dedicated support. --- # Cloud Hosted The Spice Runtime is deployed on a fully managed service within the Spice Cloud Platform, minimizing the operational burden of managing clusters, upgrades, and infrastructure. ![hosted](https://github.com/user-attachments/assets/a985527b-3481-40f4-a689-f784c893b314) **Benefits** * Reduced overhead for deployment, scaling, and maintenance. * Access to specialized hosting features and quick setup. * Helps reduce operational complexity and cost. **Considerations** * Reliance on external hosting and associated terms or limits. * Potential compliance or data residency considerations for certain industries. * May introduce latency depending on the cloud provider's infrastructure. **Use This Approach When** * Limited DevOps resources are available, or focus on application logic over infrastructure is preferred. * A fully managed environment with minimal setup time is desired. * A single, managed solution is prioritized over running own clusters. * Minimizing operational complexity and cost is the goal. **Example Use Case** A startup or team with limited DevOps support that needs a reliable, managed environment. Quick deployment and minimal in-house infrastructure responsibilities are priorities. --- # Microservice Deployment (Single or Multiple Replicas) The Spice Runtime operates as an independent microservice. Multiple replicas may be deployed behind a load balancer to achieve high availability and handle spikes in demand. ![microservice](https://github.com/user-attachments/assets/b46f050b-e500-4d53-b354-24f0ab30cad3) **Benefits** * Loose coupling between the application and the Spice Runtime. * Independent scaling and upgrades. * Can serve multiple applications or services within an organization. * Helps achieve high availability and redundancy. **Considerations** * Additional network hop introduces latency compared to sidecar. * More complex infrastructure, requiring service discovery and load balancing. * Potentially higher cost due to additional infrastructure components. **Use This Approach When** * A loosely coupled architecture and the ability to independently scale the AI service are desired. * Multiple services or teams need to share the same AI engine. * Heavy or varying traffic is anticipated, requiring independent scaling of the Spice Runtime. * Resiliency and redundancy are prioritized over simplicity. **Example Use Case** A large organization where multiple services (recommendations, analytics, etc.) need to share AI insights. A centralized Spice Runtime microservice cluster helps separate teams consume AI outputs without duplicating efforts. --- # Sharded The Spice Runtime instances can be sharded based on specific criteria, such as by customer, state, or other logical partitions. Each shard operates independently, with a 1:N Application to Spice instances ratio. ![sharded](https://github.com/user-attachments/assets/5730d108-6d22-4ea4-8c14-8e87ad6d0079) **Benefits** * Helps distribute load across multiple instances, improving performance and scalability. * Isolates failures to specific shards, enhancing resiliency. * Allows tailored configurations and optimizations for different shards. **Considerations** * More complex deployment and management due to multiple instances. * Requires effective sharding strategy to balance load and avoid hotspots. * Potentially higher cost due to multiple instances. **Use This Approach When** * Distributing load across multiple instances for better performance is needed. * Isolating failures to specific shards to improve resiliency is desired. * The application can benefit from tailored configurations for different logical partitions. * The complexity of managing multiple instances can be handled. **Example Use Case** A multi-tenant application where each customer has a dedicated Spice Runtime instance. This helps ensure that heavy usage by one customer does not impact others, and allows for customer-specific optimizations. **Sharding vs. partitioning** Sharding splits load across **multiple Spice instances**, each backing a logical slice of the system (a customer, a region, a workload). Each shard runs an independent runtime with its own datasets, accelerations, and resources. Within a single Spice instance, [acceleration partitioning](/docs/v1.10/features/data-acceleration/partitioning) splits a single dataset into multiple physical units (files, tables, or in-memory tables) so that filtered queries only read the relevant subset. The two are complementary: a shard can also use partitioning internally to keep individual datasets pruneable. | Concern | Sharded deployment | Acceleration partitioning | | ----------------- | -------------------------------------------------------- | -------------------------------------------------------------- | | Splits across… | Multiple runtimes/processes | One acceleration on one runtime | | Routing | Application picks which Spice instance to query | Spice prunes partitions automatically based on filter pushdown | | Failure isolation | Per-shard | None — single runtime | | Use when | Tenants/regions have very different load or data volumes | A single dataset is too large to scan whole on every query | --- # Sidecar Deployment Run the Spice Runtime in a separate container or process on the same machine as the main application. For example, in Kubernetes as a [Sidecar Container](https://kubernetes.io/docs/concepts/workloads/pods/sidecar-containers/). This approach minimizes communication overhead as requests to the Spice Runtime are transported over local loopback. ![sidecar](https://github.com/user-attachments/assets/716f7c23-1939-4947-85f5-b0ee2bbd63fc) **Benefits** * Low-latency communication between the application and the Spice Runtime. * Simplified lifecycle management (same pod). * Isolated environment without needing a separate microservice. * Helps ensure resiliency and redundancy by replicating data across sidecars. **Considerations** * Each application pod includes a copy of the Spice Runtime, increasing resource usage. * Updating the Spice Runtime independently requires updating each pod. * Accelerated data is replicated to each sidecar, adding resiliency and redundancy but increasing resource usage and requests to data sources. * May increase overall cost due to resource duplication. **Use This Approach When** * Fast, low-latency interactions between the application and the Spice Runtime are needed (e.g., real-time decision-making). * Scaling needs are small or moderate, making duplication of the Spice Runtime in each pod acceptable. * Keeping the architecture simple without additional services or load balancers is preferred. * Performance and latency are prioritized over cost and complexity. **Example Use Case** A real-time trading bot or a data-intensive application that relies on immediate feedback, where minimal latency is critical. Both containers in the same pod ensure very fast data exchange. --- # Tiered Deployment A hybrid approach combining sidecar deployments for performance-critical tasks and a shared microservice for batch processing or less time-sensitive workloads. ![tiered](https://github.com/user-attachments/assets/e602bad4-bd0d-4069-bc91-5b5678a10710) **Benefits** * Real-time responsiveness where needed (sidecar). * Centralized microservice handles broader or shared tasks. * Balances resource usage by limiting sidecar instances to high-priority operations. * Helps balance performance and latency with cost and complexity. **Considerations** * More complex deployment structure, mixing two patterns. * Must ensure consistent versioning between sidecar and microservice instances. * Potentially higher operational complexity and cost. **Use This Approach When** * Certain application components require ultra-low-latency responses, while others do not. * Centralized AI or analytics is needed, but localized real-time decision-making is also required. * The system can handle the operational complexity of running multiple deployment patterns. * Balancing performance and latency with cost and complexity is the goal. **Example Use Case** A logistics application that calculates routing decisions in real time (sidecar) while a microservice component processes aggregated data for periodic analysis or re-training models. --- # AWS Deployment Options Spice.ai provides multiple deployment options on Amazon Web Services (AWS), enabling data and AI applications to run on AWS's elastic infrastructure. Whether using virtual machines, container orchestration, or managed services, Spice deploys to meet requirements for performance, scalability, and cost efficiency. For a complete list of AWS-compatible data connectors, AI models, vector stores, and secret management, see [AWS Integrations](/docs/v1.10/deployment/aws/integrations). ## Benefits of Deploying on AWS[​](#benefits-of-deploying-on-aws "Direct link to Benefits of Deploying on AWS") * **Scalability**: Easily scale your Spice.ai applications with AWS's elastic infrastructure. * **Global Reach**: Deploy across AWS's [worldwide regions](https://aws.amazon.com/about-aws/global-infrastructure/) for low-latency access. * **Integration**: Connect with other AWS services like [Amazon S3](https://aws.amazon.com/s3/), [Amazon RDS](https://aws.amazon.com/rds/), and [AWS Secrets Manager](https://aws.amazon.com/secrets-manager/). * **Cost Control**: Optimize expenses with various [instance types](https://aws.amazon.com/ec2/instance-types/) and [pricing models](https://aws.amazon.com/pricing/). * **Security and Compliance**: Deploy Spice.ai within your AWS security perimeter using features like [VPC](https://aws.amazon.com/vpc/) isolation, [security groups](https://docs.aws.amazon.com/vpc/latest/userguide/vpc-security-groups.html), [IAM roles](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles.html) to meet organizational compliance requirements. ## Deployment Options[​](#deployment-options "Direct link to Deployment Options") ### Amazon EKS (Elastic Kubernetes Service)[​](#amazon-eks-elastic-kubernetes-service "Direct link to Amazon EKS (Elastic Kubernetes Service)") Leverage [Kubernetes](https://kubernetes.io/) orchestration with [Amazon EKS](https://aws.amazon.com/eks/) for containerized Spice.ai deployments. 1. **Create an EKS Cluster**: * Use the [AWS Management Console](https://console.aws.amazon.com/eks/), [AWS CLI](https://docs.aws.amazon.com/cli/latest/reference/eks/), or [eksctl](https://eksctl.io/) to create your cluster * Configure [node groups](https://docs.aws.amazon.com/eks/latest/userguide/managed-node-groups.html) according to your workload requirements * (Optional) Use [EKS Fargate profiles](https://docs.aws.amazon.com/eks/latest/userguide/fargate.html) for serverless container deployment 2. **Deploy Spice.ai on EKS**: * Apply Spice.ai Kubernetes manifests via [Helm chart](https://spiceai.org/docs/deployment/kubernetes) * Configure persistent storage using [Amazon EBS](https://aws.amazon.com/ebs/) or [Amazon EFS](https://aws.amazon.com/efs/) * Set up ingress with the [AWS Network Load Balancer (NLB)](https://docs.aws.amazon.com/eks/latest/userguide/network-load-balancing.html) * (Optional) Automate cluster and resource provisioning with Infrastructure as Code (IaC) tools such as [AWS CloudFormation](https://aws.amazon.com/cloudformation/) or [Terraform](https://www.terraform.io/) for consistent, repeatable deployments For comprehensive instructions and advanced configuration options, refer to the [Amazon EKS User Guide](https://docs.aws.amazon.com/eks/latest/userguide/what-is-eks.html), [EKS Best Practices Guide](https://aws.github.io/aws-eks-best-practices/), and [Spice.ai Kubernetes Deployment Guide](https://spiceai.org/docs/deployment/kubernetes). ### EC2 / AWS CloudFormation[​](#ec2--aws-cloudformation "Direct link to EC2 / AWS CloudFormation") Deploy Spice.ai directly on [Amazon EC2](https://aws.amazon.com/ec2/) instances for maximum control over the environment. 1. **Manual EC2 Deployment**: * Launch an [EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/EC2_GetStarted.html) with your preferred Linux distribution * Install [Docker](https://docs.docker.com/engine/install/) * Run [Spice.ai as a Docker Container](https://spiceai.org/docs/deployment/docker#running-spiceai-as-a-docker-container) on your EC2 instance * (Optional) Use Infrastructure as Code (IaC) tools like [AWS CloudFormation](https://aws.amazon.com/cloudformation/) or [Terraform](https://www.terraform.io/) to automate the provisioning, configuration, and management of EC2 resources for repeatable and consistent deployments 2. **Automated EC2 Deployment with CloudFormation**: * Define your infrastructure in a [CloudFormation template](https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGuide/template-guide.html), including EC2 instances (using a [Linux AMI](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/AMIs.html)), security groups, IAM roles, VPC, and subnets * Use EC2 [`UserData`](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html) to automate Docker installation, pull the [Spice.ai Docker image](https://hub.docker.com/r/spiceai/spiceai), retrieve configuration or secrets from [AWS Parameter Store](https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-parameter-store.html) or [Secrets Manager](https://aws.amazon.com/secrets-manager/), and run the container with required environment variables * (Optional) Add [parameters](https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGuide/parameters-section-structure.html) to your template for VPC ID, Subnet ID, [KeyPair](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-key-pairs.html), instance type, and secret names to enable flexible deployments * (Optional) Store sensitive data such as API keys in [Parameter Store](https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-parameter-store.html) or [Secrets Manager](https://aws.amazon.com/secrets-manager/) and reference them securely in `UserData` * (Optional) Deploy and manage your CloudFormation stack using the [AWS Console](https://console.aws.amazon.com/cloudformation/), [CLI](https://docs.aws.amazon.com/cli/latest/reference/cloudformation/), or [CI/CD pipelines](https://aws.amazon.com/devops/continuous-delivery/) for repeatable, version-controlled infrastructure For detailed guidance and best practices, refer to the [AWS CloudFormation User Guide](https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGuide/Welcome.html), [EC2 User Guide for Linux Instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/), and [AWS Systems Manager Parameter Store Documentation](https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-parameter-store.html). ### Amazon ECS (Elastic Container Service)[​](#amazon-ecs-elastic-container-service "Direct link to Amazon ECS (Elastic Container Service)") Deploy Spice.ai as containerized tasks on [Amazon ECS](https://aws.amazon.com/ecs/) for easy container management and flexible scaling. 1. **Create an ECS Cluster**: * Choose a launch type: [EC2](https://docs.aws.amazon.com/AmazonECS/latest/developerguide/launch_types.html#ec2-launch-type) (manage your own EC2 instances) or [Fargate](https://docs.aws.amazon.com/AmazonECS/latest/developerguide/launch_types.html#fargate-launch-type) (serverless). * Create the ECS cluster using the [AWS Console](https://console.aws.amazon.com/ecs/), [CLI](https://docs.aws.amazon.com/cli/latest/reference/ecs/), or Infrastructure as Code ([CloudFormation](https://aws.amazon.com/cloudformation/), [Terraform](https://www.terraform.io/)). 2. **Define a Task Definition**: * Specify the [Spice.ai Docker image](https://hub.docker.com/r/spiceai/spiceai), resource needs, networking, environment variables, and storage in a [Task Definition](https://docs.aws.amazon.com/AmazonECS/latest/developerguide/task_definitions.html). * (Optional) Use [AWS Secrets Manager](https://aws.amazon.com/secrets-manager/) or [Parameter Store](https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-parameter-store.html) to inject secrets securely. * Enable logging with [Amazon CloudWatch](https://aws.amazon.com/cloudwatch/). 3. **Deploy Spice.ai on ECS**: * Create an [ECS Service](https://docs.aws.amazon.com/AmazonECS/latest/developerguide/ecs_services.html) to run and manage Spice.ai tasks. * Set up load balancing with [NLB](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/introduction.html). * (Optional) Configure [auto-scaling](https://docs.aws.amazon.com/AmazonECS/latest/developerguide/service-auto-scaling.html) based on resource usage or CloudWatch metrics. * (Optional) Use [CI/CD pipelines](https://aws.amazon.com/devops/continuous-delivery/) for automated updates. Manage infrastructure with CloudFormation, Terraform, or the AWS CLI. For more details, see the [Amazon ECS Developer Guide](https://docs.aws.amazon.com/ecs/latest/developerguide/Welcome.html) and [Spice.ai Docker Deployment Guide](https://spiceai.org/docs/deployment/docker). ## Authentication[​](#authentication "Direct link to Authentication") Most AWS services that Spice connects to have explicit parameters for configuring authentication (usually by setting an `access_key_id` and `secret_access_key`). If explicit credentials are not provided, Spice follows the standard AWS SDK behavior for loading credentials from the environment based on the following sources in order: 1. **Environment Variables**: * `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` * `AWS_SESSION_TOKEN` (if using temporary credentials) 2. **Shared AWS Config/Credentials Files**: * Config file: `~/.aws/config` (Linux/Mac) or `%UserProfile%\.aws\config` (Windows) * Credentials file: `~/.aws/credentials` (Linux/Mac) or `%UserProfile%\.aws\credentials` (Windows) * The `AWS_PROFILE` environment variable can be used to specify a named profile, otherwise the `[default]` profile is used. * Supports both static credentials and SSO sessions * Example credentials file: ``` # Static credentials [default] aws_access_key_id = YOUR_ACCESS_KEY aws_secret_access_key = YOUR_SECRET_KEY # SSO profile [profile sso-profile] sso_start_url = https://my-sso-portal.awsapps.com/start sso_region = us-west-2 sso_account_id = 123456789012 sso_role_name = MyRole region = us-west-2 ``` tip To set up SSO authentication: 1. Run `aws configure sso` to configure a new SSO profile 2. Use the profile by setting `AWS_PROFILE=sso-profile` 3. Run `aws sso login --profile sso-profile` to start a new SSO session 3. **AWS STS Web Identity Token Credentials**: * Used primarily with OpenID Connect (OIDC) and OAuth * Common in Kubernetes environments using IAM roles for service accounts (IRSA) 4. **ECS Container Credentials**: * Used when running in Amazon ECS containers * Automatically uses the task's IAM role * Retrieved from the ECS credential provider endpoint * Relies on the environment variable `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` or `AWS_CONTAINER_CREDENTIALS_FULL_URI` which are automatically injected by ECS. 5. **AWS EC2 Instance Metadata Service (IMDSv2)**: * Used when running on EC2 instances. * Automatically uses the instance's IAM role. * Retrieved securely using [IMDSv2](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/configuring-instance-metadata-service.html). The connector will try each source in order until valid credentials are found. If no valid credentials are found, an authentication error will be returned. IAM Permissions Regardless of the credential source, the IAM role or user must have appropriate permissions (e.g., `s3:ListBucket`, `s3:GetObject`) to access the service. If the Spicepod connects to multiple different AWS services, the permissions should cover all of them. ## Resources[​](#resources "Direct link to Resources") ### Documentation[​](#documentation "Direct link to Documentation") * [AWS Integrations](/docs/v1.10/deployment/aws/integrations) - Complete list of AWS data connectors, AI models, vector stores, and secrets * [AWS Secrets Manager Secret Store](/docs/v1.10/components/secret-stores/aws-secrets-manager) ### AWS Blog Posts[​](#aws-blog-posts "Direct link to AWS Blog Posts") * [Architecting High-Performance AI-Driven Data Applications with Spice.ai and AWS](https://aws.amazon.com/blogs/storage/architecting-high-performance-ai-driven-data-applications-with-spice-ai-and-aws/) - AWS Storage Blog ### Spice.ai Blog Posts[​](#spiceai-blog-posts "Direct link to Spice.ai Blog Posts") * [Amazon S3 Vectors](https://spice.ai/blog/amazon-s3-vectors) - Overview of S3 Vectors integration * [Getting Started with Amazon S3 Vectors and Spice](https://spice.ai/blog/getting-started-with-amazon-s3-vectors-and-spice) - Step-by-step tutorial ### Videos[​](#videos "Direct link to Videos") * [Getting started with Amazon S3 Vectors and Spice](https://www.youtube.com/watch?v=KuWI0yDOnIU) - YouTube walkthrough ### Marketplace[​](#marketplace "Direct link to Marketplace") * [Spice.ai on AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-jmf6jskjvnq7i) - Deploy Spice.ai from AWS Marketplace --- # AWS Integrations ![Spice.ai and AWS](/assets/images/aws-spice-a361b79fcd229ce942248cd4a48f2cff.png) Spice.ai provides deep integrations with Amazon Web Services (AWS), enabling data federation, AI inference, vector search, and secure secret management across the AWS ecosystem. This page consolidates all AWS-compatible components and provides quick access to configuration guides. ## Data Connectors[​](#data-connectors "Direct link to Data Connectors") Data connectors federate SQL queries across AWS data sources without data movement. | Connector | Description | Documentation | | ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ | | **Amazon S3** | Query Parquet, CSV, and JSON files stored in S3 buckets. Supports private buckets with IAM authentication and S3-compatible storage like MinIO. | [S3 Data Connector](/docs/v1.10/components/data-connectors/s3) | | **Amazon S3 Tables** | Query Iceberg tables in [Amazon S3 Tables](https://aws.amazon.com/s3/features/tables/) using the Glue connector with S3 Tables catalog format. | [Glue Data Connector](/docs/v1.10/components/data-connectors/glue) | | **Amazon DynamoDB** | Federated SQL queries on DynamoDB tables with automatic schema inference. | [DynamoDB Data Connector](/docs/v1.10/components/data-connectors/dynamodb) | | **Amazon DynamoDB Streams** | Real-time CDC streaming of table changes via [DynamoDB Streams](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.html). | [DynamoDB Data Connector](/docs/v1.10/components/data-connectors/dynamodb) | | **Amazon Redshift** | Connect to Redshift clusters using the PostgreSQL-compatible connector. | [Redshift Data Connector](/docs/v1.10/components/data-connectors/redshift) | | **Amazon Aurora PostgreSQL** | Connect to Aurora PostgreSQL clusters using the PostgreSQL connector. | [PostgreSQL Data Connector](/docs/v1.10/components/data-connectors/postgres) | | **Amazon Aurora MySQL** | Connect to Aurora MySQL clusters using the MySQL connector. | [MySQL Data Connector](/docs/v1.10/components/data-connectors/mysql) | | **Amazon RDS PostgreSQL** | Connect to RDS PostgreSQL instances using the PostgreSQL connector. | [PostgreSQL Data Connector](/docs/v1.10/components/data-connectors/postgres) | | **Amazon RDS MySQL** | Connect to RDS MySQL instances using the MySQL connector. | [MySQL Data Connector](/docs/v1.10/components/data-connectors/mysql) | | **Amazon MSK** | Stream data from [Amazon MSK](https://aws.amazon.com/msk/) (Managed Streaming for Apache Kafka) topics using the Kafka connector. | [Kafka Data Connector](/docs/v1.10/components/data-connectors/kafka) | | **Debezium (Amazon MSK)** | Change Data Capture (CDC) from databases via Debezium running on Amazon MSK for real-time dataset updates. | [Debezium Data Connector](/docs/v1.10/components/data-connectors/debezium) | | **AWS Glue Data Catalog** | Query Iceberg tables registered in AWS Glue. | [Glue Data Connector](/docs/v1.10/components/data-connectors/glue) | | **Apache Iceberg (AWS)** | Query Iceberg tables stored in S3 with Glue or REST catalog metadata. | [Iceberg Data Connector](/docs/v1.10/components/data-connectors/iceberg) | | **Delta Lake (S3)** | Query Delta Lake tables stored in Amazon S3. | [Delta Lake Data Connector](/docs/v1.10/components/data-connectors/delta-lake) | | **AWS Athena (ODBC)** | Connect to Athena using the ODBC connector with Athena SQL dialect support. | [ODBC Data Connector](/docs/v1.10/components/data-connectors/odbc) | ### Example: Amazon S3[​](#example-amazon-s3 "Direct link to Example: Amazon S3") ``` datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet s3_region: us-east-1 s3_auth: iam_role # Uses IAM credentials from environment ``` ### Example: DynamoDB[​](#example-dynamodb "Direct link to Example: DynamoDB") ``` datasets: - from: dynamodb:users name: users params: dynamodb_aws_region: us-west-2 ``` ### Example: AWS Glue with Amazon S3 Tables[​](#example-aws-glue-with-amazon-s3-tables "Direct link to Example: AWS Glue with Amazon S3 Tables") ``` datasets: - from: glue:my_namespace.orders name: orders params: glue_catalog_id: 123635965758:s3tablescatalog/my-table-bucket glue_region: us-east-2 ``` ## Catalog Connectors[​](#catalog-connectors "Direct link to Catalog Connectors") Catalog connectors provide schema discovery and unified access to tables in AWS data catalogs. | Connector | Description | Documentation | | -------------------- | --------------------------------------------------------------------------------- | -------------------------------------------------------------- | | **AWS Glue Catalog** | Discover and query tables from AWS Glue Data Catalog with glob pattern filtering. | [Glue Catalog Connector](/docs/v1.10/components/catalogs/glue) | ### Example: Glue Catalog[​](#example-glue-catalog "Direct link to Example: Glue Catalog") ``` catalogs: - from: glue name: my_data_lake include: - '*.*' # Include all tables from all databases params: glue_region: us-east-1 ``` ## AI Models (Amazon Bedrock)[​](#ai-models-amazon-bedrock "Direct link to AI Models (Amazon Bedrock)") Spice integrates with [Amazon Bedrock](https://aws.amazon.com/bedrock/) for large language model inference, supporting Amazon Nova and other foundation models. | Provider | Supported Models | Documentation | | ------------------ | ------------------------------------------------------------------------ | ------------------------------------------------------- | | **Amazon Bedrock** | Amazon Nova (Micro, Lite, Pro, Premier), cross-region inference profiles | [Bedrock Models](/docs/v1.10/components/models/bedrock) | ### Example: Amazon Nova[​](#example-amazon-nova "Direct link to Example: Amazon Nova") ``` models: - from: bedrock:us.amazon.nova-lite-v1:0 name: nova params: aws_region: us-east-1 ``` ### Guardrails Support[​](#guardrails-support "Direct link to Guardrails Support") Bedrock Guardrails can filter model inputs and outputs: ``` models: - from: bedrock:amazon.nova-pro-v1:0 name: nova-guarded params: aws_region: us-east-1 bedrock_guardrail_identifier: arn:aws:bedrock:us-east-1:123456789012:guardrail/abc123 bedrock_guardrail_version: '1' ``` ## Embeddings (Amazon Bedrock)[​](#embeddings-amazon-bedrock "Direct link to Embeddings (Amazon Bedrock)") Generate vector embeddings using Amazon Bedrock embedding models for semantic search and RAG applications. | Provider | Supported Models | Documentation | | ------------------ | ------------------------------------------------------------------------ | --------------------------------------------------------------- | | **Amazon Bedrock** | Amazon Titan Embeddings, Amazon Nova Multimodal Embeddings, Cohere Embed | [Bedrock Embeddings](/docs/v1.10/components/embeddings/bedrock) | ### Example: Amazon Titan Embeddings[​](#example-amazon-titan-embeddings "Direct link to Example: Amazon Titan Embeddings") ``` embeddings: - from: bedrock:amazon.titan-embed-text-v2:0 name: titan params: aws_region: us-east-1 dimensions: '256' ``` ### Example: Amazon Nova Multimodal Embeddings[​](#example-amazon-nova-multimodal-embeddings "Direct link to Example: Amazon Nova Multimodal Embeddings") ``` embeddings: - from: bedrock:amazon.nova-2-multimodal-embeddings-v1:0 name: nova_embed params: dimensions: '1024' truncation_mode: START embedding_purpose: GENERIC_RETRIEVAL aws_region: us-east-1 ``` ## Vector Stores (Amazon S3 Vectors)[​](#vector-stores-amazon-s3-vectors "Direct link to Vector Stores (Amazon S3 Vectors)") [Amazon S3 Vectors](https://aws.amazon.com/s3/features/s3-vectors/) is a new S3 bucket type for storing and querying vector embeddings at scale. Spice integrates S3 Vectors as a vector index backend for hybrid search applications. | Engine | Description | Documentation | | --------------------- | ---------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- | | **Amazon S3 Vectors** | Sub-second similarity queries on billions of vectors with up to 90% cost reduction compared to traditional vector databases. | [S3 Vectors Engine](/docs/v1.10/components/vectors/s3_vectors) | ### Example: S3 Vectors with Bedrock Embeddings[​](#example-s3-vectors-with-bedrock-embeddings "Direct link to Example: S3 Vectors with Bedrock Embeddings") ``` datasets: - from: oracle:"CUSTOMER_REVIEWS" name: reviews vectors: enabled: true engine: s3_vectors params: s3_vectors_bucket: my-s3-vector-bucket s3_vectors_aws_region: us-east-1 columns: - name: body embeddings: from: bedrock_titan embeddings: - from: bedrock:amazon.titan-embed-text-v2:0 name: bedrock_titan params: aws_region: us-east-1 dimensions: '256' ``` ## Data Accelerators (S3 Express One Zone)[​](#data-accelerators-s3-express-one-zone "Direct link to Data Accelerators (S3 Express One Zone)") Spice Cayenne data accelerator supports [AWS S3 Express One Zone](https://aws.amazon.com/s3/storage-classes/express-one-zone/) for storing accelerated data with single-digit millisecond latency. This is ideal for latency-sensitive query workloads that require persistent storage while maintaining fast access. Storage Recommendation For best performance, store Cayenne data files on local NVMe storage. Use S3 Express One Zone only when persistence of accelerations is required, such as preserving accelerated data across restarts or sharing data between multiple Spice instances. | Accelerator | Description | Documentation | | ----------------- | --------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- | | **Spice Cayenne** | High-performance data accelerator using Vortex file format with S3 Express One Zone for sub-10ms latency query performance. | [Cayenne Accelerator](/docs/v1.10/components/data-accelerators/cayenne) | ### Why S3 Express One Zone?[​](#why-s3-express-one-zone "Direct link to Why S3 Express One Zone?") S3 Express One Zone directory buckets provide: * **Single-digit millisecond latency**: 10x faster than S3 Standard for first-byte latency * **High request throughput**: Up to 10x higher request rates than S3 Standard * **Cost efficiency**: Lower per-request costs for high-frequency access patterns * **Durability**: Same 99.999999999% (11 9s) durability as S3 Standard ### Example: Cayenne with S3 Express One Zone[​](#example-cayenne-with-s3-express-one-zone "Direct link to Example: Cayenne with S3 Express One Zone") ``` datasets: - from: s3://source-bucket/events/ name: analytics_events acceleration: engine: cayenne enabled: true mode: file params: # Store accelerated data in S3 Express One Zone bucket cayenne_file_path: s3://my-bucket--usw2-az1--x-s3/cayenne/ cayenne_s3_region: us-west-2 ``` ### Example: Auto-generated Bucket with IAM Role[​](#example-auto-generated-bucket-with-iam-role "Direct link to Example: Auto-generated Bucket with IAM Role") ``` datasets: - from: postgresql://db/events name: fast_events acceleration: engine: cayenne enabled: true mode: file params: # Auto-generates bucket: spice-{spicepod-name}-fast_events--usw2-az1--x-s3 cayenne_s3_zone_ids: usw2-az1 ``` ### Supported AWS Regions[​](#supported-aws-regions "Direct link to Supported AWS Regions") S3 Express One Zone is available in select regions. Spice automatically derives the region from zone IDs: | Zone ID Prefix | Region | | -------------- | -------------- | | `use1` | us-east-1 | | `use2` | us-east-2 | | `usw1` | us-west-1 | | `usw2` | us-west-2 | | `euw1` | eu-west-1 | | `euc1` | eu-central-1 | | `apne1` | ap-northeast-1 | | `apse1` | ap-southeast-1 | See AWS documentation for the complete list of [S3 Express One Zone availability zones](https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-express-Regions-and-Zones.html). ## Secret Management[​](#secret-management "Direct link to Secret Management") Securely store and retrieve credentials using AWS Secrets Manager. | Store | Description | Documentation | | ----------------------- | ----------------------------------------------------- | ------------------------------------------------------------------------------- | | **AWS Secrets Manager** | Read secrets from AWS Secrets Manager by secret name. | [AWS Secrets Manager](/docs/v1.10/components/secret-stores/aws-secrets-manager) | ### Example: Using Secrets Manager[​](#example-using-secrets-manager "Direct link to Example: Using Secrets Manager") ``` secrets: - from: aws_secrets_manager:my_database_creds name: db datasets: - from: postgres:public.users name: users params: pg_host: ${db:host} pg_user: ${db:username} pg_pass: ${db:password} ``` ## Authentication[​](#authentication "Direct link to Authentication") All AWS integrations support the standard AWS SDK credential chain. When credentials are not explicitly configured, Spice loads them from the following sources in order: 1. **Environment Variables**: `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN` 2. **Shared Credentials Files**: `~/.aws/credentials` and `~/.aws/config` 3. **AWS SSO Sessions**: Configured via `aws configure sso` 4. **Web Identity Token**: For OIDC/OAuth (common with EKS IRSA) 5. **ECS Container Credentials**: Automatic IAM role for ECS tasks 6. **EC2 Instance Metadata (IMDSv2)**: Automatic IAM role for EC2 instances ### IAM Permissions[​](#iam-permissions "Direct link to IAM Permissions") Ensure the IAM role or user has appropriate permissions for all AWS services used: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:GetObject", "s3:ListBucket", "dynamodb:Scan", "dynamodb:DescribeTable", "glue:GetTable", "glue:GetTables", "glue:GetDatabase", "glue:GetDatabases", "bedrock:InvokeModel", "secretsmanager:GetSecretValue" ], "Resource": "*" } ] } ``` ## Deployment Options[​](#deployment-options "Direct link to Deployment Options") Deploy Spice on AWS infrastructure for optimal performance and integration: | Option | Description | Documentation | | -------------- | --------------------------------------------------- | --------------------------------------------- | | **Amazon EKS** | Kubernetes orchestration with Helm chart deployment | [AWS Deployment](/docs/v1.10/deployment/aws/) | | **Amazon ECS** | Container service with Fargate or EC2 launch types | [AWS Deployment](/docs/v1.10/deployment/aws/) | | **Amazon EC2** | Direct deployment with Docker or binary | [AWS Deployment](/docs/v1.10/deployment/aws/) | ## Resources[​](#resources "Direct link to Resources") ### AWS Blog Posts[​](#aws-blog-posts "Direct link to AWS Blog Posts") * [Architecting High-Performance AI-Driven Data Applications with Spice.ai and AWS](https://aws.amazon.com/blogs/storage/architecting-high-performance-ai-driven-data-applications-with-spice-ai-and-aws/) - AWS Storage Blog ### Spice.ai Blog Posts[​](#spiceai-blog-posts "Direct link to Spice.ai Blog Posts") * [Amazon S3 Vectors](https://spice.ai/blog/amazon-s3-vectors) - Overview of S3 Vectors integration * [Getting Started with Amazon S3 Vectors and Spice](https://spice.ai/blog/getting-started-with-amazon-s3-vectors-and-spice) - Step-by-step tutorial ### Videos[​](#videos "Direct link to Videos") * [Getting started with Amazon S3 Vectors and Spice](https://www.youtube.com/watch?v=KuWI0yDOnIU) - YouTube walkthrough * [How Spice AI operationalizes data lakes for AI using Amazon S3](https://www.youtube.com/watch?v=KuWI0yDOnIU\&list=PLesJrUXEx3U-WIqfWYfha4zBkyZo9czEJ\&index=2) - Spice presentation at re:Invent ### Marketplace[​](#marketplace "Direct link to Marketplace") * [Spice.ai on AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-jmf6jskjvnq7i) - Deploy Spice.ai from AWS Marketplace ## Quick Start[​](#quick-start "Direct link to Quick Start") Get started with Spice on AWS in minutes: 1. **Install Spice CLI**: ``` curl https://install.spiceai.org | /bin/bash ``` 2. **Configure AWS credentials**: ``` aws configure ``` 3. **Create a Spicepod with S3 data**: ``` # spicepod.yaml version: v1beta1 kind: Spicepod name: aws_quickstart datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet s3_auth: iam_role ``` 4. **Start the runtime**: ``` spice run ``` 5. **Query your data**: ``` spice sql > SELECT COUNT(*) FROM taxi_trips; ``` --- # Spice Cloud Platform Deployment The Spice Cloud Platform is a managed, cloud-hosted solution designed for deploying data and AI applications and agents. It provides a secure and efficient compute environment powered by Spice.ai OSS, offering building blocks including high-speed SQL queries, LLM inference, vector search, and retrieval-augmented generation (RAG). ## Benefits of the Spice.ai Cloud Platform[​](#benefits-of-the-spiceai-cloud-platform "Direct link to Benefits of the Spice.ai Cloud Platform") * **Simplified Deployment**: Focus on creating data and AI applications without the complexity of managing infrastructure. * **High Performance**: Optimize data queries and AI workflows with cloud-scale compute resources. * **Collaboration**: Share and manage datasets, models, and tools across your team and the enterprise. * **Production-Ready**: Achieve reliability, scalability, and compliance for AI applications. ## Security and Compliance[​](#security-and-compliance "Direct link to Security and Compliance") The Spice.ai Cloud Platform prioritizes security and compliance to ensure the protection of user data and systems. It adheres to SOC 2 Type II compliance standards, providing enterprise-level security validated by third-party audits. The platform employs robust security measures, including encryption of sensitive data both in-transit and at-rest, multi-factor authentication (MFA), and role-based access control (RBAC). Access and usage are logged for auditability, and the principle of least privilege is enforced to minimize unnecessary access. For more details, visit the [Spice.ai Security Documentation](https://docs.spice.ai/security/security). ## Deployment Overview[​](#deployment-overview "Direct link to Deployment Overview") 1. **Sign Up**: Register for an account on the [Spice.ai Cloud Platform](https://spice.ai/login). 2. **Configure Your Application**: Organize datasets, models, and workflows using the platform's cloud portal. 3. **Deploy**: Launch your AI applications and agents with minimal configuration. 4. **Monitor and Scale**: Leverage built-in monitoring and observability tools to track performance and scale resources as required. ## Learn More[​](#learn-more "Direct link to Learn More") For comprehensive instructions and advanced configuration options, refer to the [Spice.ai Cloud Platform documentation](https://docs.spice.ai/). --- # Docker - Kubernetes ## Running Spice.ai as a Docker Container[​](#running-spiceai-as-a-docker-container "Direct link to Running Spice.ai as a Docker Container") For information on using Helm deployment, refer to the [Deploy Spice.ai in Kubernetes using Helm](/docs/deployment/kubernetes) section. Use the [`spiceai/spiceai` Docker image](https://hub.docker.com/r/spiceai/spiceai/tags) to run Spice.ai as a Docker container: ``` # To use AI features including embeddings, models, and search use the spiceai/spiceai:latest-models tag. See supported tags at https://hub.docker.com/r/spiceai/spiceai/tags FROM spiceai/spiceai:latest # Copy the Spicepod configuration file COPY spicepod.yaml /app/spicepod.yaml # Copy your local data COPY data /app/data # Copy the .env.local & .env files COPY .env* /app/ # Spice runtime start-up arguments CMD ["--http","0.0.0.0:8090","--metrics", "0.0.0.0:9090","--flight","0.0.0.0:50051"] EXPOSE 8090 EXPOSE 9090 EXPOSE 50051 # Start the Spicepod ``` Example Docker Compose configuration to build and start the container: ``` services: spiced: build: context: . dockerfile: Dockerfile container_name: spiced-container ports: - '50051:50051' - '8090:8090' - '9090:9090' ``` ``` docker-compose up --build ``` ``` [+] Building 0.4s (10/10) FINISHED docker:desktop-linux => [spiced internal] load build definition from Dockerfile 0.0s => => transferring dockerfile: 470B 0.0s => [spiced internal] load metadata for docker.io/spiceai/spiceai:latest 0.3s => [spiced internal] load .dockerignore 0.0s => => transferring context: 2B 0.0s => [spiced 1/4] FROM docker.io/spiceai/spiceai:latest@sha256:96bedafa678343e8acb977001ddbde335d8d66476c2ed891f9b7873d277f2b84 0.0s => [spiced internal] load build context 0.0s => => transferring context: 262B 0.0s => CACHED [spiced 2/4] COPY spicepod.yaml /app/spicepod.yaml 0.0s => CACHED [spiced 3/4] COPY data /app/data 0.0s => CACHED [spiced 4/4] COPY .env* /app/ 0.0s => [spiced] exporting to image 0.0s => => exporting layers 0.0s => => writing image sha256:7f753171835316954b7072ba9b59c388c6295d3cfc677bc2de1271ba85a1f040 0.0s => => naming to docker.io/library/accounts-spiced 0.0s => [spiced] resolving provenance for metadata file 0.0s [+] Running 1/0 ✔ Container spiced-container Recreated 0.0s Attaching to spiced-container spiced-container | 2024-12-19T00:43:13.844091Z INFO runtime::init::dataset: No datasets were configured. If this is unexpected, check the Spicepod configuration. spiced-container | 2024-12-19T00:43:13.844615Z INFO runtime::opentelemetry: Spice Runtime OpenTelemetry listening on 127.0.0.1:50052 spiced-container | 2024-12-19T00:43:13.844750Z INFO runtime::metrics_server: Spice Runtime Metrics listening on 0.0.0.0:9090 spiced-container | 2024-12-19T00:43:13.844778Z INFO runtime::flight: Spice Runtime Flight listening on 0.0.0.0:50051 spiced-container | 2024-12-19T00:43:13.844882Z INFO runtime::http: Spice Runtime HTTP listening on 0.0.0.0:8090 spiced-container | 2024-12-19T00:43:13.844899Z INFO runtime::init::results_cache: Initialized results cache; max size: 128.00 MiB, item ttl: 5s ``` --- # Docker Sandbox Guide - v1.3.0 ## v1.3.0 Docker Sandbox[​](#v130-docker-sandbox "Direct link to v1.3.0 Docker Sandbox") In the v1.3.0 release, the Docker image changed the sandbox from a script that ran at startup, to being baked into the Docker image itself. Prior to this, the Docker image would start up as a root user, and then set up a sandbox user with restricted permissions before starting the Spice runtime in that restricted context. Additionally, the Docker image includes no standard Linux tools like `bash`. Starting with v1.3.0, the sandboxing logic is baked directly into the final Docker image, and the Docker image starts up as the sandbox user. For most users, this change will be transparent. However, there are a few cases where an action is required to update. ### Building a custom Docker image based on v1.3.0[​](#building-a-custom-docker-image-based-on-v130 "Direct link to Building a custom Docker image based on v1.3.0") Building a custom Docker image based on v1.3.0 that installs additional dependencies will require using the `debian:bookworm-slim` base image and copying the `spiced` binary from the `spiceai/spiceai` image into the custom image. This approach can also be used to restore the previous behavior of including standard Linux tools like `bash`. ``` FROM debian:bookworm-slim # Copy the spiced binary from the spiceai/spiceai image into the custom image. COPY --from=spiceai/spiceai:v1.3.0 /usr/local/bin/spiced /usr/local/bin/spiced # Install any additional dependencies needed for the image. RUN apt update && apt install -y --no-install-recommends # Any other customizations needed for the image. WORKDIR /app ENTRYPOINT ["/usr/local/bin/spiced"] ``` This will restore the previous behavior of starting as the root user and including standard Linux tools, like `bash`. #### Running as a non-root user[​](#running-as-a-non-root-user "Direct link to Running as a non-root user") Spice recommends that custom Docker images based on Spice run as a non-root user. i.e. ``` RUN addgroup -g 1001 -S sandboxgroup && adduser -u 1001 -S -G sandboxgroup sandbox USER sandbox ``` This may require additional configuration of mounted volumes to ensure that the sandbox user has access to the necessary files. i.e. in Kubernetes, this requires adding a `securityContext` to the pod spec. ``` securityContext: runAsUser: 1001 runAsGroup: 1001 # Tells Kubernetes to set the group of the files in the volume to sandboxgroup, # which allows the sandbox user to access the files. fsGroup: 1001 ``` note The `fsGroup` directive does not work for all Kubernetes storage types. For example, it does not work for `hostPath` volumes. In this case, an init container can be used to set the group of the files in the volume. ### Custom Kubernetes deployments[​](#custom-kubernetes-deployments "Direct link to Custom Kubernetes deployments") Kubernetes deployments that do not use the v1.3.0 Helm chart will need to add the following `securityContext` to their pod spec: ``` securityContext: runAsUser: 65534 runAsGroup: 65534 fsGroup: 65534 ``` ### Debugging sandbox container[​](#debugging-sandbox-container "Direct link to Debugging sandbox container") To debug issues with the sandbox container, see the [Debugging Sandbox Container](/docs/v1.10/troubleshooting#debugging-sandbox-container) section of the troubleshooting guide. --- # Helm - Kubernetes ## Quickstart[​](#quickstart "Direct link to Quickstart") ``` helm repo add spiceai https://helm.spiceai.org helm repo update helm upgrade --install spiceai spiceai/spiceai ``` Deployment Architecture By default, the Spice.ai Helm chart deploys the application as a stateless Kubernetes Deployment. To persist data between restarts (e.g., for file-based acceleration), enable and configure the `stateful` section in the values file. Refer to the [Stateful Configuration](#stateful-configuration) section for details. ## What are Kubernetes and Helm?[​](#what-are-kubernetes-and-helm "Direct link to What are Kubernetes and Helm?") **Kubernetes** is an open-source platform for automating deployment, scaling, and management of containerized applications.
**Helm** is a package manager for Kubernetes that simplifies the installation and configuration of applications using reusable templates called charts. Spice publishes a Helm chart that simplifies the deployment of Spice.ai OSS on Kubernetes. ## Deploy Spice using Helm in Kubernetes[​](#deploy-spice-using-helm-in-kubernetes "Direct link to Deploy Spice using Helm in Kubernetes") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to a Kubernetes cluster. * For local testing, try running a local Kubernetes cluster using [Kind](https://kind.sigs.k8s.io/docs/user/quick-start/). * `kubectl` CLI installed and configured to interact with the target Kubernetes cluster. Visit the [Kubernetes docs](https://kubernetes.io/docs/tasks/tools/#kubectl) for installation instructions. * Helm CLI installed. Visit the [Helm docs](https://helm.sh/docs/intro/install/) for installation instructions. ### Add the Spice Helm repository[​](#add-the-spice-helm-repository "Direct link to Add the Spice Helm repository") A Helm repository (Helm repo) is a storage location where Helm charts are hosted and can be accessed for deployment in Kubernetes clusters. Add the Spice Helm repository to your local Helm client and update the index to get the latest charts: ``` helm repo add {repository-name} https://helm.spiceai.org helm repo update ``` The repository name is customizable and can be set to any preferred value. For example: ``` helm repo add spiceai https://helm.spiceai.org # or helm repo add my-spiceai-repo https://helm.spiceai.org ``` ### Install the Spice Helm chart as a new release[​](#install-the-spice-helm-chart-as-a-new-release "Direct link to Install the Spice Helm chart as a new release") Once the repository with the chart is added, install the chart into the Kubernetes cluster with the following command: ``` helm install {release-name} {repository-name}/spiceai --namespace {namespace} ``` For example: ``` helm install spiceai spiceai/spiceai --namespace default ``` Spice can be installed multiple times in the same cluster by specifying a different release name for each installation. #### Command Breakdown[​](#command-breakdown "Direct link to Command Breakdown") * `helm install`: Installs a new Helm chart. To upgrade an existing release, use `helm upgrade`. Combine both upgrade and install by specifying `helm upgrade --install`. * `spiceai`: The name of the release. This name is customizable and can be set to any preferred value, e.g., `spiceai-my-app-v1` and `spiceai-my-app-v2` are valid release names. Each Helm release is a distinct installation of the same chart. * `spiceai/spiceai`: The chart to install. The first `spiceai` is the repository name added earlier, and the second `spiceai` is the name of the chart to install. While the repository name is customizable, the chart name is not. * `--namespace default`: The Kubernetes namespace to install the chart into. This is optional and defaults to `default`. Another valid command to install the chart is (assuming the repository name is `my-spiceai-repo`): ``` helm upgrade --install spiceai-my-app-1 my-spiceai-repo/spiceai ``` ### Upgrade the Spice Helm chart[​](#upgrade-the-spice-helm-chart "Direct link to Upgrade the Spice Helm chart") To upgrade an existing release, use the `helm upgrade` command: ``` helm upgrade {release-name} {repository-name}/{chart-name} ``` For example: ``` helm upgrade spiceai-my-app-1 my-spiceai-repo/spiceai ``` ### Rollback a Helm release[​](#rollback-a-helm-release "Direct link to Rollback a Helm release") On occasion, you may need to roll back a Spice Helm release to a previous version. To do so, use the `helm rollback` command. This will notify Kubernetes to redeploy Spice back to a previous version of the Helm release: ``` helm rollback {release-name} --namespace {namespace} ``` For example: ``` helm rollback spiceai-my-app-1 --namespace default ``` ### Uninstall a Helm release[​](#uninstall-a-helm-release "Direct link to Uninstall a Helm release") To uninstall a Helm release, use the `helm uninstall` command. This will cause Kubernetes to remove the Spice deployment entirely. Note that any data stored in volumes created by configuring the `stateful` parameter will be preserved and must be manually deleted if desired: ``` helm uninstall {release-name} --namespace {namespace} ``` For example: ``` helm uninstall spiceai-my-app-1 --namespace default ``` ## Customize the Helm release[​](#customize-the-helm-release "Direct link to Customize the Helm release") By default, the Helm release installs a minimal Spice.ai setup with an empty Spicepod. To add a Spicepod and adjust other settings, customize the release as needed by creating a `values.yaml` file or by using the `--set` flag. note `values.yaml` is the configuration file used in Helm to define the user-configurable parameters of a Helm chart. The `--set` flag is used to specify individual values on the command line. Visit the [Helm docs](https://helm.sh/docs/chart_template_guide/values_files/#helm) for more information. Create a `values.yaml` file and override the default values as needed. The full list of configurable parameters and their defaults are specified in the [values.yaml file](https://github.com/spiceai/spiceai/blob/trunk/deploy/chart/values.yaml) in the `spiceai/spiceai` repository. The [Common Parameters](#common-parameters) section below lists the most commonly used configurable parameters and their descriptions. ### Spicepod[​](#spicepod "Direct link to Spicepod") To customize the Spicepod that the Spice.ai runtime will load, define a [Spicepod](https://spiceai.org/docs/getting-started/spicepods) in a new `values.yaml` file. ``` spicepod: name: app version: v1 kind: Spicepod datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet ``` Upgrade or install a new release with the custom Spicepod: ``` helm upgrade --install spiceai spiceai/spiceai -f values.yaml ``` note The Helm convention is to use a file called `values.yaml`, but any file name can be used and passed to the `-f` flag. ## Common Parameters[​](#common-parameters "Direct link to Common Parameters") | **Name** | **Description** | **Value** | | ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | | `additionalEnv` | Additional environment variables to set in the Spice.ai container. | `[]` | | `additionalLabels` | Additional labels to add to all resources. | `{}` | | `image.pullSecrets` | Specify Docker registry secret names as an array. | `[]` | | `image.repository` | The repository of the Docker image. | `spiceai` | | `image.tag` | Replace with a specific version of Spice.ai to run. | `1.3.0` | | `monitoring.podMonitor.enabled` | Enable Prometheus metrics collection for the Spice pods. Requires the [Prometheus Operator](https://prometheus-operator.dev/docs/operator/api/#monitoring.coreos.com/v1.PodMonitor) CRDs. | `false` | | `replicaCount` | Number of Spice.ai replicas to run. | `1` | | `resources` | Resource requests and limits for the Spice.ai container. See [Container resource examples](https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/#example-1). | `{}` | | `service.type` | Kubernetes service type. Can be null, ClusterIP, NodePort, or LoadBalancer. | `null` | | `serviceAccount.create` | Specifies whether a ServiceAccount should be created. | `false` | | `spicepod` | Define the [Spicepod](https://spiceai.org/docs/getting-started/spicepods) to be loaded by the Spice.ai runtime. | `{}` | | `stateful.enabled` | Use a StatefulSet with a PVC (Persistent Volume Claim) for the data volume. | `false` | | `stateful.mountPath` | Mount path in container for the persistent volume. | `/data` | | `stateful.size` | Size of each PV in the StatefulSet. | `1Gi` | | `stateful.storageClass` | Storage class for the volume claim template in the StatefulSet. | `standard` | | `tolerations` | List of node taints to tolerate. | `[]` | ## Environment Variables and Secrets[​](#environment-variables-and-secrets "Direct link to Environment Variables and Secrets") Add extra environment variables using the `additionalEnv` property. This can be useful when combining with the [Environment Secret Store](/docs/v1.10/components/secret-stores/env). ``` additionalEnv: - name: SPICED_LOG value: 'DEBUG' - name: SPICE_SECRET_SPICEAI_KEY valueFrom: secretKeyRef: name: spice-secrets key: spiceai-key ``` To create a test secret: ``` kubectl create secret generic spice-secrets --from-literal=spiceai-key="secret-value" ``` Further reading: * [Kubernetes Secrets](https://kubernetes.io/docs/concepts/configuration/secret/) * [Good practices for Kubernetes Secrets](https://kubernetes.io/docs/concepts/security/secrets-good-practices/) ## Monitoring[​](#monitoring "Direct link to Monitoring") The Spice Helm chart includes compatibility with the [Prometheus Operator](https://prometheus-operator.dev/) for collecting Prometheus metrics that can be visualized in the [Spice Grafana dashboard](/docs/monitoring/grafana). To enable this feature, set the `monitoring.podMonitor.enabled` value to `true`. This will create a `PodMonitor` resource for the Spice.ai pods that will configure Prometheus to scrape metrics from the Spice.ai pods. Install the Prometheus Operator The easiest way to install the Prometheus Operator along with Grafana is to use the [kube-prometheus-stack](https://github.com/prometheus-community/helm-charts/blob/main/charts/kube-prometheus-stack/README.md) Helm chart. ``` helm repo add prometheus-community https://prometheus-community.github.io/helm-charts helm repo update helm install prometheus-stack prometheus-community/kube-prometheus-stack \ --set prometheus.prometheusSpec.podMonitorSelectorNilUsesHelmValues=false \ --set prometheus.prometheusSpec.serviceMonitorSelectorNilUsesHelmValues=false ``` Deploy the Spice.ai Helm chart with monitoring enabled: ``` helm upgrade --install spiceai spiceai/spiceai --set monitoring.podMonitor.enabled=true ``` Once the monitoring is enabled, import the [Spice Grafana dashboard](/docs/monitoring/grafana) to visualize the Spice.ai metrics. ### Health and Readiness[​](#health-and-readiness "Direct link to Health and Readiness") Spice provides two HTTP endpoints for monitoring the runtime state: `/health` and `/v1/ready`. These endpoints are used for Kubernetes health and readiness probes in the Spice deployment. The Spice Helm chart automatically configures these probes. #### Health Probe[​](#health-probe "Direct link to Health Probe") The `/health` endpoint indicates whether the Spice process is up and running, ready to receive requests. A probe can be configured for custom deployment as follows: ``` livenessProbe: httpGet: path: /health port: 8090 ``` In Kubernetes, the pod will not be marked as healthy until the `/health` endpoint returns a `200` status. #### Readiness Probe[​](#readiness-probe "Direct link to Readiness Probe") The `/ready` endpoint indicates **whether the Spice components (datasets, models, etc) are ready**. While the `/health` endpoint might show that Spice is up and running, the `/ready` endpoint must return a `200` status to ensure that queries will return results. ``` readinessProbe: httpGet: path: /v1/ready port: 8090 ``` note For more information on how Kubernetes uses probes to determine the health of a pod, see [here](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes) ## Service Configuration[​](#service-configuration "Direct link to Service Configuration") Configure the Kubernetes service for Spice.ai: ``` service: # type can be null, ClusterIP, NodePort, or LoadBalancer type: LoadBalancer additionalAnnotations: {} # If service type is LoadBalancer, specify the source IP ranges to whitelist loadBalancerSourceRanges: - 10.0.0.0/8 - 172.16.0.0/12 # The selector to use for the service selector: custom-selector: value ``` ## Service Account[​](#service-account "Direct link to Service Account") Configure a Kubernetes ServiceAccount for Spice.ai: ``` serviceAccount: # Specifies whether a ServiceAccount should be created create: true # The name of the ServiceAccount to use. # If not set and create is true, a name is generated using the fullname template name: spice-service-account ``` ## Volumes and Volume Mounts[​](#volumes-and-volume-mounts "Direct link to Volumes and Volume Mounts") Define custom volumes and volume mounts for the Spice.ai container: ``` # Define volumes to be mounted to the container volumes: - name: custom-data persistentVolumeClaim: claimName: my-custom-pvc - name: config-volume configMap: name: custom-config # Define volume mounts for the container volumeMounts: - name: custom-data mountPath: /data - name: config-volume mountPath: /config ``` ## Stateful Configuration[​](#stateful-configuration "Direct link to Stateful Configuration") The Spice.ai Helm chart provides two deployment architectures to accommodate different persistence requirements. The default architecture deploys Spice.ai as a stateless Kubernetes Deployment, suitable for workloads that do not require data persistence between pod restarts. For workloads requiring data persistence, the chart supports deploying Spice.ai as a StatefulSet with persistent storage. This architecture becomes essential when implementing file-based acceleration for datasets or when maintaining state between pod restarts is critical. Enabling the StatefulSet architecture requires configuration of the `stateful` section: ``` # Use a StatefulSet with a PVC for the data volume stateful: enabled: true # Storage class for the volume claim template storageClass: 'standard' # Size of each PV in the StatefulSet size: 1Gi # Mount path in container mountPath: /data ``` When `stateful.enabled` is set to `true`, the Helm chart creates a StatefulSet instead of a Deployment and provisions a PersistentVolumeClaim for each replica. The persistent volume is mounted at the specified path, allowing data to persist across pod restarts and rescheduling events. ## Example values.yaml[​](#example-valuesyaml "Direct link to Example values.yaml") ``` additionalLabels: environment: production app: spice image: repository: spiceai/spiceai tag: 1.3.0 replicaCount: 1 service: type: ClusterIP additionalAnnotations: service.beta.kubernetes.io/aws-load-balancer-internal: 'true' resources: limits: cpu: 1000m memory: 2Gi requests: cpu: 500m memory: 1Gi additionalEnv: - name: SPICED_LOG value: 'INFO' - name: SPICE_SECRET_SPICEAI_KEY valueFrom: secretKeyRef: name: spice-secrets key: spiceai-key stateful: enabled: true storageClass: 'standard' size: 5Gi mountPath: /data monitoring: podMonitor: enabled: true additionalLabels: release: prometheus spicepod: name: app version: v1 kind: Spicepod datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips description: Demo taxi trips in s3 params: file_format: parquet acceleration: enabled: true engine: duckdb mode: file params: duckdb_file: /data/taxi_trips.db # Uncomment to refresh the acceleration on a schedule # refresh_check_interval: 1h # refresh_mode: full ``` ## Cookbook[​](#cookbook "Direct link to Cookbook") * [Running Spice.ai in Kubernetes](https://github.com/spiceai/cookbook/tree/trunk/kubernetes) --- # Frequently Asked Questions ## 1. What is Spice?[​](#1-what-is-spice "Direct link to 1. What is Spice?") Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven apps and agents. Spice provides four industry standard APIs in a lightweight, portable runtime (single \~140 MB binary): 1. **SQL Query APIs**: Supports HTTP, Arrow Flight, Arrow Flight SQL, ODBC, JDBC, and ADBC. 2. **OpenAI-Compatible APIs**: Provides HTTP APIs for OpenAI SDK compatibility, local model serving (CUDA/Metal accelerated), and hosted model gateway. 3. **Iceberg Catalog REST APIs**: Offers a unified API for Iceberg Catalog. 4. **MCP HTTP+SSE APIs**: Enables integration with external tools via Model Context Protocol (MCP) using HTTP and Server-Sent Events (SSE). Spice embeds [DataFusion](https://datafusion.apache.org/), the fastest single-node Parquet SQL query engine, and [DuckDB](https://duckdb.org), to serve secure, virtualized data views to data-intensive apps, AI, and agents. ## 2. Why should I use Spice?[​](#2-why-should-i-use-spice "Direct link to 2. Why should I use Spice?") Spice is primarily used for: * **Data Federation**: SQL query across any database, data warehouse, or data lake. [Learn More](/docs/v1.10/features/query-federation). * **Data Materialization and Acceleration**: Materialize, accelerate, and cache database queries. [Read the MaterializedView interview - Building a CDN for Databases](https://materializedview.io/p/building-a-cdn-for-databases-spice-ai) * **AI apps and agents**: An AI-database powering retrieval-augmented generation (RAG) and intelligent agents. [Learn More](https://github.com/spiceai/cookbook/tree/trunk/rag#readme). ## 3. How is Spice different?[​](#3-how-is-spice-different "Direct link to 3. How is Spice different?") * **Application-Centric Design:** Spice is designed for 1:1 or 1 :N mappings between applications and Spice instances, making it flexible for tenant-specific or customer-specific configurations. Unlike traditional databases designed for many applications sharing one data system, Spice often runs one instance per application or tenant. * **Dual-Engine Acceleration:** Spice supports both OLAP (DuckDB/Arrow) and OLTP (SQLite/PostgreSQL) databases at the dataset level, providing flexibility for various query workloads. * **Separation of Materialization and Storage/Compute:** Spice enables data to remain close to its source while materializing working sets for fast access, reducing data movement and query latency. * **Deployment Flexibility:** Deployable across infrastructure tiers, including edge, on-prem, and cloud environments. Spice can run as a standalone instance, sidecar, microservice, or cluster. ## 4. Can Spice handle federated queries?[​](#4-can-spice-handle-federated-queries "Direct link to 4. Can Spice handle federated queries?") Yes. Spice natively supports federated queries across disparate data sources with advanced query push-down capabilities. Spice executes portions of queries directly on source databases, reducing data transfer and improving performance. [Learn More](/docs/v1.10/features/query-federation). ## 5. Is Spice a cache?[​](#5-is-spice-a-cache "Direct link to 5. Is Spice a cache?") Not solely. Spice functions as an active cache or working dataset prefetcher. A *working dataset* is a subset of data actively used by an application or model, such as recent records or frequently accessed tables. Unlike traditional caches that fetch data reactively, Spice proactively prefetches and materializes data based on filters, intervals, triggers, or Change Data Capture (CDC), ensuring data readiness for queries. Spice also supports [results caching](/docs/v1.10/features/caching). ## 6. Is Spice a CDN for databases?[​](#6-is-spice-a-cdn-for-databases "Direct link to 6. Is Spice a CDN for databases?") Yes. Spice acts as a CDN for databases by loading and materializing datasets close to applications, reducing latency and improving query efficiency. [Read more](https://materializedview.io/p/building-a-cdn-for-databases-spice-ai). ## 7. How is Spice different from Trino/Presto and Dremio?[​](#7-how-is-spice-different-from-trinopresto-and-dremio "Direct link to 7. How is Spice different from Trino/Presto and Dremio?") Spice is purpose-built for data and AI applications and agents, designed with low-latency access, materialization, and proximity to applications. Trino/Presto and Dremio primarily target big data analytics and rely on centralized clusters. Spice's decentralized approach reduces latency, simplifies deployment, and improves efficiency. ## 8. How does Spice compare to Spark?[​](#8-how-does-spice-compare-to-spark "Direct link to 8. How does Spice compare to Spark?") Spark excels at distributed batch processing and large-scale transformations. Spice focuses on real-time, low-latency data access and AI inference. Spice materializes data locally and supports tiered storage, optimizing performance for applications requiring fast access and high concurrency. ## 9. How does Spice compare to DuckDB?[​](#9-how-does-spice-compare-to-duckdb "Direct link to 9. How does Spice compare to DuckDB?") DuckDB is an embedded analytics database optimized for OLAP queries. Spice integrates DuckDB for data acceleration, combining DuckDB's analytical capabilities with Spice's broader federation, multi-engine support, and flexible deployment. Spice can be considered an enterprise/production productization of DuckDB for data-intensive applications. ## 10. What AI capabilities does Spice provide?[​](#10-what-ai-capabilities-does-spice-provide "Direct link to 10. What AI capabilities does Spice provide?") Spice provides unified APIs for data and AI workflows, including model inference, embeddings, and an AI gateway supporting OpenAI, Anthropic, xAI, and Nvidia NIMs. Spice includes advanced LLM tools such as vector and hybrid search, text-to-SQL, SQL retrieval, data sampling, and context formatting. ## 11. What AI model providers does Spice support?[​](#11-what-ai-model-providers-does-spice-support "Direct link to 11. What AI model providers does Spice support?") Spice supports local model serving (e.g., Llama3) and gateways to hosted AI platforms including OpenAI, Anthropic, xAI, and Nvidia NIMs. [Learn More](/docs/v1.10/features/large-language-models). ## 12. What deployment options does Spice support?[​](#12-what-deployment-options-does-spice-support "Direct link to 12. What deployment options does Spice support?") Spice supports multiple deployment configurations: * Standalone binary * Sidecar or microservice * Cluster deployments * Edge, on-prem, and cloud environments Spice Cloud Platform (SCP) provides managed, SOC 2 Type II compliant deployments. [Learn More](https://spiceai.org/docs/deployment/architectures). ## 13. Where can developers find examples and recipes?[​](#13-where-can-developers-find-examples-and-recipes "Direct link to 13. Where can developers find examples and recipes?") The [Spice.ai Cookbook](https://github.com/spiceai/cookbook) provides over 65 quickstarts and examples demonstrating Spice capabilities, including federated queries, RAG, text-to-SQL, and more. ## 14. How can developers get started quickly?[​](#14-how-can-developers-get-started-quickly "Direct link to 14. How can developers get started quickly?") Visit the [Spice.ai Getting Started Guide](/docs/v1.10/getting-started) to install Spice, connect data sources, and begin querying. Spice installs the GPU-accelerated runtime by default (if supported). ## 15. What is Data-grounded AI?[​](#15-what-is-data-grounded-ai "Direct link to 15. What is Data-grounded AI?") Data-grounded AI anchors models in accurate, current, domain-specific data rather than relying solely on pre-trained knowledge. Spice unifies enterprise data across databases, data lakes, and APIs, dynamically incorporating real-world context at inference time. This helps minimize hallucinations, reduce operational risk, and build trust in AI by delivering reliable, relevant outputs. ## 16. What query engines does Spice support?[​](#16-what-query-engines-does-spice-support "Direct link to 16. What query engines does Spice support?") Spice supports multiple query engines, including Apache Arrow, Cayenne (Vortex), DuckDB, SQLite, PostgreSQL, and DataFusion. Developers can select engines based on workload requirements, balancing performance, concurrency, and latency. ## 17. Does Spice support Change Data Capture (CDC)?[​](#17-does-spice-support-change-data-capture-cdc "Direct link to 17. Does Spice support Change Data Capture (CDC)?") Yes. Spice supports CDC via Debezium, enabling real-time data ingestion and materialization from databases such as PostgreSQL and MySQL. [Learn More](/docs/v1.10/features/cdc). ## 18. Can Spice integrate with existing BI tools?[​](#18-can-spice-integrate-with-existing-bi-tools "Direct link to 18. Can Spice integrate with existing BI tools?") Yes. Spice integrates with BI tools through standard SQL interfaces (ODBC, JDBC, Arrow Flight SQL), enabling accelerated, real-time analytics for dashboards and reporting. An official [Tableau Connector](/docs/v1.10/clients/tableau) is available and a [BI Acceleration](https://www.youtube.com/watch?v=blEtLgRKu0c) demo using Apache Superset. ## 19. How does Spice handle data privacy and compliance?[​](#19-how-does-spice-handle-data-privacy-and-compliance "Direct link to 19. How does Spice handle data privacy and compliance?") Spice provides secure, auditable data access through sandboxed runtimes, secure endpoint checks, and detailed telemetry and tracing. The Spice Cloud Platform (SCP) is SOC 2 Type II compliant, meeting enterprise security and compliance requirements. ## 20. Can Spice be used for real-time analytics?[​](#20-can-spice-be-used-for-real-time-analytics "Direct link to 20. Can Spice be used for real-time analytics?") Yes. Spice accelerates data locally using Apache Arrow, Cayenne (Vortex), DuckDB, SQLite, or PostgreSQL, enabling real-time analytics and sub-second query performance for data-intensive applications and dashboards. ## 21. How can developers contribute to Spice?[​](#21-how-can-developers-contribute-to-spice "Direct link to 21. How can developers contribute to Spice?") Developers can contribute by submitting code, documentation, or raising issues on [GitHub](https://github.com/spiceai/spiceai). See [CONTRIBUTING.md](https://github.com/spiceai/spiceai/blob/trunk/CONTRIBUTING) for guidelines. --- # Features ## [📄️Query Federation](/docs/v1.10/features/query-federation) [Learn how to use federated SQL queries in Spice.ai Open Source](/docs/v1.10/features/query-federation) ## [🗃Data Acceleration](/docs/v1.10/features/data-acceleration) [6 items](/docs/v1.10/features/data-acceleration) ## [📄️Caching](/docs/v1.10/features/caching) [Learn how to use Spice in-memory caching](/docs/v1.10/features/caching) ## [📄️Distributed Query](/docs/v1.10/features/distributed-query) [Learn how to run Spice in distributed mode for larger scale queries.](/docs/v1.10/features/distributed-query) ## [📄️Change Data Capture](/docs/v1.10/features/cdc) [Learn how to use Change Data Capture (CDC) in Spice.](/docs/v1.10/features/cdc) ## [📄️Data Ingestion](/docs/v1.10/features/data-ingestion) [Learn how to ingest data in Spice.](/docs/v1.10/features/data-ingestion) ## [🗃Large Language Models](/docs/v1.10/features/large-language-models) [7 items](/docs/v1.10/features/large-language-models) ## [📄️Machine Learning Models](/docs/v1.10/features/machine-learning-models) [Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform.](/docs/v1.10/features/machine-learning-models) ## [📄️Embedding Datasets](/docs/v1.10/features/embeddings) [Learn how to define, or augment existing datasets with embedding column(s).](/docs/v1.10/features/embeddings) ## [🗃Search](/docs/v1.10/features/search) [2 items](/docs/v1.10/features/search) ## [📄️Semantic Model](/docs/v1.10/features/semantic-model) [Learn how to define and use semantic data models with Spice.](/docs/v1.10/features/semantic-model) ## [🗃Observability](/docs/v1.10/features/observability) [1 item](/docs/v1.10/features/observability) ## [📄️Web Search](/docs/v1.10/features/web-search) [Learn how Spice can perform web search](/docs/v1.10/features/web-search) --- # Caching Spice supports in-memory caching for SQL query results and search results, which are both enabled by default when querying or searching via the HTTP (`/v1/sql`, `/v1/search`) and Arrow Flight APIs. Results caching improves performance for repeated requests and non-accelerated results, such as refresh data returned [on zero results](/docs/v1.10/features/data-acceleration/data-refresh#behavior-on-zero-results). The cache uses a [least-recently-used (LRU)](https://en.wikipedia.org/wiki/Cache_replacement_policies#LRU) replacement policy. You can configure the cache to set an item expiration duration, which defaults to 1 second. ``` version: v1 kind: Spicepod name: app runtime: caching: sql_results: enabled: true max_size: 1GiB # Default 128 MiB item_ttl: 1m # Default 1s stale_while_revalidate_ttl: 30s # Default 0s (disabled) search_results: enabled: true max_size: 1GiB # Default 128 MiB item_ttl: 1m # Default 1s stale_while_revalidate_ttl: 30s # Default 0s (disabled) ``` ## `caching` Parameters[​](#caching-parameters "Direct link to caching-parameters") | Parameter name | Optional | Description | | ---------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `sql_results` | Yes | Enabled by default. Configures the Runtime cache for results from SQL queries. See the [SQL Results Parameters](#cachingsql_results-parameters) for cache parameter details. | | `search_results` | Yes | Enabled by default. Configures the Runtime cache for results from searches. See the [Common Caching Parameters](#common-caching-parameters) for cache parameter details. | | `embeddings` | Yes | Enabled by default. Configures the Runtime cache for embeddings requests. See the [Common Caching Parameters](#common-caching-parameters) for cache parameter details. | ## Common Caching Parameters[​](#common-caching-parameters "Direct link to Common Caching Parameters") Every cache type (`sql_results`, `search_results`, `embeddings`) supports the following parameters: | Parameter name | Optional | Default | Description | | ------------------- | -------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `enabled` | Yes | `true` | Defaults to `true`. | | `max_size` | Yes | `128MiB` | Maximum cache size. Defaults to `128MiB`. | | `eviction_policy` | Yes | `lru` | Cache replacement policy when the cache reaches `max_size`. Defaults to `lru`, which is currently the only supported value. | | `item_ttl` | Yes | `1s` | Cache entry expiration duration (Time to Live). Defaults to 1 second. | | `hashing_algorithm` | Yes | `xxh3` | Selects which hashing algorithm is used to hash the cache keys when storing the results. Defaults to `xxh3`. Supports `xxh3`, `ahash`, `siphash`, `blake3`, `xxh32`, `xxh64`, or `xxh128`. | ## `caching.sql_results` Parameters[​](#cachingsql_results-parameters "Direct link to cachingsql_results-parameters") In addition to the common caching parameters, `sql_results` also supports additional parameters: | Parameter name | Optional | Default | Description | | ---------------------------- | -------- | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cache_key_type` | Yes | `plan` | Determines how cache keys are generated. Defaults to `plan`. `plan` uses the query's logical plan, while `sql` uses the raw SQL query string. | | `encoding` | Yes | `none` | Compression algorithm for cached results. Defaults to `none`. Supports `none` or `zstd`. | | `stale_while_revalidate_ttl` | Yes | `0s` | Duration to serve stale cache entries while revalidating in the background. When set to a non-zero value, expired cache entries continue to be served while a background refresh occurs. Defaults to `0s` (disabled). | ### Choosing a `cache_key_type`[​](#choosing-a-cache_key_type "Direct link to choosing-a-cache_key_type") * **`plan` (Default):** Uses the query's logical plan as the cache key. This approach matches semantically equivalent queries, even if their SQL syntax differs. However, it requires query parsing, which introduces some overhead. * **`sql`:** Uses the raw SQL string as the cache key. This method provides faster lookups but requires exact string matches. Queries with dynamic functions, such as `NOW()`, may produce unexpected results because the cache key changes with each execution. Use `sql` only when query results are predictable and consistent. Use `sql` for the lowest latency with identical queries that do not include dynamic functions. Use `plan` for greater flexibility and semantic matching of queries. ## Choosing a `hashing_algorithm`[​](#choosing-a-hashing_algorithm "Direct link to choosing-a-hashing_algorithm") The hashing algorithm determines how cache keys are hashed before being stored, impacting both lookup speed and protection against potential DOS attacks. * **`xxh3` (Default):** Uses the [XXH3](https://cyan4973.github.io/xxHash/) algorithm for hashing the cache keys. XXH3 is a fast, non-cryptographic hash algorithm that provides high performance and good distribution. It is suitable for scenarios where speed is critical and cryptographic security is not required. * **`siphash`:** Uses the SipHash1-3 algorithm for hashing the cache keys, the [default hashing algorithm of Rust](https://github.com/rust-lang/rust/commit/db1b1919baba8be48d997d9f70a6a5df7e31612a). This hashing algorithm is a secure algorithm that implements verified protections against ["hash flooding"](https://v8.dev/blog/hash-flooding) denial of service (DoS) attacks. Reasonably performant, and provides a high level of security. * **`ahash`:** Uses the [AHash](https://github.com/tkaitchuck/ahash) algorithm for hashing the cache keys. The AHash algorithm is a [high quality](https://github.com/tkaitchuck/aHash/blob/master/compare/readme.md#Quality) hashing algorithm, and has claimed resistance against hashing DoS attacks. AHash has higher performance than SipHash1-3, especially when used with `cache_key_type: plan`. * **`blake3`:** Uses the [BLAKE3](https://github.com/BLAKE3-team/BLAKE3) cryptographic hash function. BLAKE3 is a fast, parallelizable hash function that provides cryptographic security while maintaining high performance. It is suitable for scenarios requiring both speed and cryptographic guarantees. * **`xxh32`, `xxh64`, `xxh128`:** Variants of the XXH hashing algorithm with different output sizes. These algorithms offer a balance between speed and collision resistance, with larger hash sizes providing better collision resistance at the cost of performance. Use `xxh3` (the default) for its superior speed in most scenarios. Use `ahash`, `xxh64` or `xxh128` for reduced collision probability when caching a large number of queries. Use `blake3` when cryptographic security is required. Use `siphash` when protection against hash flooding attacks is a priority. ### Choosing an `encoding`[​](#choosing-an-encoding "Direct link to choosing-an-encoding") The encoding algorithm determines how cached results are compressed in memory, trading CPU for memory efficiency. Currrently supported for SQL results only. * **`none` (Default):** Stores query results uncompressed. Uses more memory but has zero compression overhead. Best for small result sets or when memory is not a constraint. * **`zstd`:** Uses the [Zstandard compression algorithm](https://facebook.github.io/zstd/) to compress cached query results. Provides high compression ratios (often 50-90% reduction) with fast decompression speeds. Recommended when caching large result sets to maximize cache capacity. Use `zstd` when maximizing cache efficiency is important, especially for large queries that would otherwise quickly fill the cache. Use `none` for the lowest latency when memory is not constrained. ## Cached Responses[​](#cached-responses "Direct link to Cached Responses") Responses from HTTP APIs include a header that indicates the cache status of the applicable cache: | Cache | Header Key | | ---------------- | ----------------------------- | | `sql_results` | `Results-Cache-Status` | | `search_results` | `Search-Results-Cache-Status` | The value of the header indicates the status of the cache: | Header value | Description | | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | `HIT` | The query result was served from the cache. | | `MISS` | The cache was checked, but the result was not found. | | `BYPASS` | The cache was bypassed for this query (e.g., when `cache-control: no-cache` is specified). | | `STALE` | A stale cache entry was served while the cache is being revalidated in the background (when `stale_while_revalidate_ttl` is configured). | | *header not present* | The cache did not apply to this query (e.g., when caching is disabled or querying a system table). | ### Examples[​](#examples "Direct link to Examples") #### Cached Response[​](#cached-response "Direct link to Cached Response") ``` $ curl -XPOST -i http://localhost:8090/v1/sql -d 'select * from taxi_trips limit 1;' HTTP/1.1 200 OK content-type: text/plain; charset=utf-8 results-cache-status: HIT vary: origin, access-control-request-method, access-control-request-headers content-length: 416 date: Thu, 13 Feb 2025 03:05:39 GMT ``` #### Uncached Response[​](#uncached-response "Direct link to Uncached Response") ``` $ curl -XPOST -i http://localhost:8090/v1/sql -d 'select * from taxi_trips limit 1;' HTTP/1.1 200 OK content-type: text/plain; charset=utf-8 results-cache-status: MISS vary: origin, access-control-request-method, access-control-request-headers content-length: 416 date: Thu, 13 Feb 2025 03:13:19 GMT ``` #### Bypassed Cache with `cache-control: no-cache`[​](#bypassed-cache-with-cache-control-no-cache "Direct link to bypassed-cache-with-cache-control-no-cache") ``` $ curl -H "cache-control: no-cache" -XPOST -i http://localhost:8090/v1/sql -d 'select * from taxi_trips limit 1;' HTTP/1.1 200 OK content-type: text/plain; charset=utf-8 results-cache-status: BYPASS vary: origin, access-control-request-method, access-control-request-headers content-length: 416 date: Thu, 13 Feb 2025 03:14:00 GMT ``` #### Stale Cache Response (Stale-While-Revalidate)[​](#stale-cache-response-stale-while-revalidate "Direct link to Stale Cache Response (Stale-While-Revalidate)") ``` $ curl -XPOST -i http://localhost:8090/v1/sql -d 'select * from taxi_trips limit 1;' HTTP/1.1 200 OK content-type: text/plain; charset=utf-8 results-cache-status: STALE vary: origin, access-control-request-method, access-control-request-headers content-length: 416 date: Thu, 13 Feb 2025 03:15:30 GMT ``` ## Cache Control[​](#cache-control "Direct link to Cache Control") You can control caching behavior for specific requests using HTTP headers. The `Cache-Control` header helps skip the cache for a request while caching the results for subsequent requests. ### Stale-While-Revalidate[​](#stale-while-revalidate "Direct link to Stale-While-Revalidate") The `stale_while_revalidate_ttl` parameter configures a grace period during which stale cache entries continue to be served while a background refresh occurs. This technique reduces latency for end users by serving cached data immediately, even after `item_ttl` expires, while the system fetches fresh data asynchronously. When `stale_while_revalidate_ttl` is set to a non-zero value: 1. Cache entries are served normally until `item_ttl` expires. 2. After `item_ttl` expires but before `item_ttl + stale_while_revalidate_ttl` expires, the stale entry is served immediately with a `STALE` cache status. 3. Simultaneously, a background task refreshes the cache entry. 4. Once the background refresh completes, subsequent requests receive the fresh data with a `HIT` cache status. 5. After `item_ttl + stale_while_revalidate_ttl` expires, the entry is evicted and the next request results in a `MISS`. #### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` runtime: caching: sql_results: enabled: true item_ttl: 10s stale_while_revalidate_ttl: 10s ``` With this configuration: * Fresh cache entries are served for 10 seconds after creation. * Between 10-20 seconds after creation, stale entries are served while being refreshed in the background. * After 20 seconds, the entry is evicted if not refreshed. This approach is particularly useful for queries that take significant time to execute, providing a better user experience by reducing perceived latency while keeping data reasonably fresh. Conflict with Caching Accelerator SWR When using a dataset with `refresh_mode: caching`, you cannot configure both the results cache's `stale_while_revalidate_ttl` and the caching accelerator's `caching_stale_while_revalidate_ttl` for the same dataset. These parameters control similar behavior at different layers. Choose one approach: * **Results cache SWR**: Configure `runtime.caching.sql_results.stale_while_revalidate_ttl` for SQL query results caching * **Caching accelerator SWR**: Configure `acceleration.params.caching_stale_while_revalidate_ttl` for [HTTP-based dataset caching](/docs/v1.10/features/data-acceleration/refresh-modes/caching) ### HTTP/Flight API[​](#httpflight-api "Direct link to HTTP/Flight API") The following endpoints support the standard HTTP [`Cache-Control` header](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control): * SQL query (HTTP and Arrow Flight) * Search (HTTP) The following `Cache-Control` directives are supported: | Directive | Description | | ---------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [`no-cache`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control#no-cache) | Skips the cache for the current request but caches the results for future requests. | | [`min-fresh`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control#min-fresh) | Specifies the minimum time (in seconds) that a cached response must remain fresh. For example, `min-fresh=60` requires the cached entry to be fresh for at least 60 more seconds. | | [`max-stale`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control#max-stale) | Indicates the client will accept a stale response. An optional value in seconds specifies the maximum staleness allowed. For example, `max-stale=30` accepts responses stale for up to 30 seconds. | | [`only-if-cached`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control#only-if-cached) | Returns only cached responses. If no cached response is available, returns an error instead of fetching fresh data. | | [`stale-if-error`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control#stale-if-error) | Serves stale cached responses if an error occurs while fetching fresh data. An optional value in seconds specifies how stale the response can be. For example, `stale-if-error=600` serves responses stale for up to 10 minutes if fetching fails. | #### HTTP Example[​](#http-example "Direct link to HTTP Example") ``` # Default behavior (uses cache) curl -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' # Skip cache for this query, but cache the results for future queries curl -H "cache-control: no-cache" -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' # Only use cached response if it will be fresh for at least 30 more seconds curl -H "cache-control: min-fresh=30" -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' # Accept cached responses that are stale for up to 60 seconds curl -H "cache-control: max-stale=60" -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' # Only return cached responses, fail if cache miss curl -H "cache-control: only-if-cached" -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' # Serve stale cache (up to 300 seconds old) if fetching fresh data fails curl -H "cache-control: stale-if-error=300" -XPOST http://localhost:8090/v1/sql -d 'SELECT 1' ``` #### Arrow FlightSQL Example[​](#arrow-flightsql-example "Direct link to Arrow FlightSQL Example") The following example skips the cache for a specific query using FlightSQL in Rust: ``` let sql_command = arrow_flight::sql::CommandStatementQuery { query: "SELECT 1".to_string(), transaction_id: None, }; let sql_command_bytes = sql_command.as_any().encode_to_vec(); let mut request = FlightDescriptor::new_cmd(sql_command_bytes).into_request(); request .metadata_mut() .insert("cache-control", "no-cache"); // Send the request ``` The cache can be controlled using JDBC properties. For example, ``` Properties props = new Properties(); props.setProperty("cache-control", "no-cache"); Connection conn = DriverManager.getConnection("jdbc:arrow-flight-sql://localhost:50051", props); ``` ### `spice` CLI[​](#spice-cli "Direct link to spice-cli") The `spice sql` and `spice search` commands accept a `--cache-control` flag that supports all cache-control directives: ``` # Default behavior (use cache if available) spice sql # Same as above spice sql --cache-control cache # Skip cache for this query, but cache the results for future queries spice sql --cache-control no-cache # Only use cached response if fresh for at least 30 more seconds spice sql --cache-control min-fresh=30 # Accept cached responses stale for up to 60 seconds spice sql --cache-control max-stale=60 # Only return cached responses, fail if cache miss spice sql --cache-control only-if-cached # Serve stale cache (up to 300 seconds) if fetching fails spice sql --cache-control stale-if-error=300 # Default behavior (use cache if available) spice search # Same as above spice search --cache-control cache # Skip cache for this search, but cache the results for future searches spice search --cache-control no-cache # Accept stale search results up to 60 seconds old spice search --cache-control max-stale=60 ``` ## Custom Cache Keys[​](#custom-cache-keys "Direct link to Custom Cache Keys") Set the `Spice-Cache-Key` header to supply a custom cache key. When set, a supplied cache key takes precedence over `caching.sql_results.cache_key_type`. Info A valid cache key consists of up to 128 alphanumeric characters (and the characters `-` and `_`). ### HTTP Example[​](#http-example-1 "Direct link to HTTP Example") Consider the case of two semantically equivalent queries: ``` Time: 0.0251325 seconds. 2 rows. sql> select * from users where org_id = 1; +----+--------+-------+----------------+ | id | org_id | name | email | +----+--------+-------+----------------+ | 1 | 1 | Jane | jane@spice.ai | | 2 | 1 | Sarah | sarah@spice.ai | +----+--------+-------+----------------+ Time: 0.008993042 seconds. 2 rows. sql> select * from users where split_part(email, '@', 2) = 'spice.ai'; +----+--------+-------+----------------+ | id | org_id | name | email | +----+--------+-------+----------------+ | 1 | 1 | Jane | jane@spice.ai | | 2 | 1 | Sarah | sarah@spice.ai | +----+--------+-------+----------------+ ``` To share a cache key for these queries, set `Spice-Cache-Key`. The first request is a cache miss: ``` $ curl -i -XPOST http://localhost:8090/v1/sql -H"spice-cache-key: users_spiceai" -d "select * from users where org_id = 1;" HTTP/1.1 200 OK content-type: application/json x-cache: Miss from spiceai results-cache-status: MISS vary: Spice-Cache-Key vary: origin, access-control-request-method, access-control-request-headers content-length: 119 date: Thu, 24 Jul 2025 14:15:53 GMT [{"id":1,"org_id":1,"name":"Jane","email":"jane@spice.ai"},{"id":2,"org_id":1,"name":"Sarah","email":"sarah@spice.ai"}] ``` The subsequent request with the different (but semantically equivalent) query is a cache hit: ``` $ curl -i -XPOST http://localhost:8090/v1/sql -H"spice-cache-key: users_spiceai" -d "select * from users where split_part(email, '@', 2) = 'spice.ai';" HTTP/1.1 200 OK content-type: application/json x-cache: Hit from spiceai results-cache-status: HIT vary: Spice-Cache-Key vary: origin, access-control-request-method, access-control-request-headers content-length: 119 date: Thu, 24 Jul 2025 14:18:00 GMT ``` Note When supplying a custom cache key, **ensure the semantic equivalence of queries**. For example, this is expected behavior: ``` $ curl -i -XPOST http://localhost:8090/v1/sql -H"spice-cache-key: users_spiceai" -d "select 1" HTTP/1.1 200 OK content-type: application/json x-cache: Hit from spiceai results-cache-status: HIT vary: Spice-Cache-Key vary: origin, access-control-request-method, access-control-request-headers content-length: 119 date: Thu, 24 Jul 2025 14:21:32 GMT [{"id":1,"org_id":1,"name":"Jane","email":"jane@spice.ai"},{"id":2,"org_id":1,"name":"Sarah","email":"sarah@spice.ai"}] ``` ## Metrics[​](#metrics "Direct link to Metrics") Cache metrics can be monitored using the [Prometheus-compatible Metrics Endpoint](/docs/v1.10/features/observability). The following metrics are available for each cache type: | Metric | Type | Description | | ------------------------ | ------- | ---------------------------------------------------------- | | `*_cache_max_size_bytes` | Gauge | Maximum configured cache size in bytes. | | `*_cache_requests` | Counter | Total number of cache lookup requests. | | `*_cache_hits` | Counter | Total number of cache hits. | | `*_cache_items_count` | Gauge | Current number of items in the cache. | | `*_cache_size_bytes` | Gauge | Current cache size in bytes. | | `*_cache_evictions` | Counter | Total number of cache evictions due to size or TTL limits. | | `*_cache_hit_ratio` | Gauge | Current cache hit ratio (hits / total requests). | The `*` prefix corresponds to the cache type: * `results_*` - SQL query results cache metrics * `search_results_*` - Search results cache metrics * `embeddings_*` - Embeddings cache metrics Example metrics output: ``` # HELP results_cache_evictions Number of cache evictions. # TYPE results_cache_evictions counter results_cache_evictions 2 # HELP results_cache_hit_ratio Cache hit ratio (hits / total requests). # TYPE results_cache_hit_ratio gauge results_cache_hit_ratio 0.625 # HELP results_cache_hits Cache hit count. # TYPE results_cache_hits counter results_cache_hits 14 # HELP results_cache_items_count Number of items currently in the cache. # TYPE results_cache_items_count gauge results_cache_items_count 1 # HELP results_cache_max_size_bytes Maximum allowed size of the cache in bytes. # TYPE results_cache_max_size_bytes gauge results_cache_max_size_bytes 134217728 # HELP results_cache_misses Cache miss count. # TYPE results_cache_misses counter results_cache_misses 4 # HELP results_cache_requests Number of requests to get a key from the cache. # TYPE results_cache_requests counter results_cache_requests 18 # HELP results_cache_size_bytes Size of the cache in bytes. # TYPE results_cache_size_bytes gauge results_cache_size_bytes 7776 ``` --- # Change Data Capture (CDC) Change Data Capture (CDC) captures changed rows from a database's transaction log and delivers them to consumers with low latency. This technique enables Spice to keep [locally accelerated](/docs/v1.10/features/data-acceleration) datasets up-to-date in real time with the source data. It is efficient because it only transfers the changed rows instead of re-fetching the entire dataset. ## Benefits[​](#benefits "Direct link to Benefits") Using locally accelerated datasets configured with CDC enables Spice to provide high-performance accelerated queries and efficient real-time updates. ## Example Use Case[​](#example-use-case "Direct link to Example Use Case") Consider a fraud detection application that needs to determine whether a pending transaction is likely fraudulent. The application queries a Spice-accelerated, real-time updated table of recent transactions to check if a pending transaction resembles known fraudulent ones. With CDC, the table is kept up-to-date, allowing the application to quickly identify potential fraud. ## Considerations[​](#considerations "Direct link to Considerations") When configuring datasets to be accelerated with CDC, ensure that the [data connector](/docs/v1.10/components/data-connectors) supports CDC and can return a stream of row-level changes. See the [Supported Data Connectors](#supported-data-connectors) section for more information. The startup time for CDC-accelerated datasets may be longer than for non-CDC-accelerated datasets due to the initial synchronization. tip It is recommended to use CDC-accelerated datasets with persistent data accelerator configurations (i.e., `file` mode for [`DuckDB`](/docs/v1.10/components/data-accelerators/duckdb)/[`SQLite`](/docs/v1.10/components/data-accelerators/sqlite) or [`PostgreSQL`](/docs/v1.10/components/data-accelerators/postgres)). This ensures that when Spice restarts, it can resume from the last known state of the dataset instead of re-fetching the entire dataset. ## Supported Data Connectors[​](#supported-data-connectors "Direct link to Supported Data Connectors") Enabling CDC by setting `refresh_mode: changes` in the acceleration settings requires support from the data connector to provide a stream of row-level changes. Currently, the only supported data connector is [Debezium](/docs/v1.10/components/data-connectors/debezium). ## Example[​](#example "Direct link to Example") See an example of configuring a dataset to use CDC with Debezium by following the recipe at [Streaming changes in real-time with Debezium CDC](https://github.com/spiceai/cookbook/tree/trunk/cdc-debezium#readme). ``` version: v1 kind: Spicepod name: cdc-debezium datasets: - from: debezium:cdc.public.customer_addresses name: cdc params: debezium_transport: kafka debezium_message_format: json kafka_bootstrap_servers: localhost:19092 kafka_security_protocol: PLAINTEXT acceleration: enabled: true engine: sqlite mode: file refresh_mode: changes ``` --- # Data Acceleration Datasets and views can be locally accelerated by the Spice runtime, pulling data from any [Data Connector](/docs/v1.10/components/data-connectors) and storing it locally in a [Data Accelerator](/docs/v1.10/components/data-accelerators) for faster access. The data can be kept up-to-date in real-time or on a refresh schedule, ensuring deployments maintain the latest data locally for querying. ![Spice.ai Open Source Query Federation with Acceleration](/assets/images/data-acceleration-8a4401ca91dc42fc622704dd11d962c1.png) ## Benefits[​](#benefits "Direct link to Benefits") Local data acceleration stores data alongside the application, providing faster query times by eliminating network latency. This is especially beneficial for large query results, as data transfer over the network is avoided. Depending on the [Acceleration Engine](/docs/v1.10/components/data-accelerators) used, data can also be stored in-memory, further reducing query times. [Indexes](/docs/v1.10/features/data-acceleration/indexes) can be applied to speed up certain queries. Locally accelerated datasets can also have [primary key constraints](/docs/v1.10/features/data-acceleration/constraints) applied. This feature allows specifying actions when a constraint is violated, such as dropping the violating row or upserting it into the accelerated table. [Acceleration snapshots](/docs/v1.10/features/data-acceleration/snapshots) (preview) help file-mode accelerations become ready in seconds by bootstrapping from managed snapshots stored in object storage such as Amazon S3. ## Example Use Case[​](#example-use-case "Direct link to Example Use Case") Consider a high-volume e-trading frontend application backed by an AWS RDS database containing a table of trades. To retrieve all trades over the last 24 hours, the application would need to query the remote database and transfer the data over the network. By accelerating the trades table locally using the [AWS RDS Data Connector](https://github.com/spiceai/cookbook/tree/trunk/mysql/rds-aurora#readme), the data is brought to the application, saving round trip time and data transfer time. ## Considerations[​](#considerations "Direct link to Considerations") Data Storage: Ensure local storage has enough capacity for the accelerated data. The required storage type (Disk or RAM) and amount depend on the dataset size and the acceleration engine used. Data Security: Assess data sensitivity and secure network connections between the edge and data connector when replicating data. Secure any external Data Accelerator connected to the Spice runtime with encryption, access controls, and secure protocols. ## Example[​](#example "Direct link to Example") ### Locally Accelerating taxi\_trips[​](#locally-accelerating-taxi_trips "Direct link to Locally Accelerating taxi_trips") * Start Spice with the following dataset: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true refresh_mode: full refresh_check_interval: 10s ``` * The dataset `taxi_trips` is accelerated locally by the Spice runtime. The data refreshes every 10 seconds. * Query times can be compared against the Spice platform: ``` curl \ --url 'https://data.spiceai.io/v1/sql?api_key=[API_KEY]' \ --data 'select * from taxi_trips' ``` The locally accelerated dataset can then be queried locally: ``` spice sql select * from taxi_trips; ``` Example output: ``` +---------------+--------------+------------------+ | trip_distance | total_amount | tpep_pickup_time | +---------------+--------------+------------------+ | 1.2 | 9.80 | 2023-01-15 08:32 | | 3.4 | 18.50 | 2023-01-15 09:10 | +---------------+--------------+------------------+ Time: 0.012s. 2 rows. ``` Locally accelerated datasets provide significantly faster query times compared to remote sources. [Learn more about Data Accelerators](/docs/v1.10/components/data-accelerators) for faster access. --- # Constraints Constraints enforce data integrity in a database. Spice supports constraints on locally accelerated tables to ensure data quality and configure behavior for data updates that violate constraints. Constraints are specified using [column references](#column-references) in the Spicepod via the `primary_key` field in the acceleration configuration. Additional unique constraints are specified via the [`indexes`](/docs/v1.10/features/data-acceleration/indexes) field with the value `unique`. Data that violates these constraints will result in a [conflict](#handling-conflicts). If multiple rows in the incoming data violate any constraint, the entire incoming batch of data will be dropped. Example Spicepod: ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash # Define a primary key on the `hash` column indexes: '(number, timestamp)': unique # Add a unique index with a multicolumn key comprised of the `number` and `timestamp` columns ``` ## Column References[​](#column-references "Direct link to Column References") Column references can be used to specify which columns are part of the constraint. The column reference can be a single column name or a multicolumn key. The column reference must be enclosed in parentheses if it is a multicolumn key. Examples * `number`: Reference a constraint on the `number` column * `(hash, timestamp)`: Reference a constraint on the `hash` and `timestamp` columns ## Handling conflicts[​](#handling-conflicts "Direct link to Handling conflicts") The behavior of inserting data that violates the constraint can be configured via the `on_conflict` field to either `drop` the data that violates the constraint or `upsert` that data into the accelerated table (i.e. update all values other than the columns that are part of the constraint to match the incoming data). warning If there are multiple rows in the incoming data that violate any constraint, the entire incoming batch of data will be dropped. Example Spicepod: ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash # Define a primary key on the `hash` column indexes: '(number, timestamp)': unique # Add a unique index with a multicolumn key comprised of the `number` and `timestamp` columns on_conflict: # Upsert the incoming data when the primary key constraint on "hash" is violated, # alternatively "drop" can be used instead of "upsert" to drop the data update. hash: upsert ``` ### Advanced Upsert Options[​](#advanced-upsert-options "Direct link to Advanced Upsert Options") By default, even when `upsert` is configured, if there are constraint violations, such as duplicates within the same batch of ingested data, it will result in a constraint violation - as attempting to upsert data into the target acceleration engine results in an error if done in a single statement. (i.e. [PostgreSQL does not allow the same row to be proposed for insertion more than once](https://www.postgresql.org/docs/18/sql-insert.html)) Spice provides two `upsert` options to resolve duplicates within a single update: * `upsert_dedup`: Removes exact duplicates in the incoming batch if there is a constraint violation. (i.e. the equivalent of running `SELECT DISTINCT * FROM [batch]`) * `upsert_dedup_by_row_id`: Resolves conflicts by taking the row with the greatest row id. This is the behavior that would occur if the upsert were applied row-by-row. This guarantees that no constraint violations would result in an error, but it has the tradeoff of being effectively "random" if the incoming data is not ordered. The new behavior is only triggered when an incoming batch has a constraint violation, minimizing the effect of applying these computations to only when its necessary. However, they can have a performance impact and are not enabled by default. Full configuration example: ``` acceleration: enabled: true engine: duckdb mode: file primary_key: id on_conflict: id: upsert_dedup # upsert_dedup_by_row_id ``` Examples for advanced upsert behavior Take these two CSV files: `one.csv`: ``` foo,bar a,1 b,2 a,1 ``` Behavior on `one.csv` with a primary key on `foo` and `on_conflict` set to: * `upsert`: Will error with: `Constraint Violation: Incoming data violates uniqueness constraint on column(s): foo` * `upsert_dedup`: Will succeed in loading 2 rows, the `a,1` row is reduced to a single instance. * `upsert_dedup_by_row_id`: Same as `upsert_dedup` `two.csv`: ``` foo,bar a,1 b,2 a,10 ``` Behavior on `one.csv` with a primary key on `foo` and `on_conflict` set to: * `upsert`: Will error with: `Constraint Violation: Incoming data violates uniqueness constraint on column(s): foo` * `upsert_dedup`: Will error with: `Constraint Violation: Incoming data violates uniqueness constraint on column(s): foo` * `upsert_dedup_by_row_id`: Will succeed in loading 2 rows, `a,10` and `b,2`. The primary key violation is resolved to the row that occurred later. ## Limitations[​](#limitations "Direct link to Limitations") * **Single on\_conflict target supported**: Only a single `on_conflict` target can be specified, unless all `on_conflict` targets are specified with drop. * Examples for valid/invalid `on_conflict` targets The following Spicepod is invalid because it specifies multiple `on_conflict` targets with `upsert`: Invalid ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash indexes: '(number, timestamp)': unique on_conflict: hash: upsert '(number, timestamp)': upsert ``` The following Spicepod is valid because it specifies multiple `on_conflict` targets with `drop`, which is allowed: Valid ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash indexes: '(number, timestamp)': unique on_conflict: hash: drop '(number, timestamp)': drop ``` The following Spicepod is invalid because it specifies multiple `on_conflict` targets with `upsert` and `drop`: Invalid ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite primary_key: hash indexes: '(number, timestamp)': unique on_conflict: hash: upsert '(number, timestamp)': drop ``` * **DuckDB Limitations:** * DuckDB does not support `upsert` for datasets with List or Map types. * Standard indexes unexpectedly act like unique indexes and block updates when `upsert` is configured. * Standard indexes blocking updates The following Spicepod specifies a standard index on the `number` column, which blocks updates when `upsert` is configured for the `hash` column: ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: duckdb primary_key: hash indexes: number: enabled on_conflict: hash: upsert ``` The following error is returned when attempting to upsert data into the `eth.recent_blocks` table: ``` ERROR runtime::accelerated_table::refresh: Error adding data for eth.recent_blocks: External error: Unable to insert into duckdb table: Binder Error: Can not assign to column 'number' because it has a UNIQUE/PRIMARY KEY constraint ``` This is a limitation of DuckDB. --- # Data Refresh Acceleration data can be refreshed (updated) by: * **API**: POST to `/v1/datasets/:name/acceleration/refresh`. See [Refresh Dataset HTTP API](/docs/v1.10/api/HTTP/post-dataset-refresh). * **Interval**: Time-based refresh interval. See [Refresh Interval](#refresh-on-demand). * **Change Data Capture (CDC)**: CDC from a database using Debezium. See [Change Data Capture](/docs/v1.10/features/cdc). * **Push**: Spice-to-Spice Push over Apache Arrow DoExchange. ![Spice.ai Open Source Acceleration Refresh](/assets/images/acceleration-refresh-28183446549fbcc17f21f07881cfe7d9.png). ## Refresh Modes[​](#refresh-modes "Direct link to Refresh Modes") Spice supports four modes to refresh/update local data from a connected data source. `full` is the default mode. | Mode | Description | Example | | --------- | ---------------------------------------------------- | ---------------------------------------------------------------- | | `full` | Replace/overwrite the entire dataset on each refresh | A table of users | | `append` | Append/add data to the dataset on each refresh | Append-only, immutable datasets, such as time-series or log data | | `changes` | Apply incremental changes | Customer order lifecycle table | | `caching` | Read-through caching for SQL queries | API search results or dynamic content endpoints | Learn more about each mode: * [Caching Mode](/docs/v1.10/features/data-acceleration/refresh-modes/caching) Example: ``` datasets: - from: databricks:my_dataset name: accelerated_dataset acceleration: refresh_mode: full refresh_check_interval: 10m ``` ### Append[​](#append "Direct link to Append") Using `refresh_mode: append` requires the use of a [`time_column` dataset parameter](/docs/v1.10/reference/spicepod/datasets#time_column), specifying a column to compare the local acceleration against the remote source. Data will be incrementally refreshed where the `time_column` value in the remote source is greater-than (gt) the `max(time_column)` value in the local acceleration. E.g. ``` datasets: - from: databricks:my_dataset name: accelerated_dataset time_column: created_at acceleration: refresh_mode: append refresh_check_interval: 10m ``` Readiness with snapshots Append-mode accelerations that define a `time_column` wait to report ready until the first append refresh completes after [snapshot bootstrap](/docs/v1.10/features/data-acceleration/snapshots). This keeps the dataset out of rotation until the freshest data is available while still benefiting from the snapshot-assisted startup. If late arriving data or clock-skew needs to be accounted for, an optional overlap can also be specified. See [`acceleration.refresh_append_overlap`](/docs/v1.10/reference/spicepod/datasets#accelerationrefresh_append_overlap). #### `time_partition_column`[​](#time_partition_column "Direct link to time_partition_column") Datasets that are partitioned by a less-granular time-column (e.g. day, month, year) can also use the `time_partition_column` parameter in addition to the `time_column` parameter to specify the time-column to use for efficient partition pruning. Example: ``` datasets: - from: databricks:my_dataset name: accelerated_dataset time_column: created_at time_format: iso8601 time_partition_column: created_at_day time_partition_format: date ``` #### Append only modified files[​](#append-only-modified-files "Direct link to Append only modified files") Spice can automatically detect and append only newly created or updated files from object-store data sources. This is useful for append-only datasets where only new files are added to the source and existing files are not modified or deleted. Enable this feature by setting either `time_column` or `time_partition_column` to the special value `last_modified`. When configured this way with `refresh_mode: append`, Spice will use the file/object's metadata to determine which files are new or have been updated. This approach can drastically speed up incremental updates for large datasets, as Spice only needs to process the new files rather than scanning the entire dataset for changes to a column. This optimization is particularly valuable for datasets with many files or large file sizes. If `last_modified` already exists as a column in the parquet data, that column will take precedence over the metadata value from the file itself. Example using `time_column`: ``` datasets: - from: s3://my_bucket/my_dataset name: accelerated_dataset time_column: last_modified params: file_format: parquet acceleration: refresh_mode: append refresh_check_interval: 10m ``` Example using `time_partition_column`: ``` datasets: - from: s3://my_bucket/my_dataset name: accelerated_dataset time_column: created_at time_partition_column: last_modified params: file_format: parquet acceleration: refresh_mode: append refresh_check_interval: 10m ``` info Appending modified files is only supported for datasets that support setting the [file format parameter](/docs/v1.10/reference/file_format), such as `s3://`, `abfs://`, `file://`, etc. ### Changes (CDC)[​](#changes-cdc "Direct link to Changes (CDC)") Datasets configured with acceleration `refresh_mode: changes` requires a [Change Data Capture (CDC)](/docs/v1.10/features/cdc) supported data connector. Initial CDC support in Spice is supported by the [Debezium data connector](/docs/v1.10/components/data-connectors/debezium). ### Caching[​](#caching "Direct link to Caching") The `caching` refresh mode is designed for HTTP-based datasets where request metadata acts as cache keys. This mode is particularly useful for API responses that return multiple rows for a single request, such as search results or dynamic content endpoints. See [Caching Mode](/docs/v1.10/features/data-acceleration/refresh-modes/caching) for detailed documentation and examples. ## Ready State[​](#ready-state "Direct link to Ready State") | | | | --------------------------- | --------- | | Supported in `refresh_mode` | Any | | Required | No | | Default Value | `on_load` | By default, Spice will return an error for queries against an accelerated dataset that is still loading its initial data. The endpoint [`/v1/ready`](/docs/v1.10/api/HTTP/ready) is used in production deployments to control when queries are sent to the Spice runtime. The ready state for an accelerated dataset can be configured using the [`ready_state`](/docs/v1.10/reference/spicepod/datasets#ready_state) parameter in the dataset configuration. * `ready_state: on_load`: Default. The dataset is considered ready after the initial load of the accelerated data. For file-based accelerated datasets that have existing data, this will be ready immediately. Queries against this dataset before the data is loaded will return an error. * `ready_state: on_registration`: The dataset is considered ready when the dataset is registered in Spice, even before the initial data is loaded. Queries against this dataset before the data is loaded will automatically fallback to the federated source. Once the data is loaded, queries will be served from the acceleration. Example: ``` datasets: - from: s3://my_bucket/my_dataset name: my_dataset ready_state: on_load # or on_registration acceleration: enabled: true ``` ## Fast Cold Starts with Snapshots[​](#fast-cold-starts-with-snapshots "Direct link to Fast Cold Starts with Snapshots") File-based acceleration engines (DuckDB, SQLite, or Turso) can rely on [acceleration snapshots](/docs/v1.10/features/data-acceleration/snapshots) to download a pre-built database file on startup instead of waiting for the first refresh to finish. Configure a shared snapshot location under the top-level `snapshots` block and opt individual datasets in with `acceleration.snapshots: enabled`, `bootstrap_only`, or `create_only`. Snapshots are stored using Hive-style partitions (`month=YYYY-MM/day=YYYY-MM-DD/dataset=`) and are only supported when each dataset writes to its own acceleration file. ## Filtered Refresh[​](#filtered-refresh "Direct link to Filtered Refresh") Typically only a working subset of an entire dataset is used in an application or dashboard. Use these features to filter refresh data, creating a smaller subset for faster processing and to reduce the data transferred and stored locally. * [Refresh SQL](#refresh-sql) - Specify the filter as arbitrary SQL to be pushed down to the remote source. * [Refresh Data Window](#refresh-data-window) - Filters data from the remote source outside the specified time window. ### Refresh SQL[​](#refresh-sql "Direct link to Refresh SQL") | | | | --------------------------- | ----- | | Supported in `refresh_mode` | Any | | Required | No | | Default Value | Unset | Refresh SQL supports specifying filters for data accelerated from the connected source using arbitrary SQL. Filters will be pushed down to the remote source when possible, so only the requested data will be transferred over the network. Example: ``` datasets: - from: databricks:my_dataset name: accelerated_dataset acceleration: enabled: true refresh_mode: full refresh_check_interval: 10m refresh_sql: | SELECT * FROM accelerated_dataset WHERE city = 'Seattle' ``` The `refresh_sql` parameter can be updated at runtime on-demand using `PATCH /v1/datasets/:name/acceleration`. This change is temporary and will revert to the `spicepod.yml` definition at the next runtime restart. Columns can be selected in the query via the `SELECT` clause, but only column names are supported. Arbitrary expressions or aliases are not supported. Example: ``` curl -i -X PATCH \ -H "Content-Type: application/json" \ -d '{ "refresh_sql": "SELECT city, state FROM accelerated_dataset WHERE city = 'Bellevue'" }' \ 127.0.0.1:8090/v1/datasets/accelerated_dataset/acceleration ``` Queries that return zero results will fallback to the behavior specified by the [`on_zero_results` parameter](#behavior-on-zero-results), and will not have the `refresh_sql` applied to the results from the fallback. The `refresh_sql` only applies to acceleration refresh tasks. For the complete reference, view the `refresh_sql` section of [datasets](/docs/v1.10/reference/spicepod/datasets#accelerationrefresh_sql). Limitations * When `refresh_mode: changes` is specified, Refresh SQL can only modify the selected columns and cannot apply filters. * Running queries while using refresh SQL will not fallback to the source if any query returns more than zero rows, even when querying against columns that are not explicitly filtered by the refresh SQL. This may result in queries returning partial data, depending on the filters applied in the refresh SQL. * Refresh SQL only supports filtering data from the current dataset - joining across other datasets is not supported. * Refresh SQL modifications made via API are temporary and will revert after a runtime restart. * Not supported by [DynamoDB Streams Data Connector](/docs/v1.10/components/data-connectors/dynamodb#streams). ### Refresh Data Window[​](#refresh-data-window "Direct link to Refresh Data Window") | | | | --------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default Value | Unset | The `refresh_data_window` parameter supports refreshing data that falls within the specified time window. The `refresh_data_window` is applied cumulatively to any filters specified by the [`refresh_sql`](#refresh-sql), and applies a time filter based on `now() - refresh_data_window`. For example, the following configuration: ``` time_column: column_time acceleration: refresh_sql: "SELECT * FROM my_dataset WHERE column_one = 'value'" refresh_data_window: 1d ``` In this example, `refresh_data_window` is converted into an effective Refresh SQL of `SELECT * FROM my_dataset WHERE column_one = 'value' AND column_time > (now() - interval '1' day)`. The `time_column` column can be specified in the `refresh_sql` in conjunction with the `refresh_data_window`, and both filters are combined with `AND`. This parameter relies on the `time_column` dataset parameter specifying a column that is a timestamp type. Optionally, the `time_format` can be specified to instruct the Spice runtime on how to interpret timestamps in the `time_column`. *Example with `refresh_sql`:* ``` datasets: - from: databricks:my_dataset name: accelerated_dataset time_column: created_at acceleration: enabled: true refresh_mode: full refresh_check_interval: 10m refresh_sql: | SELECT * FROM accelerated_dataset WHERE city = 'Seattle' refresh_data_window: 1d ``` This example will only accelerate data from the federated source that matches the filter `city = 'Seattle'` and is less than 1 day old. *Example with `on_zero_results`:* ``` datasets: - from: databricks:my_dataset name: accelerated_dataset time_column: created_at acceleration: enabled: true refresh_mode: full refresh_check_interval: 10m refresh_sql: | SELECT * FROM accelerated_dataset WHERE city = 'Seattle' refresh_data_window: 1d on_zero_results: use_source ``` This example will only accelerate data from the federated source that matches the filter `city = 'Seattle'` and is less than 1 day old. If a query against the accelerated data returns zero results, the query will fallback to the source and return the direct results without any filtering. If a query against the accelerated data returns some results, the query will not fall back. For example, attempting to query for the last 2 days of data would only return the last 1 day of data without falling back. ## Behavior on Zero Results[​](#behavior-on-zero-results "Direct link to Behavior on Zero Results") | | | | --------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default Value | `return_empty` | By default, accelerated datasets only return locally materialized data. If this local data is a subset of the full dataset in the federated source—due to settings like `refresh_sql`, `refresh_data_window`, or retention policies—queries against the accelerated dataset may return zero results, even when the federated table would return results. To address this, `on_zero_results: use_source` can be configured in the acceleration configuration. Queries returning zero results will fall back to the federated source, returning results from querying the underlying data. `on_zero_results`: * `return_empty` (Default) - Return an empty result set when no data is found in the accelerated dataset. * `use_source` - Fall back to querying the federated table when no data is found in the accelerated dataset. Example: ``` datasets: - from: databricks:my_dataset name: accelerated_dataset acceleration: enabled: true refresh_sql: SELECT * FROM accelerated_dataset where city = 'Seattle' on_zero_results: use_source ``` In this example a query against `accelerated_dataset` within Spice like `SELECT * FROM accelerated_dataset WHERE city = 'Portland'` would initially query against the accelerated data, see that it returns zero results and then fallback to querying against the federated table in Databricks. warning * It is possible that even though an accelerated table returns some results, it may not contain all the data that would be returned by the federated table. `on_zero_results` only controls the behavior in the simple case where no data is returned by the acceleration for a given query. ## Refresh on Startup[​](#refresh-on-startup "Direct link to Refresh on Startup") | Parameter | Value | | --------------------------- | ------ | | Supported in `refresh_mode` | Any | | Required | No | | Default Value | `auto` | Controls the refresh behavior of an accelerated dataset across restarts. `refresh_on_startup` Options: * `auto` (Default) – Maintains refresh state across restarts: * With `refresh_check_interval`: Schedules next refresh based on last successful refresh time, triggering immediately if interval has already elapsed * Without `refresh_check_interval`: No refresh (on-demand only) * `always` – Forces a dataset refresh on every startup, regardless of the existing acceleration state. Setting `refresh_on_startup: always` ensures that accelerated data is always refreshed to match the source when the service restarts. This is useful in **development environments** or when **data consistency is critical** after deployment. Example Configuration: ``` datasets: - from: databricks:my_dataset name: accelerated_dataset acceleration: enabled: true refresh_on_startup: always ``` For the complete reference, view the `refresh_on_startup` section of [datasets](/docs/v1.10/reference/spicepod/datasets#accelerationrefresh_on_startup). ## Refresh Interval[​](#refresh-interval "Direct link to Refresh Interval") | | | | --------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default Value | Unset | The [`refresh_check_interval`](/docs/v1.10/reference/spicepod/datasets#accelerationrefresh_check_interval) parameter controls how often the accelerated dataset is refreshed. Example: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true refresh_mode: full refresh_check_interval: 10s ``` This configuration will refresh `taxi_trips` data every 10 seconds. ## Refresh On-Demand[​](#refresh-on-demand "Direct link to Refresh On-Demand") info Supported for accelerators with a `refresh_mode` of `full` or `append`. Accelerated datasets can be refreshed on-demand via the `refresh` CLI command or `POST /v1/datasets/:name/acceleration/refresh` API endpoint. CLI example: ``` spice refresh eth_recent_blocks ``` API example using cURL: ``` curl -i -XPOST 127.0.0.1:8090/v1/datasets/eth_recent_blocks/acceleration/refresh ``` with response: ``` HTTP/1.1 201 Created content-type: application/json content-length: 55 date: Thu, 11 Apr 2024 20:11:18 GMT {"message":"Dataset refresh triggered for eth_recent_blocks."} ``` Note On-demand refresh always initiates a new refresh, terminating any in-progress refresh for the dataset. ## Refresh Schedules[​](#refresh-schedules "Direct link to Refresh Schedules") | | | | --------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default Value | Unset | The [`refresh_cron`](/docs/v1.10/reference/spicepod/datasets#accelerationrefresh_cron) parameter supports specifying a cron schedule which controls when datasets refresh. Example: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true refresh_mode: full refresh_cron: '0 12 * * 1-5' ``` This configuration will refresh `taxi_trips` data at midday every weekday. For more information about cron schedules, see the [cron schedule reference](/docs/v1.10/reference/cron). The `refresh_cron` parameter cannot be specified in conjunction with a `refresh_check_interval` parameter. ## Refresh Retries[​](#refresh-retries "Direct link to Refresh Retries") | | | | ------------------------------------ | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default `refresh_retry_enabled` | `false` | | Default `refresh_retry_max_attempts` | Unset | By default, data refreshes for accelerated datasets are retried on transient errors (connectivity issues, compute warehouse goes idle, etc.) using a [Fibonacci](https://en.wikipedia.org/wiki/Fibonacci_sequence) backoff strategy. Retry behavior can be configured using the [`acceleration.refresh_retry_enabled`](/docs/v1.10/reference/spicepod/datasets#accelerationrefresh_retry_enabled) and [`acceleration.refresh_retry_max_attempts`](/docs/v1.10/reference/spicepod/datasets#accelerationrefresh_retry_max_attempts) parameters. Example: Disable retries ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: refresh_retry_enabled: false refresh_check_interval: 30s ``` Example: Limit retries to a maximum of 10 attempts ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: refresh_retry_max_attempts: 10 refresh_check_interval: 30s ``` ## Retention Policy[​](#retention-policy "Direct link to Retention Policy") | | | | ---------------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default `retention_check_enabled` | `false` | | Default `retention_period` | Unset | | Default `retention_sql` | Unset | | Default `retention_check_interval` | Unset | Accelerated datasets can be configured to automatically evict data using two different retention strategies: ### Time-based Retention[​](#time-based-retention "Direct link to Time-based Retention") Automatically evict time-series data exceeding a retention period by setting a retention policy based on the configured `time_column` and `acceleration.retention_period`. The policy is set using the [`acceleration.retention_check_enabled`](/docs/v1.10/reference/spicepod/datasets#accelerationretention_check_enabled), [`acceleration.retention_period`](/docs/v1.10/reference/spicepod/datasets#accelerationretention_period) and [`acceleration.retention_check_interval`](/docs/v1.10/reference/spicepod/datasets#accelerationretention_check_interval) parameters, along with the [`time_column`](/docs/v1.10/reference/spicepod/datasets#time_column) and [`time_format`](/docs/v1.10/reference/spicepod/datasets#time_format) dataset parameters. When `retention_check_enabled` is set to `true`, `retention_check_interval` and `retention_period` are required parameters. Example: ``` datasets: - from: mysql:user_events name: user_events time_column: created_at acceleration: enabled: true refresh_mode: append retention_check_enabled: true retention_period: 30d retention_check_interval: 1h ``` ### Custom SQL-based Retention[​](#custom-sql-based-retention "Direct link to Custom SQL-based Retention") Evict data from an acceleration based on custom filter predicates using the [`acceleration.retention_sql`](/docs/v1.10/reference/spicepod/datasets#accelerationretention_sql) parameter. This is useful for scenarios like soft-deleting rows in append datasets or removing data based on complex business logic. The `retention_sql` parameter takes the form of a `DELETE FROM
WHERE ` statement. Example - Soft delete retention: ``` datasets: - from: mysql:user_events name: user_events time_column: created_at acceleration: enabled: true refresh_mode: append primary_key: user_id on_conflict: user_id: upsert retention_check_enabled: true retention_check_interval: 5m retention_sql: DELETE FROM user_events WHERE status = 'archived' ``` note * Time-based retention (`retention_period`) and custom SQL retention (`retention_sql`) can be used independently or together. When both are configured, both retention policies will be applied during each retention check. ## Refresh Jitter[​](#refresh-jitter "Direct link to Refresh Jitter") | | | | -------------------------------- | ---------------- | | Supported in `refresh_mode` | `full`, `append` | | Required | No | | Default `refresh_jitter_enabled` | `false` | | Default `refresh_jitter_max` | Unset | Accelerated datasets can include a random jitter in their refresh interval to prevent the [Thundering herd problem](https://en.wikipedia.org/wiki/Thundering_herd_problem), where multiple datasets refresh simultaneously. The jitter is a random value between 0 and `refresh_jitter_max`, which is added to or subtracted from the base `refresh_check_interval`. If `refresh_jitter_max` is not specified, it defaults to 10% of `refresh_check_interval`. Refresh Jitter applies to the initial dataset load. If multiple similarly configured Spice instances are restarted at the same time, they will load with a jitter between 0 and `refresh_jitter_max`. Example: ``` datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: refresh_check_interval: 10s refresh_jitter_enabled: true refresh_jitter_max: 1s ``` In the configuration above: 1. The initial load will include a random delay between **0** and **1 second**. 2. Subsequent refresh intervals will vary randomly between **9 seconds** and **11 seconds**. Refresh jitter configuration: * [`refresh_jitter_enabled`](/docs/v1.10/reference/spicepod/datasets#accelerationrefresh_jitter_enabled) * [`refresh_jitter_max`](/docs/v1.10/reference/spicepod/datasets#accelerationrefresh_jitter_max) ## Configuration Examples[​](#configuration-examples "Direct link to Configuration Examples") ### Accelerating a full set of data that sometimes changes[​](#accelerating-a-full-set-of-data-that-sometimes-changes "Direct link to Accelerating a full set of data that sometimes changes") In this example, Spice connects with a dataset that changes infrequently and is not configured for CDC. For example, a list of product categories. ``` datasets: - from: mysql:product_categories name: product_categories acceleration: refresh_mode: full refresh_check_interval: 8h ``` In this scenario, Spice uses a simple acceleration configuration - full refreshes on an 8 hour schedule. No additional behaviors are enabled, so queries matching for new product codes will return no results until the next refresh cycle. ### Accelerating a subset of data that changes frequently[​](#accelerating-a-subset-of-data-that-changes-frequently "Direct link to Accelerating a subset of data that changes frequently") In this example, Spice connects with a dataset that has frequently changing data that is not configured for CDC. For example, user's posts on a social media platform. ``` datasets: - from: mysql:posts name: posts acceleration: refresh_mode: full refresh_check_interval: 10m refresh_sql: "SELECT * FROM posts WHERE updated_at > now() - interval '1' day" on_zero_results: use_source ``` With this configuration, Spice will refresh every 10 minutes accelerating posts that have been updated in the last day. When querying for posts by direct ID, if a post is not accelerated Spice will fallback to retrieving the post from the non-accelerated source due to the behavior of `on_zero_results: use_source`. However, if querying for a range of posts that includes some which have updated in the last day Spice will only return those results without falling back to the source. This could result in queries for a range of posts excluding posts that exist in the non-accelerated source because they have been filtered out due to their `updated_at` value. ### Accelerating application logs[​](#accelerating-application-logs "Direct link to Accelerating application logs") In this example, Spice connects to a data source that is immutable, receives new rows, and is not configured for CDC. For example, a database that contains some application logs. ``` datasets: - from: duckdb:logs name: logs time_column: created_at params: duckdb_open: logs.duckdb acceleration: refresh_mode: append refresh_check_interval: 10m refresh_sql: "SELECT * FROM logs WHERE asset = 'asset_id'" refresh_data_window: 1d on_zero_results: use_source retention_check_enabled: true retention_period: 7d retention_check_interval: 10m ``` This acceleration configuration applies a number of different behaviors: 1. A `refresh_data_window` was specified. When Spice starts, it will apply this `refresh_data_window` to the `refresh_sql`, and retrieve only the last day's worth of logs with an `asset = 'asset_id'`. 2. Because a `refresh_sql` is specified, every refresh (including initial load) will have the filter applied to the refresh query. 3. 10 minutes after loading, as specified by the `refresh_check_interval`, the first refresh will occur - retrieving new rows where `asset = 'asset_id'`. 4. Running a query to retrieve logs with an `asset` that is *not* `asset_id` will fall back to the source, because of the `on_zero_results: use_source` parameter. 5. Running a query to retrieve a log longer than 1 day ago will fall back to the source, because of the `on_zero_results: use_source` parameter. 6. Running a query to retrieve logs within a range of now to longer than 1 day ago will only return logs from the last day. This is due to the `refresh_data_window` only accelerating the last day's worth of logs, which will return some results. Because results are returned, Spice will not fall back to the source even though `on_zero_results: use_source` is specified. 7. Spice will retain newly appended log rows for 7 days before discarding them, as specified by the `retention_*` parameters. ## Cookbook[​](#cookbook "Direct link to Cookbook") * Configure accelerated dataset retention policy. [Accelerated Dataset Retention Policy](https://github.com/spiceai/cookbook/tree/trunk/retention#readme) * Dynamically refresh specific data at runtime by programmatically updating refresh\_sql and triggering data refreshes. [Advanced Data Refresh](https://github.com/spiceai/cookbook/tree/trunk/acceleration/data-refresh#readme) * Configure `refresh_data_window` to filter refreshed data to recent data [Refresh Data Window](https://github.com/spiceai/cookbook/tree/trunk/refresh-data-window#readme) --- # Indexes Database indexes are essential for optimizing query performance. This document explains how to add indexes to tables created by Spice for local data acceleration. Example Spicepod: ``` datasets: - from: spice.ai/eth.recent_blocks name: eth.recent_blocks acceleration: enabled: true engine: sqlite indexes: number: enabled # Index the `number` column '(hash, timestamp)': unique # Add a unique index with a multicolumn key comprised of the `hash` and `timestamp` columns ``` ## Column References[​](#column-references "Direct link to Column References") Column references can be used to specify which columns to index. The column reference can be a single column name or a multicolumn key. The column reference must be enclosed in parentheses if it is a multicolumn key. Examples * `number`: Index the `number` column * `(hash, timestamp)`: Index the `hash` and `timestamp` columns ## Index Types[​](#index-types "Direct link to Index Types") There are two types of indexes that can be specified in a Spicepod: * `enabled`: Creates a standard index on the specified column(s). * Similar to specifying `CREATE INDEX my_index ON my_table (my_column)`. * `unique`: Creates a unique index on the specified column(s). See [Constraints](/docs/v1.10/features/data-acceleration/constraints) for more information on working with unique constraints on locally accelerated tables. * Similar to specifying `CREATE UNIQUE INDEX my_index ON my_table (my_column)`. Limitations Traditional indexes are not supported for the in-memory Arrow or [Spice Cayenne](/docs/v1.10/components/data-accelerators/cayenne) acceleration engines. Use [DuckDB](/docs/v1.10/components/data-accelerators/duckdb), [SQLite](/docs/v1.10/components/data-accelerators/sqlite), or [PostgreSQL](/docs/v1.10/components/data-accelerators/postgres) as the acceleration engine to enable indexing. Spice Cayenne Point Lookup Performance While Spice Cayenne does not support traditional indexes, [Vortex](https://github.com/vortex-data/vortex) provides [100x faster random access reads](https://bench.vortex.dev) compared to Parquet through segment statistics (similar to zone-maps), fast random access encodings ([FSST](https://www.vldb.org/pvldb/vol13/p2649-boncz.pdf), [FastLanes](https://www.vldb.org/pvldb/vol16/p2132-afroozeh.pdf)), and compute push-down on compressed data. For many point lookup workloads, Spice Cayenne matches or exceeds indexed query performance without requiring explicit index configuration. See the [Spice Cayenne documentation](/docs/v1.10/components/data-accelerators/cayenne) for details. --- # Partitioning Accelerations can be partitioned using an arbitrary expression to group rows together into separate files. This allows Spice to avoid reading unnecessary partitions, making particular queries faster. To partition your accelerations, add the `partition_by` acceleration parameter: ``` datasets: - from: s3://spiceai-demo-datasets/taxi_trips/2024/ name: taxi_trips params: file_format: parquet acceleration: enabled: true engine: duckdb mode: file partition_by: - bucket(50, PULocationID) ``` This example uses a `bucket` user-defined function (UDF) to hash the `PULocationID` column and put each row into one of 50 partition files. This allows partition pruning for queries that filter on the column referenced in the `partition_by` expression: ``` SELECT * FROM taxi_trips WHERE PULocationID IN (1, 2, 3, 4, 5) ``` This will result in a scan plan that only reads from the partitions that contain the values from the `IN` list. Limitations * Partitioning is currently limited to `engine: duckdb` and `mode: file`. * `partition_by` must have only 1 expression. * Expression must reference exactly one column from the dataset. * Expression must produce a scalar value * Expression cannot contain a subquery * Partition pruning is limited to specific filter expressions such as: * `WHERE foo = bar` * `WHERE foo IN (bar, baz, ...)` * `WHERE foo NOT IN (bar, baz, ...)` --- # Caching Refresh Mode The `caching` refresh mode provides intelligent caching for HTTP-based datasets where multiple result rows can share the same request metadata. This mode is specifically designed for scenarios like API responses where the same request parameters can return different content over time or multiple rows of data. The caching mode supports two key paradigms: 1. **Stale-While-Revalidate (SWR)** - Serves cached data immediately while refreshing in the background, optimizing for low latency and reduced API costs 2. **Cache Persistence** - Stores cached data to disk using file-based accelerators (DuckDB, SQLite, or Cayenne) for fast cold starts and durability ## Overview[​](#overview "Direct link to Overview") Unlike traditional refresh modes that treat datasets as single sources of truth, the `caching` mode treats HTTP request metadata (path, query parameters, and body) as cache keys. This approach is particularly useful for: * REST API responses that return multiple records for a single request * Search API results where the same query may return different results over time * Dynamic content APIs where responses change based on server state Future Enhancement While currently designed for HTTP-based datasets, future versions of Spice will extend the `caching` mode to support arbitrary queries from any data source, enabling flexible caching strategies across all connector types. ## How It Works[​](#how-it-works "Direct link to How It Works") The `caching` mode uses HTTP request filter values as cache keys rather than enforcing primary key constraints. When a refresh occurs: 1. **Cache Key Generation**: By default, the combination of `request_path`, `request_query`, and `request_body` acts as the cache key. If a `primary_key` is explicitly specified in the acceleration configuration, it will be used instead of the metadata fields. 2. **Row Replacement**: All existing rows matching the cache key are removed before inserting new data 3. **Multiple Results**: Multiple rows with identical request metadata can coexist, representing different content items from the same API response 4. **Timestamp Tracking**: Each row includes a `fetched_at` timestamp indicating when the data was retrieved ## Schema[​](#schema "Direct link to Schema") Datasets using `caching` mode include the following metadata fields in addition to the content data: | Field Name | Type | Description | | --------------- | --------- | ------------------------------------------------------------------- | | `request_path` | String | The URL path used for the request | | `request_query` | String | The query parameters used for the request | | `request_body` | String | The request body (for POST requests) | | `content` | String | The response content | | `fetched_at` | Timestamp | The timestamp when the data was fetched (based on HTTP Date header) | The `fetched_at` timestamp uses the HTTP `Date` response header when available, falling back to the current system time if not present. ## Configuration[​](#configuration "Direct link to Configuration") To use `caching` mode, configure an `HTTP`/`HTTPS` dataset with `refresh_mode: caching`: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_cache params: file_format: json allowed_request_paths: '/search/shows,/shows/*' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file refresh_check_interval: 30s params: caching_ttl: 10s # How long a cache entry is considered fresh caching_stale_while_revalidate_ttl: 30s # How long after the `caching_ttl` to serve stale data while refreshing in the background ``` ## Use Cases[​](#use-cases "Direct link to Use Cases") ### Caching TV Show Search Results[​](#caching-tv-show-search-results "Direct link to Caching TV Show Search Results") Cache TV show search API results where the same query may return different results over time: ``` datasets: - from: https://api.tvmaze.com name: tv_search_cache params: file_format: json allowed_request_paths: '/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 15s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 30s refresh_sql: | SELECT * FROM tv_search_cache WHERE request_path = '/search/shows' AND request_query = 'q=game+of+thrones' ``` This configuration: * Fetches search results for "game of thrones" every 30 seconds * Stores all result items with the same request metadata * Replaces all previous results for this query on each refresh * Preserves the timestamp of when results were fetched ### Caching TV Show Episodes[​](#caching-tv-show-episodes "Direct link to Caching TV Show Episodes") Cache responses from a TV show episodes API: ``` datasets: - from: https://api.tvmaze.com name: episodes_cache params: file_format: json allowed_request_paths: '/shows/*/episodes' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 10s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 20s refresh_sql: | SELECT * FROM episodes_cache WHERE request_path = '/shows/82/episodes' AND request_query = 'season=1' ``` ### Multi-Endpoint Caching[​](#multi-endpoint-caching "Direct link to Multi-Endpoint Caching") Cache responses from multiple TVMaze API endpoints: ``` datasets: - from: https://api.tvmaze.com name: multi_endpoint_cache params: file_format: json allowed_request_paths: '/shows/*,/search/shows,/people/*' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 10s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 30s refresh_sql: | SELECT * FROM multi_endpoint_cache WHERE (request_path = '/shows/82' OR request_path = '/shows/169') OR (request_path = '/search/shows' AND request_query = 'q=breaking+bad') ``` ## Querying Cached Data[​](#querying-cached-data "Direct link to Querying Cached Data") Query cached data using standard SQL, filtering by request metadata or content: ``` -- Get all cached search results for a specific TV show query SELECT content, fetched_at FROM tv_search_cache WHERE request_query = 'q=game+of+thrones' ORDER BY fetched_at DESC; -- Find the most recent cache entry for each unique request SELECT request_path, request_query, MAX(fetched_at) as last_fetched FROM tv_shows_cache GROUP BY request_path, request_query; -- Get cached results fetched within the last hour SELECT * FROM episodes_cache WHERE fetched_at > NOW() - INTERVAL '1 hour'; -- Parse JSON to extract show information SELECT json_get_str(content, 'name') as show_name, json_get_str(content, 'type') as show_type, fetched_at FROM tv_shows_cache WHERE request_path = '/shows/82'; ``` ## Stale-While-Revalidate Pattern[​](#stale-while-revalidate-pattern "Direct link to Stale-While-Revalidate Pattern") The `caching` mode supports the Stale-While-Revalidate (SWR) pattern for acceleration, providing optimal performance by serving cached data immediately while refreshing in the background. ### How SWR Works[​](#how-swr-works "Direct link to How SWR Works") When configured with background refresh, the caching mode: 1. **Serves stale data immediately** - Returns cached results without waiting for a refresh 2. **Triggers background refresh** - Initiates an asynchronous refresh of the cache 3. **Updates cache transparently** - Subsequent queries receive fresh data once the refresh completes This pattern reduces query latency by eliminating wait times for data fetches while keeping the cache reasonably fresh. ### Background Refresh Configuration[​](#background-refresh-configuration "Direct link to Background Refresh Configuration") Configure background refresh using `refresh_check_interval` to specify how frequently the cache should be updated: ``` datasets: - from: https://api.tvmaze.com name: shows_cache params: file_format: json allowed_request_paths: '/shows/*' acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file # Persist cache to disk params: caching_ttl: 10s # Data is fresh for 10 seconds caching_stale_while_revalidate_ttl: 10s # Serve stale data for 10 seconds while refreshing refresh_check_interval: 30s # Refresh every 30 seconds in background refresh_sql: | SELECT * FROM shows_cache WHERE request_path = '/shows/82' ``` ### SWR Benefits for API Caching[​](#swr-benefits-for-api-caching "Direct link to SWR Benefits for API Caching") The SWR pattern is particularly valuable for caching API responses: * **Reduced latency** - Queries return immediately from the cache without waiting for HTTP requests * **Lower API costs** - Fewer requests to external APIs reduce usage and costs * **Improved reliability** - Cached data remains available even if the API is temporarily unavailable * **Better user experience** - Consistent fast response times improve application performance ### Example: SWR with On-Demand Refresh[​](#example-swr-with-on-demand-refresh "Direct link to Example: SWR with On-Demand Refresh") Combine background refresh with on-demand refresh for maximum flexibility: ``` datasets: - from: https://api.tvmaze.com name: tv_search_swr params: file_format: json allowed_request_paths: '/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file # Persist cache to disk params: caching_ttl: 15s # Cache data is fresh for 15 seconds caching_stale_while_revalidate_ttl: 10s # Serve stale data for 10 seconds while refreshing refresh_check_interval: 30s # Background refresh every 30 seconds refresh_on_startup: always # Always refresh on startup refresh_sql: | SELECT * FROM tv_search_swr WHERE request_path = '/search/shows' AND request_query = 'q=breaking+bad' ``` With this configuration: * The cache refreshes every 30 seconds automatically * Queries are served immediately from the cache * Manual refresh is available via `/v1/datasets/tv_search_swr/acceleration/refresh` * Cache is guaranteed fresh on application startup ## Cache Persistence[​](#cache-persistence "Direct link to Cache Persistence") The `caching` mode supports persisting cached data to disk using file-based acceleration engines, enabling the cache to survive application restarts and reducing cold start times. ### File-Based Accelerators[​](#file-based-accelerators "Direct link to File-Based Accelerators") Three acceleration engines support file persistence for caching mode: * **DuckDB** - High-performance analytical database with excellent compression * **SQLite** - Lightweight, reliable database ideal for embedded scenarios * **Cayenne** - Spice's native accelerator optimized for analytical workloads ### Configuring File Persistence[​](#configuring-file-persistence "Direct link to Configuring File Persistence") Enable file persistence by setting `acceleration.mode: file` and specifying an acceleration engine: ``` datasets: - from: https://api.tvmaze.com name: shows_persistent_cache params: file_format: json allowed_request_paths: '/shows/*,/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb # or sqlite, cayenne mode: file # Enable file persistence params: caching_ttl: 10s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 30s refresh_sql: | SELECT * FROM shows_persistent_cache WHERE request_path IN ('/shows/82', '/shows/169') OR (request_path = '/search/shows' AND request_query = 'q=game+of+thrones') ``` ### DuckDB Persistence Example[​](#duckdb-persistence-example "Direct link to DuckDB Persistence Example") DuckDB provides excellent performance for analytical queries on cached data: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_duckdb params: file_format: json allowed_request_paths: '/shows/*,/shows/*/episodes' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 10s caching_stale_while_revalidate_ttl: 10s duckdb_file: tv_shows_cache.db # Specify custom file location refresh_check_interval: 30s refresh_sql: | SELECT * FROM tv_shows_duckdb WHERE request_path IN ('/shows/82', '/shows/169') OR (request_path = '/shows/82/episodes' AND request_query = 'season=1') ``` ### SQLite Persistence Example[​](#sqlite-persistence-example "Direct link to SQLite Persistence Example") SQLite is ideal for lightweight caching scenarios: ``` datasets: - from: https://api.tvmaze.com name: tv_search_sqlite params: file_format: json allowed_request_paths: '/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: sqlite mode: file params: caching_ttl: 15s caching_stale_while_revalidate_ttl: 10s sqlite_file: tv_search_cache.db refresh_check_interval: 30s refresh_sql: | SELECT * FROM tv_search_sqlite WHERE request_path = '/search/shows' AND request_query IN ('q=breaking+bad', 'q=game+of+thrones') ``` ### Cayenne Persistence Example[​](#cayenne-persistence-example "Direct link to Cayenne Persistence Example") Cayenne provides optimized performance for Spice workloads: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_cayenne params: file_format: json allowed_request_paths: '/shows/*' acceleration: enabled: true refresh_mode: caching engine: cayenne mode: file params: caching_ttl: 10s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 30s refresh_sql: | SELECT * FROM tv_shows_cayenne WHERE request_path IN ('/shows/82', '/shows/169', '/shows/73') ``` ### Benefits of File Persistence[​](#benefits-of-file-persistence "Direct link to Benefits of File Persistence") Persisting the cache to disk provides several advantages: * **Fast cold starts** - Cache is immediately available on application restart without fetching from APIs * **Reduced API load** - No need to refetch all data after restarts * **Cost savings** - Fewer API requests reduce metered API costs * **Offline capability** - Cached data remains queryable even when the API is unavailable * **Data durability** - Cache survives application crashes and restarts ### Memory vs. File Mode[​](#memory-vs-file-mode "Direct link to Memory vs. File Mode") Choose between in-memory and file-based caching based on your requirements: | Aspect | Memory Mode (`mode: memory`) | File Mode (`mode: file`) | | --------------- | --------------------------------- | ----------------------------- | | **Performance** | Fastest - all data in RAM | Fast with disk I/O | | **Persistence** | Lost on restart | Survives restarts | | **Capacity** | Limited by available memory | Limited by disk space | | **Cold start** | Slow - must refetch all data | Fast - loads from disk | | **Best for** | Small, frequently changing caches | Large, stable caches | | **Engines** | `arrow` (default) | `duckdb`, `sqlite`, `cayenne` | ### Combining SWR and Persistence[​](#combining-swr-and-persistence "Direct link to Combining SWR and Persistence") For optimal performance, combine the SWR pattern with file persistence: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_optimized params: file_format: json allowed_request_paths: '/shows/*,/search/shows' request_query_filters: enabled acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file # Persist to disk params: caching_ttl: 15s # Cache data is fresh for 15 seconds caching_stale_while_revalidate_ttl: 10s # Serve stale data for 10 seconds while refreshing refresh_check_interval: 30s # Background refresh (SWR) refresh_on_startup: auto # Use persisted cache on startup refresh_sql: | SELECT * FROM tv_shows_optimized WHERE request_path IN ('/shows/82', '/shows/169') OR (request_path = '/search/shows' AND request_query = 'q=game+of+thrones') ``` This configuration provides: * Immediate query response from persisted cache on startup * Background refresh every 30 seconds without blocking queries * Durable cache that survives application restarts * Reduced API requests and costs ## Behavior and Characteristics[​](#behavior-and-characteristics "Direct link to Behavior and Characteristics") ### Row-Level Replacement[​](#row-level-replacement "Direct link to Row-Level Replacement") The `caching` mode uses `InsertOp::Replace` to handle data updates. When new data is fetched for a given cache key (request metadata): 1. All existing rows matching that cache key are removed 2. All new rows are inserted 3. This operation is atomic, ensuring consistent cache state This behavior differs from other modes: * **`full` mode**: Replaces the entire dataset * **`append` mode**: Adds new rows without removing existing ones * **`changes` mode**: Applies CDC events * **`caching` mode**: Replaces rows matching the specific cache key ### Cache Key Behavior[​](#cache-key-behavior "Direct link to Cache Key Behavior") The `caching` mode determines cache keys based on the acceleration configuration: **Default (No Primary Key Specified)**: * Uses HTTP request metadata fields as the cache key: `request_path`, `request_query`, and `request_body` * Multiple result rows can share the same request metadata * The cache key serves as the logical grouping mechanism for row replacement * Content within a response may have duplicate values across different requests **With Primary Key Specified**: * Uses the explicitly configured `primary_key` columns as the cache key * Provides fine-grained control over cache key composition * Useful when caching requires uniqueness based on response content fields rather than request metadata Example with custom primary key: ``` datasets: - from: https://api.tvmaze.com name: tv_episodes_custom_key params: file_format: json allowed_request_paths: '/shows/*/episodes' acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file primary_key: [id, season, number] # Use episode fields as cache key params: caching_ttl: 15s caching_stale_while_revalidate_ttl: 10s refresh_check_interval: 30s refresh_sql: | SELECT * FROM tv_episodes_custom_key WHERE request_path = '/shows/82/episodes' ``` ### HTTP Date Header[​](#http-date-header "Direct link to HTTP Date Header") The `fetched_at` timestamp respects the HTTP `Date` response header when present. This provides: * Accurate server-side timestamps for cached responses * Consistency across distributed systems * Proper cache age calculation based on server time If the `Date` header is not present, the system falls back to using the current system time. ## Refresh Configuration[​](#refresh-configuration "Direct link to Refresh Configuration") The `caching` mode supports standard refresh configuration options. See [Stale-While-Revalidate Pattern](#stale-while-revalidate-pattern) for background refresh details and [Cache Persistence](#cache-persistence) for file-based caching configuration. | Parameter | Description | Default | | ------------------------ | --------------------------------------------------------------------- | -------------- | | `refresh_check_interval` | How often to refresh cached data in the background | None | | `refresh_sql` | SQL query defining what data to cache | None | | `refresh_on_startup` | Whether to refresh on startup (`auto` or `always`) | `auto` | | `on_zero_results` | Behavior when cache returns no results (`return_empty`, `use_source`) | `return_empty` | | `engine` | Acceleration engine (`arrow`, `duckdb`, `sqlite`, `cayenne`) | `arrow` | | `mode` | Persistence mode (`memory` or `file`) | `memory` | ### Cache TTL (Time-to-Live)[​](#cache-ttl-time-to-live "Direct link to Cache TTL (Time-to-Live)") The caching mode provides parameters to control cache freshness and staleness behavior: | Parameter | Description | Default | | ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | | `caching_ttl` | Duration that cached data is considered fresh. After this period, data becomes stale and triggers a background refresh. | `30s` | | `caching_stale_while_revalidate_ttl` | Duration after `caching_ttl` expires during which stale data is served while refreshing in the background. After this period, queries wait for fresh data. | None | | `caching_stale_if_error` | When set to `enabled`, serves expired cached data if the upstream source returns an error. Valid values: `enabled`, `disabled`. | `disabled` | The `caching_ttl` parameter defines how long cached data is considered fresh before it becomes stale. Once cached data exceeds this age, the SWR pattern triggers background refresh to update the cache while continuing to serve the stale data during the `caching_stale_while_revalidate_ttl` window. If `caching_stale_while_revalidate_ttl` is omitted, cached data becomes rotten immediately after `caching_ttl` expires, and queries will wait for fresh data rather than returning stale results. When a value is specified, stale data is served during that window after `caching_ttl` expires while a background refresh occurs. Once the combined `caching_ttl + caching_stale_while_revalidate_ttl` period has passed, queries will wait for fresh data instead of returning stale results. **Configuring Cache TTL**: ``` datasets: - from: https://api.tvmaze.com name: tv_shows_ttl params: file_format: json allowed_request_paths: '/shows/*' acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file # Persist cache to disk refresh_check_interval: 30s # Periodic check for stale data params: caching_ttl: 15s # Cache data is fresh for 15 seconds caching_stale_while_revalidate_ttl: 10s # Serve stale data for 10 seconds while refreshing ``` **How Cache TTL Works**: 1. When data is fetched, the `fetched_at` timestamp is recorded 2. On subsequent queries, the system checks `now - fetched_at > caching_ttl` 3. If data is within TTL, it is served immediately (fresh) 4. If data exceeds TTL, it becomes stale: * Stale data is served immediately (no query delay) if within `caching_stale_while_revalidate_ttl` * Background refresh is triggered to update the cache * Next query receives the refreshed data **TTL Format**: Duration strings support common units: * Seconds: `30s`, `90s` * Minutes: `5m`, `15m` * Hours: `1h`, `24h` * Mixed: `1h30m`, `2h15m30s` **Default Behavior**: When `caching_ttl` is not specified, it defaults to `30s` (30 seconds). This provides a reasonable balance between freshness and cache efficiency for most use cases. When `caching_stale_while_revalidate_ttl` is not specified, stale data is not served after the TTL expires, and queries will wait for fresh data. ### Stale-If-Error Behavior[​](#stale-if-error-behavior "Direct link to Stale-If-Error Behavior") The `caching_stale_if_error` parameter controls whether expired cached data is served when the upstream data source returns an error during a refresh attempt. This provides fault tolerance by returning stale data instead of failing the query when the upstream source is temporarily unavailable. ``` datasets: - from: https://api.tvmaze.com name: tv_shows_resilient acceleration: enabled: true refresh_mode: caching engine: duckdb mode: file params: caching_ttl: 15s caching_stale_while_revalidate_ttl: 30s caching_stale_if_error: enabled # Serve stale data on upstream errors ``` When `caching_stale_if_error: enabled`: * If the upstream source returns an error during refresh, expired cached data is served instead of failing * Queries continue to return data even when the upstream API is unavailable * Useful for APIs with intermittent availability or rate limits When `caching_stale_if_error: disabled` (default): * Errors from the upstream source propagate to the query * Queries fail when fresh data cannot be fetched and cached data has expired Stale-While-Revalidate Configuration Conflict When using `refresh_mode: caching`, you cannot configure both the caching accelerator's `caching_stale_while_revalidate_ttl` and the [results cache](/docs/v1.10/features/caching)'s `stale_while_revalidate_ttl` for the same dataset. These parameters control similar behavior at different layers, and having both enabled creates a conflict. Choose one approach: * **Caching accelerator SWR**: Use `acceleration.params.caching_stale_while_revalidate_ttl` for HTTP-based dataset caching * **Results cache SWR**: Use `runtime.caching.sql_results.stale_while_revalidate_ttl` for SQL query results caching **TTL Considerations**: * **Shorter TTL** (e.g., `10s`, `30s`): More frequent refresh, higher data freshness, more API requests * **Longer TTL** (e.g., `10m`, `1h`): Fewer refreshes, lower API costs, potentially stale data * **Matching patterns**: Set `caching_ttl` shorter than `refresh_check_interval` to define the staleness window ## Limitations[​](#limitations "Direct link to Limitations") * Currently only available for HTTP-based datasets using the [HTTPS connector](/docs/v1.10/components/data-connectors/https). Future releases will extend support to arbitrary queries from any data source. * Requires `acceleration.enabled: true` * When no `primary_key` is specified, cache keys default to request metadata fields (`request_path`, `request_query`, `request_body`) * On-demand refresh via `/v1/datasets/:name/acceleration/refresh` API triggers a new refresh for all cache keys defined in `refresh_sql` ## Related Documentation[​](#related-documentation "Direct link to Related Documentation") * [HTTPS Data Connector](/docs/v1.10/components/data-connectors/https) - Detailed HTTP connector configuration * [Data Refresh](/docs/v1.10/features/data-acceleration/data-refresh) - Overview of all refresh modes * [Refresh SQL](/docs/v1.10/features/data-acceleration/data-refresh#refresh-sql) - Using SQL to control refresh behavior * [Special Metadata Fields](/docs/v1.10/components/data-connectors/https#special-metadata-fields) - HTTP request metadata fields * [Data Accelerators](/docs/v1.10/components/data-accelerators) - Acceleration engines for cache persistence * [DuckDB Accelerator](/docs/v1.10/components/data-accelerators/duckdb) - DuckDB acceleration engine * [SQLite Accelerator](/docs/v1.10/components/data-accelerators/sqlite) - SQLite acceleration engine --- # Snapshots ## Spicepod Example[​](#spicepod-example "Direct link to Spicepod Example") ``` snapshots: enabled: true location: s3://some_bucket/some_folder/ bootstrap_on_failure_behavior: warn params: s3_auth: iam_role datasets: - name: some_table acceleration: engine: duckdb mode: file snapshots: enabled params: duckdb_file: /nvme/some_table.db ``` ## Overview[​](#overview "Direct link to Overview") Acceleration snapshots let Spice reuse a pre-built acceleration file on startup instead of waiting for a full refresh. When a dataset uses a file-mode acceleration engine (DuckDB, SQLite, or Turso) and the local file is missing (for example on first boot or when using ephemeral NVMe storage), Spice downloads the most recent snapshot from object storage and moves the dataset straight to a ready state. Preview Acceleration snapshots are available in preview. ## How it works[​](#how-it-works "Direct link to How it works") * On startup, Spice checks whether the file supplied in `acceleration.params` (for example `duckdb_file`) exists. * If the file is missing and snapshots are enabled, Spice looks under the configured snapshot location and downloads the newest snapshot for that dataset. * If no snapshot is available, the acceleration boots empty and refreshes from the source. * After each refresh, Spice writes a new snapshot unless the dataset is configured in `bootstrap_only` mode. Snapshots are organized with Hive-style partitioning so they are easy to retain and prune. For a dataset named `my_dataset`, Spice writes files such as: ``` s3://some_bucket/some_folder/month=2025-09/day=2025-09-30/dataset=my_dataset/my_dataset_20250919T134522Z.db ``` The timestamp is recorded in UTC using ISO 8601 without punctuation. Dedicated files only Every accelerated dataset must write to its own file (for example, `/nvme/my_dataset.db`). Sharing a single file across multiple datasets is not supported. ## Configure snapshot storage[​](#configure-snapshot-storage "Direct link to Configure snapshot storage") Snapshots are controlled with a new top-level `snapshots` block in the Spicepod. The location must point to a folder on S3 or the local filesystem. When the location is an S3 bucket, the configuration accepts any S3 dataset parameters under `params`. ``` snapshots: enabled: true location: s3://some_bucket/some_folder/ # Folder where snapshots are written bootstrap_on_failure_behavior: warn # retry | fallback | warn params: s3_auth: iam_role # Defaults to iam_role for snapshots ``` ### Failure behavior[​](#failure-behavior "Direct link to Failure behavior") `bootstrap_on_failure_behavior` controls what Spice does when it cannot load the most recent snapshot. * `retry` – keep retrying the newest snapshot until it succeeds. * `fallback` – try older snapshot files until one loads successfully. * `warn` – log a warning and continue with an empty acceleration. (Default.) ## Enable snapshots per dataset[​](#enable-snapshots-per-dataset "Direct link to Enable snapshots per dataset") Each dataset opts into snapshotting through the `acceleration.snapshots` field. Four modes are available: * `enabled` – download snapshots on startup and write a new snapshot after each refresh. * `bootstrap_only` – only download snapshots; never write new ones. * `create_only` – write new snapshots after refreshes, but never download them on startup. * `disabled` – disable snapshot usage for this dataset. (Default.) Example Spicepod configuration: ``` snapshots: enabled: true location: s3://some_bucket/some_folder/ bootstrap_on_failure_behavior: warn params: s3_auth: iam_role datasets: - from: s3://some_bucket/some_table/ name: some_table params: file_format: parquet s3_auth: iam_role acceleration: enabled: true engine: duckdb mode: file snapshots: enabled params: duckdb_file: /nvme/some_table.db ``` Readiness with append refreshes Append-mode accelerations that define a `time_column` wait to report ready until the first append refresh completes after snapshot bootstrap. This keeps the dataset out of rotation until the freshest data is available while still benefiting from the snapshot-assisted startup. See [Fast Cold Starts](/docs/v1.10/features/data-acceleration/data-refresh#fast-cold-starts-with-snapshots) for additional context. ## Best practices[​](#best-practices "Direct link to Best practices") * **Pair with ephemeral storage:** Deployments commonly place the acceleration file on fast ephemeral disks (such as NVMe instance storage) while relying on snapshots for persistence across restarts. * **Align retention policies:** Apply an object storage lifecycle rule that mirrors the desired snapshot retention policy. * **Monitor bootstraps:** Track warning logs emitted when Spice falls back to an empty acceleration so operators can respond quickly if snapshot loading fails. For the full reference, see [`snapshots` in the Spicepod specification](/docs/v1.10/reference/spicepod#snapshots) and [`acceleration.snapshots`](/docs/v1.10/reference/spicepod/datasets#accelerationsnapshots). --- # Data Ingestion Data can be ingested by the Spice runtime into a Data Connector using the following methods: 1. **SQL Statements** – Write data directly to [write-capable connectors](/docs/v1.10/tags/write) using standard SQL syntax. 2. **OpenTelemetry (OTEL) Ingestion** – Stream OTEL metrics for real-time processing and acceleration. Data ingestion is useful for scenarios such as collecting metrics from edge devices, writing application events for later analysis, or populating datasets from external sources. ## SQL Statements[​](#sql-statements "Direct link to SQL Statements") Spice supports writing data to **compatible data connectors** using standard SQL `INSERT INTO` syntax. ### Write-Capable Connectors[​](#write-capable-connectors "Direct link to Write-Capable Connectors") Data connectors that support write operations are tagged as [write](/docs/v1.10/tags/write): * **[Apache Iceberg](/docs/v1.10/components/data-connectors/iceberg)** - Write to Iceberg tables via data connector or [catalog connector](/docs/v1.10/components/catalogs/iceberg) * **[AWS Glue](/docs/v1.10/components/data-connectors/glue)** - Write to Glue Data Catalog tables via data connector or [catalog connector](/docs/v1.10/components/catalogs/glue) ### Configuration for Write Operations[​](#configuration-for-write-operations "Direct link to Configuration for Write Operations") To enable write operations, configure your dataset or catalog with [read\_write access](/docs/v1.10/reference/spicepod/datasets#access): ``` datasets: - from: glue:my_catalog.my_schema.my_table name: my_table access: read_write params: # ... connector-specific parameters ``` ### Example SQL[​](#example-sql "Direct link to Example SQL") ``` INSERT INTO my_table (column1, column2) VALUES ('value1', 'value2'); INSERT INTO my_table (column1, column2) SELECT source_column1, source_column2 FROM source_table WHERE condition = 'filter'; ``` For more details on the `INSERT` statement syntax, see the [SQL INSERT documentation](/docs/v1.10/reference/sql/dml#insert). ## OpenTelemetry Data Ingestion[​](#opentelemetry-data-ingestion "Direct link to OpenTelemetry Data Ingestion") By default, the runtime exposes an [OpenTelemetry](https://opentelemetry.io) (OTEL) endpoint at grpc://127.0.0.1:50052 for the OTEL data ingestion. OTEL metrics will be inserted into datasets with matching names (metric name = dataset name) and optionally replicated to the dataset source. ## Benefits[​](#benefits "Direct link to Benefits") Spice.ai OSS includes built-in data ingestion support, allowing the collection of the latest data from edge nodes for use in subsequent queries. This feature eliminates the need for additional ETL pipelines and enhances the speed of the feedback loop. For example, consider CPU usage anomaly detection. When CPU metrics are sent to the Spice OpenTelemetry endpoint, the loaded machine learning model can use the most recent observations for inferencing and provide recommendations to the edge node. This process occurs quickly on the edge itself, within milliseconds, and without generating network traffic. Additionally, Spice will periodically replicate the data to the data connector for further use. ## Considerations[​](#considerations "Direct link to Considerations") Data Quality: Use Spice SQL capabilities to transform and cleanse ingested edge data, ensuring high-quality inputs. Data Security: Evaluate data sensitivity and secure network connections between the edge and data connector when replicating data for further use. Implement encryption, access controls, and secure protocols. ## Example[​](#example "Direct link to Example") ### [Disk SMART](https://en.wikipedia.org/wiki/Self-Monitoring,_Analysis_and_Reporting_Technology)[​](#disk-smart "Direct link to disk-smart") Start Spice with the following dataset: ``` datasets: - from: spice.ai/coolorg/smart/datasets/drive_stats name: smart_attribute_raw_value access: read_write replication: enabled: true acceleration: enabled: true ``` Start telegraf with the following config: ``` [[inputs.smart]] attributes = true [[outputs.opentelemetry]] service_address = "localhost:50052" [agent] interval = "1s" flush_interval = "1s" ``` SMART data will be available in the `smart_attribute_raw_value` dataset in Spice.ai OSS and replicated to the `coolorg.smart.drive_stats` dataset in Spice.ai Cloud. ## Limitations[​](#limitations "Direct link to Limitations") Current Limitations * Write Support: Only selected [write-capable connectors and catalogs](/docs/v1.10/tags/write) support write operations. * Only Spice.ai replication is supported for OpenTelemetry ingestion --- # Distributed Query Learn how to configure and run Spice in distributed mode to handle larger scale queries across multiple nodes. Preview Multi-node distributed query execution based on Apache Ballista is available as a preview feature in Spice `v1.9.0`. ## Overview[​](#overview "Direct link to Overview") Spice integrates [Apache Ballista](https://github.com/apache/datafusion-ballista) to schedule and coordinate distributed queries across multiple executor nodes. This integration enables distributed execution when running large queries over partitioned data lake formats such as Parquet, Delta Lake, or Iceberg. ## Architecture[​](#architecture "Direct link to Architecture") A distributed Spice cluster consists of two components: * **Scheduler** – Plans distributed queries and manages the work queue for the executor fleet. Single instance per cluster. * **Executors** – One or more nodes responsible for executing physical query plans. The scheduler holds the cluster-wide configuration for a Spicepod, while executors connect to the scheduler to receive work. ## Getting Started[​](#getting-started "Direct link to Getting Started") Cluster deployment typically starts with a scheduler instance, followed by one or more executors that register with the scheduler. ### Start the Scheduler[​](#start-the-scheduler "Direct link to Start the Scheduler") The scheduler is the only `spiced` process that needs to be configured (i.e. have a `spicepod.yaml` in the current dir). Override the Flight bind address when it must be reachable outside of `localhost`: ``` # Start scheduler spiced --cluster-mode scheduler --flight 0.0.0.0:50051 ``` ### Start Executors[​](#start-executors "Direct link to Start Executors") Executors need the scheduler's Flight URI to register and pull work. The executors do not require a `spicepod.yaml` to be present, it will fetch the configuration from the coordinator. Each executor automatically selects a free port if the default is unavailable: ``` # Start executor spiced --cluster-mode executor --scheduler-url spiced://localhost:50051 ``` ## Query Execution[​](#query-execution "Direct link to Query Execution") Queries run against the scheduler endpoint. The `EXPLAIN` output confirms that distributed planning is active—Spice includes a `distributed_plan` section showing how stages are split across executors: ``` EXPLAIN SELECT count(id) FROM my_dataset; ``` Limitations * Accelerated datasets are not yet supported; distributed query currently targets partitioned data lake sources. * As a preview feature, clusters may encounter stability or performance issues. * Accelerator support is planned for future releases; follow release notes for updates. --- # Embedding Datasets Learn how to define and augment datasets with embedding columns for advanced search capabilities. ## Overview[​](#overview "Direct link to Overview") Spice provides three distinct methods for handling embedding columns in datasets: 1. **[Just-in-Time (JIT) Embeddings](/docs/v1.10/components/embeddings#jit-embeddings)**: Dynamically computes embeddings, on-demand, during query execution, without precomputing data. 2. **[Accelerated Embeddings](/docs/v1.10/components/embeddings#accelerated-embeddings)**: Precomputes embeddings by transforming and augmenting the source dataset for faster query and search performance. 3. **[Passthrough Embeddings](/docs/v1.10/components/embeddings#passthrough-embeddings)**: Utilizes pre-existing embeddings directly from the underlying source datasets, bypassing any additional computation. ## Configuring Embedding Models[​](#configuring-embedding-models "Direct link to Configuring Embedding Models") Before configuring dataset embeddings define the embedding models in the `spicepod.yaml`, for example: ``` embeddings: - name: local_embedding_model from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 - from: openai name: remote_service params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } ``` See [Embedding components](/docs/v1.10/components/embeddings) for more information on embedding models. ## Vector Searches[​](#vector-searches "Direct link to Vector Searches") Spice supports complex searches by utilizing embeddings. Both local and remote embedding models can be used for vector searches. To run a vector search, embeddings must be defined for the relevant columns in your dataset. Once configured, similarity searches can be performed using the defined embeddings. For detailed instructions and examples on running vector searches, refer to the [Vector-Based Search documentation](/docs/v1.10/features/search/vector-search). ## Generating Embeddings in Queries[​](#generating-embeddings-in-queries "Direct link to Generating Embeddings in Queries") The [`embed()` scalar function](/docs/v1.10/reference/sql/scalar_functions#embed) allows you to generate embeddings directly within SQL queries. This function can process both single text strings and arrays of text, making it useful for ad-hoc embedding generation and comparison operations. --- # Large Language Models Spice provides a high-performance, OpenAI API-compatible AI Gateway optimized for managing and scaling large language models (LLMs). It offers tools for Enterprise Retrieval-Augmented Generation (RAG), such as SQL query across federated datasets and an advanced search feature (see [Search](/docs/v1.10/features/search)). ![ai-gateway](https://github.com/user-attachments/assets/4a45cd62-ebfc-4a73-956d-661f1ab44cd8) Spice supports **full OpenTelemetry observability**, helping with detailed tracking of model tool use, recursion, data flows and requests for full transparency and easier debugging. ## Configuring Language Models[​](#configuring-language-models "Direct link to Configuring Language Models") Spice supports a variety of LLMs (see [Model Providers](/docs/v1.10/components/models)). ### Core Features[​](#core-features "Direct link to Core Features") * **SQL Integration**: Invoke LLMs directly within SQL queries using the `ai()` function for text generation tasks. See [SQL Reference: ai function](/docs/v1.10/reference/sql/scalar_functions#ai). * **Custom Tools**: Provide models with tools to interact with the Spice runtime. See [Tools](/docs/v1.10/features/large-language-models/tools). * **System Prompts**: Customize system prompts and override defaults for [`v1/chat/completion`](/docs/v1.10/api/HTTP/post-chat-completions). See [Parameter Overrides](/docs/v1.10/features/large-language-models/parameter_overrides). * **Memory**: Provide LLMs with memory persistence tools to store and retrieve information across conversations. See [Memory](/docs/v1.10/features/large-language-models/memory). * **Vector Search**: Perform advanced vector-based searches using embeddings. See [Vector Search](/docs/v1.10/features/search/vector-search). * **Evals**: Evaluate, track, compare, and improve language model performance for specific tasks. See [Evals](/docs/v1.10/features/large-language-models/evals). * **Local Models**: Load and serve models locally from various sources, including local filesystems and Hugging Face. See [Local Models](/docs/v1.10/features/large-language-models/serving). For API usage, refer to the [API Documentation](/docs/v1.10/api). ## [📄️Tools](/docs/v1.10/features/large-language-models/tools) [Learn how LLMs interact with the Spice runtime.](/docs/v1.10/features/large-language-models/tools) ## [📄️MCP](/docs/v1.10/features/large-language-models/mcp) [Learn how to use the Model Context Protocol (MCP) with Spice.](/docs/v1.10/features/large-language-models/mcp) ## [📄️Memory](/docs/v1.10/features/large-language-models/memory) [Learn how to provide LLMs with memory](/docs/v1.10/features/large-language-models/memory) ## [📄️Evals](/docs/v1.10/features/large-language-models/evals) [Learn how Spice evaluates, tracks, compares, and improves language model performance for specific tasks](/docs/v1.10/features/large-language-models/evals) ## [📄️Parameter Overrides](/docs/v1.10/features/large-language-models/parameter_overrides) [Learn how to override default LLM hyperparameters in Spice.](/docs/v1.10/features/large-language-models/parameter_overrides) ## [📄️Local Models](/docs/v1.10/features/large-language-models/serving) [Learn how to load and serve large learning models.](/docs/v1.10/features/large-language-models/serving) ## [📄️Parameterized Prompts](/docs/v1.10/features/large-language-models/parameterized_prompts) [Learn how to update system prompts for each request with Jinja-styled templating.](/docs/v1.10/features/large-language-models/parameterized_prompts) --- # Evaluating Language Models Language models can perform complex tasks. Evals help measure a model's ability to perform a specific task. Evals are defined as Spicepod components and can evaluate any Spicepod model's performance. Refer to the [Cookbook](https://github.com/spiceai/cookbook/tree/trunk/evals) for related examples. ## Overview[​](#overview "Direct link to Overview") In Spice, an eval consists of the following core components: * **Evals**: A defined task for a model to perform and a method to measure its performance. * **Eval Run**: An single evaluation of a specific model. * **Eval Result**: The model output and score for a single input task within an eval run. * **Eval Scorer**: A method to score the model's performance on an eval result. ### Eval Components[​](#eval-components "Direct link to Eval Components") An eval component is defined as follows: ``` evals: - name: australia description: Make sure the model understands Aussies, and importantly Cricket. dataset: cricket_questions scorers: - match datasets: - name: cricket_questions from: https://github.com/openai/evals/raw/refs/heads/main/evals/registry/data/cricket_situations/samples.jsonl ``` Where: * `name` is a unique identifier for this eval (like `models`, `datasets`, etc.). * `dataset` is a dataset component. * `scorers` is a list of scoring methods. For complete details on the `evals` component, see the [Spicepod reference](/docs/v1.10/reference/spicepod/evals). ## Running an Eval[​](#running-an-eval "Direct link to Running an Eval") To run an eval, ensure 1. Define an `eval` component (and it's associated `dataset`). 2. Add a language model to the spicepod (this is the model that will be evaluated). An eval can be started via the HTTP API: ``` curl -XPOST http://localhost:8090/v1/evals/australia \ -H 'Content-Type: application/json' \ -d '{ "model": "my_model", }' ``` Depending on the dataset and model, the eval run can take some time to complete. On completion, results will be available in two tables: * `eval.runs`: Summarises the status and scores from the eval run. * `eval.results`: Contains the input, expected output, and actual output for each eval run, and the score from each scorer. ## Dataset Formats[​](#dataset-formats "Direct link to Dataset Formats") Datasets are used to define the input and expected output for an eval. Evals expect a particular format, * `input`: The input to the model. It should be either: * A plain string (e.g., `"Hello, how are you?"`), interpreted as a single user message. * A JSON array is interpreted as multiple OpenAI-compatible [messages](https://platform.openai.com/docs/api-reference/chat/create#chat-create-messages) (e.g., `[{"role":"system","content":"You are a helpful assistant."}, ...]`). * For the `ideal` column: * A plain string (e.g., `"I'm doing well, thanks!"`), interpreted as a single assistant response. * A JSON array is interpreted as multiple OpenAI-compatible [choices](https://platform.openai.com/docs/api-reference/chat/object#chat/object-choices) (e.g., `[{"index":0,"message":{"role":"assistant","content":"Sure!"}, ...}]`). To use a dataset with a different format, use a `view`. For example: ``` views: # This view defines an eval dataset containing previous ai completion tasks from the `runtim.task_history` table. - name: user_queries sql: | SELECT json_get_json(input, 'messages') AS input, json_get_str((captured_output -> 0), 'content') as ideal FROM runtime.task_history WHERE task='ai_completion' ``` ## Eval Scorers[​](#eval-scorers "Direct link to Eval Scorers") An eval scorer is a method to score the model's performance on a single eval case. A scorer is given the input given to the model, the models output and the expected output and produces an associated score. Spice has several out of the box scorers: * `match`: Checks for an exact match between the expected and actual outputs. * `json_match`: Checks for an equivalent JSON between expected and actual outputs. * `includes`: Checks for the actual output to include the expected output. * `fuzzy_match`: Checks whether a normalised version (ignoring casing, punctuation, articles (e.g. a, the), excess whitespace) of either the expected and actual outputs are a subset of the other. * `levenshtein`: Computes the Levenshtein distance between the two output strings, normalised to the string length. The Levenshtein distance between two words is the minimum number of single-character edits (insertions, deletions or substitutions) required to change one word into the other. Spice has two other methods to define new scorers based on other spicepod components: * Embedding models can be used to compute the similarity between the expected and actual output from the model being evaluated. Any `embeddings` model defined in the `spicepod.yaml` is automatically available as a scorer. * Other language models can be used to judge the model being evaluated. This is often called an LLM-as-a-judge. Any `models` model defined in the `spicepod.yaml` is automatically available as a scorer. Note, however, that these models should generally be configured purposefully to be a judge. There are also constraints the model must satisfy, see [below](#llm-judge). Below is an example of an eval that uses all three: a builtin scorer, an embedding model scorer and an LLM judge. ``` evals: - name: australia description: Make sure the model understands Aussies, and importantly Cricket. dataset: cricket_questions scorers: - hf_minilm - judge - match embeddings: - name: hf_minilm from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 models: - name: judge from: openai:gpt-4o params: openai_api_key: ${ secrets:OPENAI_API_KEY } parameterized_prompt: enabled system_prompt: | Score these two stories between 0.0 and 1.0 based on how similar their moral lesson is. Story A: {{ actual }} Story B: {{ ideal }} openai_response_format: type: json_schema json_schema: name: judge schema: type: object properties: score: type: number format: float additionalProperties: true required: - score strict: false ``` ### LLM-as-a-Judge[​](#llm-judge "Direct link to LLM-as-a-Judge") Spicepod models can be used to provide eval scores for other models. To do so in Spice, the LLM must: 1. Return a valid JSON as the response. The JSON must have at least a single number field `.score`. e.g. ``` { "score": 0.42, "rationale": "It was a good story, they both are about love." } ``` 2. Use [Parameterized prompts](/docs/v1.10/features/large-language-models/parameterized_prompts) to provide details about the eval step. When used as an eval scorer, the model will be provided with the following variables: `input`, `actual` & `ideal`. The type of these variables will depend on the dataset, as per the [dataset format](#dataset-formats). --- # Model Context Protocol (MCP) The Model Context Protocol (MCP) helps integrate external tools and services into the Spice runtime. MCP tools can be run internally or connected over HTTP using the Server-Sent Events (SSE) protocol. ![Spice.ai Open Source Model-Context-Protocol (MCP) support](/assets/images/mcp-ea1443c9eb03219835e6945023b5b036.png) ## Overview[​](#overview "Direct link to Overview") MCP enables Spice to: 1. Run stdio-based MCP servers internally. 2. Connect to external MCP servers over SSE. This flexibility helps extend the capabilities of language models by providing access to external tools and services. ## Configuring MCP Tools[​](#configuring-mcp-tools "Direct link to Configuring MCP Tools") To configure MCP tools, define them in the `tools` section of your `spicepod.yaml` file. The `from` field specifies the transport mechanism (e.g., `mcp:npx` for stdio or an HTTP URL for SSE). ### Example: Adding an MCP Tool[​](#example-adding-an-mcp-tool "Direct link to Example: Adding an MCP Tool") ``` tools: - name: google_maps from: mcp:npx params: mcp_args: -y @modelcontextprotocol/server-google-maps ``` ### Example: Connecting to an External MCP Server[​](#example-connecting-to-an-external-mcp-server "Direct link to Example: Connecting to an External MCP Server") ``` tools: - name: external_mcp_server from: mcp:http://example.com/v1/mcp/sse ``` ## Using MCP Tools with Models[​](#using-mcp-tools-with-models "Direct link to Using MCP Tools with Models") Once configured, MCP tools can be assigned to models via the `tools` parameter. ``` models: - name: model_with_mcp from: openai:gpt-4o params: tools: google_maps ``` ## Spice as an MCP Server[​](#spice-as-an-mcp-server "Direct link to Spice as an MCP Server") Spice can also act as an MCP server, exposing its tools over SSE. This enables other Spice instances or external systems to connect and use the tools. ### Example: Connecting to another Spice instance via MCP[​](#example-connecting-to-another-spice-instance-via-mcp "Direct link to Example: Connecting to another Spice instance via MCP") ``` tools: - name: spice_instance from: mcp:http://localhost:8090/v1/mcp/sse ``` ## Additional Configuration Options[​](#additional-configuration-options "Direct link to Additional Configuration Options") ### `from`[​](#from "Direct link to from") The `from` field specifies the transport mechanism for the MCP tool: * **SSE**: Use an HTTP URL ending with `/sse` (e.g., `http://localhost:8090/v1/mcp/sse`). * **Stdio**: Use commands like `mcp:npx` or `mcp:docker`. Additional arguments can be passed via `params.mcp_args`. ### `params`[​](#params "Direct link to params") The `params` field provides additional configuration for MCP tools. For stdio-based tools, use `mcp_args` to specify command-line arguments. ``` tools: - name: custom_tool from: mcp:npx params: mcp_args: -y @custom/tool ``` ### `env`[​](#env "Direct link to env") For stdio-based MCP tools, environment variables can be set using the `env` field. ``` tools: - name: tool_with_env from: mcp:docker env: API_KEY: your_api_key ``` For more details, see the [MCP Tools Reference](/docs/v1.10/components/tools/mcp). --- # Language Model Memory Spice provides memory persistence tools that help language models store and retrieve information across conversations. These tools are available through the `memory` tool group. ## Enabling Memory Tools[​](#enabling-memory-tools "Direct link to Enabling Memory Tools") To enable memory tools for Spice models, define a `store` [memory](/docs/v1.10/components/data-connectors/memory) dataset and specify `memory` in the model's `tools` parameter. ### Example: Enabling Memory Tools[​](#example-enabling-memory-tools "Direct link to Example: Enabling Memory Tools") ``` datasets: - from: memory:store name: llm_memory access: read_write models: - name: memory-enabled-model from: openai:gpt-4o params: tools: memory, sql # Can be combined with other tool groups ``` For more information on tools, see [Tool components](/docs/v1.10/components/tools). --- # Language Model Overrides ### Chat Completion Parameter Overrides[​](#chat-completion-parameter-overrides "Direct link to Chat Completion Parameter Overrides") The [`v1/chat/completion`](/docs/v1.10/api/HTTP/post-chat-completions) endpoint is compatible with OpenAI's API. It supports a subset of request body parameters defined in the [OpenAI reference documentation](https://platform.openai.com/docs/api-reference/chat/create). Spice helps configure different defaults for these request parameters. Supported parameters: * [`frequency_penalty`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-frequency_penalty) * [`logit_bias`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-logit_bias) * [`logprobs`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-logprobs) * [`max_completion_tokens`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-max_completion_tokens) * [`metadata`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-metadata) * [`n`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-n) * [`parallel_tool_calls`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-parallel_tool_calls) * [`presence_penalty`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-presence_penalty) * [`response_format`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format) * [`seed`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-seed) * [`stop`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-stop) * [`store`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-store) * [`stream`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-stream) * [`stream_options`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-stream_options) * [`temperature`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-temperature) * [`tool_choice`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-tool_choice) * [`tools`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-tools) * [`top_logprobs`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-top_logprobs) * [`top_p`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-top_p) * [`user`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-user) ### Example: Setting Default Overrides[​](#example-setting-default-overrides "Direct link to Example: Setting Default Overrides") Deprecated Default Overrides Parameters The `openai_` prefix is deprecated for non-OpenAI model providers. Use the [model provider prefix](/docs/v1.10/components/models#model-provider-prefix) instead. To specify a default override for a parameter, use the [model provider prefix](/docs/v1.10/components/models#model-provider-prefix) followed by the parameter name. For example, to set the `temperature` parameter to `0.1` for all requests with this model for Hugging Face model, use `hf_temperature: 0.1`. A `temperature` parameter in the request body will still override the default. ``` models: - name: pirate-haikus from: openai:gpt-4o params: openai_temperature: 0.1 openai_response_format: { 'type': 'json_object' } ``` When sending this payload to spice `/v1/chat/completions`: ``` { "model": "pirate-haikus", "messages": [ { "role": "user", "content": "What is the capital of France?" } ], "temperature": 0.5 } ``` Will be passed to the OpenAI API as: ``` { "model": "gpt-4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ], "temperature": 0.5, // temperature overriden by value in request body "response_format": { "type": "json_object" } // default response format from model configuration } ``` ### System Prompt[​](#system-prompt "Direct link to System Prompt") In addition to any system prompts provided in message dialogue, or added by model providers, Spice can configure an additional system prompt. ``` models: - name: pirate-haikus from: openai:gpt-4o params: system_prompt: | Write everything in Haiku like a pirate ``` Any request to [HTTP `v1/chat/completion`](/docs/v1.10/api/HTTP/post-chat-completions) will include the configured system prompt. ### Example: Enforcing default structured output and using system prompt[​](#example-enforcing-default-structured-output-and-using-system-prompt "Direct link to Example: Enforcing default structured output and using system prompt") This example demonstrates how to create a specialized math tutoring model by combining system prompts with structured JSON output. The configuration ensures consistent, step-by-step mathematical solutions in a machine-readable format. ``` models: - name: math-tutor from: openai:gpt-4o params: system_prompt: | You are a helpful math tutor. Guide the user through the solution step by step. openai_response_format: type: json_schema json_schema: name: math_reasoning schema: type: object properties: steps: type: array items: type: object properties: explanation: type: string output: type: string required: - explanation - output additionalProperties: false final_answer: type: string required: - steps - final_answer additionalProperties: false strict: true ``` To use the configured math tutor, send a simple request to the chat completions endpoint: ``` curl -s -XPOST http://localhost:8090/v1/chat/completions -H "Content-Type: application/json" -d \ '{ "model": "math-tutor", "messages": [{ "role": "user", "content" :"how can I solve 8x + 7 = -23" }] }' \ | jq '.choices[0].message.content | fromjson' ``` Example response: ``` { "final_answer": "x = -3.75", "steps": [ { "explanation": "We start with the given equation that we need to solve.", "output": "8x + 7 = -23" }, { "explanation": "Our goal is to solve for x. We can start by isolating the term with x on one side of the equation. To do this, we need to eliminate the constant term (7) on the left side. We subtract 7 from both sides of the equation in order to keep it balanced.", "output": "8x + 7 - 7 = -23 - 7" }, { "explanation": "Subtracting 7 from both sides simplifies the equation. On the left side, the +7 and -7 cancel out, leaving just the term with the variable.", "output": "8x = -30" }, { "explanation": "Now, we have 8 times x equals -30. To solve for x, we divide both sides of the equation by the coefficient of x, which is 8.", "output": "8x / 8 = -30 / 8" }, { "explanation": "Dividing both sides results in x on the left side and simplifies the fraction on the right side. The fraction -30/8 can be simplified further by dividing both the numerator and the denominator by their greatest common divisor, which is 2.", "output": "x = -3.75" }, { "explanation": "The solution has been simplified completely, giving us the value of x.", "output": "x = -3.75" } ] } ``` Visit [OpenAI Structured Outputs](https://platform.openai.com/docs/guides/structured-outputs) for more information on how to use structured output formats. --- # System Prompt parameterization Spice supports defining system prompts for Large Language Models (LLM)s in the [spicepod](/docs/v1.10/features/large-language-models/parameter_overrides#system-prompt). **Example**: ``` models: - name: advice from: openai:gpt-4o params: system_prompt: | Write everything in Haiku like a pirate from Australia ``` More than this, system prompts can use Jinja syntax to allow system prompts to be altered on each [v1/chat/completion](/docs/v1.10/api/HTTP/post-chat-completions) request. This involves three steps: 1. Add `parameterized_prompt: enabled` to the model. 2. Use Jinja syntax in the `system_prompt` parameter for the model in the spicepods. ``` models: - name: advice from: openai:gpt-4o params: parameterized_prompt: enabled system_prompt: | Write everything in {{ form }} like a {{ user.character }} from {{ user.country }} ``` 3. Provide the required variables in [v1/chat/completion](/docs/v1.10/api/HTTP/post-chat-completions) via the `.metadata` field. ``` curl -X POST http://localhost:8090/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "advice", "messages": [ {"role": "user", "content": "Where should I visit in San Francisco?"} ], "metadata": { "form": "haiku", "user": { "character": "pirate", "country": "australia" } } }' ``` --- # Load and Serve Models Locally Spice supports loading and serving LLMs from various sources for embeddings and inference, including local filesystems and Hugging Face. ### Example: Loading a LLM from Hugging Face[​](#example-loading-a-llm-from-hugging-face "Direct link to Example: Loading a LLM from Hugging Face") ``` models: - name: llama_3.2_1B from: huggingface:huggingface.co/meta-llama/Llama-3.2-1B params: hf_token: ${ secrets:HF_TOKEN } ``` ## Filesystem[​](#filesystem "Direct link to Filesystem") Models can be hosted on a local filesystem and referenced directly in the configuration. For more details, see the [Filesystem Model Component](/docs/v1.10/components/models/filesystem). ## Hugging Face[​](#hugging-face "Direct link to Hugging Face") Spice integrates with Hugging Face, enabling you to use a wide range of pre-trained models. For more information, see the [Hugging Face Model Component](/docs/v1.10/components/models/huggingface). --- # Language Models Tools Spice provides tools that help LLMs interact with the runtime. To specify these tools for a Spice model, include them in its `params.tools`. For a list of available tools, or how to define additional tools, see [Tool Components](/docs/v1.10/components/tools). ### Example: Specifying Tools for a Model[​](#example-specifying-tools-for-a-model "Direct link to Example: Specifying Tools for a Model") ``` models: - name: sql-model from: openai:gpt-4o params: tools: list_datasets, sql, table_schema ``` ### Example: Specifying tools via a Tool Group[​](#example-specifying-tools-via-a-tool-group "Direct link to Example: Specifying tools via a Tool Group") ``` - name: full-runtime from: openai:gpt-4o params: tools: auto # Use all default tools ``` For details on tool groups, see [Tool Components](/docs/v1.10/components/tools#tool-groups). ### Example: Specifying tools and tool groups[​](#example-specifying-tools-and-tool-groups "Direct link to Example: Specifying tools and tool groups") ``` models: - name: full-runtime from: openai:gpt-4o params: tools: memory, sql ``` ### Tool Recursion Limit[​](#tool-recursion-limit "Direct link to Tool Recursion Limit") When a model requests to call a runtime tool, Spice runs the tool internally and feeds it back to the model. The `tool_recursion_limit` parameter limits the depth of internal recursion Spice will undertake. By default, this limit is set to 10. ``` models: - name: my-model from: openai params: tool_recursion_limit: 3 ``` --- # Machine Learning Models Spice supports loading and serving ONNX models for inference, from sources including local filesystems, Hugging Face, and the Spice.ai Cloud platform. Example `spicepod.yml` loading an ONNX model from HuggingFace: ``` models: - from: huggingface:huggingface.co/spiceai/darts:latest name: hf_model files: - path: model.onnx datasets: - taxi_trips ``` ## Filesystem[​](#filesystem "Direct link to Filesystem") Models can be hosted on a local filesystem and referenced directly in the configuration. For more details, see the [Filesystem Model Component](/docs/v1.10/components/models/filesystem). ## Hugging Face[​](#hugging-face "Direct link to Hugging Face") Spice integrates with Hugging Face, enabling you to use a wide range of pre-trained models. For more information, see the [Hugging Face Model Component](/docs/v1.10/components/models/huggingface). ## Spice Cloud Platform[​](#spice-cloud-platform "Direct link to Spice Cloud Platform") The Spice Cloud platform provides a scalable environment for training, hosting, and managing your models. For further details, see the [Spice Cloud Platform Model Component](/docs/v1.10/components/models/spiceai). --- # Observability & Monitoring Spice provides monitoring and observability through three mechanisms: * **Prometheus-compatible metrics endpoint**: Exposes metrics in the [Prometheus exposition format](https://prometheus.io/docs/instrumenting/exposition_formats/#basic-info) for scraping by monitoring systems like [Datadog](https://www.datadoghq.com/), [New Relic](https://newrelic.com/), and [Chronosphere](https://chronosphere.io/). * **OpenTelemetry metrics export**: Pushes metrics to an [OpenTelemetry](https://opentelemetry.io/) collector using gRPC. * **Distributed tracing**: Integrates with [Zipkin](https://zipkin.io/) and compatible tracing systems for request tracing. ![observability](https://github.com/user-attachments/assets/2468e3e7-4fb4-4a74-8b26-45eeeee90310) ### Monitoring Integrations[​](#monitoring-integrations "Direct link to Monitoring Integrations") * [Datadog](/docs/v1.10/monitoring/datadog) * [Grafana & Prometheus](/docs/v1.10/monitoring/grafana) * [Zipkin](/docs/v1.10/monitoring/zipkin) ## Prometheus Metrics Endpoint[​](#prometheus-metrics-endpoint "Direct link to Prometheus Metrics Endpoint") Spice exposes a Prometheus-compatible metrics endpoint that monitoring systems can scrape. The endpoint serves metrics in the [Prometheus exposition format](https://prometheus.io/docs/instrumenting/exposition_formats/), which is supported by most enterprise monitoring platforms including Datadog, New Relic, Chronosphere, Grafana Cloud, and others. ### Default Configuration[​](#default-configuration "Direct link to Default Configuration") The metrics endpoint listens on port `9090` by default. The endpoint address is logged at startup: ``` 2024-11-28T19:48:10.942003Z INFO runtime::metrics_server: Spice Runtime Metrics listening on 127.0.0.1:9090 ``` ### Custom Port Binding[​](#custom-port-binding "Direct link to Custom Port Binding") Use the `--metrics` flag to bind to a specific address and port: ``` spiced --metrics 0.0.0.0:9091 ``` For Docker deployments: ``` FROM spiceai/spiceai:latest CMD ["--metrics", "0.0.0.0:9090"] EXPOSE 9090 ``` ### Verifying the Endpoint[​](#verifying-the-endpoint "Direct link to Verifying the Endpoint") Verify the metrics endpoint is working with a GET request: ``` curl http://localhost:9090/metrics # HELP runtime_flight_server_started Indicates the runtime Flight server has started. # TYPE runtime_flight_server_started counter runtime_flight_server_started 1 # HELP runtime_http_server_started Indicates the runtime HTTP server has started. # TYPE runtime_http_server_started counter runtime_http_server_started 1 # HELP dataset_load_state Status of the dataset. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. # TYPE dataset_load_state gauge dataset_load_state{dataset="taxi_trips"} 2 dataset_load_state{dataset="taxi_trips_accelerated"} 2 # HELP dataset_active_count Number of currently loaded datasets. # TYPE dataset_active_count gauge dataset_active_count{engine="None"} 1 dataset_active_count{engine="duckdb"} 1 ... ``` ## OpenTelemetry Metrics Exporter[​](#opentelemetry-metrics-exporter "Direct link to OpenTelemetry Metrics Exporter") Spice can push metrics to an [OpenTelemetry](https://opentelemetry.io/) collector, enabling integration with platforms such as [Jaeger](https://www.jaegertracing.io/), [New Relic](https://newrelic.com/), [Honeycomb](https://www.honeycomb.io/), and other OpenTelemetry-compatible backends. ### Configuration[​](#configuration "Direct link to Configuration") Configure the OpenTelemetry exporter in `spicepod.yaml` under `runtime.telemetry.otel_exporter`: | Parameter | Required | Default | Description | | --------------- | -------- | ------- | --------------------------------------------------------------------- | | `enabled` | No | `true` | Whether the OpenTelemetry exporter is enabled. | | `endpoint` | Yes | - | The OpenTelemetry collector endpoint. | | `push_interval` | No | `60s` | How frequently metrics are pushed to the collector. | | `metrics` | No | `[]` | List of metric names to export. When empty, all metrics are exported. | ### Protocol[​](#protocol "Direct link to Protocol") Spice currently supports only the gRPC protocol for OpenTelemetry metrics export. Specify the collector `endpoint` as a host and port (e.g., `localhost:4317`). ### Examples[​](#examples "Direct link to Examples") **gRPC (default port 4317):** ``` runtime: telemetry: enabled: true otel_exporter: endpoint: 'localhost:4317' push_interval: '30s' ``` ### Metric Filtering[​](#metric-filtering "Direct link to Metric Filtering") To export only specific metrics, use the `metrics` parameter: ``` runtime: telemetry: enabled: true otel_exporter: endpoint: 'localhost:4317' metrics: - query_duration_ms - query_executions - dataset_load_state ``` When `metrics` is empty or omitted, all available metrics are exported. For full configuration details, see the [runtime.telemetry reference](/docs/v1.10/reference/spicepod/runtime#runtimetelemetry). ## Available Metrics[​](#available-metrics "Direct link to Available Metrics") Spice exposes the following metrics. All metrics include relevant labels (dimensions) for filtering and aggregation. | Metric | Description | | --------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `accelerated_ready_state_federated_fallback`
*(count)* | Number of times the federated table was queried due to the accelerated table loading the initial data. | | `accelerated_zero_results_federated_fallback`
*(count)* | Number of times the federated table was queried due to the accelerated table returning zero results. | | `ai_inferences_with_spice_count`
*(count)* | AI Inferences with Spice count. | | `catalog_load_errors`
*(count)* | Number of errors loading the catalog provider. | | `catalog_load_state`
*(gauge)* | Status of the catalog provider. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | | `component_metric_registered_count`
*(gauge)* | Number of currently registered component metrics. | | `dataset_acceleration_ingestion_lag_ms`
*(gauge)* | Lag between the current wall-clock time and the maximum time\_column value after the refresh operation, in milliseconds. [Disabled by default](/docs/v1.10/reference/spicepod/runtime#runtimemetrics) | | `dataset_acceleration_last_refresh_time_ms`
*(gauge)* | Unix timestamp in milliseconds when the last refresh completed. [Disabled by default](/docs/v1.10/reference/spicepod/runtime#runtimemetrics) | | `dataset_acceleration_max_timestamp_after_refresh_ms`
*(gauge)* | Maximum value of the dataset's time\_column after the refresh operation, in milliseconds. [Disabled by default](/docs/v1.10/reference/spicepod/runtime#runtimemetrics) | | `dataset_acceleration_max_timestamp_before_refresh_ms`
*(gauge)* | Maximum value of the dataset's time\_column before the refresh operation, in milliseconds. [Disabled by default](/docs/v1.10/reference/spicepod/runtime#runtimemetrics) | | `dataset_acceleration_refresh_data_fetches_skipped`
*(count)* | Number of refresh data fetches skipped due to unchanged file metadata. | | `dataset_acceleration_refresh_duration_ms`
*(histogram)* | Duration in milliseconds to load a full or appended refresh data. | | `dataset_acceleration_refresh_errors`
*(count)* | Number of errors refreshing the dataset. | | `dataset_acceleration_refresh_lag_ms`
*(gauge)* | Difference between the maximum time\_column value after and before the refresh operation, in milliseconds. | | `dataset_acceleration_refresh_worker_panics`
*(count)* | Number of times a refresh worker panicked while refreshing a dataset. | | `dataset_acceleration_snapshot_bootstrap_bytes`
*(gauge)* | Number of bytes downloaded when bootstrapping the acceleration from a snapshot. | | `dataset_acceleration_snapshot_bootstrap_checksum`
*(gauge)* | Checksum of the snapshot downloaded during bootstrap (emitted with `checksum` attribute). | | `dataset_acceleration_snapshot_bootstrap_duration_ms`
*(count)* | Time in milliseconds taken to download the snapshot used to bootstrap acceleration. | | `dataset_acceleration_snapshot_failure_count`
*(count)* | Number of failures encountered while writing snapshots. | | `dataset_acceleration_snapshot_write_bytes`
*(gauge)* | Number of bytes written for the most recent snapshot. | | `dataset_acceleration_snapshot_write_checksum`
*(gauge)* | Checksum of the most recent snapshot write (emitted with `checksum` attribute). | | `dataset_acceleration_snapshot_write_duration_ms`
*(histogram)* | Time in milliseconds taken to write the latest snapshot to object storage. | | `dataset_acceleration_snapshot_write_timestamp`
*(gauge)* | Unix timestamp (seconds) when the most recent snapshot write completed. | | `dataset_active_count`
*(gauge)* | Number of currently loaded datasets. | | `dataset_load_errors`
*(count)* | Number of errors loading the dataset. | | `dataset_load_state`
*(gauge)* | Status of the dataset. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | | `dataset_unavailable_time_ms`
*(gauge)* | Time dataset went offline in milliseconds. | | `embeddings_active_count`
*(gauge)* | Number of currently loaded embeddings. | | `embeddings_cache_evictions`
*(count)* | Number of cache evictions. | | `embeddings_cache_hit_ratio`
*(gauge)* | Cache hit ratio (hits / total requests). | | `embeddings_cache_hits`
*(count)* | Cache hit count. | | `embeddings_cache_items_count`
*(gauge)* | Number of items currently in the cache. | | `embeddings_cache_max_size_bytes`
*(gauge)* | Maximum allowed size of the cache in bytes. | | `embeddings_cache_misses`
*(count)* | Cache miss count. | | `embeddings_cache_requests`
*(count)* | Number of requests to get a key from the cache. | | `embeddings_cache_size_bytes`
*(gauge)* | Size of the cache in bytes. | | `embeddings_cache_stale_swr_count`
*(count)* | Number of stale-while-revalidate background refreshes skipped due to existing in-flight revalidation. | | `embeddings_cache_swr_background_query_count`
*(count)* | Number of background queries triggered for stale-while-revalidate cache refreshes. | | `embeddings_failures`
*(count)* | Number of embedding failures. | | `embeddings_internal_request_duration_ms`
*(histogram)* | The duration of running an embedding(s) internally. | | `embeddings_load_errors`
*(count)* | Number of errors loading the embedding. | | `embeddings_load_state`
*(gauge)* | Status of the embedding. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | | `embeddings_requests`
*(count)* | Number of embedding requests. | | `flight_do_exchange_data_updates_sent`
*(count)* | Number of data updates sent via DoExchange. | | `flight_request_duration_ms`
*(histogram)* | Measures the duration of Flight requests in milliseconds. | | `flight_requests`
*(count)* | Total number of Flight requests. | | `http_requests`
*(count)* | Number of HTTP requests. | | `http_requests_duration_ms`
*(histogram)* | Measures the duration of HTTP requests in milliseconds. | | `llm_failures`
*(count)* | Number of LLM failures. | | `llm_internal_request_duration_ms`
*(histogram)* | The duration of running an LLM request internally. | | `llm_load_state`
*(gauge)* | Status of the LLM model. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | | `llm_requests`
*(count)* | Number of LLM requests. | | `model_active_count`
*(gauge)* | Number of currently loaded models. | | `model_load_duration_ms`
*(histogram)* | Duration in milliseconds to load the model. | | `model_load_errors`
*(count)* | Number of errors loading the model. | | `model_load_state`
*(gauge)* | Status of the model. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | | `query_active_count`
*(histogram)* | Number of concurrent top-level queries actively being processed in the runtime. Includes the `protocol` dimension (`http`, `flight`, `flightsql`, `internal`) to indicate the query type. | | `query_duration_ms`
*(histogram)* | The total amount of time spent planning and executing queries in milliseconds. | | `query_execution_duration_ms`
*(histogram)* | The total amount of time spent only executing queries (0 for cached queries). | | `query_executions`
*(count)* | Number of query executions. | | `query_failures`
*(count)* | Number of query failures. | | `query_processed_bytes`
*(count)* | Number of bytes processed by the runtime. | | `query_produced_spills`
*(count)* | Number of spills produced by the query. | | `query_returned_bytes`
*(count)* | Number of bytes returned to query clients. | | `query_returned_rows`
*(histogram)* | Number of rows returned to query clients. | | `query_spilled_bytes`
*(count)* | Number of spilled bytes produced by the query. | | `query_spilled_rows`
*(count)* | Number of spilled rows produced by the query. | | `results_cache_evictions`
*(count)* | Number of cache evictions. | | `results_cache_hit_ratio`
*(gauge)* | Cache hit ratio (hits / total requests). | | `results_cache_hits`
*(count)* | Cache hit count. | | `results_cache_items_count`
*(gauge)* | Number of items currently in the cache. | | `results_cache_max_size_bytes`
*(gauge)* | Maximum allowed size of the cache in bytes. | | `results_cache_misses`
*(count)* | Cache miss count. | | `results_cache_requests`
*(count)* | Number of requests to get a key from the cache. | | `results_cache_size_bytes`
*(gauge)* | Size of the cache in bytes. | | `results_cache_stale_swr_count`
*(count)* | Number of stale-while-revalidate background refreshes skipped due to existing in-flight revalidation. | | `results_cache_swr_background_query_count`
*(count)* | Number of background queries triggered for stale-while-revalidate cache refreshes. | | `runtime_flight_server_started`
*(count)* | Indicates the runtime Flight server has started. | | `runtime_http_server_started`
*(count)* | Indicates the runtime HTTP server has started. | | `search_results_cache_evictions`
*(count)* | Number of cache evictions. | | `search_results_cache_hit_ratio`
*(gauge)* | Cache hit ratio (hits / total requests). | | `search_results_cache_hits`
*(count)* | Search cache hit count. | | `search_results_cache_items_count`
*(gauge)* | Number of items currently in the search cache. | | `search_results_cache_max_size_bytes`
*(gauge)* | Maximum allowed size of the search cache in bytes. | | `search_results_cache_misses`
*(count)* | Cache miss count. | | `search_results_cache_requests`
*(count)* | Number of requests to get a key from the search cache. | | `search_results_cache_size_bytes`
*(gauge)* | Size of the search cache in bytes. | | `search_results_cache_stale_swr_count`
*(count)* | Number of stale-while-revalidate background refreshes skipped due to existing in-flight revalidation. | | `search_results_cache_swr_background_query_count`
*(count)* | Number of background queries triggered for stale-while-revalidate cache refreshes. | | `secrets_store_load_duration_ms`
*(histogram)* | Duration in milliseconds to load the secret stores. | | `tool_active_count`
*(gauge)* | Number of currently loaded LLM tools. | | `tool_load_errors`
*(count)* | Number of errors loading the LLM tool. | | `tool_load_state`
*(gauge)* | Status of the LLM tools. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | | `view_load_errors`
*(count)* | Number of errors loading the view. | | `view_load_state`
*(gauge)* | Status of the views. 0=Initializing, 1=Ready, 2=Disabled, 3=Error, 4=Refreshing, 5=ShuttingDown. | | `worker_active_count`
*(gauge)* | Number of currently loaded workers. | | `workers_load_duration_ms`
*(histogram)* | Duration in milliseconds to load the worker. | Component Metrics In addition to these core metrics, individual components can expose their own metrics. For example, the MySQL data connector exposes [connection pool metrics](/docs/v1.10/components/data-connectors/mysql#metrics). See [Component Metrics](/docs/v1.10/features/observability/component_metrics) for more information. --- # Component Metrics Component metrics provide detailed insights into the internal state and performance of individual components in Spice. Each component can expose its own set of metrics that can be enabled selectively to monitor specific aspects of its operation. ## Enabling Component Metrics[​](#enabling-component-metrics "Direct link to Enabling Component Metrics") Component metrics are disabled by default and can be enabled by adding a `metrics` section to the component configuration. Each metric can be enabled individually by specifying its name in the metrics list. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") ``` datasets: - from: some_component:my_resource name: my_resource metrics: - name: metric_one enabled: true - name: metric_two enabled: true - name: metric_three enabled: false params: param_one: value_one param_two: value_two ``` ## Available Metrics[​](#available-metrics "Direct link to Available Metrics") Each component defines its own set of available metrics. These metrics are exposed in the Prometheus format with the following naming convention: ``` {component_type}_{component_name}_{metric_name} ``` For example, a MySQL dataset component's metrics would be prefixed with `dataset_mysql_`. ## Monitoring Component Metrics[​](#monitoring-component-metrics "Direct link to Monitoring Component Metrics") Component metrics are exposed through the same Prometheus-compatible metrics endpoint as other Spice metrics. These metrics can be accessed using standard Prometheus tools or any monitoring system that supports Prometheus metrics. To view the metrics, make a GET request to the metrics endpoint: ``` curl http://localhost:9090/metrics ``` The response will include all enabled component metrics in Prometheus format, with proper HELP and TYPE annotations. ## Component-Specific Metrics[​](#component-specific-metrics "Direct link to Component-Specific Metrics") For detailed information about metrics available for specific components, view all [components that expose metrics](/docs/v1.10/tags/component-metrics). --- # Query Federation Spice supports query federation, enabling you to join, combine, and query data using SQL from multiple sources, including databases (PostgreSQL, MySQL), data warehouses (Databricks, Snowflake, BigQuery), and data lakes (S3, MinIO). ![Spice.ai Open Source Query Federation](/assets/images/query-federation-d1077a7b9c335e975aeec36b9dffd8ae.png) For a full list of supported sources, see [Data Connectors](/docs/v1.10/components/data-connectors). ## Getting Started[​](#getting-started "Direct link to Getting Started") To start using federated queries in Spice, follow these steps: **Step 1.** Install Spice by following the [installation instructions](/docs/v1.10/getting-started). **Step 2.** Clone the Spice Cookbook repository and navigate to the `federation` directory. ``` git clone https://github.com/spiceai/cookbook.git cd cookbook/federation ``` **Step 3.** Login to the demo Dremio. ``` spice login dremio -u demo -p demo1234 ``` **Step 4.** Create a new Spice app called `demo`. ``` # Create Spice app "demo" spice init demo # Change to demo directory. cd demo ``` **Step 5.** Add the `spiceai/fed-demo` Spicepod. ``` # Change to demo directory. cd demo spice add spiceai/fed-demo ``` Note in the Spice runtime output several datasets are loaded. **Step 6.** Start the Spice runtime. ``` spice run ``` **Step 7.** Show available tables and query them, regardless of source. ``` # Start the Spice SQL REPL. spice sql ``` Show the available tables: ``` show tables; ``` Execute the queries: ``` -- Query S3 (Parquet) SELECT * FROM s3_source LIMIT 10; -- Query S3 (Parquet) accelerated SELECT * FROM s3_source_accelerated LIMIT 10; -- Query Dremio SELECT * FROM dremio_source LIMIT 10; -- Query Dremio accelerated SELECT * FROM dremio_source_accelerated LIMIT 10; ``` **Step 8.** Join tables across remote sources and locally accelerated source ``` -- Query across S3 and Dremio WITH all_sales AS ( SELECT sales FROM s3_source UNION ALL select fare_amount+tip_amount as sales from dremio_source ) SELECT SUM(sales) as total_sales, COUNT(*) AS total_transactions, MAX(sales) AS max_sale, AVG(sales) AS avg_sale FROM all_sales; +--------------------+--------------------+----------+--------------------+ | total_sales | total_transactions | max_sale | avg_sale | +--------------------+--------------------+----------+--------------------+ | 11501140.079999998 | 102823 | 14082.8 | 111.85376890384445 | +--------------------+--------------------+----------+--------------------+ Time: 1.079320792 seconds. 1 rows. ``` **Step 9.** Join tables across locally accelerated sources and query ``` -- Query across S3 accelerated and Dremio accelerated WITH all_sales AS ( SELECT sales FROM s3_source_accelerated UNION ALL select fare_amount+tip_amount as sales from dremio_source_accelerated ) SELECT SUM(sales) as total_sales, COUNT(*) AS total_transactions, MAX(sales) AS max_sale, AVG(sales) AS avg_sale FROM all_sales; +-------------+--------------------+----------+--------------------+ | total_sales | total_transactions | max_sale | avg_sale | +-------------+--------------------+----------+--------------------+ | 11501140.08 | 102823 | 14082.8 | 111.85376890384447 | +-------------+--------------------+----------+--------------------+ Time: 0.011524375 seconds. 1 rows. ``` ### Acceleration[​](#acceleration "Direct link to Acceleration") While the query in step 8 successfully returned results from federated remote data sources, the performance was suboptimal due to data transfer overhead. Step 9 demonstrates the same query executed against locally materialized datasets using [Data Accelerators](/docs/v1.10/components/data-accelerators). By storing data locally, queries avoid network round-trips and achieve significantly faster response times. Limitations * **Query Performance:** Without acceleration, federated queries will be slower than local queries due to network latency and data transfer. * **Query Capabilities:** Not all SQL features and data types are supported across all data sources. More complex data type queries may not work as expected. ## Related Topics[​](#related-topics "Direct link to Related Topics") * [Distributed Query](/docs/v1.10/features/distributed-query) - Scale queries across multiple nodes * [Results Caching](/docs/v1.10/features/caching) - Cache query results for improved performance * [Arrow Flight SQL API](/docs/v1.10/api/arrow-flight-sql) - High-performance query protocol * [ADBC](/docs/v1.10/api/adbc) - Arrow Database Connectivity --- # Search Functionality > 🎓 For a practical walkthrough, see the: [Amazon S3 Vectors with Spice](https://spiceai.org/blog/amazon-s3-vectors-with-spice) engineering blog post. Spice provides robust search capabilities enabling developers to query datasets beyond traditional SQL, including semantic (vector-based) search, full-text keyword search, and hybrid search methods. ## Search Methods Overview[​](#search-methods-overview "Direct link to Search Methods Overview") Spice supports multiple search methods: * **Vector Search**: Semantic search using embeddings to retrieve data by meaning and similarity. * **Full-Text Search**: Keyword-driven search optimized for text data retrieval. * **Hybrid Search**: Combine multiple search methods using Reciprocal Rank Fusion (RRF) for improved relevance. * **SQL Search**: Traditional SQL queries for precise and structured searches. ### Vector Search[​](#vector-search "Direct link to Vector Search") Vector search uses embeddings—numerical representations of data—to identify similar or related content based on semantic meaning. **Requirements:** * Configured data connectors or accelerators * Defined embeddings for datasets **Getting Started:** * [Configure Embeddings](/docs/v1.10/components/embeddings) * [Performing Vector Search](/docs/v1.10/features/search/vector-search) **Example SQL Vector Search:** ``` SELECT id, extra_column, score FROM vector_search(my_table, 'search query') WHERE date_published > '2021-01-01' ORDER BY score DESC LIMIT 5 ``` For complete SQL UDTF specifications, see [Vector-Based Search SQL UDTF](/docs/v1.10/features/search/vector-search#sql-udtf). ### Full-Text Search[​](#full-text-search "Direct link to Full-Text Search") Full-text search efficiently retrieves records matching specific keywords. **Requirements:** * Indexed columns within datasets **Getting Started:** * [Full-Text Search Docs](/docs/v1.10/features/search/full-text) **Example SQL Full-Text Search:** ``` SELECT id, extra_column, score FROM text_search(my_table, 'search terms') WHERE date_published > '2021-01-01' ORDER BY score DESC LIMIT 5 ``` For detailed SQL UDTF instructions, see [Full-Text Search SQL UDTF](/docs/v1.10/features/search/full-text#searching-with-sql). ### Hybrid Search with RRF[​](#hybrid-search-with-rrf "Direct link to Hybrid Search with RRF") Reciprocal Rank Fusion (RRF) combines results by merging rankings from multiple search methods to improve relevance. **Requirements:** * Multiple search methods configured (vector, full-text, etc.) **Example SQL Hybrid Search:** ``` SELECT id, title, content, fused_score FROM rrf( vector_search(documents, 'machine learning algorithms'), text_search(documents, 'neural networks deep learning', content), join_key => 'id' -- join key for optimal performance ) ORDER BY fused_score DESC LIMIT 5 ``` For complete RRF syntax and parameters, see [Search SQL Reference](/docs/v1.10/reference/sql/search#reciprocal-rank-fusion-rrf). ## [📄️Vector Search](/docs/v1.10/features/search/vector-search) [Learn how Spice can perform searches using vector-based methods.](/docs/v1.10/features/search/vector-search) ## [📄️Full-text Search](/docs/v1.10/features/search/full-text) [Learn how Spice can perform full text search](/docs/v1.10/features/search/full-text) --- # Full-Text Search Spice provides full text search functionality with BM25 scoring. Datasets can be augmented with a full-text search index that enables efficient search. Dataset columns are included in the full-text index based on the column configuration. ## Enabling Full-Text Search[​](#enabling-full-text-search "Direct link to Enabling Full-Text Search") To enable full-text search, configure your dataset columns within your dataset definition as follows: ``` datasets: - from: github:github.com/spiceai/docs/pulls name: doc.pulls params: github_token: ${secrets:GITHUB_TOKEN} acceleration: enabled: true columns: - name: title full_text_search: enabled: true row_id: - id - name: body full_text_search: enabled: true ``` In this example, full-text search indexing is enabled on both the `title` and `body` columns. The `row_id` specifies a unique identifier for referencing search results and retrieving additional data. ## Searching with the HTTP API[​](#searching-with-the-http-api "Direct link to Searching with the HTTP API") After enabling indexing, you can perform searches using the HTTP API endpoint `/v1/search`. Results will be ranked based on the relevance to your keyword query across indexed columns (`title` and `body` in this example). For details on using this endpoint, see the \[API reference for `/v1/search`(../../api/HTTP/post-search). ## Searching with SQL[​](#searching-with-sql "Direct link to Searching with SQL") Spice also provides full-text search through SQL using a user-defined table function (UDTF), `text_search()`. ### Example SQL Query[​](#example-sql-query "Direct link to Example SQL Query") Here's how you can query using SQL: ``` SELECT id, title, score FROM text_search(doc.pulls, 'search keywords', body) ORDER BY score DESC LIMIT 5; ``` This returns the top 5 results from the `doc.pulls` dataset that best match your search keywords within the `body` column. ### Function Signature[​](#function-signature "Direct link to Function Signature") The `text_search()` function has the following signature: ``` text_search( table STRING, -- Dataset name (required) query STRING, -- Keyword or phrase to search (required) col STRING, -- Specific column to search (required if dataset has multiple indexed columns) limit INTEGER, -- Maximum results returned (optional, defaults to 1000) include_score BOOLEAN -- Include relevance scores in results (optional, defaults to TRUE) ) RETURNS TABLE -- Original table columns plus an optional FLOAT column `score` ``` By default, `text_search` retrieves up to 1000 results. To adjust this, specify the `limit` parameter in the function call. Use this function to integrate robust full-text search directly into your data workflows with minimal setup. --- # Vector-Based Search > 🎓 Learn how it works with the [Amazon S3 Vectors with Spice](https://spiceai.org/blog/amazon-s3-vectors-with-spice) engineering blog post. Spice provides advanced vector-based search capabilities, enabling more nuanced and intelligent searches. ## Embedding Models[​](#embedding-models "Direct link to Embedding Models") Spice supports two types of embedding providers: * **Local embedding models** e.g., [sentence-transformers/all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2). * **Remote embedding services** e.g., [OpenAI Embeddings API](https://platform.openai.com/docs/api-reference/embeddings/create). Embedding models are defined in the `spicepod.yaml` file as top-level components. ``` embeddings: - name: openai_embeddings from: openai params: openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY } - name: local_embedding_model from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2 ``` ## Configuring Datasets for Embeddings[​](#configuring-datasets-for-embeddings "Direct link to Configuring Datasets for Embeddings") To enable vector search, specify embeddings for the dataset columns in `spicepod.yaml`: ``` datasets: - from: github:github.com/spiceai/spiceai/issues name: spiceai.issues params: github_token: ${ secrets:GITHUB_TOKEN } acceleration: enabled: true columns: - name: body embeddings: - from: local_embedding_model ``` This configuration instructs Spice to create embeddings from the `body` column, enabling similarity searches on body content. ## Performing a Vector Search[​](#performing-a-vector-search "Direct link to Performing a Vector Search") Execute similarity searches using Spice's HTTP API: ``` curl -X POST http://localhost:8090/v1/search \ -H 'Content-Type: application/json' \ -d '{ "datasets": ["spiceai.issues"], "text": "cutting edge AI", "where": "author=\"jeadie\"", "additional_columns": ["title", "state"], "limit": 2 }' ``` For detailed API documentation, see [Search API Reference](/docs/v1.10/api/HTTP/post-search). ## Retrieving Full Documents[​](#retrieving-full-documents "Direct link to Retrieving Full Documents") If the dataset uses chunking, Spice returns relevant chunks. To retrieve entire documents, include the embedding column in `additional_columns`: ``` curl -X POST http://localhost:8090/v1/search \ -H 'Content-Type: application/json' \ -d '{ "datasets": ["spiceai.issues"], "text": "cutting edge AI", "where": "array_has(assignees, \"jeadie\")", "additional_columns": ["title", "state", "body"], "limit": 2 }' ``` Response: ```` { "matches": [ { "value": "implements a scalar UDF `array_distance`:\n```\narray_distance(FixedSizeList[Float32], FixedSizeList[Float32])", "dataset": "spiceai.issues", "metadata": { "title": "Improve scalar UDF array_distance", "state": "Closed", "body": "## Overview\n- Previous PR https://github.com/spiceai/spiceai/pull/1601 implements a scalar UDF `array_distance`:\n```\narray_distance(FixedSizeList[Float32], FixedSizeList[Float32])\narray_distance(FixedSizeList[Float32], List[Float64])\n```\n\n### Changes\n - Improve using Native arrow function, e.g. `arrow_cast`, [`sub_checked`](https://arrow.apache.org/rust/arrow/array/trait.ArrowNativeTypeOp.html#tymethod.sub_checked)\n - Support a greater range of array types and numeric types\n - Possibly create a sub operator and UDF, e.g.\n\t- `FixedSizeList[Float32] - FixedSizeList[Float32]`\n\t- `Norm(FixedSizeList[Float32])`" }, "score": 0.66, }, { "value": "est external tools being returned for toolusing models", "dataset": "spiceai.issues", "metadata": { "title": "Automatic NSQL retries in /v1/nsql ", "state": "Open", "body": "To mimic our ability for LLMs to repeatedly retry tools based on errors, the `/v1/nsql`, which does not use this same paradigm, should retry internally.\n\nIf possible, improve the structured output to increase the likelihood of valid SQL in the response. Currently we just inforce JSON like this\n```json\n{\n "sql": "SELECT ..."\n}\n```" }, "score": 0.52, } ], "duration_ms": 45 } ```` ## SQL UDTF[​](#sql-udtf "Direct link to SQL UDTF") The embedding index can also be used to perform search in SQL, via a user-defined table function (UDTF). ``` SELECT id, title, score FROM vector_search('sales', 'cutting edge AI') ORDER BY score DESC LIMIT 5; ``` **SQL Function Signature of `vector_search`:** ``` vector_search( table STRING, -- Dataset name (required) query STRING, -- Search text (required) col STRING, -- Column name (optional if single embedding column) limit INTEGER, -- Results limit (default: 1000) include_score BOOLEAN -- Include relevance scores (default: TRUE) ) RETURNS TABLE -- The original table and: -- - A FLOAT column `score` (if `include_score`). ``` By default, `vector_search` retrieves up to 1000 results. To adjust this limit, specify the `limit` parameter in the function call. When using a specific vector engine, such as `s3_vectors` the limit defaults to that of the vector engine. ``` SELECT id, title, score FROM vector_search('sales', 'cutting edge AI', 1500) ORDER BY score DESC; ``` Limitations * `vector_search` UDTF does not yet support chunked embedding columns. Chunking support is on the roadmap. ## Using Existing Embeddings[​](#using-existing-embeddings "Direct link to Using Existing Embeddings") Spice supports vector searches on datasets with pre-existing embeddings. Ensure the dataset meets these requirements: 1. **Column Naming**: The embedding column name must be `_embedding`. 2. **Data Types**: Embedding columns must use Arrow types: * Non-chunked: `FixedSizeList[Float32|Float64, N]` * Chunked: `List[FixedSizeList[Float32|Float64, N]]` 3. **Offset Columns**: For chunked embeddings, an additional offset column (`_offsets`) is required: * Type: `List[FixedSizeList[Int32, 2]]`, indicating chunk boundaries. Example dataset structure (`sales` table): Non-chunked: ``` sql> describe sales; +-------------------+-----------------------------------------+-------------+ | column_name | data_type | is_nullable | +-------------------+-----------------------------------------+-------------+ | order_number | Int64 | YES | | quantity_ordered | Int64 | YES | | price_each | Float64 | YES | | order_line_number | Int64 | YES | | address | Utf8 | YES | | address_embedding | FixedSizeList( | NO | | | Field { | | | | name: "item", | | | | data_type: Float32, | | | | nullable: false, | | | | dict_id: 0, | | | | dict_is_ordered: false, | | | | metadata: {} | | | | }, | | | | 384 | | +-------------------+-----------------------------------------+-------------+ ``` Chunked: ``` sql> describe sales; +-------------------+-----------------------------------------+-------------+ | column_name | data_type | is_nullable | +-------------------+-----------------------------------------+-------------+ | order_number | Int64 | YES | | quantity_ordered | Int64 | YES | | price_each | Float64 | YES | | order_line_number | Int64 | YES | | address | Utf8 | YES | | address_embedding | List(Field { | NO | | | name: "item", | | | | data_type: FixedSizeList( | | | | Field { | | | | name: "item", | | | | data_type: Float32, | | | | }, | | | | 384 | | | | ), | | | | }) | | +-------------------+-----------------------------------------+-------------+ | address_offset | List(Field { | NO | | | name: "item", | | | | data_type: FixedSizeList( | | | | Field { | | | | name: "item", | | | | data_type: Int32, | | | | }, | | | | 2 | | | | ), | | | | }) | | +-------------------+-----------------------------------------+-------------+ ``` ### Constraints[​](#constraints "Direct link to Constraints") 1. **Underlying Column Presence:** * The underlying column must exist in the table, and be of `string` [Arrow data type](/docs/v1.10/reference/datatypes/accelerators) . 2. **Embeddings Column Naming Convention:** * For each underlying column, the corresponding embeddings column must be named as `_embedding`. For example, a `customer_reviews` table with a `review` column must have a `review_embedding` column. 3. **Embeddings Column Data Type:** * The embeddings column must have the following [Arrow data type](/docs/v1.10/reference/datatypes/accelerators) when loaded into Spice: 1. `FixedSizeList[Float32 or Float64, N]`, where `N` is the dimension (size) of the embedding vector. `FixedSizeList` is used for efficient storage and processing of fixed-size vectors. 2. If the column is \[**chunked**(../../components/embeddings#chunking), use `List[FixedSizeList[Float32 or Float64, N]]`. 4. **Offset Column for Chunked Data:** * If the underlying column is chunked, there must be an additional offset column named `_offsets` with the following Arrow data type: 1. `List[FixedSizeList[Int32, 2]]`, where each element is a pair of integers `[start, end]` representing the start and end indices of the chunk in the underlying text column. This offset column maps each chunk in the embeddings back to the corresponding segment in the underlying text column. * *For instance, `[[0, 100], [101, 200]]` indicates two chunks covering indices 0–100 and 101–200, respectively.* By following these guidelines, you can ensure that your dataset with pre-existing embeddings is fully compatible with the vector search and other embedding functionalities provided by Spice. ### Example[​](#example "Direct link to Example") A table `sales` with an `address` column and corresponding embedding column(s). --- # Semantic Model Semantic data models in Spice are defined using the `datasets[*].columns` configuration. These models provide structured and meaningful data representations, which are beneficial for both AI large language models (LLMs) and traditional data analysis. ## Use-Cases[​](#use-cases "Direct link to Use-Cases") ### Large Language Models (LLMs)[​](#large-language-models-llms "Direct link to Large Language Models (LLMs)") The semantic model is automatically used by [Spice Models](/docs/v1.10/reference/spicepod/models) as context to produce more accurate and context-aware AI responses. ## Defining a Semantic Model[​](#defining-a-semantic-model "Direct link to Defining a Semantic Model") Semantic data models are defined within the `spicepod.yaml` file, specifically under the `datasets` section. Each dataset supports `description`, `metadata`, and a `columns` field where individual columns are described with metadata and features for utility and clarity. ### Example Configuration[​](#example-configuration "Direct link to Example Configuration") Example `spicepod.yaml`: ``` datasets: - name: taxi_trips description: NYC taxi trip rides metadata: instructions: Always provide citations with reference URLs. reference_url_template: https://d37ci6vzurychx.cloudfront.net/trip-data/yellow_tripdata_.parquet columns: - name: tpep_pickup_time description: 'The time the passenger was picked up by the taxi' - name: notes description: 'Optional notes about the trip' embeddings: - from: hf_minilm # A defined Spice Model chunking: enabled: true target_chunk_size: 512 overlap_size: 128 trim_whitespace: true ``` ## Dataset Metadata[​](#dataset-metadata "Direct link to Dataset Metadata") Datasets can be defined with the following metadata: * `instructions`: Optional. Instructions to provide to a language model when using this dataset. * `reference_url_template`: Optional. A URL template for citation links. For detailed `metadata` configuration, see the [Dataset Reference](/docs/v1.10/reference/spicepod/datasets#metadata) ## Column Definitions[​](#column-definitions "Direct link to Column Definitions") Each column in the dataset can be defined with the following attributes: * `description`: Optional. A description of the column's contents and purpose. * `embeddings`: Optional. Vector embeddings configuration for this column. For detailed `columns` configuration, see the [Dataset Reference](/docs/v1.10/reference/spicepod/datasets#columns) --- # Web Search Spice provides web search functionality through LLMs and tools, enabling access to recent information and relevant context. Spice supports two ways of using web search in the runtime, namely through tools and through specific model providers. ## Web Search Through LLM Tools[​](#web-search-through-llm-tools "Direct link to Web Search Through LLM Tools") One way of using web search with Spice is through the dedicated web search tool configured to use Perplexity as the engine. Sample Spicepod configuration: ``` tools: - name: the_internet from: websearch description: 'Search the web for information.' params: engine: perplexity perplexity_auth_token: ${ secrets:SPICE_PERPLEXITY_AUTH_TOKEN } ``` This tool can then be provided to any configured models like so: ``` models: - from: openai:gpt-4.1 name: my-model params: openai_api_key: ${secrets:OPENAI_API_KEY} tools: websearch # configure the model to use the websearch tool. - from: anthropic:claude-3-5-sonnet-latest name: claude_3_5_sonnet params: anthropic_api_key: ${ secrets:SPICE_ANTHROPIC_API_KEY } tools: websearch # configure the model to use the websearch tool. - from: xai:grok4 name: xai params: xai_api_key: ${secrets:SPICE_GROK_API_KEY} tools: websearch # configure the model to use the websearch tool. ``` These models can then be invoked via an interactive REPL through [`spice chat`](/docs/v1.10/cli/reference/chat) or via the OpenAI-compatible `/v1/chat/completions` HTTP endpoint. To learn more about the `websearch` tool, refer to [the reference](/docs/v1.10/components/tools/websearch). ## Web Search Through OpenAI Hosted Tools[​](#web-search-through-openai-hosted-tools "Direct link to Web Search Through OpenAI Hosted Tools") Spice also supports web search using [OpenAI's hosted web search tool](https://platform.openai.com/docs/guides/tools-web-search?api-mode=responses) when using [OpenAI's Responses API](https://platform.openai.com/docs/api-reference/responses). Sample Spicepod configuration: ``` models: - from: openai:gpt-4o-mini # Or any other model supported by OpenAI's Responses API name: openai_model params: openai_api_key: ${secrets:OPENAI_API_KEY} tools: auto responses_api: enabled # Required for using web search openai_responses_tools: web_search # Allowlist the web search tool via OpenAI's Responses API ``` **Sample Usage (with the above configuration)** Start Spice with `spice run`. Then, execute the following command, which makes a request to the `/v1/responses` endpoint. ``` curl -s -H "Content-Type: application/json" -X POST "http://localhost:8090/v1/responses" -d '{"model": "openai_model", "input": "what is the latest news today? use the web search tool"}' | jq -r '.output[] | select(.type=="message") | .content[] | select(.type=="output_text") | .text' ``` Output: ``` Here are some of the latest news updates as of August 30, 2025: **International Affairs** - **Thailand's Prime Minister Removed from Office**: The Constitutional Court of Thailand has dismissed Prime Minister Paetongtarn Shinawatra from office due to ethical misconduct related to leaked phone calls with former Cambodian Prime Minister Hun Sen. ([en.wikipedia.org](https://en.wikipedia.org/wiki/2025?utm_source=openai)) - **Armenia and Azerbaijan Sign Peace Deal**: On August 8, Armenia and Azerbaijan signed a peace agreement mediated by the United States, ending 37 years of hostilities over the Nagorno-Karabakh region. ([en.wikipedia.org](https://en.wikipedia.org/wiki/2025?utm_source=openai)) **Science and Technology** - **OpenAI Releases GPT-5**: OpenAI has unveiled GPT-5, an upgraded version of its language model, featuring "PhD-level intelligence." ([en.wikipedia.org](https://en.wikipedia.org/wiki/2025_in_science?utm_source=openai)) - **NASA-ISRO Synthetic Aperture Radar (NISAR) Satellite Launched**: A joint project between NASA and the Indian Space Research Organisation (ISRO), the NISAR satellite was launched on July 30, 2025. It is the first dual-band radar imaging satellite, designed for remote sensing of Earth's surface. ([en.wikipedia.org](https://en.wikipedia.org/wiki/2025_in_spaceflight?utm_source=openai)) **United States** - **Federal Reserve Governor Lisa Cook Dismissed**: President Donald Trump has removed Federal Reserve Governor Lisa Cook from her position, citing "false statements" on "one or more mortgage agreements." This decision has led to a weakening of the U.S. dollar. ([cnbc.com](https://www.cnbc.com/2025/08/25/stock-market-today-live-updates.html?utm_source=openai)) - **Stock Market Performance**: On August 26, 2025, major U.S. stock indices closed with gains. The Nasdaq Composite advanced 0.44%, the S&P 500 gained 0.41%, and the Dow Jones Industrial Average climbed 0.30%. ([cnbc.com](https://www.cnbc.com/2025/08/25/stock-market-today-live-updates.html?utm_source=openai)) **Space Exploration** - **SpaceX's Seventh Starship Test Flight**: On January 16, 2025, SpaceX launched its seventh test flight of the Starship launch vehicle from its Starbase site in Boca Chica, Texas. The first stage was successfully recovered, but the second stage broke up shortly before engine shutdown. ([en.wikipedia.org](https://en.wikipedia.org/wiki/2025_in_Texas?utm_source=openai)) Please note that news developments are ongoing, and it's advisable to consult multiple sources for the most current information. ``` To invoke this model, use [`spice chat --responses`](/docs/v1.10/cli/reference/chat) for an interactive REPL or the OpenAI-compatible `/v1/responses` HTTP endpoint in the runtime. To learn more about configuring models provided by OpenAI, view [the reference](/docs/v1.10/components/models/openai). ## References[​](#references "Direct link to References") * [Perplexity Search Documentation](https://docs.perplexity.ai/getting-started/overview) * [OpenAI Web Search Tool Documentation](https://platform.openai.com/docs/guides/tools-web-search?api-mode=responses) * [OpenAI Responses API Documentation](https://platform.openai.com/docs/api-reference/responses) --- # Getting started with Spice.ai OSS []() ### Follow these steps to get started with Spice[​](#follow-these-steps-to-get-started-with-spice "Direct link to Follow these steps to get started with Spice") Download the latest version of Spice, connect to a dataset in S3, and ask questions about the data using AI, in less than 5 minutes. **Step 1.** Install the Spice CLI: * macOS, Linux, and WSL * Windows ### Install Script[​](#install-script "Direct link to Install Script") ``` curl https://install.spiceai.org | /bin/bash ``` ### Homebrew[​](#homebrew "Direct link to Homebrew") ``` brew install spiceai/spiceai/spice ``` ### PowerShell Install Script[​](#powershell-install-script "Direct link to PowerShell Install Script") ``` iex ((New-Object System.Net.WebClient).DownloadString("https://install.spiceai.org/Install.ps1")) ``` **Step 2.** Initialize a new Spice app with the `spice init` command: ``` spice init spice_qs ``` A `spicepod.yaml` file is created in the `spice_qs` directory. Change to that directory: ``` cd spice_qs ``` **Step 3.** Add the `spiceai/quickstart` Spicepod. A Spicepod is a package of configuration defining datasets and ML models. ``` spice add spiceai/quickstart ``` The `spicepod.yaml` file will be updated with the `spiceai/quickstart` dependency. ``` version: v1 kind: Spicepod name: spice_qs dependencies: - spiceai/quickstart ``` **Step 4.** Add an [OpenAI](/docs/components/models/openai) model to the Spicepod, with tools enabled to enable the model to access the data: ``` models: - from: openai:gpt-4o-mini name: openai_model params: openai_api_key: ${ env:OPENAI_API_KEY } tools: auto ``` Add your OpenAI API key to a `.env` file that will be automatically loaded: ``` echo "OPENAI_API_KEY=sk-..." > .env ``` Start the Spice runtime: ``` spice run ``` Example output: ``` Spice.ai runtime starting... 2024-08-05T13:02:40.247484Z INFO runtime::flight: Spice Runtime Flight listening on 127.0.0.1:50051 2024-08-05T13:02:40.247490Z INFO runtime::metrics_server: Spice Runtime Metrics listening on 127.0.0.1:9090 2024-08-05T13:02:40.247949Z INFO runtime: Initialized results cache; max size: 128.00 MiB, item ttl: 1s 2024-08-05T13:02:40.248611Z INFO runtime::http: Spice Runtime HTTP listening on 127.0.0.1:8090 2024-08-05T13:02:40.252356Z INFO runtime::opentelemetry: Spice Runtime OpenTelemetry listening on 127.0.0.1:50052 ``` **Step 5.** Start the Spice Chat REPL and ask a question: ``` $ spice chat Using model: openai_model chat> How many taxi trips were taken? A total of 2,964,624 trips were taken according to the dataset. ``` The OpenAI model was automatically provided with the `taxi_trips` dataset and was able to answer the question. **Step 6.** Start the Spice SQL REPL: ``` spice sql ``` The SQL REPL interface will be shown: ``` Welcome to the Spice.ai SQL REPL! Type 'help' for help. show tables; -- list available tables sql> ``` Enter `show tables;` to display the available tables for query: ``` sql> show tables +---------------+--------------+---------------+------------+ | table_catalog | table_schema | table_name | table_type | +---------------+--------------+---------------+------------+ | spice | public | taxi_trips | BASE TABLE | | spice | runtime | task_history | BASE TABLE | | spice | runtime | metrics | BASE TABLE | +---------------+--------------+---------------+------------+ Time: 0.022671708 seconds. 3 rows. ``` Enter a query to display the longest taxi trips: ``` sql> SELECT trip_distance, total_amount FROM taxi_trips ORDER BY trip_distance DESC LIMIT 10; ``` Output: ``` +---------------+--------------+ | trip_distance | total_amount | +---------------+--------------+ | 312722.3 | 22.15 | | 97793.92 | 36.31 | | 82015.45 | 21.56 | | 72975.97 | 20.04 | | 71752.26 | 49.57 | | 59282.45 | 33.52 | | 59076.43 | 23.17 | | 58298.51 | 18.63 | | 51619.36 | 24.2 | | 44018.64 | 52.43 | +---------------+--------------+ Time: 0.045150667 seconds. 10 rows. ``` ## Next Steps[​](#next-steps "Direct link to Next Steps") ## [🔗Cookbook](https://github.com/spiceai/cookbook) [Spice.ai Cookbook Recipes.](https://github.com/spiceai/cookbook) --- # Community Data The [Spice.ai Cloud Platform](https://docs.spice.ai) includes a comprehensive set of free, ready-to-query [sample datasets](https://spicerack.org/). The Spice runtime can query these datasets using the [Spice.ai Data Connector](/docs/v1.10/components/data-connectors/spiceai). ## Quickstart[​](#quickstart "Direct link to Quickstart") To access these community datasets, navigate to [spice.ai](https://spice.ai/), and create a new account by clicking Try for Free. ![spiceai\_try\_for\_free-1](https://github.com/spiceai/spiceai/assets/112157037/27fb47ed-4825-4fa8-94bd-48197406cfaa) After logging in, create an app in order to get an API key. ![create\_app-1](https://github.com/spiceai/spiceai/assets/112157037/d2446406-1f06-40fb-8373-1b6d692cb5f7) This quickstart will use the `taxi_trips` dataset from Spice.ai app. **Step 1.** Initialize a new project: ``` # Initialize a new Spice app spice init spice_app # Change to app directory cd spice_app ``` **Step 2.** Log in to the Spice Cloud Platform from the command line using the `spice login` command. A pop up browser window will prompt you to authenticate: ``` spice login ``` Logging in will create or update a `.env` file in the project directory with the API key. **Step 3.** Start the runtime: ``` # Start the runtime spice run ``` **Step 4.** Configure the dataset: In a new terminal window, configure a new dataset using the `spice dataset configure` command: ``` spice dataset configure ``` Enter a dataset name that will be used to reference the dataset in queries. This name does not need to match the name in the dataset source. ``` dataset name: (spice_app) taxi_trips ``` Enter the description of the dataset: ``` description: Taxi trips dataset ``` Enter the location of the dataset: ``` from: spice.ai/spiceai/quickstart/datasets/taxi_trips ``` Select `y` when prompted whether to accelerate the data: ``` Locally accelerate (y/n)? y ``` You should see the following output from your runtime terminal: ``` 2024-12-16T05:12:45.803694Z INFO runtime::init::dataset: Dataset taxi_trips registered (spice.ai/spiceai/quickstart/datasets/taxi_trips), acceleration (arrow, 10s refresh), results cache enabled. 2024-12-16T05:12:45.805494Z INFO runtime::accelerated_table::refresh_task: Loading data for dataset taxi_trips 2024-12-16T05:13:24.218345Z INFO runtime::accelerated_table::refresh_task: Loaded 2,964,624 rows (8.41 GiB) for dataset taxi_trips in 38s 412ms. ``` **Step 5.** In a new terminal window, use the Spice SQL REPL to query the dataset ``` spice sql ``` ``` SELECT tpep_pickup_datetime, passenger_count, trip_distance from taxi_trips LIMIT 10; ``` The output displays the results of the query along with the query execution time: ``` +----------------------+-----------------+---------------+ | tpep_pickup_datetime | passenger_count | trip_distance | +----------------------+-----------------+---------------+ | 2024-01-11T12:55:12 | 1 | 0.0 | | 2024-01-11T12:55:12 | 1 | 0.0 | | 2024-01-11T12:04:56 | 1 | 0.63 | | 2024-01-11T12:18:31 | 1 | 1.38 | | 2024-01-11T12:39:26 | 1 | 1.01 | | 2024-01-11T12:18:58 | 1 | 5.13 | | 2024-01-11T12:43:13 | 1 | 2.9 | | 2024-01-11T12:05:41 | 1 | 1.36 | | 2024-01-11T12:20:41 | 1 | 1.11 | | 2024-01-11T12:37:25 | 1 | 2.04 | +----------------------+-----------------+---------------+ Time: 0.00538925 seconds. 10 rows. ``` You can experiment with the time it takes to generate queries when using non-accelerated datasets. You can change the acceleration setting from `true` to `false` in the datasets.yaml file. ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` -- Query to display the average trip distance SELECT AVG(trip_distance) FROM taxi_trips; ``` The output displays the average gas used: ``` +-------------------------------+ | avg(taxi_trips.trip_distance) | +-------------------------------+ | 3.652169178958276 | +-------------------------------+ Time: 0.031145625 seconds. 1 rows. ``` --- # Spicepods ## Overview[​](#overview "Direct link to Overview") A **Spicepod** is a package that encapsulates application-centric datasets and machine learning (ML) models. Spicepods are analogous to code packaging systems, like NPM, however differ by expanding the concepts to data and ML models. ## Structure[​](#structure "Direct link to Structure") A Spicepod is described by a YAML manifest file, typically named `spicepod.yaml`, which includes the following key sections: * **Metadata:** Basic information about the Spicepod, such as its name and version. * **Datasets:** Definitions of datasets that are used or produced within the Spicepod. * **Catalogs:** Definitions of catalogs that are used within the Spicepod. * **Models:** Definitions of language or traditional ML models that the Spicepod manages, including their sources and associated datasets. * **Secrets:** Configuration for any secret stores used within the Spicepod. ## Example Manifest[​](#example-manifest "Direct link to Example Manifest") ``` version: v1 kind: Spicepod name: my_spicepod datasets: - from: spice.ai/spiceai/quickstart/datasets/taxi_trips name: taxi_trips acceleration: enabled: true models: - from: openai:gpt-4o-mini name: openai_model params: openai_api_key: ${ env:OPENAI_API_KEY } tools: auto secrets: - from: env name: env ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` version: v1 kind: Spicepod name: another_spicepod datasets: - from: databricks:spiceai_demo.public.dataset name: sample_ds params: mode: delta_lake databricks_endpoint: dbc-a1b2345c-d6e7.cloud.databricks.com databricks_token: ${secrets:my_token} databricks_aws_access_key_id: ${secrets:aws_access_key_id} databricks_aws_secret_access_key: ${secrets:aws_secret_access_key} acceleration: enabled: true refresh_mode: full models: - from: huggingface.co/microsoft/Phi-3.5-mini-instruct name: phi secrets: - from: env name: env ``` ## Key Components[​](#key-components "Direct link to Key Components") ### Datasets[​](#datasets "Direct link to Datasets") Datasets in a Spicepod can be sourced from various locations, including local files or remote databases. They can be materialized and accelerated using different engines such as DuckDB, SQLite, or PostgreSQL to optimize performance. Learn more at [Datasets](/docs/v1.10/reference/spicepod/datasets). ### Catalogs[​](#catalogs "Direct link to Catalogs") Catalogs in a Spicepod can contain multiple schemas. Each schema, in turn, contains multiple tables where the actual data is stored. Learn more at [Catalogs](/docs/v1.10/reference/spicepod/catalogs). ### Models[​](#models "Direct link to Models") ML models are integrated into the Spicepod similarly to datasets. The models can be specified using paths to local files or remote locations. ML inference can be performed using the models and datasets defined within the Spicepod. Learn more at [Models](/docs/v1.10/reference/spicepod/models). ### Secrets[​](#secrets "Direct link to Secrets") Spice.ai supports various secret stores to manage sensitive information such as API keys or database credentials. Supported secret store types include environment variables, files, AWS Secrets Manager, Kubernetes secrets, and keyrings. Learn more at [Secret Stores](/docs/v1.10/components/secret-stores) --- # Telemetry Spice collects anonymous telemetry data, which is used to help understand how to improve the product in future versions. Data collected includes: * The version of Spice being used (i.e. `v1.0.0`) * An anonymous identifier for the Spice instance, computed as `sha256(hostname + spicepod.name)`. * An anonymous identifier for the Spicepod, computed as `sha256(spicepod.name)`. * The code to calculate these identifiers is here: * Various metrics related to usage of features of the runtime, see the full list here: Data collected is sent to `https://telemetry.spiceai.org` once every hour. ## Disabling Telemetry[​](#disabling-telemetry "Direct link to Disabling Telemetry") Telemetry can be disabled in one of three ways: 1. Running the Spice runtime with the CLI flag `--telemetry-enabled false`: ``` spice run -- --telemetry-enabled false ``` or ``` spiced --telemetry-enabled false ``` 2. Add the following configuration to the Spicepod configuration file (`spicepod.yaml`): ``` runtime: telemetry: enabled: false ``` 3. Compile the Spice runtime without the `anonymous_telemetry` default feature: ``` cargo build --release --no-default-features --features "" ``` i.e. ``` cargo build --release --no-default-features --features "duckdb,postgres,sqlite,mysql,flightsql,delta_lake,databricks,dremio,clickhouse,spark,snowflake,ftp,debezium" ``` --- # Spice.ai OSS Installation Installation options for Spice.ai OSS For deployment options, such as to Kubernetes, see [`Deployment`](/docs/v1.10/deployment). * macOS, Linux, and WSL * Windows ### Install Script[​](#install-script "Direct link to Install Script") ``` curl https://install.spiceai.org | /bin/bash ``` ### Homebrew[​](#homebrew "Direct link to Homebrew") ``` brew install spiceai/spiceai/spice ``` ### PowerShell Install Script[​](#powershell-install-script "Direct link to PowerShell Install Script") ``` iex ((New-Object System.Net.WebClient).DownloadString("https://install.spiceai.org/Install.ps1")) ``` ## Direct Download[​](#direct-download "Direct link to Direct Download") Binaries for Linux, Windows, and macOS are available for download from GitHub at [github.com/spiceai/spiceai/releases](https://github.com/spiceai/spiceai/releases). ## Building Spice from Source[​](#building-spice-from-source "Direct link to Building Spice from Source") ### Build prerequisites[​](#build-prerequisites "Direct link to Build prerequisites") * macOS * Linux (Ubuntu) 1. If you don't already have it installed, install Homebrew: ``` # Note: Be sure to follow the steps in the Homebrew installation output to add Homebrew to your PATH. /bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)" ``` 2. Install the Xcode Command Line tools ``` xcode-select --install ``` 3. Install dependencies ``` brew install rust brew install go brew install cmake brew install protobuf ``` 1. Install system dependencies ``` sudo apt update sudo apt install build-essential curl openssl libssl-dev pkg-config protobuf-compiler cmake ``` 2. Install Go ``` export GO_VERSION="1.22.4" rm -rf /tmp/spice mkdir -p /tmp/spice cd /tmp/spice wget https://go.dev/dl/go$GO_VERSION.linux-amd64.tar.gz tar xvfz go$GO_VERSION.linux-amd64.tar.gz sudo mv ./go /usr/local/go echo 'export PATH=$PATH:/usr/local/go/bin' >> $HOME/.profile source $HOME/.profile cd $HOME rm -rf /tmp/spice ``` 3. Install Rust ``` curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y # install unattended source $HOME/.cargo/env ``` ### Build Spice OSS[​](#build-spice-oss "Direct link to Build Spice OSS") ``` # Clone SpiceAI OSS Repo git clone https://github.com/spiceai/spiceai.git cd spiceai # Build and install OSS project in release mode make install # Build and install OSS project in dev mode make install-dev # Build and install OSS project with models make install-with-models # Also you can specify specific features SPICED_CUSTOM_FEATURES="postgres sqlite" make install # Run the following to temporarily add spice to your PATH. # Add it to the end of your .bashrc or .zshrc to permanently add spice to your PATH. export PATH="$PATH:$HOME/.spice/bin" ``` ### Build with Hardware Acceleration[​](#build-with-hardware-acceleration "Direct link to Build with Hardware Acceleration") Spice OSS supports running both local language models and embedding models on dedicated hardware. This is for models from either Huggingface or models located locally. The Spice CLI will automatically detect and download the appropriate runtime binary with hardware acceleration if available. #### CUDA Support[​](#cuda-support "Direct link to CUDA Support") **Steps**: 1. GPUs with Cuda compute capabilities < 7.5 are not supported (V100, Titan V, GTX 1000 series, ...). 2. Ensure both Cuda and associated Nvidia drivers are installed. Requires CUDA version 12.2 or higher 3. Ensure Nvidia binaries are in your path: ``` export PATH=$PATH:/usr/local/cuda/bin ``` 1. Build Spice OSS with CUDA support: ``` make install-with-models-cuda ``` This ensures CUDA devices are selected on model load, and CUDA-specifiy kernels used when possible. #### Metal Support[​](#metal-support "Direct link to Metal Support") **Steps**: 1. Ensure you have Xcode installed. ``` xcode-select --install ``` 1. Build Spice OSS with Metal support: ``` make install-with-models-metal ``` Similarily, this ensures Metal devices are selected on model load, and Metal-specific kernels used when possible. --- # Intelligent Applications ## Building Data-Driven AI Applications with Spice.ai[​](#building-data-driven-ai-applications-with-spiceai "Direct link to Building Data-Driven AI Applications with Spice.ai") Spice.ai represents a paradigm shift in how intelligent applications are developed, deployed, and managed. As outlined in the blog post [Making Apps That Learn and Adapt](https://blog.spiceai.org/posts/2021/11/05/making-apps-that-learn-and-adapt/), the goal of Spice.ai is to eliminate the technical complexity that often hampers developers when building AI-powered solutions. By colocating federated data and machine learning models with applications, Spice.ai provides lightweight, high-performance, and highly scalable AI copilot sidecars. These sidecars streamline application workflows, significantly enhancing both speed and efficiency. At its core, Spice.ai addresses the fragmented nature of traditional AI infrastructure. Data often resides in multiple systems: modern cloud-based warehouses, legacy databases, or unstructured formats like files on FTP servers. Integrating these disparate sources into a unified application pipeline typically requires extensive engineering effort. Spice.ai simplifies this process by federating data across all these sources, materializing it locally for low-latency access, and offering a unified SQL API. This eliminates the need for complex and costly ETL pipelines or federated query engines that operate with high latency. Spice.ai also colocates machine learning models with the application runtime. This approach reduces the data transfer overhead that occurs when sending data to external inference services. By performing inference locally, applications can respond faster and operate more reliably, even in environments with intermittent network connectivity. The result is an infrastructure that enables developers to focus on building value-driven features rather than wrestling with data and deployment complexities. *** ## The Intelligent Application Workflow[​](#the-intelligent-application-workflow "Direct link to The Intelligent Application Workflow") The workflow for creating intelligent applications with Spice.ai is designed to provide developers with a straightforward, efficient path from data to decision-making. It begins with the creation of `Spicepods`, self-contained packages that define datasets and machine learning models. These packages can be distributed through [Spicerack.org](https://spicerack.org), a registry that allows developers to publish, share, and reuse datasets and models for various applications. Once deployed, federated datasets are materialized locally within the Spice runtime. Materialization involves prefetching and precomputing data, storing it in high-performance local stores like DuckDB or SQLite. This approach ensures that queries are executed with minimal latency, offering high concurrency and predictable performance. Accelerated access is made possible through advanced caching and query optimization techniques, enabling applications to perform even complex operations without relying on remote databases. Applications interact with the Spice runtime through high-performance APIs, calling machine learning models for inference tasks such as predictions, recommendations, or anomaly detection. These models are colocated with the runtime, allowing them to leverage the same locally materialized datasets. For example, an e-commerce application could use this infrastructure to provide real-time product recommendations based on user behavior, or a manufacturing system could detect equipment failures before they happen by analyzing time-series sensor data. As the application runs, contextual and environmental data—such as user actions or external sensor readings are ingested into the runtime. This data is replicated back to centralized compute clusters where machine learning models are retrained and fine-tuned to improve accuracy and performance. The updated models are automatically versioned and deployed to the runtime, where they can be A/B tested in real time. This continuous feedback loop ensures that applications evolve and improve without manual intervention, reducing time to value while maintaining model relevance. ![Spice.ai Intelligent Application Workflow](https://github.com/spiceai/docs/assets/80174/22b02c5e-5fcb-4856-b79d-911ac5d084c6) *** ## Why Spice.ai Is the Future of Intelligent Applications[​](#why-spiceai-is-the-future-of-intelligent-applications "Direct link to Why Spice.ai Is the Future of Intelligent Applications") Spice.ai introduces an entirely new way of thinking about application infrastructure by making intelligent applications accessible to developers of all skill levels. It replaces the need for custom integrations and fragmented tools with a unified runtime optimized for AI and data-driven applications. Unlike traditional architectures that rely heavily on centralized databases or cloud-based inference engines, Spice.ai focuses on tier-optimized deployments. This means that data and computation are colocated wherever the application runs—whether in the cloud, on-prem, or at the edge. Federation and materialization are at the heart of Spice.ai’s architecture. Instead of querying remote data sources directly, Spice.ai materializes working datasets locally. For example, a logistics application might materialize only the last seven days of shipment data from a cloud data lake, ensuring that 99% of queries are served locally while retaining the ability to fall back to the full dataset as needed. This reduces both latency and costs while improving the user experience. Machine learning models benefit from the same localized efficiency. Because the models are colocated with the runtime, inference happens in milliseconds rather than seconds, even for complex operations. This is critical for use cases like fraud detection, where split-second decisions can save businesses millions, or real-time personalization, where user engagement depends on instant feedback. Spice.ai also shines in its ability to integrate with diverse infrastructures. It supports modern cloud-native systems like Snowflake and Databricks, legacy databases like SQL Server, and even unstructured sources like files stored on FTP servers. With support for industry-standard APIs like JDBC, ODBC, and Arrow Flight, it integrates seamlessly into existing applications without requiring extensive refactoring. In addition to data acceleration and model inference, Spice.ai provides comprehensive observability and monitoring. Every query, inference, and data flow can be tracked and audited, ensuring that applications meet enterprise standards for security, compliance, and reliability. This makes Spice.ai particularly well-suited for industries such as healthcare, finance, and manufacturing, where data privacy and traceability are paramount. *** ## Getting Started with Spice.ai[​](#getting-started-with-spiceai "Direct link to Getting Started with Spice.ai") Developers can start building intelligent applications with Spice.ai by installing the open-source runtime from [GitHub](https://github.com/spiceai/spiceai). The installation process is simple, and the runtime can be deployed across cloud, on-premises, or edge environments. Once installed, developers can create and manage `Spicepods` to define datasets and machine learning models. These `Spicepods` serve as the building blocks for their applications, streamlining data and model integration. For those looking to accelerate development, [Spicerack.org](https://spicerack.org) provides a curated library of reusable datasets and models. By using these pre-built components, developers can reduce time to deployment while focusing on the unique features of their applications. The Spice.ai community is an essential resource for new and experienced developers alike. Through forums, documentation, and hands-on support, the community helps developers unlock the full potential of intelligent applications. Whether you’re building real-time analytics systems, AI-enhanced enterprise tools, or edge-based IoT applications, Spice.ai provides the infrastructure you need to succeed. In a world where intelligent applications are increasingly becoming the norm, Spice.ai stands out as the definitive platform for building fast, scalable, and secure AI-driven solutions. Its unified approach to data, computation, and machine learning sets a new standard for how applications are developed and deployed. --- # Monitoring ![](/assets/images/observability-9095d33215e242807146ca3255f60915.png) Learn how to monitor Spice.ai deployments. * [Datadog](/docs/monitoring/datadog) * [Grafana & Prometheus](/docs/monitoring/grafana) --- # Datadog Spice can be monitored with [Datadog](https://www.datadoghq.com/) using the [Spice Metrics Endpoint](/docs/features/observability) and pre-built dashboards available in the [Spice repository](https://github.com/spiceai/spiceai/tree/trunk/monitoring). ## Datadog Agent Configuration[​](#datadog-agent-configuration "Direct link to Datadog Agent Configuration") Prerequisite: [Datadog Agent version 6.5.0 or later is installed](https://docs.datadoghq.com/getting_started/agent/). Configure the Datadog Agent to scrape the Spice metrics endpoint: 1. Edit the `openmetrics.d/conf.yaml` file in the `conf.d/` folder at the root of your [Agent’s configuration directory](https://docs.datadoghq.com/agent/guide/agent-configuration-files/#agent-configuration-directory): ``` init_config: instances: - prometheus_url: SPICE-METRICS-ENDPOINT>/metrics # for example http://localhost:9090/metrics namespace: spice metrics: - '*' ``` 1. [Restart the Agent](https://docs.datadoghq.com/agent/guide/agent-commands/#start-stop-and-restart-the-agent) to start collecting Spice metrics. 2. Refer to [Prometheus and OpenMetrics metrics collection from a host](https://docs.datadoghq.com/integrations/guide/prometheus-host-collection/) for all available configuration options and supported parameters. 3. Open Datadog Metrics Explorer and type `spice` to confirm Spice telemetry information is successfully collected. ![](/img/datadog/spice_datadog_metrics_explorer.png) ## Import the Spice Datadog Dashboard[​](#import-the-spice-datadog-dashboard "Direct link to Import the Spice Datadog Dashboard") 1. Create [New Datadog Dashboard](https://docs.datadoghq.com/dashboards/#get-started) ![](/img/datadog/spice_datadog_dashboard_new.png) 2. Click **Import dashboard JSON** and drag and drop [monitoring/datadog-dashboard.json](https://raw.githubusercontent.com/spiceai/spiceai/trunk/monitoring/datadog-dashboard.json) file ![](/img/datadog/spice_datadog_dashboard_import.png) 3. Dashboard is now configured to display Spice.ai OSS key performance metrics ![](/img/datadog/spice_datadog_dashboard.png) --- # Grafana & Prometheus Spice can be monitored with [Grafana](https://grafana.com/grafana/) using the [Spice Metrics Endpoint](/docs/v1.10/features/observability) and pre-built dashboards available in the [Spice repository](https://github.com/spiceai/spiceai/tree/trunk/monitoring). ## Import Grafana Dashboard[​](#import-grafana-dashboard "Direct link to Import Grafana Dashboard") Navigate to the Dashboards section in Grafana and click "New" > "Import". ![](/img/grafana/import-dashboard-button.png) Copy the dashboard JSON from [monitoring/grafana-dashboard.json](https://github.com/spiceai/spiceai/blob/trunk/monitoring/grafana-dashboard.json) into the Grafana import box. ![](/img/grafana/import-dashboard.png) Click "Load". ## Kubernetes[​](#kubernetes "Direct link to Kubernetes") View the [Kubernetes](/docs/v1.10/deployment/kubernetes) deployment guide for configuring the Prometheus Operator to scrape metrics from the Spice instances in Kubernetes. ## Prometheus[​](#prometheus "Direct link to Prometheus") Configure a Prometheus instance to scrape metrics from the Spice runtimes. ``` global: scrape_interval: 1s scrape_configs: - job_name: spiceai static_configs: - targets: ['127.0.0.1:9090'] # Change to your Spice runtime endpoint + port ``` ## Local Quickstart[​](#local-quickstart "Direct link to Local Quickstart") This tutorial creates and configures Grafana and Prometheus locally to scrape and display metrics from several Spice instances. It assumes: * Two Spice runtimes, `spiced-main` and `spiced-edge`, are running on `127.0.0.1:9091` and `127.0.0.1:9092` respectively. 1. Create a `compose.yaml`: ``` version: '3' services: prometheus: image: prom/prometheus:latest volumes: - ./prometheus.yaml:/etc/prometheus/prometheus.yml ports: - 9090:9090 network_mode: 'host' grafana: image: grafana/grafana:latest volumes: - ./.grafana/provisioning:/etc/grafana/provisioning ports: - 3000:3000 network_mode: 'host' ``` 2. Create a `prometheus.yaml` to ``` global: scrape_interval: 1s scrape_configs: - job_name: spiced-main static_configs: - targets: ['127.0.0.1:9091'] - job_name: spiced-edge static_configs: - targets: ['127.0.0.1:9092'] ``` 3. Add a prometheus as a source to grafana. Create a `.grafana/provisioning/datasources/prometheus.yml` ``` apiVersion: 1 datasources: - name: Prometheus type: prometheus access: proxy url: http://localhost:9090 isDefault: true ``` 4. Run the Docker Compose ``` docker-compose up ``` 5. Go to `http://localhost:3000/dashboard/import` and add the JSON from [monitoring/grafana-dashboard.json](https://github.com/spiceai/spiceai/blob/trunk/monitoring/grafana-dashboard.json). 6. The dashboard will have data from the Spice runtimes. ![](/img/grafana/screenshot.png) --- # Zipkin Integration Spice supports distributed tracing by integrating with [Zipkin](https://zipkin.io/) and compatible tracing systems. ![zipkin](https://github.com/user-attachments/assets/b2763cdf-1ec9-4b85-9a88-16a11a98aaf6) ![zipkin](https://github.com/user-attachments/assets/c28e042e-96d8-4aab-8da4-d1f493642e02) ## Configuration[​](#configuration "Direct link to Configuration") Enable Zipkin tracing by configuring the `runtime.tracing` section in `spicepod.yaml`: ``` runtime: tracing: zipkin_enabled: true zipkin_endpoint: "http://your_zipkin_host:9411/api/v2/spans" ``` Trace data will be available in the Zipkin UI at `http://your_zipkin_host:9411`. See the [Zipkin Quickstart](https://zipkin.io/pages/quickstart) to run a test server. ## Trace Information[​](#trace-information "Direct link to Trace Information") A trace in Spice represents a completed task (such as a SQL query, AI chat completion, or tool call). Each trace is a unique span, recording execution time, inputs, outputs, and errors. ``` +----------------------------------+------------------+---------------------+----------------------------+----------------------------+-----------------------+---------------------------------------------------------------------------------------------+ | trace_id | span_id | task | start_time | end_time | execution_duration_ms | error_message | +----------------------------------+------------------+---------------------+----------------------------+----------------------------+-----------------------+---------------------------------------------------------------------------------------------+ | 687e0970f8c49d19c5a08764ea2d4dc1 | f4f52ed29db8b151 | text_embed | 2024-11-25T05:39:37.444749 | 2024-11-25T05:39:53.577195 | 16132.446000000002 | | | 1e881188e5fd252b26adb8a8d838efb8 | 532b0019ad778094 | sql_query | 2024-11-25T05:40:38.864982 | 2024-11-25T05:40:38.871090 | 6.108 | | | ac5abd8bfec7e5aa7c19fc84772c55f1 | 316622ac359e3c00 | sql_query | 2024-11-25T05:39:39.872946 | 2024-11-25T05:39:39.872994 | 0.048 | This feature is not implemented: The context currently only supports a single SQL statement | | 701874d7282dd47791e7519b343a9694 | 5dacf75c4537ee0e | accelerated_refresh | 2024-11-25T05:39:30.452534 | 2024-11-25T05:39:30.452900 | 0.366 | | | 18d76b6389898cc5253a49294607477d | cc0d06a4e69cbcd5 | health | 2024-11-25T05:39:30.451626 | 2024-11-25T05:39:31.563876 | 1112.25 | | | 3c75d16b6b4b8da98c551d115e1c049c | 9a16dc065a95236a | sql_query | 2024-11-25T05:42:27.386754 | 2024-11-25T05:42:27.386859 | 0.10500000000000001 | SQL error: ParserError("Expected: an SQL statement, found: ELECT") | +----------------------------------+------------------+---------------------+----------------------------+----------------------------+-----------------------+---------------------------------------------------------------------------------------------+ ``` For more details, see [task\_history](/docs/reference/task_history). --- # Spice.ai OSS Reference Docs ## [🗃Spicepod specification](/docs/v1.10/reference/spicepod) [10 items](/docs/v1.10/reference/spicepod) ## [🗃SQL Reference](/docs/v1.10/reference/sql) [12 items](/docs/v1.10/reference/sql) ## [🗃Data Types Reference](/docs/v1.10/reference/datatypes) [2 items](/docs/v1.10/reference/datatypes) ## [📄️Models Grade Report](/docs/v1.10/reference/models) [Spice AI graded Large-Language-Model (LLM) evaluation report](/docs/v1.10/reference/models) ## [📄️Task History](/docs/v1.10/reference/task_history) [The Spice runtime stores information about completed tasks in the spice.runtime.task\_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span.](/docs/v1.10/reference/task_history) ## [📄️Duration](/docs/v1.10/reference/duration) [Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted.](/docs/v1.10/reference/duration) ## [📄️File Formats](/docs/v1.10/reference/file_format) [Spice currently supports CSV, JSON, and Parquet data file-formats for data connectors that can read files from a file system or cloud object storage (i.e. s3//, file://, etc.). Support for Iceberg and other file-formats are on the roadmap.](/docs/v1.10/reference/file_format) ## [📄️Cron Schedules](/docs/v1.10/reference/cron) [The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc).](/docs/v1.10/reference/cron) ## [📄️System Requirements](/docs/v1.10/reference/system_requirements) [System requirements for running Spice.ai Open Source](/docs/v1.10/reference/system_requirements) ## [📄️Memory](/docs/v1.10/reference/memory) [Guidelines and best practices for managing memory usage and optimizing performance in Spice.ai Open Source deployments.](/docs/v1.10/reference/memory) --- # Cron Schedules The Runtime supports cron expressions with optional seconds, like `*/10 * * * * *` which evaluates to every 10th second (10, 20, 30, etc). Cron expressions in the Runtime evaluate according to the systems local time where the Runtime is running. The Spice Runtime uses the [`croner` Rust crate](https://github.com/hexagon/croner-rust?tab=readme-ov-file#pattern) for parsing cron expressions. ## Examples[​](#examples "Direct link to Examples") ### At 1am every Monday[​](#at-1am-every-monday "Direct link to At 1am every Monday") ``` 0 1 * * 1 ``` ### At midday every weekday (Monday-Friday)[​](#at-midday-every-weekday-monday-friday "Direct link to At midday every weekday (Monday-Friday)") ``` 0 12 * * 1-5 ``` ### Every hour, at 5 minutes past the hour[​](#every-hour-at-5-minutes-past-the-hour "Direct link to Every hour, at 5 minutes past the hour") ``` 5 * * * * ``` ### Every 10 minutes[​](#every-10-minutes "Direct link to Every 10 minutes") ``` */10 * * * * ``` ### Every 5 minutes at 30 seconds[​](#every-5-minutes-at-30-seconds "Direct link to Every 5 minutes at 30 seconds") ``` 30 */5 * * * * ``` --- # Data Types Reference ## [📄️Accelerator Data Types](/docs/v1.10/reference/datatypes/accelerators) [Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator.](/docs/v1.10/reference/datatypes/accelerators) ## [📄️Object Store Data Types](/docs/v1.10/reference/datatypes/object_store) [Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice.](/docs/v1.10/reference/datatypes/object_store) --- # Accelerator Data Types Spice adheres to Apache Arrow data [types](https://docs.rs/arrow/latest/arrow/datatypes/index.html). Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. | Arrow Type | Description | [Cayenne (Vortex)](https://github.com/vortex-data/vortex) | [DuckDB](https://duckdb.org/docs/sql/data_types/overview) | [SQLite](https://sqlite.org/datatype3.html) | [Postgres](https://www.postgresql.org/docs/current/datatype.html#DATATYPE-TABLE) | | -------------------------- | ---------------------------------------------------------------------------- | --------------------------------------------------------- | ---------------------------------------------------------- | ------------------------------------------- | -------------------------------------------------------------------------------- | | na | A NULL type having no physical storage. | `Null` | | | | | bool | Boolean as 1 bit, LSB bit-packed ordering. | `Bool` | `BOOLEAN` | `BOOL` | `BOOL` | | uint8 | Unsigned 8-bit little-endian integer. | `Primitive(U8)` | `UTINYINT` | `TINYINT` | `SMALLINT` | | int8 | Signed 8-bit little-endian integer. | `Primitive(I8)` | `TINYINT` | `TINYINT` | `SMALLINT` | | uint16 | Unsigned 16-bit little-endian integer. | `Primitive(U16)` | `USMALLINT` | `SMALLINT` | `SMALLINT` | | int16 | Signed 16-bit little-endian integer. | `Primitive(I16)` | `SMALLINT` | `SMALLINT` | `SMALLINT` | | uint32 | Unsigned 32-bit little-endian integer. | `Primitive(U32)` | `UINTEGER` | `INT` | `INTEGER` | | int32 | Signed 32-bit little-endian integer. | `Primitive(I32)` | `INTEGER` | `INT` | `INTEGER` | | uint64 | Unsigned 64-bit little-endian integer. | `Primitive(U64)` | `UBIGINT` | `BIGINT` | `BIGINT` | | int64 | Signed 64-bit little-endian integer. | `Primitive(I64)` | `BIGINT` | `BIGINT` | `BIGINT` | | half\_float | 2-byte floating point value | `Primitive(F16)` | | | | | float | 4-byte floating point value | `Primitive(F32)` | `FLOAT` | `FLOAT` | `REAL` | | double | 8-byte floating point value | `Primitive(F64)` | `DOUBLE` | `DOUBLE` | `DOUBLE PRECISION` | | string | UTF8 variable-length string as List\ | `Utf8` | `VARCHAR` | `TEXT` | `TEXT` | | binary | Variable-length bytes (no guarantee of UTF8-ness) | `Binary` | `BLOB` | | | | fixed\_size\_binary | Each value has equal bytes of binary. | | `BLOB` | | | | date32 | int32\_t days since the UNIX epoch | `Extension(Date)` | `DATE` | `DATE` | | | date64 | int64\_t milliseconds since the UNIX epoch | `Extension(Date)` | `DATE` | `TIMESTAMP` | | | timestamp | Exact timestamp encoded with int64 since UNIX epoch, seconds or milliseconds | `Extension(Timestamp)` | `TIMESTAMP_S`, `TIMESTAMP_MS`, `TIMESTAMP`, `TIMESTAMP_NS` | `TIMESTAMP` | `TIMESTAMP` | | time32 | Time as signed 32-bit integer, seconds or milliseconds since midnight. | `Extension(Time)` | `TIME` | `TIME` | | | time64 | Time as signed 64-bit integer, microseconds or nanoseconds since midnight. | `Extension(Time)` | `TIME` | `TIME` | | | interval\_months | YEAR\_MONTH interval in SQL style. | | | | | | interval\_day\_time | DAY\_TIME interval in SQL style. | | | | | | decimal128 | Precision- and scale-based decimal type with 128 bits. | `Decimal` | `DECIMAL(P, S)` | `DECIMAL(38, 10)` | `DECIMAL(38, 10)` | | decimal | Defined for backward-compatibility. | `Decimal` | | | | | decimal256 | Precision- and scale-based decimal type with 256 bits. | `Decimal` | | | | | list | A list of some logical data type. | `List` | `TYPE[]` | | `TYPE[]` | | struct | Struct of logical types. | `Struct` | `STRUCT(...)` | | | | sparse\_union | Sparse unions of logical types. | | | | | | dense\_union | Dense unions of logical types. | | | | | | dictionary | Dictionary-encoded type, | | | | | | map | Map, a repeated struct logical type. | | | | | | extension | Custom data type, implemented by user. | `Extension` | | | | | fixed\_size\_list | Fixed size list of some logical type. | `FixedSizeList` | `TYPE[N]` | | | | duration | Elapsed time in seconds, milliseconds, microseconds or nanoseconds. | | | | | | large\_string | Like STRING, but with 64-bit offsets. | `Utf8` | `VARCHAR` | | | | large\_binary | Like BINARY, but with 64-bit offsets. | `Binary` | `BLOB` | | | | large\_list | Like LIST, but with 64-bit offsets. | `List` | `TYPE[]` | | | | interval\_month\_day\_nano | Calendar interval type with three fields. | | `INTERVAL` | | | | run\_end\_encoded | Run-end encoded data. | | | | | | string\_view | UTF8 view type with 4-byte prefix & inline small string optimization. | `Utf8` | `VARCHAR` | | | | binary\_view | Bytes view type with 4-byte prefix and inline small string optimization. | `Binary` | `BLOB` | | | | list\_view | A list of some logical data type represented by offset and size. | | | | | | large\_list\_view | Like LIST\_VIEW, but with 64-bit offsets and sizes. | | | | | Note: Where `TYPE` is used (e.g. `TYPE[]`), it refers an established supported type for the specific data accelerator (e.g. `INTEGER[]`). Cayenne (Vortex) provides zero-copy compatibility with Apache Arrow and supports most Arrow types through its logical type system. For DuckDB, the accelerator hands DuckDB the Arrow schema and uses the `CREATE TABLE` statement DuckDB derives from it, so unsigned Arrow integers keep their unsignedness (`uint32` becomes `UINTEGER`, not `INTEGER`) and `decimal128` keeps its precision and scale. `TYPE[N]` denotes a fixed-length list (e.g. `INTEGER[4]`), and a timezone-aware `timestamp` becomes `TIMESTAMP WITH TIME ZONE`. --- # Object Store Data Types Spice adheres to Apache Arrow data [types](https://docs.rs/arrow/latest/arrow/datatypes/index.html). The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. ## Parquet[​](#parquet "Direct link to Parquet") | Parquet Physical Type | Arrow Type | | ---------------------- | -------------------------------------------------------------------------------------------- | | `BOOLEAN` | `Boolean` | | `INT32` | `Int8`/`Int16`/`Int32`/`UInt8`/`UInt16`/`UInt32`/`Date32`/`Time32`/`Decimal128` | | `INT64` | `Int64`/`UInt64`/`Time64`/`Timestamp(Millisecond/Microsecond/Nanosecond, None)`/`Decimal128` | | `INT96` | `Timestamp(Nanosecond, None)` | | `FLOAT` | `Float32` | | `DOUBLE` | `Float64` | | `BYTE_ARRAY` | `Utf8`/`Binary`/`Decimal128`/`Decimal256` | | `FIXED_LEN_BYTE_ARRAY` | `Decimal128`/`Decimal256`/`Interval`/`Float16`/`FixedSizeBinary` | --- # Duration Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. `1.5h`) are accepted. Supported time units are: | Time Unit | Identifier | Calculation | | ------------- | ---------- | ----------- | | `Nanosecond` | ns | `1` | | `Millisecond` | ms | `1000000` | | `Second` | s | `1s` | | `Minute` | m | `60s` | | `Hour` | h | `60m` | | `Day` | d | `24h` | | `Week` | w | `7d` | Microseconds are also supported, spelled `Ms` (capital `M`, lowercase `s`). Month and year units are **not** supported — their length is ambiguous, so express longer intervals in weeks or days. ## Example[​](#example "Direct link to Example") ``` # 1 second 1s # 250 milliseconds 250ms # 3 minutes 3m # 1 hour 1h ``` ### Additional Example[​](#additional-example "Direct link to Additional Example") ``` # 2 days 2d # 1 week 1w # 90 minutes 1.5h # 30 seconds (no suffix) 30 ``` --- # File Formats Spice currently supports CSV, JSON, and Parquet data file-formats for data connectors that can read files from a file system or cloud object storage (i.e. [`s3://`](/docs/components/data-connectors/s3), [`abfs://`](/docs/components/data-connectors/abfs), [`file://`](/docs/components/data-connectors/file), etc.). Support for Iceberg and other file-formats are on the roadmap. The parameters supported for specific file-formats are detailed on this page. ## Parquet[​](#parquet "Direct link to Parquet") Spice automatically supports reading any Parquet file, regardless of the compression codec or data encoding used. Compression codecs: * [`UNCOMPRESSED`](https://parquet.apache.org/docs/file-format/data-pages/compression/#uncompressed) * [`SNAPPY`](https://parquet.apache.org/docs/file-format/data-pages/compression/#snappy) * [`GZIP`](https://parquet.apache.org/docs/file-format/data-pages/compression/#gzip) * [`LZO`](https://parquet.apache.org/docs/file-format/data-pages/compression/#lzo) * [`BROTLI`](https://parquet.apache.org/docs/file-format/data-pages/compression/#brotli) * [`LZ4`](https://parquet.apache.org/docs/file-format/data-pages/compression/#lz4) (deprecated in favor of `LZ4_RAW`) * [`LZ4_RAW`](https://parquet.apache.org/docs/file-format/data-pages/compression/#lz4_raw) * [`ZSTD`](https://parquet.apache.org/docs/file-format/data-pages/compression/#zstd) Data encodings: * [`PLAIN`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#plain-plain--0) * [`PLAIN_DICTIONARY` / `RLE_DICTIONARY`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#dictionary-encoding-plain_dictionary--2-and-rle_dictionary--8) * [`RLE`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#run-length-encoding--bit-packing-hybrid-rle--3) * [`BIT_PACKED`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#bit-packed-deprecated-bit_packed--4) (deprecated in favor of `RLE`) * [`DELTA_BINARY_PACKED`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#delta-binary-packing-delta_binary_packed--5) * [`DELTA_LENGTH_BYTE_ARRAY`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#delta-length-byte-array-delta_length_byte_array--6) * [`DELTA_BYTE_ARRAY`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#delta-strings-delta_byte_array--7) * [`BYTE_STREAM_SPLIT`](https://parquet.apache.org/docs/file-format/data-pages/encodings/#byte-stream-split-byte_stream_split--9) ## CSV[​](#csv "Direct link to CSV") ### Parameters[​](#parameters "Direct link to Parameters") * `csv_has_header`: Optional. Indicate if the CSV file has header row. Defaults to `true` * `csv_quote`: Optional. A one-character string used to quote fields containing special characters. Defaults to `"` * `csv_escape`: Optional. A one-character string used to represent special characters or to include characters that would normally be interpreted as delimiters or new line characters within a field value. Defaults to `null` * `csv_schema_infer_max_records`: Optional. A number used to set the limit in terms of records to scan to infer the schema. Defaults to `1000` * `csv_delimiter`: Optional. A one-character string used to separate individual fields. Defaults to `,` ## JSON[​](#json "Direct link to JSON") ### Parameters[​](#parameters-1 "Direct link to Parameters") * `json_format`: Optional. Specifies the JSON format to parse. Valid values are `array`, `ndjson`, and `jsonl`. Defaults to `jsonl` --- # Managing Memory Usage Effective memory management is essential for maintaining optimal performance and stability in Spice deployments. This guide outlines recommendations and best practices for managing memory usage across different [Data Accelerators](/docs/components/data-accelerators). ## General Memory Recommendations[​](#general-memory-recommendations "Direct link to General Memory Recommendations") Memory requirements vary based on workload characteristics, dataset sizes, query complexity, and refresh modes. Recommended allocations include: * **Typical workloads**: At least 8 GB RAM. * **Larger datasets**: * `refresh_mode: full`: 2.5x dataset size. * `refresh_mode: append`: 1.5x dataset size. * `refresh_mode: changes`: Primarily influenced by CDC event volume and frequency; 1.5x dataset size is a reasonable estimate. Memory requirements can be reduced by using file-based acceleration with [DuckDB](/docs/components/data-accelerators/duckdb), [SQLite](/docs/components/data-accelerators/sqlite), or [Spice Cayenne](/docs/components/data-accelerators/cayenne), which store data on disk and support spilling. ## Accelerator-Specific Memory Management[​](#accelerator-specific-memory-management "Direct link to Accelerator-Specific Memory Management") Different acceleration engines have distinct memory characteristics and tuning options. ### Arrow (In-Memory)[​](#arrow-in-memory "Direct link to Arrow (In-Memory)") The default Arrow accelerator stores all data in memory uncompressed. Datasets must fit entirely in available RAM. * Data is stored uncompressed in Apache Arrow format * No configuration options for memory limits * Best for smaller datasets requiring maximum query speed * Consider switching to file-based accelerators for datasets exceeding available memory **Hash Index Memory (Experimental, v1.11.0-rc.2+):** When using the optional hash index, additional memory is required: | Component | Memory per Row | | ------------ | -------------- | | Hash slot | 16 bytes | | Bloom filter | \~1.25 bytes | | **Total** | \~17.25 bytes | For a 10 million row dataset with hash index enabled, expect \~165 MB additional memory overhead. ### Spice Cayenne[​](#spice-cayenne "Direct link to Spice Cayenne") [Spice Cayenne](/docs/components/data-accelerators/cayenne) stores data on disk using the [Vortex](https://github.com/vortex-data/vortex) columnar format, with configurable caches for metadata and frequently accessed data segments. The caches can be configured to reside either in memory or on disk, which impacts overall memory behavior. Spice Cayenne is DataFusion query-native, meaning all query execution adheres to the `runtime.query.memory_limit` setting. When query memory is exhausted, DataFusion spills intermediate results to disk. This architecture provides predictable memory usage while maintaining high query performance. **Memory Configuration Parameters:** | Parameter | Default | Description | | -------------------------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------- | | `cayenne_footer_cache_mb` | `128` | Size of the in-memory Vortex footer cache in megabytes. Larger values improve query performance for repeated scans by caching file metadata. | | `cayenne_segment_cache_mb` | `256` | Size of the in-memory Vortex segment cache in megabytes. Caches decompressed data segments for improved query performance. | **Memory Usage Guidelines:** * Base memory: \~500 MB for runtime overhead * Footer cache: 128 MB default, increase for datasets with many files * Segment cache: 256 MB default, increase for workloads with repeated scans on the same data * Query execution memory: Depends on query complexity and concurrency **Example Configuration:** ``` datasets: - from: s3://my-bucket/large-dataset/ name: large_dataset acceleration: engine: cayenne mode: file params: cayenne_footer_cache_mb: 256 cayenne_segment_cache_mb: 512 ``` ### DuckDB[​](#duckdb "Direct link to DuckDB") [DuckDB](/docs/components/data-accelerators/duckdb) manages memory through streaming execution, intermediate spilling, and buffer management. By default, each DuckDB instance uses up to 80% of available system memory. **Memory Configuration Parameters:** | Parameter | Default | Description | | --------------------- | ----------------- | -------------------------------------- | | `duckdb_memory_limit` | 80% of system RAM | Maximum memory for the DuckDB instance | **Memory Usage Guidelines:** * Set `duckdb_memory_limit` to control memory per DuckDB instance * DuckDB indexes do not support spilling and may consume significant memory * Allocate at least 30% additional container/machine memory for the runtime process **Example Configuration:** ``` datasets: - from: postgres:analytics.orders name: orders acceleration: engine: duckdb mode: file params: duckdb_memory_limit: 4GB ``` ### SQLite[​](#sqlite "Direct link to SQLite") [SQLite](/docs/components/data-accelerators/sqlite) is lightweight and efficient for smaller datasets but does not support intermediate spilling. Datasets must fit in memory or use application-level paging. ## Refresh Modes and Memory Implications[​](#refresh-modes-and-memory-implications "Direct link to Refresh Modes and Memory Implications") Refresh modes affect memory usage as follows: * **Full Refresh**: Temporarily loads data into a new table before replacing the existing table so that it can be atomically swapped and maintain consistency. This requires memory for both tables simultaneously, resulting in higher usage. * **Append Refresh**: Incrementally inserts or upserts data, using memory only for the incremental data, which reduces usage. * **Changes Refresh**: Applies CDC events incrementally. Memory usage depends on event volume and frequency, typically resulting in lower and predictable usage. ## DataFusion Memory Management[​](#datafusion-memory-management "Direct link to DataFusion Memory Management") Spice.ai uses DataFusion as its query execution engine. By default, DataFusion does not enforce strict memory limits, which can lead to unbounded usage. Spice.ai addresses this through: * **Memory Limit**: The `runtime.query.memory_limit` parameter defines the maximum memory available for query execution. Once the memory limit is reached, supported query operations spill data to disk, helping prevent out-of-memory errors and maintain query stability. See [Spicepod Configuration](/docs/v1.10/reference/spicepod/runtime#runtimequerymemory_limit) for details. * **Memory Budgeting**: Limits memory per query execution. Queries exceeding the limit return an error. See [Spicepod Configuration](/docs/v1.10/reference/spicepod) for details. * **Spill-to-Disk**: Operators such as Sort, Join, and GroupByHash spill intermediate results to disk when memory limits are exceeded, preventing out-of-memory errors. * **Spill File Compression**: The `runtime.query.spill_compression` parameter controls how spill files are compressed during query execution. By default, Spice.ai uses Zstandard (`zstd`) compression, which offers a high compression ratio and helps reduce disk usage when queries spill intermediate data. See [Spicepod Configuration](/docs/v1.10/reference/spicepod/runtime#runtimequeryspill_compression) for details. DataFusion supports spilling for several operators, but not all operations are currently supported. Notably, the following operations do not support spilling: * HashJoin ([tracking issue](https://github.com/apache/arrow-datafusion/issues/1047)) * ExternalSorterMerge (no current tracking issue; previously discussed in the context of SortMergeJoin) * RepartitionMerge (spilling is suggested to be supported, but may depend on HashJoin support; see [issue](https://github.com/apache/arrow-datafusion/issues/1047)) ## Embedded Data Accelerators[​](#embedded-data-accelerators "Direct link to Embedded Data Accelerators") Spice.ai integrates with embedded accelerators like [SQLite](/docs/components/data-accelerators/sqlite) and [DuckDB](/docs/components/data-accelerators/duckdb), each with unique memory considerations: Operators such as Sort, Join, and GroupByHash spill intermediate results to disk when memory limits are exceeded, preventing out-of-memory errors. DataFusion writes spill files using the [Arrow IPC Stream format](https://arrow.apache.org/docs/format/Columnar.html#ipc-streaming-format). **Spill Compression:** The `runtime.query.spill_compression` parameter controls how spill files are compressed: | Value | Description | | ---------------- | ---------------------------------------------- | | `z