# Spice.ai OSS > A portable SQL query and AI compute engine, written in Rust, for data-grounded apps and agents. - [Spice.ai OSS](/) ## cookbook A collection of guides and samples to help you build data-grounded AI apps and agents with Spice.ai Open-Source. Find ready-to-use examples for data acceleration, AI agents, LLM memory, and more. - [🧑‍🍳 Spice.ai OSS Cookbook](/cookbook): A collection of guides and samples to help you build data-grounded AI apps and agents with Spice.ai Open-Source. Find ready-to-use examples for data acceleration, AI agents, LLM memory, and more. ## docs Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven applications and AI agents. Learn about data federation, acceleration, RAG, and building intelligent apps. - [Spice.ai Open Source](/docs): Spice is an open-source SQL query and AI compute engine, written in Rust, for data-driven applications and AI agents. Learn about data federation, acceleration, RAG, and building intelligent apps. ### tags - [Tags](/docs/tags) - [One doc tagged with "Acknowledgements"](/docs/tags/acknowledgements): Project acknowledgements, credits, and attributions. - [3 docs tagged with "ADBC"](/docs/tags/adbc): Arrow Database Connectivity driver configuration and usage. - [4 docs tagged with "API"](/docs/tags/api): HTTP, Arrow Flight SQL, ODBC, JDBC, and ADBC API reference. - [One doc tagged with "Argo CD"](/docs/tags/argocd): Argo CD declarative GitOps continuous delivery for Kubernetes. - [2 docs tagged with "Arrow"](/docs/tags/arrow): Apache Arrow columnar data format integration. - [One doc tagged with "Arrow Flight SQL"](/docs/tags/arrow-flight-sql): Arrow Flight SQL protocol implementation and configuration. - [One doc tagged with "Auth"](/docs/tags/auth): Authentication and authorization mechanisms. - [2 docs tagged with "Authentication"](/docs/tags/authentication): User authentication methods and security protocols. - [5 docs tagged with "Azure"](/docs/tags/azure): Microsoft Azure cloud services and integrations. - [2 docs tagged with "Blob Storage"](/docs/tags/blob-storage): Object storage services and blob data access. - [One doc tagged with "Caching"](/docs/tags/caching): Data caching strategies and performance optimization. - [13 docs tagged with "Catalogs"](/docs/tags/catalogs): Data catalog connectors and metadata management. - [3 docs tagged with "Cayenne"](/docs/tags/cayenne): Cayenne (Vortex) data accelerator built on Vortex columnar format. - [2 docs tagged with "CLI"](/docs/tags/cli): Spice command-line interface commands and usage. - [4 docs tagged with "Cloud Connect"](/docs/tags/cloud-connect): Connect a self-hosted Spice instance to Spice Cloud. - [5 docs tagged with "Component Metrics"](/docs/tags/component-metrics): Runtime component performance metrics and monitoring. - [One doc tagged with "Components"](/docs/tags/components): Runtime components including catalogs, data connectors, and models. - [2 docs tagged with "Configuration"](/docs/tags/configuration): System configuration files and runtime settings. - [2 docs tagged with "Cosmos DB"](/docs/tags/cosmosdb): Azure Cosmos DB (NoSQL / Core SQL) data connector. - [One doc tagged with "C#"](/docs/tags/csharp): C# programming language topics. - [8 docs tagged with "Data Accelerators"](/docs/tags/data-accelerators): Data acceleration engines for high-performance query execution. - [48 docs tagged with "Data Connectors"](/docs/tags/data-connectors): Data source connectors and integration patterns. - [One doc tagged with "Data Lake"](/docs/tags/data-lake): Data lake architectures and storage solutions. - [3 docs tagged with "Databricks"](/docs/tags/databricks): Databricks platform integration and Unity Catalog support. - [2 docs tagged with "Datasets"](/docs/tags/datasets): Dataset definitions and data source configurations. - [One doc tagged with "Debezium"](/docs/tags/debezium): Debezium change data capture integration. - [One doc tagged with "Debugging"](/docs/tags/debugging): Debugging techniques and troubleshooting tools. - [3 docs tagged with "Delta Lake"](/docs/tags/delta-lake): Delta Lake open table format support. - [One doc tagged with "Dependencies"](/docs/tags/dependencies): Third-party dependencies and software requirements. - [13 docs tagged with "Deployment"](/docs/tags/deployment): Production deployment using Docker and Kubernetes. - [3 docs tagged with "Docker"](/docs/tags/docker): Docker containerization and deployment configurations. - [One doc tagged with ".NET"](/docs/tags/dotnet): .NET framework topics. - [One doc tagged with "Dremio"](/docs/tags/dremio): Dremio data lake engine integration. - [2 docs tagged with "DuckDB"](/docs/tags/duckdb): DuckDB embedded analytical database integration. - [2 docs tagged with "DuckLake"](/docs/tags/ducklake): DuckLake catalog integration. - [2 docs tagged with "DynamoDB"](/docs/tags/dynamodb): Amazon DynamoDB NoSQL database integration. - [2 docs tagged with "Elasticsearch"](/docs/tags/elasticsearch): Elasticsearch data connector and vector engine integration. - [7 docs tagged with "Embeddings"](/docs/tags/embeddings): Vector embeddings and semantic similarity operations. - [6 docs tagged with "Features"](/docs/tags/features): Core platform features including acceleration, caching, and search. - [One doc tagged with "Federation"](/docs/tags/federation): Cross-database queries and federated data access. - [One doc tagged with "File"](/docs/tags/file): File-based data connectors and local file access. - [One doc tagged with "Flux"](/docs/tags/flux): Flux CD GitOps toolkit for Kubernetes. - [One doc tagged with "Flux CD"](/docs/tags/fluxcd): Flux CD GitOps toolkit for Kubernetes. - [2 docs tagged with "Functions"](/docs/tags/functions): User-defined SQL scalar functions and remote function endpoints. - [One doc tagged with "Getting Started"](/docs/tags/getting-started): Installation guides and quickstart tutorials. - [3 docs tagged with "GitHub"](/docs/tags/github): GitHub repository data connector and integration. - [3 docs tagged with "GitOps"](/docs/tags/gitops): GitOps continuous delivery patterns for Kubernetes. - [2 docs tagged with "Glue"](/docs/tags/glue): AWS Glue ETL service and data catalog integration. - [One doc tagged with "Go"](/docs/tags/go): Go programming language topics. - [One doc tagged with "Golang"](/docs/tags/golang): Go (Golang) programming language topics. - [One doc tagged with "GraphQL"](/docs/tags/graphql): GraphQL API data connectors and query language support. - [3 docs tagged with "Helm"](/docs/tags/helm): Helm package manager for deploying Spice.ai on Kubernetes. - [One doc tagged with "HTTPS"](/docs/tags/https): HTTPS data connectors and secure web API access. - [2 docs tagged with "Hugging Face"](/docs/tags/huggingface): Hugging Face model hub and transformer model integration. - [2 docs tagged with "Iceberg"](/docs/tags/iceberg): Apache Iceberg open table format support. - [One doc tagged with "In-Memory"](/docs/tags/in-memory): In-memory data processing and temporary storage. - [One doc tagged with "Integration"](/docs/tags/integration): Third-party system integrations and connectivity patterns. - [One doc tagged with "Java"](/docs/tags/java): Java programming language topics. - [One doc tagged with "JavaScript"](/docs/tags/javascript): JavaScript programming language topics. - [One doc tagged with "JDBC"](/docs/tags/jdbc): Java Database Connectivity driver configuration. - [One doc tagged with "Kafka"](/docs/tags/kafka): Apache Kafka event streaming platform integration. - [8 docs tagged with "Kubernetes"](/docs/tags/kubernetes): Kubernetes orchestration and secret management. - [One doc tagged with "launchd"](/docs/tags/launchd): Run Spice as a macOS service. - [One doc tagged with "libSQL"](/docs/tags/libsql): libSQL database topics. - [One doc tagged with "Local"](/docs/tags/local): Local model execution and on-premise services. - [One doc tagged with "Logging"](/docs/tags/logging): Application logging configuration and log management. - [One doc tagged with "Login"](/docs/tags/login): User login processes and session management. - [One doc tagged with "Manifest"](/docs/tags/manifest): Configuration manifest files and schema definitions. - [One doc tagged with "MCP"](/docs/tags/mcp): Model Context Protocol for AI tool integration. - [2 docs tagged with "Memory"](/docs/tags/memory): In-memory data connectors and volatile storage. - [16 docs tagged with "Models"](/docs/tags/models): Machine learning models and AI inference engines. - [One doc tagged with "MongoDB"](/docs/tags/mongodb): MongoDB NoSQL document database integration. - [One doc tagged with "MSSQL"](/docs/tags/mssql): Microsoft SQL Server database integration. - [3 docs tagged with "MySQL"](/docs/tags/mysql): MySQL relational database integration. - [One doc tagged with "Node.js"](/docs/tags/nodejs): Node.js runtime environment topics. - [3 docs tagged with "NoSQL"](/docs/tags/nosql): NoSQL database systems and document stores. - [29 docs tagged with "Observability"](/docs/tags/observability): Monitoring, tracing, and metrics for system visibility. - [One doc tagged with "ODBC"](/docs/tags/odbc): Open Database Connectivity driver configuration. - [One doc tagged with "Open Source"](/docs/tags/open-source): Open source licenses and community contributions. - [3 docs tagged with "OpenAI"](/docs/tags/openai): OpenAI API integration and GPT model access. - [2 docs tagged with "Oracle"](/docs/tags/oracle): Documentation related to Oracle database integration. - [2 docs tagged with "Overrides"](/docs/tags/overrides): Parameter overrides and configuration customization. - [2 docs tagged with "Overview"](/docs/tags/overview): High-level component and feature overviews. - [2 docs tagged with "Parameters"](/docs/tags/parameters): Configuration parameters and runtime options. - [2 docs tagged with "Performance"](/docs/tags/performance): Performance optimization and benchmarking. - [One doc tagged with "Persistence"](/docs/tags/persistence): Data persistence and durable storage mechanisms. - [4 docs tagged with "Postgres"](/docs/tags/postgres): PostgreSQL database integration and configuration. - [One doc tagged with "Python"](/docs/tags/python): Python programming language topics. - [3 docs tagged with "Query"](/docs/tags/query): SQL query execution, parameterized queries, and prepared statements. - [4 docs tagged with "Reference"](/docs/tags/reference): API reference, CLI commands, and configuration syntax. - [One doc tagged with "Relational"](/docs/tags/relational): Relational database systems and SQL operations. - [3 docs tagged with "Runtime"](/docs/tags/runtime): Runtime behavior and execution environment. - [One doc tagged with "Rust"](/docs/tags/rust): Rust programming language topics. - [2 docs tagged with "S3"](/docs/tags/s3): Amazon S3 object storage integration. - [One doc tagged with "S3 Express"](/docs/tags/s3-express): Amazon S3 Express One Zone storage. - [One doc tagged with "Sandbox"](/docs/tags/sandbox): Sandbox environments and isolated testing. - [One doc tagged with "ScyllaDB"](/docs/tags/scylladb): ScyllaDB database integration. - [6 docs tagged with "SDK"](/docs/tags/sdk): Software Development Kit topics and usage. - [10 docs tagged with "Search"](/docs/tags/search): Vector search, semantic search, and ranking capabilities. - [2 docs tagged with "Security"](/docs/tags/security): Security features and data protection mechanisms. - [One doc tagged with "Snowflake"](/docs/tags/snowflake): Snowflake cloud data warehouse integration. - [12 docs tagged with "SpiceAI"](/docs/tags/spiceai): SpiceAI cloud platform and managed services. - [5 docs tagged with "Spicepod"](/docs/tags/spicepod): Spicepod configuration files, manifest syntax, and package management. - [6 docs tagged with "SQL"](/docs/tags/sql): SQL query language support and database operations. - [One doc tagged with "systemd"](/docs/tags/systemd): Run Spice as a Linux service. - [5 docs tagged with "Tools"](/docs/tags/tools): Development tools and utility integrations. - [2 docs tagged with "Tracing"](/docs/tags/tracing): Distributed tracing and request monitoring. - [One doc tagged with "Troubleshooting"](/docs/tags/troubleshooting): Problem diagnosis and resolution guides. - [One doc tagged with "Turso"](/docs/tags/turso): Turso data accelerator integration. - [2 docs tagged with "UDF"](/docs/tags/udf): User-defined functions registered with the SQL engine. - [2 docs tagged with "Unity Catalog"](/docs/tags/unity-catalog): Databricks Unity Catalog data governance integration. - [One doc tagged with "Views"](/docs/tags/views): Virtual views and data transformation layers. - [One doc tagged with "Vortex"](/docs/tags/vortex): Vortex columnar file format and storage engine. - [10 docs tagged with "Write"](/docs/tags/write): Data connectors and catalogs that support write operations. - [One doc tagged with "YAML"](/docs/tags/yaml): YAML configuration syntax and file formats. - [One doc tagged with "Zipkin"](/docs/tags/zipkin): Zipkin distributed tracing system integration. ### acknowledgements Spice AI acknowledges the following open source projects for making this project possible: - [Open Source Acknowledgements](/docs/acknowledgements): Spice AI acknowledges the following open source projects for making this project possible: ### api Spice.ai API reference including HTTP REST API, Arrow Flight SQL, JDBC, ODBC, ADBC connectors, authentication, and TLS configuration. - [Spice.ai API Reference](/docs/api): Spice.ai API reference including HTTP REST API, Arrow Flight SQL, JDBC, ODBC, ADBC connectors, authentication, and TLS configuration. - [ADBC: Arrow Database Connectivity](/docs/api/adbc): ADBC API Documentation - [Arrow Flight SQL API](/docs/api/arrow-flight-sql): Query Spice using JDBC/ODBC/ADBC - [Authentication](/docs/api/auth): Authentication documentation - [Generate Package](/docs/api/HTTP/generate-package): This endpoint generates a zip package from a specified GitHub source. - [List Catalogs](/docs/api/HTTP/get-catalogs): Returns a list of all registered catalogs (data sources). Catalogs provide metadata about schemas and tables available from external data sources. - [Get Iceberg API config](/docs/api/HTTP/get-config): This endpoint returns the Iceberg Catalog API configuration, including details about overrides, defaults, and available endpoints. - [List Datasets](/docs/api/HTTP/get-datasets): This endpoint returns a list of configured datasets. The response can be formatted as **JSON** or **CSV**, - [List Iceberg namespaces](/docs/api/HTTP/get-iceberg-namespaces): This endpoint retrieves namespaces available in the Iceberg catalog. - [List Models](/docs/api/HTTP/get-models): List all models, both machine learning and language models, available in the runtime. - [Check if a namespace exists.](/docs/api/HTTP/get-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [List Spicepods](/docs/api/HTTP/get-spicepods): Get a list of spicepods and their summary details. - [Check Runtime Status](/docs/api/HTTP/get-status): Return the status of all connections (http, flight, metrics, opentelemetry) in the runtime. - [Get a table.](/docs/api/HTTP/get-table): This endpoint returns the table if it exists, otherwise it returns a 404 Not Found response. - [List Workers](/docs/api/HTTP/get-workers): Returns a list of all registered workers in the runtime. Workers are configurable processing units that can perform tasks like load balancing between models or implementing fallback strategies. - [Check Namespace exists](/docs/api/HTTP/head-namespace): This endpoint returns a 200 OK response if the namespace exists, otherwise it returns a 404 Not Found response. - [Check if a table exists.](/docs/api/HTTP/head-table): This endpoint returns a 200 OK response if the table exists, otherwise it returns a 404 Not Found response. - [list_tables](/docs/api/HTTP/list-tables): list_tables - [List Tools](/docs/api/HTTP/list-tools): Returns a list of all available tools in the Spice runtime. Tools provide reusable functionality that can be invoked programmatically or by AI agents. - [Send a Model Context Protocol message](/docs/api/HTTP/mcp-message): Send a JSON-RPC message to the Spice MCP server using the MCP Streamable HTTP transport. The response is either a single JSON-RPC response (`application/json`) or an SSE stream (`text/event-stream`), selected via the `Accept` header. Session continuity is carried via the `Mcp-Session-Id` header. - [Open an MCP server-to-client SSE stream](/docs/api/HTTP/mcp-stream): Open a long-lived server-to-client SSE stream for the current MCP session as defined by the Streamable HTTP transport. The `Mcp-Session-Id` header must identify an existing session created via `POST /v1/mcp`. - [Terminate an MCP Streamable HTTP session](/docs/api/HTTP/mcp-terminate-session): Terminate the MCP session identified by the `Mcp-Session-Id` header. Subsequent requests bearing the same session id will receive `404 Not Found`. - [Update Refresh SQL](/docs/api/HTTP/patch-dataset-acceleration): Update the refresh SQL for a dataset's acceleration. - [Create Chat Completion](/docs/api/HTTP/post-chat-completions): Creates a model response for the given chat conversation. - [Refresh Dataset](/docs/api/HTTP/post-dataset-refresh): Trigger an on-demand refresh for an accelerated dataset. - [Create Embeddings](/docs/api/HTTP/post-embeddings): Creates an embedding vector representing the input text. - [Text-to-SQL (NSQL)](/docs/api/HTTP/post-nsql): Generate and optionally execute a natural-language text-to-SQL (NSQL) query. - [post_responses](/docs/api/HTTP/post-responses): post_responses - [Search](/docs/api/HTTP/post-search): Perform a vector similarity search (VSS) operation on a dataset. - [SQL Query](/docs/api/HTTP/post-sql): Execute a SQL query and return the results. - [Check Readiness](/docs/api/HTTP/ready): Check the runtime status of all the components of the runtime. If the service is ready, it returns an HTTP 200 status with the message 'ready'. If not, it returns a 503 status with the message 'not ready'. - [Run Tool](/docs/api/HTTP/run-tool): Execute a specific tool by name. The request body schema and response format are defined by each individual tool's specification. Use `GET /v1/tools` to discover available tools and their parameter schemas. - [runtime](/docs/api/HTTP/runtime): The spiced runtime - [JDBC: Java Database Connectivity](/docs/api/jdbc): JDBC API Documentation - [ODBC: Open Database Connectivity](/docs/api/odbc): ODBC API Documentation - [API Overview](/docs/api/overview): Spice.ai API overview, including SQL query interfaces, OpenAI-compatible endpoints, Iceberg catalog REST APIs, and the Model Context Protocol (MCP) for integrating external tools. - [TLS: Transport Layer Security](/docs/api/tls): Encryption in transit with TLS documentation ### cli Complete CLI reference for Spice.ai including commands to create, manage Spicepods, run queries, and interact with the Spice runtime. - [Spice.ai CLI Reference](/docs/cli): Complete CLI reference for Spice.ai including commands to create, manage Spicepods, run queries, and interact with the Spice runtime. - [Spice.ai OSS CLI command reference](/docs/cli/reference): Spice CLI command reference - [add](/docs/cli/reference/add): Add a Spicepod to the project. - [catalogs](/docs/cli/reference/catalogs): List catalogs currently loaded by the Spice runtime. - [chat](/docs/cli/reference/chat): spice chat CLI documentation - [cloud](/docs/cli/reference/cloud): Manage Spice Cloud: sign in, enroll a directory, create and inspect projects, manage secrets, read logs, and deploy. - [completions](/docs/cli/reference/completions): Generate shell completions for the Spice CLI. - [connect](/docs/cli/reference/connect): spice connect / is deprecated. Use spice add /. - [dataset](/docs/cli/reference/dataset): Add or configure dataset entries in spicepod.yaml. - [datasets](/docs/cli/reference/datasets): Lists datasets loaded by the Spice runtime - [feedback](/docs/cli/reference/feedback): Open the Spice.ai community Slack in the default browser to share feedback. - [init](/docs/cli/reference/init): Initialize Spice app in the current working directory. - [install](/docs/cli/reference/install): Download and install the latest version of the Spice runtime. - [login](/docs/cli/reference/login): Login to the Spice.ai Platform, or other services with sub-commands. - [models](/docs/cli/reference/models): Lists models loaded by the Spice runtime - [nsql](/docs/cli/reference/nsql): spice nsql CLI documentation - [pods](/docs/cli/reference/pods): Lists Spicepods loaded by the Spice runtime - [query](/docs/cli/reference/query): Submit an async query or start an interactive async query REPL against the Spice runtime's distributed query engine. - [refresh](/docs/cli/reference/refresh): Refreshes an accelerated dataset loaded by the Spice runtime - [run](/docs/cli/reference/run): Run Spice - starts the Spice runtime, installing if necessary. - [search](/docs/cli/reference/search): Performs embeddings-based searches across search-configured datasets. - [spiced](/docs/cli/reference/spiced): Command-line reference for the spiced runtime binary — flags, defaults, and differences from spice run. - [sql](/docs/cli/reference/sql): Start an interactive SQL query session against the Spice runtime - [status](/docs/cli/reference/status): Spice runtime status - [trace](/docs/cli/reference/trace): Provides a user-friendly trace stack into an operation that occurred in Spice. This command retrieves and displays task execution traces from the runtime.task_history table. - [upgrade](/docs/cli/reference/upgrade): Upgrades the Spice CLI and runtime to the latest or specified version - [validate](/docs/cli/reference/validate): Validate a spicepod.yaml without starting the runtime. - [version](/docs/cli/reference/version): Outputs the current version of the Spice CLI and runtime - [Configuring Trace Levels](/docs/cli/tracing): Configuring Spice.ai OSS trace output verbosity levels ### clients Client and tools for connecting to Spice - [Clients and Tools](/docs/clients): Client and tools for connecting to Spice - [DBeaver](/docs/clients/dbeaver): Configure DBeaver to query Spice via JDBC - [JetBrains DataGrip](/docs/clients/jetbrains-datagrip): Configure JetBrains Datagrip to query Spice via JDBC - [Microsoft Power BI Connector](/docs/clients/powerbi): Use Microsoft Power BI to access, visualize and analyze Spice datasets. - [Apache Superset](/docs/clients/superset): Use Apache Superset to query and visualize datasets loaded in Spice. - [Tableau](/docs/clients/tableau): Use Tableau to to access, visualise and analyse datasets loaded in Spice. ### comparison Compare Spice.ai with data platforms (Databricks, Snowflake), query engines (Trino, Dremio, ClickHouse), vector databases (Turbopuffer, LanceDB), search engines (Elasticsearch), and AI frameworks (LangChain, LlamaIndex, Ollama). - [How Spice Compares](/docs/comparison): Compare Spice.ai with data platforms (Databricks, Snowflake), query engines (Trino, Dremio, ClickHouse), vector databases (Turbopuffer, LanceDB), search engines (Elasticsearch), and AI frameworks (LangChain, LlamaIndex, Ollama). ### components Configure Spice.ai runtime components including data connectors, data accelerators, catalog connectors, model providers, embedding models, and secret stores. - [Spice.ai Runtime Components](/docs/components): Configure Spice.ai runtime components including data connectors, data accelerators, catalog connectors, model providers, embedding models, and secret stores. - [Catalog Connectors](/docs/components/catalogs): Connect to external catalog providers like Unity Catalog, Databricks, Iceberg, AWS Glue, Snowflake, ADBC, PostgreSQL, MySQL, MSSQL, Oracle, and more for federated SQL query in Spice. - [ADBC Catalog Connector](/docs/components/catalogs/adbc): Connect to databases via ADBC for automatic schema and table discovery. - [Databricks Catalog Connector](/docs/components/catalogs/databricks): Connect to a Databricks Unity Catalog provider. - [DuckLake Catalog Connector](/docs/components/catalogs/ducklake): Connect to a DuckLake catalog for federated SQL query. - [Glue Catalog Connector](/docs/components/catalogs/glue): Connect to an AWS Glue Data Catalog. - [Iceberg Catalog Connector](/docs/components/catalogs/iceberg): Connect to an Iceberg catalog provider. - [Microsoft SQL Server Catalog Connector](/docs/components/catalogs/mssql): Connect to a Microsoft SQL Server database as a catalog provider for federated SQL query. - [MySQL Catalog Connector](/docs/components/catalogs/mysql): Connect to a MySQL database as a catalog provider for federated SQL query. - [Oracle Catalog Connector](/docs/components/catalogs/oracle): Connect to an Oracle database as a catalog provider for federated SQL query. - [PostgreSQL Catalog Connector](/docs/components/catalogs/postgres): Connect to a PostgreSQL database as a catalog provider for federated SQL query. - [Snowflake Catalog Connector](/docs/components/catalogs/snowflake): Connect to a Snowflake database as a catalog provider for federated SQL query. - [Spice.ai Catalog Connector](/docs/components/catalogs/spiceai): Connect to the Spice.ai built-in catalog. - [Unity Catalog Catalog Connector](/docs/components/catalogs/unity-catalog): Connect to a Unity Catalog provider. - [Unity Catalog Catalog Connector Deployment Guide](/docs/components/catalogs/unity-catalog/deployment): Operating guide for the Unity Catalog catalog connector in production: workspace authentication, table-type filtering, effective-permissions flow, and observability. - [Data Accelerators](/docs/components/data-accelerators): Data acceleration engines for local materialization and query acceleration in Spice - [In-Memory Arrow Data Accelerator](/docs/components/data-accelerators/arrow): In-Memory Arrow Data Accelerator Documentation - [Arrow Data Accelerator Deployment Guide](/docs/components/data-accelerators/arrow/deployment): Operating guide for the Arrow (in-memory) data accelerator in production: memory sizing, indexes, and observability. - [Spice Cayenne Data Accelerator](/docs/components/data-accelerators/cayenne): Spice Cayenne Data Accelerator (Vortex) Documentation - [Cayenne Data Accelerator Deployment Guide](/docs/components/data-accelerators/cayenne/deployment): Operating guide for Spice Cayenne in production: footer and segment caches, S3 Express, metastore durability, and observability. - [Spice Cayenne Performance Tuning](/docs/components/data-accelerators/cayenne/performance): Tuning Spice Cayenne: storage tier and NVMe placement, self-tuning and goal-driven SLOs, cache sizing, compression strategy, sorted data, write-path and compaction tuning, and file-size tuning. - [DuckDB Data Accelerator](/docs/components/data-accelerators/duckdb): DuckDB Data Accelerator Documentation - [DuckDB Data Accelerator Deployment Guide](/docs/components/data-accelerators/duckdb/deployment): Operating guide for the DuckDB data accelerator in production: memory vs file mode, checkpointing, spill, pool sizing, and observability. - [PostgreSQL Data Accelerator](/docs/components/data-accelerators/postgres): PostgreSQL Data Accelerator Documentation - [PostgreSQL Data Accelerator Deployment Guide](/docs/components/data-accelerators/postgres/deployment): Operating guide for the PostgreSQL data accelerator in production: authentication, connection pooling, and observability. - [SQLite Data Accelerator](/docs/components/data-accelerators/sqlite): SQLite Data Accelerator Documentation - [SQLite Data Accelerator Deployment Guide](/docs/components/data-accelerators/sqlite/deployment): Operating guide for the SQLite data accelerator in production: file mode, busy timeout, pool, and observability. - [Turso Data Accelerator](/docs/components/data-accelerators/turso): Turso (libSQL) Data Accelerator Documentation - [Data Connectors](/docs/components/data-connectors): Learn how to use Data Connector to query external data. - [Azure BlobFS Data Connector](/docs/components/data-connectors/abfs): Azure BlobFS Data Connector Documentation - [ADBC Data Connector](/docs/components/data-connectors/adbc): ADBC Data Connector Documentation - [ClickHouse Data Connector](/docs/components/data-connectors/clickhouse): ClickHouse Data Connector Documentation - [Azure Cosmos DB Data Connector](/docs/components/data-connectors/cosmosdb): Query Azure Cosmos DB (NoSQL / Core SQL API) containers as SQL tables in Spice. Read-only scan with schema inferred from a sample of documents. - [Azure Cosmos DB Data Connector Deployment Guide](/docs/components/data-connectors/cosmosdb/deployment): Operating guide for the Azure Cosmos DB data connector in production: authentication, RU sizing, resilience, metrics, and observability. - [Databricks Data Connector](/docs/components/data-connectors/databricks): Databricks Data Connector Documentation - [Databricks Deployment Guide](/docs/components/data-connectors/databricks/deployment): Operating guide for the Databricks connector in production: resilience controls, Unity Catalog behavior, metrics, and observability. - [Debezium Data Connector](/docs/components/data-connectors/debezium): Debezium Data Connector Documentation - [Delta Lake Data Connector](/docs/components/data-connectors/delta-lake): Delta Lake Data Connector Documentation - [Delta Lake Data Connector Deployment Guide](/docs/components/data-connectors/delta-lake/deployment): Operating guide for the Delta Lake data connector in production: object store auth, metadata caching, metrics, and observability. - [Dremio Data Connector](/docs/components/data-connectors/dremio): Dremio Data Connector Documentation - [Dremio Data Connector Deployment Guide](/docs/components/data-connectors/dremio/deployment): Operating guide for the Dremio data connector in production: authentication, Flight SQL transport, and observability. - [DuckDB Data Connector](/docs/components/data-connectors/duckdb): DuckDB Data Connector Documentation - [DuckDB Data Connector Deployment Guide](/docs/components/data-connectors/duckdb/deployment): Operating guide for the DuckDB data connector in production: file vs memory mode, durability, and observability. - [DuckLake Data Connector](/docs/components/data-connectors/ducklake): DuckLake Data Connector Documentation - [DynamoDB Data Connector](/docs/components/data-connectors/dynamodb): DynamoDB Data Connector Documentation - [DynamoDB Data Connector Deployment Guide](/docs/components/data-connectors/dynamodb/deployment): Operating guide for the DynamoDB data connector in production: IAM, streams, checkpointing, lag behavior, and observability. - [Elasticsearch Data Connector](/docs/components/data-connectors/elasticsearch): Query Elasticsearch indexes as SQL tables in Spice, including kNN vector search, full-text search, and hybrid search. - [Elasticsearch Data Connector Deployment Guide](/docs/components/data-connectors/elasticsearch/deployment): Operating guide for the Elasticsearch data connector in production: authentication, TLS, resilience, and operational tuning. - [File Data Connector](/docs/components/data-connectors/file): File Data Connector Documentation - [File Data Connector Deployment Guide](/docs/components/data-connectors/file/deployment): Operating guide for the File data connector in production: permissions, formats, performance, and observability. - [Flight SQL Data Connector](/docs/components/data-connectors/flightsql): Flight SQL Data Connector Documentation - [FTP/SFTP Data Connector](/docs/components/data-connectors/ftp): FTP/SFTP Data Connector Documentation - [GCS Data Connector](/docs/components/data-connectors/gcs): GCS (Google Cloud Storage) Data Connector Documentation - [GitHub Data Connector](/docs/components/data-connectors/github): GitHub Data Connector Documentation - [GitHub Data Connector Deployment Guide](/docs/components/data-connectors/github/deployment): Operating guide for the GitHub data connector in production: PATs, rate limits, pagination, and observability. - [Glue Data Connector](/docs/components/data-connectors/glue): Connect to and query tables in an AWS Glue Data Catalog - [GraphQL Data Connector](/docs/components/data-connectors/graphql): GraphQL Data Connector Documentation - [GraphQL Data Connector Deployment Guide](/docs/components/data-connectors/graphql/deployment): Operating guide for the GraphQL data connector in production: authentication, pagination, rate limits, and observability. - [HTTP(s) Data Connector](/docs/components/data-connectors/https): HTTP(s) Data Connector Documentation - [HTTP(s) Data Connector Deployment Guide](/docs/components/data-connectors/https/deployment): Operating guide for the HTTP(s) data connector in production: authentication, rate control, retries, and observability. - [Iceberg Data Connector](/docs/components/data-connectors/iceberg): Connect to and query Apache Iceberg tables - [IMAP Data Connector](/docs/components/data-connectors/imap): IMAP Data Connector Documentation - [Kafka Data Connector](/docs/components/data-connectors/kafka): Kafka Data Connector Documentation - [Localpod Data Connector](/docs/components/data-connectors/localpod): Localpod Data Connector Documentation - [Memory Data Connector](/docs/components/data-connectors/memory): Memory Data Connector Documentation - [MongoDB Data Connector](/docs/components/data-connectors/mongodb): MongoDB Data Connector Documentation - [Microsoft SQL Server Data Connector](/docs/components/data-connectors/mssql): Microsoft SQL Server Data Connector - [Microsoft SQL Server Performance](/docs/components/data-connectors/mssql/performance): Performance tuning for the Microsoft SQL Server connector: TopK / ORDER BY ... LIMIT pushdown and NULL ordering, read-only routing, and acceleration on local NVMe. - [MySQL Data Connector](/docs/components/data-connectors/mysql): MySQL Data Connector Documentation - [MySQL Data Connector Deployment Guide](/docs/components/data-connectors/mysql/deployment): Operating guide for the MySQL data connector in production: authentication, connection pooling, TLS, metrics, and observability. - [NFS Data Connector](/docs/components/data-connectors/nfs): NFS Data Connector Documentation - [ODBC Data Connector](/docs/components/data-connectors/odbc): ODBC Data Connector Documentation - [Oracle Data Connector](/docs/components/data-connectors/oracle): Oracle Data Connector Documentation - [PostgreSQL Data Connector](/docs/components/data-connectors/postgres): PostgreSQL Data Connector Documentation - [PostgreSQL Data Connector Deployment Guide](/docs/components/data-connectors/postgres/deployment): Operating guide for the PostgreSQL data connector in production: authentication, connection pooling, TLS, metrics, and observability. - [Amazon Redshift Data Connector](/docs/components/data-connectors/redshift): Connect to Amazon Redshift using the PostgreSQL connector in Spice. - [S3 Data Connector](/docs/components/data-connectors/s3): S3 Data Connector Documentation - [S3 Data Connector Deployment Guide](/docs/components/data-connectors/s3/deployment): Operating guide for the S3 data connector in production: IAM, credential chains, file formats, metrics, and observability. - [ScyllaDB Data Connector](/docs/components/data-connectors/scylladb): ScyllaDB Data Connector Documentation - [ScyllaDB Performance](/docs/components/data-connectors/scylladb/performance): Performance considerations for the ScyllaDB connector: partition/clustering key filters, acceleration and where to store it, datacenter locality, and connection timeouts. - [SharePoint Data Connector](/docs/components/data-connectors/sharepoint): SharePoint Data Connector Documentation - [SMB Data Connector](/docs/components/data-connectors/smb): SMB Data Connector Documentation - [Snowflake Data Connector](/docs/components/data-connectors/snowflake): Snowflake Data Connector Documentation - [Apache Spark Connector](/docs/components/data-connectors/spark): Apache Spark Connector Documentation - [Spice.ai Data Connector](/docs/components/data-connectors/spiceai): Federated SQL across Spice runtimes — Spice Cloud Platform datasets and self-hosted Spice instances (cluster-sidecar pattern). - [Spice.ai Data Connector Deployment Guide](/docs/components/data-connectors/spiceai/deployment): Operating guide for the Spice.ai connector in production: API keys, Flight endpoints, message sizing, sidecar topology, and observability. - [Embedding Models](/docs/components/embeddings): Describes how embedding models are used in Spice to convert text into numerical vectors for machine learning and search applications. - [Azure OpenAI Embedding Models](/docs/components/embeddings/azure): To use an embedding model hosted on Azure OpenAI, specify the azure path in the from field and the following parameters from the Azure OpenAI Model Deployment page: - [Amazon Bedrock Model Provider](/docs/components/embeddings/bedrock): Instructions for using Amazon Bedrock embedding models - [Databricks Model Provider](/docs/components/embeddings/databricks): Instructions for using Databricks Mosaic AI Models - [Google Vertex AI Embedding Models](/docs/components/embeddings/google): from selects a Vertex AI embedding model. A model ID, GCP project, location, - [HuggingFace Text Embedding Models](/docs/components/embeddings/huggingface): To use an embedding model from HuggingFace with Spice, specify the huggingface path in the from field of your configuration. The model and its related files will be automatically downloaded, loaded, and served locally by Spice. - [Hugging Face Embedding Deployment Guide](/docs/components/embeddings/huggingface/deployment): Operating guide for Hugging Face embeddings in production: tokens, download cache, pooling, device selection, and observability. - [Local Filesystem Embedding Models](/docs/components/embeddings/local): Embedding models can be run with files stored locally. This method is useful for using models that are not hosted on remote services. - [Local Embedding Deployment Guide](/docs/components/embeddings/local/deployment): Operating guide for filesystem-loaded embedding models in production: formats, pooling, device selection, and observability. - [Model2Vec Embedding Models](/docs/components/embeddings/model2vec): Model2Vec embedding models help generate efficient static word embeddings from sentence transformer models for use in Spice, supporting local and Hugging Face sources with options for private models and performance tuning. - [OpenAI (or Compatible) Embedding Models](/docs/components/embeddings/openai): To use a hosted OpenAI (or compatible) embedding model, specify the openai path in the from field of your configuration. - [OpenAI Embedding Deployment Guide](/docs/components/embeddings/openai/deployment): Operating guide for the OpenAI embedding provider in production: API keys, usage tiers, batching, retries, metrics, and observability. - [Model Providers](/docs/components/models): Overview of supported model providers for LLMs in Spice. - [Anthropic Models](/docs/components/models/anthropic): Instructions for using language models hosted on Anthropic with Spice. - [Azure OpenAI Models](/docs/components/models/azure): Instructions for using Azure OpenAI models - [Amazon Bedrock Models](/docs/components/models/bedrock): How to use Amazon Bedrock models with Spice. - [Databricks Model Provider](/docs/components/models/databricks): Instructions for using Databricks Mosaic AI Models - [Filesystem Hosted Models](/docs/components/models/filesystem): Instructions for using models hosted on a filesystem with Spice. - [Filesystem Model Deployment Guide](/docs/components/models/filesystem/deployment): Operating guide for filesystem-loaded models in production: formats, device selection, memory footprint, and observability. - [Google Vertex AI Models](/docs/components/models/google): Instructions for using language models hosted on Google Vertex AI with Spice. - [HuggingFace](/docs/components/models/huggingface): Instructions for using machine learning models hosted on HuggingFace with Spice. - [Hugging Face Model Deployment Guide](/docs/components/models/huggingface/deployment): Operating guide for the Hugging Face model in production: tokens, download cache, device selection, local inference footprint, and observability. - [OpenAI (or Compatible) Language Models](/docs/components/models/openai): Instructions for using language models hosted on OpenAI or compatible services with Spice. - [OpenAI Model Deployment Guide](/docs/components/models/openai/deployment): Operating guide for the OpenAI model in production: API keys, usage tiers, rate limiting, Responses API, metrics, and observability. - [Perplexity Models (Deprecated)](/docs/components/models/perplexity): Perplexity model support is no longer supported in Spice. - [Spice Cloud Platform](/docs/components/models/spiceai): Instructions for using language models served by the Spice.ai Cloud Platform with Spice. - [xAI Models](/docs/components/models/xai): Instructions for using xAI models - [Secret Stores](/docs/components/secret-stores): Configure secret stores to manage sensitive data like passwords, tokens, and API keys. - [AWS Secrets Manager Secret Store](/docs/components/secret-stores/aws-secrets-manager): AWS Secrets Manager Secret Store Documentation - [Azure Key Vault Secret Store](/docs/components/secret-stores/azure-keyvault): Azure Key Vault Secret Store Documentation - [Environment Secret Store](/docs/components/secret-stores/env): Environment Variables Secret Store Documentation - [HashiCorp Vault Secret Store](/docs/components/secret-stores/hashicorp-vault): HashiCorp Vault Secret Store Documentation - [Keyring Secret Store](/docs/components/secret-stores/keyring): Keyring Secret Store Documentation - [Kubernetes Secret Store](/docs/components/secret-stores/kubernetes): Kubernetes Secret Store Documentation - [LLM Tools (Function Calling)](/docs/components/tools): Overview of supported LLM tools (function calling) and how to define new tools - [Model Context Protocol Tools](/docs/components/tools/mcp): Spice integrates with tools and services using the Model Context Protocol (MCP). MCP tools can be configured to run internally or connect to external servers over HTTP using the Streamable HTTP transport. - [Web Search Tool (Deprecated)](/docs/components/tools/websearch): The websearch tool is no longer supported in Spice. - [Vector Engines](/docs/components/vectors): Configure vector engines for efficient embedding storage and similarity search in Spice. - [DuckDB Vector Engine](/docs/components/vectors/duckdb): Use DuckDB as a vector engine in Spice for HNSW-based vector search via the DuckDB VSS extension. - [Elasticsearch Vector Engine](/docs/components/vectors/elasticsearch): Use Elasticsearch as a vector engine in Spice for kNN vector search, full-text search, and hybrid search. - [Amazon S3 Vectors Engine](/docs/components/vectors/s3_vectors): Amazon S3 Vectors Engine Documentation ### deployment Deploy Spice.ai in your environment using Docker, Kubernetes, AWS, Azure, or the Spice Cloud Platform. Learn about sidecar, microservice, tiered, and cluster deployment architectures. - [Spice.ai Deployment Guide](/docs/deployment): Deploy Spice.ai in your environment using Docker, Kubernetes, AWS, Azure, or the Spice Cloud Platform. Learn about sidecar, microservice, tiered, and cluster deployment architectures. - [Deployment Architectures](/docs/deployment/architectures): Explore Spice deployment architectures including sidecar, microservice, tiered, sharded, and cluster configurations. - [Cluster-Based Deployment (Spice.ai Enterprise)](/docs/deployment/architectures/cluster): Deploying Spice as a cluster - [Cluster-Sidecar Deployment](/docs/deployment/architectures/cluster-sidecar): Deploy Spice with application-local sidecars for localhost query, search, and inference, backed by a centralized cluster for ingest, acceleration, and distributed query. - [Cloud Hosted](/docs/deployment/architectures/hosted): Deploying Spice cloud hosted in the Spice Cloud Platform - [Microservice Deployment (Single or Multiple Replicas)](/docs/deployment/architectures/microservice): Deploying Spice as a microservice - [Sharded](/docs/deployment/architectures/sharded): Deploying Spice with shards - [Sidecar Deployment](/docs/deployment/architectures/sidecar): Deploying Spice as a application sidecar - [Tiered Deployment](/docs/deployment/architectures/tiered): Deploying Spice in tiers - [AWS Deployment Options](/docs/deployment/aws): Guide to deploying Spice.ai applications on Amazon Web Services (AWS) - [AWS Integrations](/docs/deployment/aws/integrations): Complete guide to Spice.ai integrations with Amazon Web Services, including data connectors, AI models, vector stores, and secret management. - [Azure Deployment Options](/docs/deployment/azure): Guide to deploying Spice.ai applications on Microsoft Azure - [Azure Integrations](/docs/deployment/azure/integrations): Spice.ai integrations with Microsoft Azure, including data connectors, AI models, embeddings, and authentication. - [CI/CD Deployment](/docs/deployment/ci-cd): Deploy Spice.ai applications using continuous integration and delivery pipelines, including Helm, Kubernetes GitOps with Argo CD or Flux, GitHub Actions, and the Spice Cloud deploy action. - [Spice Cloud Platform](/docs/deployment/cloud): Managed Spice.ai hosting, and connecting a self-hosted instance to it with Cloud Connect. - [Cloud Connect](/docs/deployment/cloud/cloud-connect): Connect a self-hosted Spice instance to Spice Cloud for remote management. - [Cloud Connect on a Development Machine](/docs/deployment/cloud/cloud-connect/development): Connect a development machine to Spice Cloud and run the instance in the foreground. - [Headless Cloud Connect](/docs/deployment/cloud/cloud-connect/headless): Connect a Spice instance to Spice Cloud without an interactive terminal. - [Cloud Connect as a Service](/docs/deployment/cloud/cloud-connect/service): Run a connected Spice instance as a Linux or macOS service. - [Docker](/docs/deployment/docker): Run Spice.ai as a Docker container. - [Docker Sandbox Guide - v1.3.0](/docs/deployment/docker/sandbox): Migrating to v1.3.0 - [Google Cloud Deployment Options](/docs/deployment/gcp): Guide to deploying Spice.ai applications on Google Cloud Platform (GCP). - [GCP Integrations](/docs/deployment/gcp/integrations): Spice.ai integrations with Google Cloud Platform, including data connectors, AI models, embeddings, and authentication. - [Kubernetes Deployment](/docs/deployment/kubernetes): Deploy Spice.ai on Kubernetes using Helm, Argo CD, or Flux. - [Kubernetes - Argo CD](/docs/deployment/kubernetes/argocd): Deploy Spice.ai on self-hosted Kubernetes using Argo CD and the Spice Helm chart. - [Kubernetes - Flux](/docs/deployment/kubernetes/flux): Deploy Spice.ai on Kubernetes using Flux CD and the Spice Helm chart. - [Kubernetes - Helm](/docs/deployment/kubernetes/helm): Deploy Spice.ai in Kubernetes using Helm. - [Kubernetes - Local NVMe Storage](/docs/deployment/kubernetes/local-nvme): Step-by-step guide to giving Spice acceleration files and query spill a local NVMe volume on Kubernetes — EKS, GKE, AKS, and self-hosted clusters. - [Read/Write Separation](/docs/deployment/read-write-separation): Separate write/ingest workloads (cluster) from read workloads (application sidecars, agents) using shared snapshots and live query delegation. ### faq Answers to frequently asked questions about Spice.ai including features, use cases, differences from Trino/Presto/Dremio, federated queries, caching, and AI capabilities. - [Spice.ai FAQ](/docs/faq): Answers to frequently asked questions about Spice.ai including features, use cases, differences from Trino/Presto/Dremio, federated queries, caching, and AI capabilities. ### features Explore Spice.ai features including data federation, data acceleration, caching, search, LLM integration, embeddings, observability, and more for building data-driven AI applications. - [Spice.ai Features](/docs/features): Explore Spice.ai features including data federation, data acceleration, caching, search, LLM integration, embeddings, observability, and more for building data-driven AI applications. - [Caching](/docs/features/caching): Learn how to use Spice in-memory caching - [Change Data Capture (CDC)](/docs/features/cdc): Learn how to use Change Data Capture (CDC) in Spice. - [Debezium (CDC over Kafka)](/docs/features/cdc/debezium): Consume Debezium change events from Kafka into a Spice-accelerated dataset for sources without a native Spice CDC path. - [Debezium Push Ingest (CDC without Kafka)](/docs/features/cdc/debezium-ingest): Stream Debezium change events directly into a Spice-accelerated dataset over HTTP — any Debezium source plugin, no Kafka bus required. - [DynamoDB Streams (Native CDC)](/docs/features/cdc/dynamodb-streams): Stream INSERT, UPDATE, and DELETE events from Amazon DynamoDB directly into a Spice-accelerated dataset using DynamoDB Streams. - [MongoDB Change Streams (Native CDC)](/docs/features/cdc/mongodb-streams): Stream insert, update, replace, and delete events from MongoDB directly into a Spice-accelerated dataset using MongoDB Change Streams. - [MySQL Binlog Replication (Native CDC)](/docs/features/cdc/mysql-replication): Stream INSERT, UPDATE, and DELETE events from MySQL directly into a Spice-accelerated dataset using native binary log (binlog) replication. - [PostgreSQL Logical Replication (Native CDC)](/docs/features/cdc/postgres-replication): Stream INSERT, UPDATE, and DELETE events from PostgreSQL directly into a Spice-accelerated dataset using native logical replication. - [Data Acceleration](/docs/features/data-acceleration): Learn how to use local data acceleration in Spice. - [Constraints](/docs/features/data-acceleration/constraints): Learn how to add/configure constraints on local acceleration tables in Spice. - [Data Refresh](/docs/features/data-acceleration/data-refresh): Data refresh for accelerated datasets - [Hash Index for Arrow Acceleration](/docs/features/data-acceleration/hash-index): Learn how to use hash indexes for O(1) point lookups on Arrow-accelerated datasets. - [Indexes](/docs/features/data-acceleration/indexes): Learn how to add indexes to local acceleration tables in Spice. - [Partitioning](/docs/features/data-acceleration/partitioning): Partition accelerated datasets to make filtered queries faster by reading only the relevant partitions. - [Refresh Modes](/docs/features/data-acceleration/refresh-modes): Refresh modes for accelerated datasets in Spice. - [Append Refresh Mode](/docs/features/data-acceleration/refresh-modes/append): Incrementally append new rows to an accelerated dataset. - [Caching Refresh Mode](/docs/features/data-acceleration/refresh-modes/caching): Learn how to use caching refresh mode for HTTP-based datasets - [Changes Refresh Mode](/docs/features/data-acceleration/refresh-modes/changes): Apply incremental inserts, updates, and deletes via Change Data Capture. - [Full Refresh Mode](/docs/features/data-acceleration/refresh-modes/full): Replace the entire accelerated dataset on each refresh. - [Snapshot Refresh Mode](/docs/features/data-acceleration/refresh-modes/snapshot): Reload acceleration data exclusively from the snapshot store. - [Snapshots](/docs/features/data-acceleration/snapshots): Bootstrap file-mode accelerations from managed snapshots to eliminate cold starts. - [Data Ingestion](/docs/features/data-ingestion): Learn how to ingest data in Spice. - [Distributed Query](/docs/features/distributed-query): Learn how to run Spice in distributed mode for larger scale queries, including the async queries API. - [Embedding Datasets](/docs/features/embeddings): Learn how to define, or augment existing datasets with embedding column(s). - [Functions](/docs/features/functions): Define custom scalar and table SQL functions inline (SQL tier) or by calling remote HTTP services (Remote tier), automatically exposed as SQL functions and LLM tools. - [Large Language Models](/docs/features/large-language-models): Learn how to configure large language models (LLMs) - [Evaluating Language Models (Deprecated)](/docs/features/large-language-models/evals): Language model evals are no longer supported in Spice. - [Model Context Protocol (MCP)](/docs/features/large-language-models/mcp): Learn how to use the Model Context Protocol (MCP) with Spice. - [Language Model Memory](/docs/features/large-language-models/memory): Learn how to provide LLMs with memory - [Language Model Overrides](/docs/features/large-language-models/parameter_overrides): Learn how to override default LLM hyperparameters in Spice. - [System Prompt parameterization](/docs/features/large-language-models/parameterized_prompts): Learn how to update system prompts for each request with Jinja-styled templating. - [Load and Serve Models Locally](/docs/features/large-language-models/serving): Learn how to load and serve large learning models. - [Language Models Tools](/docs/features/large-language-models/tools): Learn how LLMs interact with the Spice runtime. - [Machine Learning Models](/docs/features/machine-learning-models): Removed in v2.2: Support for loading and serving traditional machine learning (ONNX) models for inference was removed in v2.2, along with the /v1/predict and /v1/models//predict prediction endpoints. See the v2.1 docs for documentation of this feature. - [Observability & Monitoring](/docs/features/observability): Monitor Spice with Prometheus metrics, OpenTelemetry, and distributed tracing. - [Component Metrics](/docs/features/observability/component_metrics): Learn how to enable optional component metrics. - [Query Federation](/docs/features/query-federation): Learn how to use federated SQL queries in Spice.ai Open Source - [Parameterized Queries](/docs/features/query-federation/parameterized-queries): Learn how to use prepared statements and parameterized queries in Spice for improved security and performance. - [URL Tables](/docs/features/query-federation/url-tables): Query object store files directly using URLs without pre-registering datasets - [Search Functionality](/docs/features/search): Learn how Spice can search across datasets using database-native and vector-search methods. - [Full-Text Search](/docs/features/search/full-text): Learn how Spice can perform full text search - [Multi-Vector Search](/docs/features/search/multi-vector): Embed list-of-strings columns as a column of vectors and use ColBERT-style late-interaction search in Spice. - [Reranking](/docs/features/search/rerank): Rerank search results using dedicated reranker models or LLM-as-reranker for improved relevance. - [Vector-Based Search](/docs/features/search/vector-search): Learn how Spice can perform searches using vector-based methods. - [Semantic Model](/docs/features/semantic-model): Attach descriptions and metadata to datasets, views, and columns in Spice so LLMs, SQL functions, and humans share the same understanding of your data. - [Tool Registry](/docs/features/tool-registry): Reduce per-turn token cost and improve LLM tool selection accuracy by replacing individual tool definitions with searchable tool_search and tool_invoke meta-tools backed by hybrid full-text, keyword, schema, and vector search. - [Views](/docs/features/views): Documentation for defining Views in Spice - [Web Search](/docs/features/web-search): Learn how Spice can perform web search - [Workers](/docs/features/workers): Configure workers in the Spice runtime to coordinate interactions between LLMs and tools, with load-balancing, round-robin, and fallback strategies. ### getting-started Get started with Spice.ai in 5 minutes. Install the CLI, connect to datasets, run SQL queries, and use AI models with OpenAI-compatible APIs. - [Getting Started with Spice.ai OSS](/docs/getting-started): Get started with Spice.ai in 5 minutes. Install the CLI, connect to datasets, run SQL queries, and use AI models with OpenAI-compatible APIs. - [Community Data](/docs/getting-started/spiceai): Connect to the Spice.ai Cloud Platform to access community datasets. - [Spicepods](/docs/getting-started/spicepods): An introduction to Spicepods - [Telemetry](/docs/getting-started/telemetry): Learn how Spice AI uses anonymous telemetry. ### installation Install Spice.ai OSS on macOS, Linux, Windows, or WSL using the install script, Homebrew, PowerShell, or direct download from GitHub releases. - [Install Spice.ai OSS](/docs/installation): Install Spice.ai OSS on macOS, Linux, Windows, or WSL using the install script, Homebrew, PowerShell, or direct download from GitHub releases. ### intelligent-applications Learn how to build intelligent, data-driven AI applications and agents with Spice.ai. Explore patterns for RAG, LLM integration, and real-time AI inference. - [Building Intelligent AI Applications with Spice.ai](/docs/intelligent-applications): Learn how to build intelligent, data-driven AI applications and agents with Spice.ai. Explore patterns for RAG, LLM integration, and real-time AI inference. ### monitoring Monitor Spice.ai deployments with Datadog, Grafana, New Relic, Prometheus, and Zipkin integrations. - [Monitoring](/docs/monitoring): Monitor Spice.ai deployments with Datadog, Grafana, New Relic, Prometheus, and Zipkin integrations. - [Datadog](/docs/monitoring/datadog): Monitoring Spice with Datadog - [Grafana & Prometheus](/docs/monitoring/grafana): Monitoring Spice instances with Grafana & Prometheus - [New Relic](/docs/monitoring/new-relic): Monitoring Spice with New Relic - [Spice Cloud Platform](/docs/monitoring/spice-cloud): Connect a self-hosted Spice runtime to the Spice Cloud Platform to centralize task history and runtime observability across deployments. - [Zipkin Integration](/docs/monitoring/zipkin): Learn how to integrate Spice with Zipkin tracing. ### reference Reference documentation for Spice.ai including API reference, CLI commands, Spicepod configuration syntax, SQL reference, and data type specifications. - [Spice.ai OSS Reference Docs](/docs/reference): Reference documentation for Spice.ai including API reference, CLI commands, Spicepod configuration syntax, SQL reference, and data type specifications. - [Cron Schedules](/docs/reference/cron): The Runtime supports cron expressions with optional seconds, like /10 which evaluates to every 10th second (10, 20, 30, etc). - [Data Types Reference](/docs/reference/datatypes): Spice uses Apache Arrow data types internally, providing consistent type handling across different data sources and accelerators. This section documents how Arrow types map to specific accelerators and object store formats. - [Accelerator Data Types](/docs/reference/datatypes/accelerators): Spice adheres to Apache Arrow data types. Data accelerators do not support all Arrow data types. The table below outlines the data type compatibility for each accelerator, and datatype used within the accelerator. - [Object Store Data Types](/docs/reference/datatypes/object_store): Spice adheres to Apache Arrow data types. The table below lists the types of supported file type from object stores and their corresponding Apache Arrow type mappings in Spice. - [Spice Runtime Distributions](/docs/reference/distributions): Distribution variants of the Spice runtime for different use cases and deployment scenarios, including data-only, GPU-accelerated, NAS, and allocator variants. - [Duration](/docs/reference/duration): Durations are represented as a number with a time unit suffix. A value without a suffix is interpreted as seconds, and fractional values (e.g. 1.5h) are accepted. - [File Formats](/docs/reference/file_format): File-based data connectors — including s3//, file//, sftp://, and others — support multiple structured and document file formats. This page details the format-specific parameters available for each. - [Managing Memory Usage](/docs/reference/memory): Guidelines and best practices for managing memory usage and optimizing performance in Spice deployments. - [Models Grade Report](/docs/reference/models): Spice AI graded Large-Language-Model (LLM) evaluation report - [Performance Tuning](/docs/reference/performance-tuning): Comprehensive guide to optimizing query performance, acceleration, storage, spill-to-disk, and resource utilization in Spice deployments. - [YAML syntax for Spicepod manifests](/docs/reference/spicepod): Detailed documentation on the Spicepod manifest syntax (spicepod.yaml) - [catalogs](/docs/reference/spicepod/catalogs): Catalogs YAML reference - [Datasets](/docs/reference/spicepod/datasets): Datasets YAML reference - [Embeddings](/docs/reference/spicepod/embeddings): Embeddings YAML reference - [Evals (Deprecated)](/docs/reference/spicepod/evals): The evals Spicepod component is no longer supported in Spice. - [Functions (User-Defined Functions)](/docs/reference/spicepod/functions): User-defined functions YAML reference - [Reserved Keywords](/docs/reference/spicepod/keywords): Reserved keywords for datasets - [Models](/docs/reference/spicepod/models): Models YAML reference - [Runtime](/docs/reference/spicepod/runtime): Runtime YAML reference - [Tools (Function Calling)](/docs/reference/spicepod/tools): Tools YAML reference - [Views](/docs/reference/spicepod/views): Views YAML reference - [Workers](/docs/reference/spicepod/workers): Workers YAML reference - [SQL Reference](/docs/reference/sql): Complete SQL reference for Spice.ai including SELECT syntax, subqueries, DML statements, aggregate functions, AI functions, JSON operators, and search capabilities. - [Aggregate Functions](/docs/reference/sql/aggregate_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [AI Functions](/docs/reference/sql/ai): AI functions in Spice provide direct integration with large language models (LLMs) and embedding models within SQL queries. These functions process text through configured model providers and return generated responses or vector embeddings. - [DML (Data Manipulation Language)](/docs/reference/sql/dml): Data Manipulation Language (DML) statements for inserting and modifying data in Spice. - [Explain](/docs/reference/sql/explain): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Information Schema](/docs/reference/sql/information_schema): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [JSON Functions and Operators](/docs/reference/sql/json): Reference for JSON functions and operators in Spice SQL - [Operators](/docs/reference/sql/operators): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Prepared Statements](/docs/reference/sql/prepared_statements): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Scalar Functions](/docs/reference/sql/scalar_functions): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. When using a data accelerator like DuckDB, function support is specific to each acceleration engine, and not all functions are supported by all acceleration engines. - [Search in SQL](/docs/reference/sql/search): Reference for search functions and filtering in Spice SQL. - [SELECT](/docs/reference/sql/select): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Subqueries](/docs/reference/sql/subqueries): Spice is built on Apache DataFusion and uses the PostgreSQL dialect, even when querying datasources with different SQL dialects. - [Spice.ai Open Source System Requirements](/docs/reference/system_requirements): System requirements for running Spice.ai Open Source - [Task History](/docs/reference/task_history): The Spice runtime stores information about completed tasks in the spice.runtime.task_history table. Each task represents a single unit of execution within the runtime, such as a SQL query or an AI chat completion, and is represented by a unique span. ### sdks Connect to Spice using official SDKs - [SDKs](/docs/sdks): Connect to Spice using official SDKs - [Dotnet SDK](/docs/sdks/dotnet): Connect to Spice using the Dotnet SDK - [Go SDK](/docs/sdks/golang): Connect to Spice using the Go SDK - [Java SDK](/docs/sdks/java): Connect to Spice using the Java SDK - [JavaScript SDK](/docs/sdks/javascript): Connect to Spice using the JavaScript SDK - [Python SDK](/docs/sdks/python): Connect to Spice using the Python SDK - [Rust SDK](/docs/sdks/rust): Connect to Spice using the Rust SDK ### troubleshooting Review and debug runtime tasks, logs, and diagnostic steps in Spice. - [Troubleshooting Spice](/docs/troubleshooting): Review and debug runtime tasks, logs, and diagnostic steps in Spice. ### use-cases Discover how to use Spice.ai for data federation, reverse-ETL, database CDN, enterprise search, RAG, and building AI-powered applications and agents. - [Spice.ai Use Cases](/docs/use-cases): Discover how to use Spice.ai for data federation, reverse-ETL, database CDN, enterprise search, RAG, and building AI-powered applications and agents. - [AI Applications and Agents](/docs/use-cases/ai): AI Applications and Agents - [Agentic AI Applications and Agents](/docs/use-cases/ai/agentic-apps): Spice.ai builds intelligent, autonomous agents for SaaS applications, enabling context-aware automation and decision-making. - [Edge-Enabled AI Applications and Agents](/docs/use-cases/ai/edge-ai): Spice.ai deploys AI applications and agents across cloud and edge for low-latency decisions in security IoT use cases. - [Federated MCP Client for Distributed Tool Ecosystems](/docs/use-cases/ai/federated-mcp-server): Spice.ai federates external MCP servers for scalable, tool-driven AI applications in security, improving threat analysis. - [Multi-Tenant AI Agents](/docs/use-cases/ai/multi-tenant-agents): Deploy AI agents across many SaaS tenants with strict isolation and no per-tenant ETL pipelines. - [Object-Store Based SQL Query, Search, and LLM Inference Engine](/docs/use-cases/ai/object-store-ai-engine): Spice.ai enables SQL queries, hybrid search, and LLM inference on object-store data for security applications, delivering real-time insights. - [Real-Time Decision-Making for Intelligent Applications](/docs/use-cases/ai/real-time-decision-making): Spice.ai powers instant, context-aware decisions for applications like security recommendations by grounding AI in federated, low-latency datasets. - [Tool-Augmented AI with Model Context Protocol Server](/docs/use-cases/ai/tool-calling-ai): Spice.ai extends AI with custom tools via MCP server in finserv, integrating domain-specific APIs for enhanced functionality. - [Caching](/docs/use-cases/caching): Caching use cases with Spice.ai, including write-through, read-through, SQL, S3, and HTTP caching. - [HTTP Cache](/docs/use-cases/caching/http-cache): Use Spice.ai to cache HTTP API responses locally, reducing API call frequency and providing fast, SQL-queryable access to external API data. - [Read-Through Cache](/docs/use-cases/caching/read-through-cache): Use Spice.ai as a read-through cache with the SQL results cache for federated data sources and HTTP APIs. - [S3 Cache](/docs/use-cases/caching/s3-cache): Use Spice.ai to cache S3 and object store data locally for fast, repeatable SQL queries without re-reading from remote storage. - [SQL and Database Cache](/docs/use-cases/caching/sql-database-cache): Use Spice.ai to cache SQL database tables and query results locally for low-latency access and reduced load on upstream databases. - [Write-Through Cache](/docs/use-cases/caching/write-through-cache): Use Spice.ai as a write-through cache that writes data to both a local accelerator and the upstream data source. - [Data Federation, Acceleration, and SQL Query](/docs/use-cases/data): Data Federation, Acceleration, and SQL Query - [Application Resilience and Performance Optimization](/docs/use-cases/data/application-resilence-and-acceleration): Spice.ai colocates dynamic data with SaaS applications as a database CDN, ensuring resilience and high performance. - [Data Mesh for Unified Data Access](/docs/use-cases/data/data-mesh): Give domain teams decentralized, real-time data ownership and access across disparate sources through a unified SQL interface. - [Database CDN for Enhanced Performance](/docs/use-cases/data/database-cdn): Spice.ai acts as a database CDN for SaaS applications, caching dynamic data to ensure high performance and resilience. - [ETL-free Workflows and Data Migrations](/docs/use-cases/data/etl-free-workflows): Federate legacy and modern data systems without ETL for faster migrations, lower overhead, and zero application downtime. - [Object-Store Data Engine](/docs/use-cases/data/object-store-data-engine): Spice.ai federates, accelerates, and queries object-store data for finserv applications, enabling real-time data access without centralized warehouses. - [Reverse-ETL for Operational Workflows](/docs/use-cases/data/reverse-etl): Serve enriched data from warehouses and data lakes to operational systems and applications, eliminating complex pipelines. - [Spice for Retrieval-Augmented-Generation (RAG)](/docs/use-cases/rag): Use Spice for Retrieval-Augmented-Generation (RAG) - [RAG for Contextual Applications](/docs/use-cases/rag/applications): Build context-rich AI applications using Spice for Retrieval-Augmented Generation (RAG). - [Retrieval-Augmented Generation for AI-Powered Reporting](/docs/use-cases/rag/reporting): Spice.ai generates dynamic, context-aware AI-driven reports for operational insights in health-tech, ensuring compliance and precision. - [Search & Retrieval](/docs/use-cases/search): Search & Retrieval - [Simplifying Real-Time Data Collection and Search](/docs/use-cases/search/data-collection-and-search): Spice.ai processes streaming and static data with integrated search for real-time insights, focusing on application logic. - [Enterprise Search and Retrieval](/docs/use-cases/search/enterprise-search): Spice.ai powers semantic and precise search for finserv knowledge bases with hybrid vector and keyword capabilities. - [Object-Store Native Search Engine](/docs/use-cases/search/object-store-search-engine): Spice.ai powers a cloud-native embedded search engine on object-store data for security applications, enabling semantic and precise search. ## Optional - [GitHub Repository](https://github.com/spiceai/spiceai): Spice.ai OSS source code and issue tracker - [Cookbook](https://github.com/spiceai/cookbook): Ready-to-use recipes and examples for Spice.ai - [Spice Cloud Platform](https://spice.ai): Managed cloud platform for Spice.ai