Published Apr 2024

Cookbook recipes to get up and running with Spice.ai quickly ๐
spice add spiceai/cookbook-cayenne-recipespice connect spiceai/cookbook-cayenne-recipedependencies:
- spiceai/cookbook-cayenne-recipetaxi trips in s3
version: v2
kind: Spicepod
name: cayenne-acceleration
datasets:
- from: s3://spiceai-demo-datasets/taxi_trips/2024/
name: taxi_trips
description: taxi trips in s3
params:
file_format: parquet
acceleration:
enabled: true
engine: cayenne
mode: file
Works with v2.0+
This recipe will walkthrough how to accelerate a local copy of the taxi trips dataset stored in S3 using Cayenne as the data accelerator engine.
Step 1. Initialize a new Spice app.
spice init cayenne-acceleration-qs
cd cayenne-acceleration-qs
Step 2. Configure s3 dataset: copy and paste the YAML below to spicepod.yaml in the Spice app.
version: v2
kind: Spicepod
name: cayenne-acceleration-qs
datasets:
- from: s3://spiceai-demo-datasets/taxi_trips/2024/
name: taxi_trips
description: taxi trips in s3
params:
file_format: parquet
Step 3. Start the Spice runtime.
spice run
Confirm in the terminal output the taxi_trips dataset has been loaded:
Spice.ai runtime starting...
2026-08-24T20:57:43.603253Z INFO spiced: Starting runtime v2.1.5+models.metal
2026-08-24T20:57:43.687793Z INFO runtime::init::caching: Initialized sql results cache; max size: 128.00 MiB, item ttl: 1s, hashing algorithm: XXH3, encoding: none
2026-08-24T20:57:43.968078Z INFO runtime::flight: Spice Runtime Flight listening on 127.0.0.1:50051
2026-08-24T20:57:43.972502Z INFO runtime::http: Spice Runtime HTTP listening on 127.0.0.1:8090
2026-08-24T20:57:49.910117Z INFO runtime::init::dataset: Dataset taxi_trips registered (s3://spiceai-demo-datasets/taxi_trips/2024/), results cache enabled. duration_ms=18
2026-08-24T20:57:50.043569Z INFO runtime: All components are loaded. Spice runtime is ready!
Step 4. Run queries against the dataset using the Spice SQL REPL.
In a new terminal, start the Spice SQL REPL
spice sql
Query the taxi_trips dataset, observing the long query time.
select "VendorID", tpep_pickup_datetime, tpep_dropoff_datetime, passenger_count from taxi_trips limit 10;
+----------+----------------------+-----------------------+-----------------+
| VendorID | tpep_pickup_datetime | tpep_dropoff_datetime | passenger_count |
| int32 | timestamp[us] | timestamp[us] | int64 |
+----------+----------------------+-----------------------+-----------------+
| 1 | 2024-01-02T18:04:31 | 2024-01-02T18:11:49 | 0 |
| 1 | 2024-01-02T18:13:28 | 2024-01-02T18:34:45 | 0 |
| 1 | 2024-01-02T18:52:21 | 2024-01-02T18:57:43 | 0 |
| 1 | 2024-01-02T18:37:05 | 2024-01-02T18:51:38 | 0 |
| 1 | 2024-01-02T18:46:54 | 2024-01-02T18:53:18 | 0 |
| 1 | 2024-01-02T18:18:22 | 2024-01-02T18:24:30 | 0 |
| 1 | 2024-01-02T18:06:50 | 2024-01-02T18:25:04 | 0 |
| 1 | 2024-01-02T18:45:31 | 2024-01-02T18:58:08 | 0 |
| 1 | 2024-01-02T18:33:29 | 2024-01-02T18:42:09 | 0 |
| 1 | 2024-01-02T18:05:25 | 2024-01-02T18:17:25 | 0 |
+----------+----------------------+-----------------------+-----------------+
Time: 1.256426375 seconds. 10 rows.
Step 5. Update the spicepod.yaml to enable Cayenne acceleration.
version: v2
kind: Spicepod
name: cayenne-acceleration-qs
datasets:
- from: s3://spiceai-demo-datasets/taxi_trips/2024/
name: taxi_trips
description: taxi trips in s3
params:
file_format: parquet
acceleration:
enabled: true
engine: cayenne
mode: file
Step 6. Restart the Spice app and observe the dataset loading and accelerating.
Spice.ai runtime starting...
2026-08-24T20:58:40.483388Z INFO spiced: Starting runtime v2.1.5+models.metal
2026-08-24T20:58:40.631669Z INFO runtime::init::caching: Initialized sql results cache; max size: 128.00 MiB, item ttl: 1s, hashing algorithm: XXH3, encoding: none
2026-08-24T20:58:40.897943Z INFO runtime::flight: Spice Runtime Flight listening on 127.0.0.1:50051
2026-08-24T20:58:40.905264Z INFO runtime::http: Spice Runtime HTTP listening on 127.0.0.1:8090
2026-08-24T20:58:43.033807Z INFO runtime::init::dataset: Dataset taxi_trips registered (s3://spiceai-demo-datasets/taxi_trips/2024/), acceleration (cayenne:file), results cache enabled. duration_ms=477
2026-08-24T20:58:43.035086Z INFO runtime_table::accelerated::refresh_task: Loading data for dataset taxi_trips
2026-08-24T20:58:53.353320Z INFO runtime_table::accelerated::refresh_task: Loaded 2,964,624 rows (399.38 MiB) for dataset taxi_trips in 10s 302ms.
2026-08-24T20:58:53.400890Z INFO runtime: All components are loaded. Spice runtime is ready!
Step 7. Run a query against the taxi_trips dataset again, observing the fast query time.
select "VendorID", tpep_pickup_datetime, tpep_dropoff_datetime, passenger_count from taxi_trips limit 10;
+----------+----------------------+-----------------------+-----------------+
| VendorID | tpep_pickup_datetime | tpep_dropoff_datetime | passenger_count |
| int32 | timestamp[us] | timestamp[us] | int64 |
+----------+----------------------+-----------------------+-----------------+
| 1 | 2024-01-04T18:01:18 | 2024-01-04T18:11:46 | 1 |
| 1 | 2024-01-04T18:13:06 | 2024-01-04T18:27:17 | 1 |
| 1 | 2024-01-04T18:29:48 | 2024-01-04T18:52:12 | 1 |
| 2 | 2024-01-04T18:19:18 | 2024-01-04T18:49:18 | 1 |
| 2 | 2024-01-04T18:52:38 | 2024-01-04T19:12:11 | 1 |
| 2 | 2024-01-04T18:29:18 | 2024-01-04T18:34:46 | 1 |
| 2 | 2024-01-04T18:36:24 | 2024-01-04T18:55:09 | 1 |
| 2 | 2024-01-04T18:26:17 | 2024-01-04T18:35:32 | 1 |
| 2 | 2024-01-04T18:51:48 | 2024-01-04T19:05:35 | 1 |
| 2 | 2024-01-04T18:06:09 | 2024-01-04T18:30:35 | 1 |
+----------+----------------------+-----------------------+-----------------+
Time: 0.064644125 seconds. 10 rows.
Cayenne Data Accelerator Documentation (opens in a new tab).
For using spice sql, see the CLI reference (opens in a new tab).
See the datasets reference (opens in a new tab) for additional dataset configuration options.
Published Apr 2024
Published Sep 2024
Published Sep 2024
Published Sep 2024