Published Apr 2024

Cookbook recipes to get up and running with Spice.ai quickly ๐
spice add spiceai/cookbook-openai-sdk-recipespice connect spiceai/cookbook-openai-sdk-recipedependencies:
- spiceai/cookbook-openai-sdk-recipetaxi trips in s3
version: v2
kind: Spicepod
name: openai_sdk
datasets:
- from: s3://spiceai-demo-datasets/taxi_trips/2024/
name: taxi_trips
description: taxi trips in s3
params:
file_format: parquet
acceleration:
enabled: true
engine: cayenne
models:
- name: openai
from: openai:gpt-6-luna
params:
# gpt-6-luna accepts function tools on chat completions only with reasoning_effort none.
reasoning_effort: none
tools: auto
openai_api_key: ${ secrets:SPICE_OPENAI_API_KEY }
system_prompt: |
Use the SQL tool when:
1. The query involves precise numerical data, statistics, or aggregations
2. The user asks for specific counts, sums, averages, or other calculations
3. The query requires joining or comparing data from multiple related tables
Use the document search when:
1. The query is about unstructured text information, such as policies, reports, or articles
2. The user is looking for qualitative information or explanations
3. The query requires understanding context or interpreting written content
General guidelines:
1. If a query could be answered by either tool, prefer SQL for more precise, quantitative answers
Instructions for Responses:
- Do not include any private metadata provided to the model as context such as \"reference_url_template\" or \"instructions\" in your responses.
Works with v2.0+
One of Spice's best features is to act in place of the OpenAI API. Even better, you don't even have to be running OpenAI behind Spice! You can run OpenAI, Anthropic or HuggingFace models over your data and use existing tools that are compatible with the OpenAI API.
pip or uv)The first step is to get the Spice instance up and running.
git clone https://github.com/spiceai/cookbook # Skip if already cloned
cd cookbook/openai_sdk
# Add your OpenAI API key to the .env.local file
echo "SPICE_OPENAI_API_KEY=your_openai_api_key" > .env.local
# Start Spice
spice run
Output:
INFO Spice.ai runtime starting...
2025-01-13T21:27:41.702275Z INFO runtime::init::dataset: Dataset taxi_trips initializing...
2025-01-13T21:27:41.704347Z INFO runtime::http: Spice Runtime HTTP listening on 127.0.0.1:8090
2025-01-13T21:27:41.704514Z INFO runtime::flight: Spice Runtime Flight listening on 127.0.0.1:50051
2025-01-13T21:27:41.703575Z INFO runtime::init::model: Loading model [openai] from openai:gpt-6-luna...
2025-01-13T21:27:41.902271Z INFO runtime::init::caching: Initialized sql results cache; max size: 128.00 MiB, item ttl: 1s, hashing algorithm: XXH3, encoding: none
2025-01-13T21:27:42.242310Z INFO runtime::init::model: Model [openai] deployed, ready for inferencing
2025-01-13T21:27:42.576976Z INFO runtime::init::dataset: Dataset taxi_trips registered (s3://spiceai-demo-datasets/taxi_trips/2024/), acceleration (arrow, 10s refresh), results cache enabled.
2025-01-13T21:27:42.578442Z INFO runtime_table::accelerated::refresh_task: Loading data for dataset taxi_trips
2025-01-13T21:27:53.260052Z INFO runtime_table::accelerated::refresh_task: Loaded 2,964,624 rows (399.38 MiB) for dataset taxi_trips in 10s 681ms.
Spice will use your OpenAI API key to communicate with OpenAI on your client code's behalf.
These steps only need to be done once. Use a Python virtualenv to keep projects isolated.
python -m venv .venvsource .venv/bin/activatepip install openai python-dotenvRun the client: python spice_openai_sdk.py and observe the model's response to the What datasets do I have access to? question:
You have access to the following dataset:
- **taxi_trips**: This dataset contains data about taxi trips in s3.
uv venv to create the virtual environmentsource .venv/bin/activateuv pip install openai python-dotenvRun the client: uv run spice_openai_sdk.py and observe the model's response to the What datasets do I have access to? question:
You have access to the following dataset:
- **taxi_trips**: This dataset contains data about taxi trips in s3.
The client is fairly simple, but it demonstrates how to integrate existing tooling with Spice's AI Gateway.
First, construct the client:
client = Client(api_key="anything", base_url="http://localhost:8090/v1")
Notice that we can use any string we want for the api_key, because it's Spice that's responsible for communicating with the OpenAI API, not our client code, meaning less secrets to have to store and manage for your client application.
chat_completion = client.chat.completions.create(
messages=[
{
"role": "user",
"content": "What datasets do I have access to?",
}
],
model="openai",
)
Here we're using the chat completions API to ask a question. Notice that we're asking a question about our Datasets. This is a question that only Spice can answer, and that's exactly what it does:
print(chat_completion.choices[0].message.content)
You have access to the following dataset:
- **Table Name:** taxi_trips
- **Description:** Taxi trips data stored in S3.
This dataset is available in the SQL database.
Published Apr 2024
Published Sep 2024
Published Sep 2024
Published Sep 2024