Venice improvements

This commit is contained in:
Aine
2026-06-22 17:18:08 +01:00
parent 2ae641109f
commit 9e5f6de965
12 changed files with 668 additions and 34 deletions

View File

@@ -101,6 +101,8 @@ Prompts may contain the following **placeholder variables** which will be replac
💡 `{{ baibot_now_utc }}` changes as time goes on, which prevents [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) from working. It's better to use `{{ baibot_conversation_start_time_utc }}` in prompts, as its value doesn't change yet still orients the bot to the current date/time.
💡 On the [Venice provider](../providers.md#venice), baibot derives the prompt-cache key from the system prompt and the conversation start time, both stable for the life of a conversation, and ships `prompt_cache_retention: 24h` by default. A stable system prompt then stays cached across the whole conversation instead of being reprocessed (and re-billed) on every turn.
Here's a prompt that combines some of the above variables:
> You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."

View File

@@ -26,7 +26,7 @@ Text Generation is the bot's ability to **respond to users' messages with text**
![Screenshot of Text Generation - a user sends a message and the bot replies in a new conversation thread](./screenshots/text-generation.webp)
Some models also support vision and document understanding, so you may be able to mix text, images, and files (PDFs, text documents, etc.) in the same conversation. Note that certain providers may not support all file types or may have issues with specific files (e.g. scanned/image-based PDFs). If a file is rejected by the provider, the conversation thread may become unusable — start a new thread to work around this.
Some models also support vision and document understanding, so you may be able to mix text, images, and files (PDFs, text documents, etc.) in the same conversation. File inputs (documents such as PDFs) are currently accepted only by the OpenAI and Venice providers; the others skip them. Note that certain providers may not support all file types or may have issues with specific files (e.g. scanned/image-based PDFs). If a file is rejected by the provider, the conversation thread may become unusable — start a new thread to work around this.
In multi-user (group) rooms, to avoid disturbing the normal conversation between people, the bot is auto-configured to only respond to messages starting with the command prefix (`!bai`) or direct mentions via the [💬 Text Generation / 🗟 Prefix Requirement Type](./configuration/text-generation.md#-prefix-requirement-type) setting.

View File

@@ -180,7 +180,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- 🆔 Identifier: `venice`
- 🔗 Links: [🏠 Home page](https://venice.ai), [👤 Sign up](https://venice.ai), [📋 Models list](https://api.venice.ai/api/v1/models)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation) (incl. editing, via the native knob-rich `/image/generate` and `/image/edit` endpoints), [💬 text-generation](./features.md#-text-generation) (incl. vision; native web search via the `venice_parameters` config), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation) (incl. editing, via the native knob-rich `/image/generate` and `/image/edit` endpoints), [💬 text-generation](./features.md#-text-generation) (incl. vision, file inputs like PDF and DOCX, and prompt caching; native web search via the `venice_parameters` config), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local venice my-venice-agent`
- create a global agent: `!bai agent create-global venice my-venice-agent`
@@ -193,7 +193,19 @@ Unlike the [OpenAI Compatible](#openai-compatible) provider (which can talk to V
Every parameter below is optional unless marked otherwise. Omitting a knob lets Venice apply its own server-side default; this is **not** the same as setting it to `false`, which actively sends `false`.
**`text_generation.venice_parameters`** — Venice-specific request knobs sent in the `venice_parameters` bag (alongside the standard `model_id`, `prompt`, `temperature`, `max_response_tokens`, and `max_context_tokens` fields). Set any of them to override Venice's behavior. The `Default` column shows the value baibot's sample config ships; a `—` means the knob is left unset, so Venice's own default applies.
**`text_generation`** (top-level knobs) — sampling, caching, and reasoning controls that sit directly on `text_generation`, next to `model_id`, `prompt`, `temperature`, `max_response_tokens`, and `max_context_tokens`. They map to top-level fields on Venice's request, separate from the `venice_parameters` bag below.
| Knob | What it does | Default |
|------|--------------|---------|
| `top_p` | Nucleus sampling, `0.0`–`1.0`. An alternative to `temperature`. | — |
| `frequency_penalty` | Penalize tokens by how often they have already appeared, `-2.0`–`2.0`. | — |
| `presence_penalty` | Penalize tokens that have appeared at all, `-2.0`–`2.0`. | — |
| `repetition_penalty` | Penalize repetition. Values above `1.0` discourage repeats. | — |
| `reasoning_effort` | Reasoning budget for models that support it: `low`, `medium`, `high`. | — |
| `prompt_cache_retention` | How long Venice keeps the prompt prefix cached: `default`, `extended`, or `24h`. `24h` is the lever that makes a long, stable system prompt cheap across a day of conversations. | `24h` |
| `show_reasoning` | Append the model's reasoning (its `reasoning_content`) below the answer. Reads a field separate from the answer text, so it works regardless of `strip_thinking_response`. | `false` |
**`text_generation.venice_parameters`** — Venice-specific request knobs sent in the `venice_parameters` bag. Set any of them to override Venice's behavior. The `Default` column shows the value baibot's sample config ships; a `—` means the knob is left unset, so Venice's own default applies.
| Knob | What it does | Default |
|------|--------------|---------|
@@ -208,6 +220,7 @@ Every parameter below is optional unless marked otherwise. Omitting a knob lets
| `strip_thinking_response` | Strip `<think></think>` blocks from reasoning models so the user sees only the answer. | `true` |
| `disable_thinking` | Disable the model's reasoning step entirely. | — |
| `enable_e2ee` | Run in end-to-end-encrypted mode rather than the default TEE-only mode. | `false` |
| `verbosity` | Response verbosity for models that support it: `low`, `medium`, `high`. | — |
**`text_to_speech`**:

View File

@@ -6,6 +6,23 @@ text_generation:
temperature: 1.0
max_response_tokens: 4096
max_context_tokens: 128000
# Prompt caching: how long Venice keeps the prompt prefix cached. "default", "extended", or "24h".
# "24h" (shipped by default) makes a long, stable system prompt cheap across a day of conversations.
prompt_cache_retention: 24h
# Top-level sampling and reasoning knobs (uncomment to override Venice's default):
# Nucleus sampling, 0.0-1.0 (an alternative to temperature).
# top_p: 0.9
# Penalize tokens by how often they have already appeared, -2.0-2.0.
# frequency_penalty: 0.0
# Penalize tokens that have appeared at all, -2.0-2.0.
# presence_penalty: 0.0
# Penalize repetition; values above 1.0 discourage repeats.
# repetition_penalty: 1.0
# Reasoning budget for models that support it: low, medium, high.
# reasoning_effort: medium
# Append the model's reasoning below the answer. Reads a field separate from the answer text, so
# it works alongside strip_thinking_response (which only strips <think> blocks from the answer).
# show_reasoning: true
# Venice-specific request parameters. Only the keys present below are sent to Venice; omit a
# key to fall back to Venice's own default. Omitting a knob is NOT the same as setting it to
# `false` — `false` actively sends `false`.
@@ -24,6 +41,8 @@ text_generation:
# return_search_results_as_documents: true
# enable_x_search: true
# disable_thinking: true
# Response verbosity for models that support it: low, medium, high.
# verbosity: medium
# character_slug: public-character-id
speech_to_text:
model_id: nvidia/parakeet-tdt-0.6b-v3