Compare commits
10 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
982ebc6657 | ||
|
|
4a75be742f | ||
|
|
aae3c4c08f | ||
|
|
08cf50885d | ||
|
|
1eebb1a55b | ||
|
|
2888cb9450 | ||
|
|
5de497482d | ||
|
|
9e5f6de965 | ||
|
|
2ae641109f | ||
|
|
105ca7b506 |
22
CHANGELOG.md
22
CHANGELOG.md
@@ -1,3 +1,25 @@
|
||||
# (2026-06-24) Version 1.23.1
|
||||
|
||||
- (**Bugfix**) The [Venice](https://venice.ai) provider now auto-recovers when a model rejects an optional knob it does not support. Venice's request body is strict (`additionalProperties: false`), so a model that lacks `prompt_cache_retention`, `reasoning_effort`, or `prompt_cache_key` rejected the whole request with a `400 Bad Request` — breaking agent creation and every reply. baibot now drops the unsupported field and retries, remembering the rejection per model so later requests skip it without a wasted round-trip. Only these meaning-preserving fields are dropped; sampling knobs that change the output (`temperature`, `top_p`, the penalties) are never silently removed and still surface as an error.
|
||||
|
||||
- (**Improvement**) When Venice rejects a request with a `400 Bad Request`, baibot now surfaces Venice's actual error message (e.g. `Extra inputs are not permitted, field: 'prompt_cache_retention'`) instead of a generic "configuration does not result in a working agent". This makes agent-creation failures self-explanatory. Other error statuses keep their bodies redacted, since those can carry account or rate-limit details.
|
||||
|
||||
|
||||
# (2026-06-23) Version 1.23.0
|
||||
|
||||
- (**Feature**) The [Venice](https://venice.ai) provider now accepts file inputs (PDF, DOCX, and other documents, up to 25MB), the same way it already handled images. This makes Venice the second provider after OpenAI to accept files; the others (Anthropic and the OpenAI-compatible providers) skip them. See the [text-generation feature docs](./docs/features.md#-text-generation).
|
||||
|
||||
- (**Feature**) Add prompt caching to the Venice provider, on by default (`prompt_cache_retention: 24h`). baibot derives the cache key from the system prompt and the conversation start time (both fixed for the life of a conversation), so a long, stable system prompt stays cached across the day instead of being reprocessed and re-billed on every turn. See [Text Generation / Prompt Override](./docs/configuration/text-generation.md#️-prompt-override).
|
||||
|
||||
- (**Feature**) Wire up the rest of Venice's sampling and reasoning controls: top-level `top_p`, `frequency_penalty`, `presence_penalty`, `repetition_penalty`, and `reasoning_effort`; `verbosity` in the `venice_parameters` bag; and a `show_reasoning` toggle that appends the model's reasoning to the reply as a collapsible, folded-by-default `💭 Reasoning` block (off by default). See the [Venice configuration reference](./docs/providers.md#venice).
|
||||
|
||||
- (**Feature**) Render Venice web-search citations as readable `[n]` references with a `Sources:` list of links, instead of leaving Venice's raw `^n^` superscripts in the reply.
|
||||
|
||||
- (**Security**) Escape citation titles and validate citation URLs before rendering them, and drop user-supplied filenames from error messages, so a hostile web page or a crafted filename cannot inject a spoofed link into the bot's reply.
|
||||
|
||||
- (**Bugfix**) The [OpenAI-compatible](./docs/providers.md#openai-compatible) provider now trusts the system CA store (honoring `SSL_CERT_FILE`), so endpoints served behind a private/internal CA (FreeIPA, organization PKI) no longer fail the TLS handshake with `invalid peer certificate: UnknownIssuer`. Fixed upstream in `etke_openai_api_rust` 0.1.10. Thanks to [@shaba](https://github.com/shaba) for the report in [#188](https://github.com/etkecc/baibot/pull/188).
|
||||
|
||||
|
||||
# (2026-06-21) Version 1.22.0
|
||||
|
||||
- (**Feature**) Add a native [Venice](https://venice.ai) provider with [🖌️ image-generation](./docs/features.md#️-image-creation) (incl. editing), [💬 text-generation](./docs/features.md#-text-generation) (incl. vision), [🗣️ text-to-speech](./docs/features.md#️-text-to-speech), [🦻 speech-to-text](./docs/features.md#-speech-to-text), and Venice's native web search via the full `venice_parameters` knob set. Unlike the [OpenAI-compatible](./docs/providers.md#openai-compatible) path (which drops images and can't reach Venice's audio or native image endpoints), it talks to Venice's API directly, using the knob-rich native `/image/generate` and `/image/edit` endpoints. See the [Venice provider docs](./docs/providers.md#venice).
|
||||
|
||||
91
Cargo.lock
generated
91
Cargo.lock
generated
@@ -196,7 +196,7 @@ dependencies = [
|
||||
"futures",
|
||||
"getrandom 0.3.4",
|
||||
"rand 0.9.4",
|
||||
"reqwest 0.13.3",
|
||||
"reqwest 0.13.4",
|
||||
"secrecy",
|
||||
"serde",
|
||||
"serde_json",
|
||||
@@ -315,7 +315,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "baibot"
|
||||
version = "1.22.0"
|
||||
version = "1.23.1"
|
||||
dependencies = [
|
||||
"anthropic",
|
||||
"anyhow",
|
||||
@@ -329,7 +329,7 @@ dependencies = [
|
||||
"mxlink",
|
||||
"quick_cache",
|
||||
"regex",
|
||||
"reqwest 0.12.28",
|
||||
"reqwest 0.13.4",
|
||||
"serde",
|
||||
"serde_json",
|
||||
"serde_yaml_ng",
|
||||
@@ -644,6 +644,16 @@ version = "0.4.2"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "3d52eff69cd5e647efe296129160853a42795992097e8af39800e1060caeea9b"
|
||||
|
||||
[[package]]
|
||||
name = "core-foundation"
|
||||
version = "0.9.4"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "91e195e091a93c46f7102ec7818a2aa394e1e1771c3ab4825963fa03e45afb8f"
|
||||
dependencies = [
|
||||
"core-foundation-sys",
|
||||
"libc",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "core-foundation"
|
||||
version = "0.10.1"
|
||||
@@ -1078,9 +1088,9 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "etke_openai_api_rust"
|
||||
version = "0.1.9"
|
||||
version = "0.1.10"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "c1499ac61bf9b9a5973a5830894cbf2016a96a4dab9eb2dab4a526c720632c93"
|
||||
checksum = "e9600ff550c189fe1bcc291def9057e17147ea9413c201d6747da32e1851865d"
|
||||
dependencies = [
|
||||
"log",
|
||||
"mime",
|
||||
@@ -1587,11 +1597,10 @@ dependencies = [
|
||||
"hyper",
|
||||
"hyper-util",
|
||||
"rustls",
|
||||
"rustls-native-certs",
|
||||
"rustls-native-certs 0.8.3",
|
||||
"tokio",
|
||||
"tokio-rustls",
|
||||
"tower-service",
|
||||
"webpki-roots 1.0.7",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
@@ -2173,10 +2182,10 @@ dependencies = [
|
||||
"oauth2-reqwest",
|
||||
"percent-encoding",
|
||||
"pin-project-lite",
|
||||
"reqwest 0.13.3",
|
||||
"reqwest 0.13.4",
|
||||
"ruma",
|
||||
"rustls",
|
||||
"rustls-native-certs",
|
||||
"rustls-native-certs 0.8.3",
|
||||
"rustls-pki-types",
|
||||
"serde",
|
||||
"serde_html_form",
|
||||
@@ -2568,7 +2577,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "234fb5c965bbce983ee5de636a7a51d6a3223da8067ea02f9ab2d2d78ac08be2"
|
||||
dependencies = [
|
||||
"oauth2",
|
||||
"reqwest 0.13.3",
|
||||
"reqwest 0.13.4",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
@@ -2583,6 +2592,12 @@ version = "0.3.1"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "c08d65885ee38876c4f86fa503fb49d7b507c2b62552df7c70b2fce627e06381"
|
||||
|
||||
[[package]]
|
||||
name = "openssl-probe"
|
||||
version = "0.1.6"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "d05e27ee213611ffe7d6348b942e8f942b37114c00cc03cec254295a4a17852e"
|
||||
|
||||
[[package]]
|
||||
name = "openssl-probe"
|
||||
version = "0.2.1"
|
||||
@@ -3076,12 +3091,11 @@ dependencies = [
|
||||
"hyper-util",
|
||||
"js-sys",
|
||||
"log",
|
||||
"mime_guess",
|
||||
"percent-encoding",
|
||||
"pin-project-lite",
|
||||
"quinn",
|
||||
"rustls",
|
||||
"rustls-native-certs",
|
||||
"rustls-native-certs 0.8.3",
|
||||
"rustls-pki-types",
|
||||
"serde",
|
||||
"serde_json",
|
||||
@@ -3098,14 +3112,13 @@ dependencies = [
|
||||
"wasm-bindgen-futures",
|
||||
"wasm-streams 0.4.2",
|
||||
"web-sys",
|
||||
"webpki-roots 1.0.7",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "reqwest"
|
||||
version = "0.13.3"
|
||||
version = "0.13.4"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "62e0021ea2c22aed41653bc7e1419abb2c97e038ff2c33d0e1309e49a97deec0"
|
||||
checksum = "219c5811de6525e5416c7d5d53bb656d3afdbc6c5af816e0802bcfa42dbdc1c3"
|
||||
dependencies = [
|
||||
"base64 0.22.1",
|
||||
"bytes",
|
||||
@@ -3396,16 +3409,38 @@ dependencies = [
|
||||
"zeroize",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "rustls-native-certs"
|
||||
version = "0.7.3"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "e5bfb394eeed242e909609f56089eecfe5fda225042e8b171791b9c95f5931e5"
|
||||
dependencies = [
|
||||
"openssl-probe 0.1.6",
|
||||
"rustls-pemfile",
|
||||
"rustls-pki-types",
|
||||
"schannel",
|
||||
"security-framework 2.11.1",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "rustls-native-certs"
|
||||
version = "0.8.3"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "612460d5f7bea540c490b2b6395d8e34a953e52b491accd6c86c8164c5932a63"
|
||||
dependencies = [
|
||||
"openssl-probe",
|
||||
"openssl-probe 0.2.1",
|
||||
"rustls-pki-types",
|
||||
"schannel",
|
||||
"security-framework",
|
||||
"security-framework 3.7.0",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "rustls-pemfile"
|
||||
version = "2.2.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "dce314e5fee3f39953d46bb63bb8a46d40c2f8fb7cc5a3b6cab2bde9721d6e50"
|
||||
dependencies = [
|
||||
"rustls-pki-types",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
@@ -3424,16 +3459,16 @@ version = "0.7.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "26d1e2536ce4f35f4846aa13bff16bd0ff40157cdb14cc056c7b14ba41233ba0"
|
||||
dependencies = [
|
||||
"core-foundation",
|
||||
"core-foundation 0.10.1",
|
||||
"core-foundation-sys",
|
||||
"jni",
|
||||
"log",
|
||||
"once_cell",
|
||||
"rustls",
|
||||
"rustls-native-certs",
|
||||
"rustls-native-certs 0.8.3",
|
||||
"rustls-platform-verifier-android",
|
||||
"rustls-webpki",
|
||||
"security-framework",
|
||||
"security-framework 3.7.0",
|
||||
"security-framework-sys",
|
||||
"webpki-root-certs",
|
||||
"windows-sys 0.61.2",
|
||||
@@ -3514,6 +3549,19 @@ dependencies = [
|
||||
"zeroize",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "security-framework"
|
||||
version = "2.11.1"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "897b2245f0b511c87893af39b033e5ca9cce68824c4d7e7630b5a1d339658d02"
|
||||
dependencies = [
|
||||
"bitflags 2.11.1",
|
||||
"core-foundation 0.9.4",
|
||||
"core-foundation-sys",
|
||||
"libc",
|
||||
"security-framework-sys",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "security-framework"
|
||||
version = "3.7.0"
|
||||
@@ -3521,7 +3569,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "b7f4bc775c73d9a02cde8bf7b2ec4c9d12743edf609006c7facc23998404cd1d"
|
||||
dependencies = [
|
||||
"bitflags 2.11.1",
|
||||
"core-foundation",
|
||||
"core-foundation 0.10.1",
|
||||
"core-foundation-sys",
|
||||
"libc",
|
||||
"security-framework-sys",
|
||||
@@ -4301,6 +4349,7 @@ dependencies = [
|
||||
"log",
|
||||
"once_cell",
|
||||
"rustls",
|
||||
"rustls-native-certs 0.7.3",
|
||||
"rustls-pki-types",
|
||||
"serde",
|
||||
"serde_json",
|
||||
|
||||
@@ -7,7 +7,7 @@ license = "AGPL-3.0-or-later"
|
||||
readme = "README.md"
|
||||
keywords = ["matrix", "chat", "bot", "AI", "LLM"]
|
||||
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
|
||||
version = "1.22.0"
|
||||
version = "1.23.1"
|
||||
edition = "2024"
|
||||
|
||||
[lib]
|
||||
@@ -28,10 +28,9 @@ mxlink = ">=1.15.0"
|
||||
etke_openai_api_rust = "0.1.*"
|
||||
quick_cache = "0.6.*"
|
||||
regex = "1.12.*"
|
||||
# Direct dep for the native `venice` provider's HTTP client. Pinned to 0.12 (the version
|
||||
# async-openai 0.41 already resolves) with rustls only and default-features off, so we ride
|
||||
# the existing reqwest+rustls copy instead of pulling a second TLS stack (native-tls/openssl).
|
||||
reqwest = { version = "0.12.*", default-features = false, features = ["json", "multipart", "rustls-tls"] }
|
||||
# HTTP client for the native `venice` provider. rustls only (no extra TLS stack), matching the
|
||||
# reqwest copy async-openai/matrix-sdk/mxlink already use.
|
||||
reqwest = { version = "0.13.*", default-features = false, features = ["json", "multipart", "rustls"] }
|
||||
serde = { version = "1.0.*", features = ["derive"], default-features = false }
|
||||
serde_json = "1.0.*"
|
||||
serde_yaml_ng = "0.10.*"
|
||||
|
||||
@@ -101,6 +101,8 @@ Prompts may contain the following **placeholder variables** which will be replac
|
||||
|
||||
💡 `{{ baibot_now_utc }}` changes as time goes on, which prevents [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) from working. It's better to use `{{ baibot_conversation_start_time_utc }}` in prompts, as its value doesn't change yet still orients the bot to the current date/time.
|
||||
|
||||
💡 On the [Venice provider](../providers.md#venice), baibot derives the prompt-cache key from the system prompt and the conversation start time, both stable for the life of a conversation, and ships `prompt_cache_retention: 24h` by default. A stable system prompt then stays cached across the whole conversation instead of being reprocessed (and re-billed) on every turn.
|
||||
|
||||
Here's a prompt that combines some of the above variables:
|
||||
|
||||
> You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||
|
||||
@@ -26,7 +26,7 @@ Text Generation is the bot's ability to **respond to users' messages with text**
|
||||
|
||||

|
||||
|
||||
Some models also support vision and document understanding, so you may be able to mix text, images, and files (PDFs, text documents, etc.) in the same conversation. Note that certain providers may not support all file types or may have issues with specific files (e.g. scanned/image-based PDFs). If a file is rejected by the provider, the conversation thread may become unusable — start a new thread to work around this.
|
||||
Some models also support vision and document understanding, so you may be able to mix text, images, and files (PDFs, text documents, etc.) in the same conversation. File inputs (documents such as PDFs) are currently accepted only by the OpenAI and Venice providers; the others skip them. Note that certain providers may not support all file types or may have issues with specific files (e.g. scanned/image-based PDFs). If a file is rejected by the provider, the conversation thread may become unusable — start a new thread to work around this.
|
||||
|
||||
In multi-user (group) rooms, to avoid disturbing the normal conversation between people, the bot is auto-configured to only respond to messages starting with the command prefix (`!bai`) or direct mentions via the [💬 Text Generation / 🗟 Prefix Requirement Type](./configuration/text-generation.md#-prefix-requirement-type) setting.
|
||||
|
||||
|
||||
@@ -180,7 +180,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
|
||||
|
||||
- 🆔 Identifier: `venice`
|
||||
- 🔗 Links: [🏠 Home page](https://venice.ai), [👤 Sign up](https://venice.ai), [📋 Models list](https://api.venice.ai/api/v1/models)
|
||||
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation) (incl. editing, via the native knob-rich `/image/generate` and `/image/edit` endpoints), [💬 text-generation](./features.md#-text-generation) (incl. vision; native web search via the `venice_parameters` config), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
||||
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation) (incl. editing, via the native knob-rich `/image/generate` and `/image/edit` endpoints), [💬 text-generation](./features.md#-text-generation) (incl. vision, file inputs like PDF and DOCX, and prompt caching; native web search via the `venice_parameters` config), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
||||
- 🗲 Quick start:
|
||||
- create a room-local agent: `!bai agent create-room-local venice my-venice-agent`
|
||||
- create a global agent: `!bai agent create-global venice my-venice-agent`
|
||||
@@ -193,7 +193,19 @@ Unlike the [OpenAI Compatible](#openai-compatible) provider (which can talk to V
|
||||
|
||||
Every parameter below is optional unless marked otherwise. Omitting a knob lets Venice apply its own server-side default; this is **not** the same as setting it to `false`, which actively sends `false`.
|
||||
|
||||
**`text_generation.venice_parameters`** — Venice-specific request knobs sent in the `venice_parameters` bag (alongside the standard `model_id`, `prompt`, `temperature`, `max_response_tokens`, and `max_context_tokens` fields). Set any of them to override Venice's behavior. The `Default` column shows the value baibot's sample config ships; a `—` means the knob is left unset, so Venice's own default applies.
|
||||
**`text_generation`** (top-level knobs) — sampling, caching, and reasoning controls that sit directly on `text_generation`, next to `model_id`, `prompt`, `temperature`, `max_response_tokens`, and `max_context_tokens`. They map to top-level fields on Venice's request, separate from the `venice_parameters` bag below.
|
||||
|
||||
| Knob | What it does | Default |
|
||||
|------|--------------|---------|
|
||||
| `top_p` | Nucleus sampling, `0.0`–`1.0`. An alternative to `temperature`. | — |
|
||||
| `frequency_penalty` | Penalize tokens by how often they have already appeared, `-2.0`–`2.0`. | — |
|
||||
| `presence_penalty` | Penalize tokens that have appeared at all, `-2.0`–`2.0`. | — |
|
||||
| `repetition_penalty` | Penalize repetition. Values above `1.0` discourage repeats. | — |
|
||||
| `reasoning_effort` | Reasoning budget for models that support it: `low`, `medium`, `high`. | — |
|
||||
| `prompt_cache_retention` | How long Venice keeps the prompt prefix cached: `default`, `extended`, or `24h`. `24h` is the lever that makes a long, stable system prompt cheap across a day of conversations. | `24h` |
|
||||
| `show_reasoning` | Append the model's reasoning (its `reasoning_content`) below the answer, as a collapsible `💭 Reasoning` block that stays folded until clicked. Reads a field separate from the answer text, so it works regardless of `strip_thinking_response`. | `false` |
|
||||
|
||||
**`text_generation.venice_parameters`** — Venice-specific request knobs sent in the `venice_parameters` bag. Set any of them to override Venice's behavior. The `Default` column shows the value baibot's sample config ships; a `—` means the knob is left unset, so Venice's own default applies.
|
||||
|
||||
| Knob | What it does | Default |
|
||||
|------|--------------|---------|
|
||||
@@ -208,6 +220,7 @@ Every parameter below is optional unless marked otherwise. Omitting a knob lets
|
||||
| `strip_thinking_response` | Strip `<think></think>` blocks from reasoning models so the user sees only the answer. | `true` |
|
||||
| `disable_thinking` | Disable the model's reasoning step entirely. | — |
|
||||
| `enable_e2ee` | Run in end-to-end-encrypted mode rather than the default TEE-only mode. | `false` |
|
||||
| `verbosity` | Response verbosity for models that support it: `low`, `medium`, `high`. | — |
|
||||
|
||||
**`text_to_speech`**:
|
||||
|
||||
|
||||
@@ -6,6 +6,24 @@ text_generation:
|
||||
temperature: 1.0
|
||||
max_response_tokens: 4096
|
||||
max_context_tokens: 128000
|
||||
# Prompt caching: how long Venice keeps the prompt prefix cached. "default", "extended", or "24h".
|
||||
# "24h" (shipped by default) makes a long, stable system prompt cheap across a day of conversations.
|
||||
prompt_cache_retention: 24h
|
||||
# Top-level sampling and reasoning knobs (uncomment to override Venice's default):
|
||||
# Nucleus sampling, 0.0-1.0 (an alternative to temperature).
|
||||
# top_p: 0.9
|
||||
# Penalize tokens by how often they have already appeared, -2.0-2.0.
|
||||
# frequency_penalty: 0.0
|
||||
# Penalize tokens that have appeared at all, -2.0-2.0.
|
||||
# presence_penalty: 0.0
|
||||
# Penalize repetition; values above 1.0 discourage repeats.
|
||||
# repetition_penalty: 1.0
|
||||
# Reasoning budget for models that support it: low, medium, high.
|
||||
# reasoning_effort: medium
|
||||
# Append the model's reasoning below the answer as a collapsible "💭 Reasoning" block (folded by
|
||||
# default). Reads a field separate from the answer text, so it works alongside
|
||||
# strip_thinking_response (which only strips <think> blocks from the answer).
|
||||
# show_reasoning: true
|
||||
# Venice-specific request parameters. Only the keys present below are sent to Venice; omit a
|
||||
# key to fall back to Venice's own default. Omitting a knob is NOT the same as setting it to
|
||||
# `false` — `false` actively sends `false`.
|
||||
@@ -24,6 +42,8 @@ text_generation:
|
||||
# return_search_results_as_documents: true
|
||||
# enable_x_search: true
|
||||
# disable_thinking: true
|
||||
# Response verbosity for models that support it: low, medium, high.
|
||||
# verbosity: medium
|
||||
# character_slug: public-character-id
|
||||
speech_to_text:
|
||||
model_id: nvidia/parakeet-tdt-0.6b-v3
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
services:
|
||||
element-web:
|
||||
image: ghcr.io/element-hq/element-web:v1.12.21
|
||||
image: ghcr.io/element-hq/element-web:v1.12.22
|
||||
user: "${UID}:${GID}"
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
|
||||
@@ -184,9 +184,7 @@ impl ControllerTrait for ControllerType {
|
||||
ControllerType::Anthropic(controller) => {
|
||||
controller.generate_image(prompt, params).await
|
||||
}
|
||||
ControllerType::Venice(controller) => {
|
||||
controller.generate_image(prompt, params).await
|
||||
}
|
||||
ControllerType::Venice(controller) => controller.generate_image(prompt, params).await,
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -1,3 +1,9 @@
|
||||
use std::collections::hash_map::DefaultHasher;
|
||||
use std::hash::{Hash, Hasher};
|
||||
use std::sync::OnceLock;
|
||||
|
||||
use regex::Regex;
|
||||
|
||||
use crate::agent::AgentPurpose;
|
||||
use crate::agent::provider::entity::{TextGenerationParams, TextGenerationResult};
|
||||
use crate::conversation::llm::{
|
||||
@@ -6,13 +12,14 @@ use crate::conversation::llm::{
|
||||
};
|
||||
use crate::strings;
|
||||
|
||||
use super::config::Config;
|
||||
use super::config::{Config, WebSearchMode};
|
||||
use super::utils::convert_llm_messages_to_venice;
|
||||
use super::wire::{ChatCompletionRequest, ChatCompletionResponse};
|
||||
use super::wire::{ChatCompletionRequest, ChatCompletionResponse, WebSearchCitation};
|
||||
|
||||
pub async fn generate_text(
|
||||
config: &Config,
|
||||
http: &reqwest::Client,
|
||||
unsupported: &super::recovery::UnsupportedFieldsCache,
|
||||
conversation: LLMConversation,
|
||||
params: TextGenerationParams,
|
||||
) -> anyhow::Result<TextGenerationResult> {
|
||||
@@ -31,6 +38,17 @@ pub async fn generate_text(
|
||||
.trim(),
|
||||
);
|
||||
|
||||
// Prompt-cache routing key. Hash ONLY conversation-stable inputs: the rendered system prompt
|
||||
// and the conversation start time. Folding in anything per-turn (message content, the current
|
||||
// time, the message count) would mint a fresh key every turn, miss the cache every lookup, and
|
||||
// pay full price plus the hashing cost. The start time is rendered explicitly here so the key
|
||||
// stays stable even when the user's prompt template never mentions the time variable; an
|
||||
// unknown start time renders "unknown" and simply keys on the prompt alone.
|
||||
let conversation_start_time = params
|
||||
.prompt_variables
|
||||
.format("{{ baibot_conversation_start_time_utc }}");
|
||||
let prompt_cache_key = derive_prompt_cache_key(&prompt_text, &conversation_start_time);
|
||||
|
||||
let prompt_message = if prompt_text.is_empty() {
|
||||
None
|
||||
} else {
|
||||
@@ -58,54 +76,132 @@ pub async fn generate_text(
|
||||
conversation_messages.insert(0, prompt_message);
|
||||
}
|
||||
|
||||
let messages = convert_llm_messages_to_venice(conversation_messages);
|
||||
let messages = convert_llm_messages_to_venice(conversation_messages)?;
|
||||
|
||||
let temperature = params
|
||||
.temperature_override
|
||||
.unwrap_or(text_generation_config.temperature);
|
||||
|
||||
let request = ChatCompletionRequest {
|
||||
// When web search is active, ask Venice to return structured search results so we can render
|
||||
// readable citations from them. Respect an explicit user choice and only fill the flag when
|
||||
// the user left it unset.
|
||||
let venice_parameters = text_generation_config
|
||||
.venice_parameters
|
||||
.clone()
|
||||
.map(|mut vp| {
|
||||
let web_search_active = matches!(
|
||||
vp.enable_web_search,
|
||||
Some(WebSearchMode::On | WebSearchMode::Auto)
|
||||
);
|
||||
if web_search_active && vp.return_search_results_as_documents.is_none() {
|
||||
vp.return_search_results_as_documents = Some(true);
|
||||
}
|
||||
vp
|
||||
});
|
||||
|
||||
let mut request = ChatCompletionRequest {
|
||||
model: text_generation_config.model_id.clone(),
|
||||
messages,
|
||||
temperature: Some(temperature),
|
||||
// Web search rides entirely inside `venice_parameters`; there is no `tools` array here.
|
||||
// `max_tokens` is deprecated on Venice in favor of `max_completion_tokens`.
|
||||
max_completion_tokens: text_generation_config.max_response_tokens,
|
||||
venice_parameters: text_generation_config.venice_parameters.clone(),
|
||||
top_p: text_generation_config.top_p,
|
||||
frequency_penalty: text_generation_config.frequency_penalty,
|
||||
presence_penalty: text_generation_config.presence_penalty,
|
||||
repetition_penalty: text_generation_config.repetition_penalty,
|
||||
reasoning_effort: text_generation_config.reasoning_effort.clone(),
|
||||
prompt_cache_key: Some(prompt_cache_key),
|
||||
prompt_cache_retention: text_generation_config.prompt_cache_retention.clone(),
|
||||
venice_parameters,
|
||||
};
|
||||
|
||||
let url = format!(
|
||||
"{}/chat/completions",
|
||||
config.base_url.trim_end_matches('/')
|
||||
);
|
||||
let url = format!("{}/chat/completions", config.base_url.trim_end_matches('/'));
|
||||
|
||||
tracing::trace!(
|
||||
model = text_generation_config.model_id,
|
||||
messages_count = request.messages.len(),
|
||||
"Sending Venice chat completion API request"
|
||||
);
|
||||
let model_id = text_generation_config.model_id.clone();
|
||||
|
||||
let response = http
|
||||
.post(&url)
|
||||
.bearer_auth(&config.api_key)
|
||||
.json(&request)
|
||||
.send()
|
||||
.await?;
|
||||
// Proactively drop fields this model has already rejected earlier in this process, so a known
|
||||
// mismatch costs zero wasted round-trips after the first discovery. Venice's body is
|
||||
// `additionalProperties: false`, so sending a known-unsupported field would 400 again.
|
||||
for field in unsupported.known_for(&model_id) {
|
||||
super::recovery::strip_droppable_field(&mut request, &field);
|
||||
}
|
||||
|
||||
let status = response.status();
|
||||
if !status.is_success() {
|
||||
// Log the body server-side for debugging (Venice explains a rejected strict body there),
|
||||
// but keep it OUT of the returned error: that error surfaces in the Matrix room, and the
|
||||
// body can carry account / rate-limit details that shouldn't reach room members.
|
||||
// Send with bounded auto-recovery. When Venice 400s because a model does not support an optional
|
||||
// knob, it names the field (`field: '...'`); if that field is one we may safely drop, we strip
|
||||
// it, remember the rejection for this model, and retry. The loop is bounded: each retry clears a
|
||||
// distinct droppable field (strip returns false once it is gone), so after at most
|
||||
// `DROPPABLE_FIELDS.len()` retries the request either succeeds or surfaces the error.
|
||||
let response = loop {
|
||||
tracing::trace!(
|
||||
model = model_id,
|
||||
messages_count = request.messages.len(),
|
||||
"Sending Venice chat completion API request"
|
||||
);
|
||||
|
||||
let response = http
|
||||
.post(&url)
|
||||
.bearer_auth(&config.api_key)
|
||||
.json(&request)
|
||||
.send()
|
||||
.await?;
|
||||
|
||||
let status = response.status();
|
||||
if status.is_success() {
|
||||
break response;
|
||||
}
|
||||
|
||||
// Always log the body server-side: Venice explains a rejected strict body there.
|
||||
let body = response.text().await.unwrap_or_default();
|
||||
tracing::warn!(%status, body, "Venice chat completion request failed");
|
||||
|
||||
// Recover from a strict-body 400 over an unsupported optional knob: strip the named field
|
||||
// and retry. Only fields in `DROPPABLE_FIELDS` are eligible, so a meaning-bearing knob (a
|
||||
// sampling parameter) is never silently dropped; that case falls through to surface below.
|
||||
if status == reqwest::StatusCode::BAD_REQUEST
|
||||
&& let Some(field) = super::recovery::parse_rejected_field(&body)
|
||||
&& super::recovery::strip_droppable_field(&mut request, &field)
|
||||
{
|
||||
unsupported.record(&model_id, &field);
|
||||
tracing::info!(
|
||||
model = model_id,
|
||||
field,
|
||||
"Venice rejected an unsupported field; dropping it and retrying"
|
||||
);
|
||||
continue;
|
||||
}
|
||||
|
||||
// A 413 almost always means an attached file pushed the request past Venice's size limit.
|
||||
if status == reqwest::StatusCode::PAYLOAD_TOO_LARGE {
|
||||
return Err(anyhow::anyhow!(
|
||||
"The request was too large for Venice, most likely an attached file over the 25MB limit."
|
||||
));
|
||||
}
|
||||
|
||||
// A 400 is a complaint about the request baibot built, so the body is safe and useful to
|
||||
// surface: it tells the operator (e.g. at agent-create time) exactly which field or value
|
||||
// Venice rejected, instead of an opaque status. Other statuses keep the body OUT of the
|
||||
// returned error, since it can carry account / rate-limit details that shouldn't reach the
|
||||
// room.
|
||||
if status == reqwest::StatusCode::BAD_REQUEST {
|
||||
return Err(anyhow::anyhow!(
|
||||
"Venice rejected the request (400 Bad Request): {}",
|
||||
super::recovery::extract_error_message(&body)
|
||||
));
|
||||
}
|
||||
|
||||
return Err(anyhow::anyhow!(
|
||||
"Venice chat completion request failed with status {status}"
|
||||
));
|
||||
}
|
||||
};
|
||||
|
||||
let response: ChatCompletionResponse = response.json().await?;
|
||||
|
||||
let citations = response
|
||||
.venice_parameters
|
||||
.map(|vp| vp.web_search_citations)
|
||||
.unwrap_or_default();
|
||||
|
||||
let Some(choice) = response.choices.into_iter().next() else {
|
||||
return Err(anyhow::anyhow!(
|
||||
"No choices were returned from the Venice chat completion API"
|
||||
@@ -118,5 +214,138 @@ pub async fn generate_text(
|
||||
));
|
||||
};
|
||||
|
||||
Ok(TextGenerationResult { text: content })
|
||||
let text = render_with_citations(content, &citations);
|
||||
let text = append_reasoning(
|
||||
text,
|
||||
choice.message.reasoning_content,
|
||||
text_generation_config.show_reasoning,
|
||||
);
|
||||
|
||||
Ok(TextGenerationResult { text })
|
||||
}
|
||||
|
||||
/// Builds the prompt-cache routing key from conversation-stable inputs. `DefaultHasher::new()` is a
|
||||
/// fixed-seed SipHasher (keys 0,0), so it is deterministic across processes and restarts: identical
|
||||
/// inputs always produce the same key, which is what lets a restarted bot keep hitting the warm
|
||||
/// cache. The algorithm is not guaranteed stable across Rust std versions, so a rebuild on a new
|
||||
/// toolchain can shift every key once, a one-time cache warm-up with no correctness effect.
|
||||
pub(super) fn derive_prompt_cache_key(prompt_text: &str, conversation_start_time: &str) -> String {
|
||||
let mut hasher = DefaultHasher::new();
|
||||
prompt_text.hash(&mut hasher);
|
||||
conversation_start_time.hash(&mut hasher);
|
||||
format!("{:016x}", hasher.finish())
|
||||
}
|
||||
|
||||
/// Appends the model's thinking to the reply only when the deployment opts in via `show_reasoning`.
|
||||
/// `reasoning_content` is a field separate from the answer `content` (it is unaffected by
|
||||
/// `strip_thinking_response`, which only strips inline `<think>` blocks from `content`), so reading
|
||||
/// it here is independent of that knob. Default-off matches today's behavior: thinking never reaches
|
||||
/// a room that did not ask for it.
|
||||
///
|
||||
/// The thinking renders as a Matrix-native collapsible `<details>` block: folded by default, one
|
||||
/// click to expand, so it stays out of the way of the answer instead of dumping a wall of reasoning
|
||||
/// inline. This survives the send path: the reply goes through markdown (`send_text_markdown`),
|
||||
/// whose pulldown-cmark pass writes raw HTML verbatim rather than escaping it, and ruma's HTML
|
||||
/// sanitizer allow-lists `<details>`/`<summary>`. Clients that do not render `<details>` degrade to
|
||||
/// showing the summary and reasoning inline, so nothing is lost there either.
|
||||
pub(super) fn append_reasoning(
|
||||
text: String,
|
||||
reasoning_content: Option<String>,
|
||||
show_reasoning: bool,
|
||||
) -> String {
|
||||
if !show_reasoning {
|
||||
return text;
|
||||
}
|
||||
|
||||
match reasoning_content {
|
||||
Some(reasoning) if !reasoning.trim().is_empty() => {
|
||||
// The blank lines around the trimmed reasoning keep it a separate markdown block from
|
||||
// the surrounding `<details>`/`</details>` HTML blocks, so the reasoning itself still
|
||||
// renders as markdown (lists, code, emphasis) inside the collapsible.
|
||||
let reasoning = reasoning.trim();
|
||||
format!(
|
||||
"{text}\n\n<details><summary>💭 Reasoning</summary>\n\n{reasoning}\n\n</details>"
|
||||
)
|
||||
}
|
||||
_ => text,
|
||||
}
|
||||
}
|
||||
|
||||
/// Rewrites Venice's inline `^n^` citation superscripts into readable `[n]` references and appends
|
||||
/// a `Sources:` list of markdown links, one per citation in order. Returns the content unchanged
|
||||
/// when web search returned no citations, so non-search replies are never touched.
|
||||
///
|
||||
/// Citation `title` and `url` come from scraped web pages, so they are attacker-influenced. The
|
||||
/// title is escaped so it cannot break out of the markdown link label, and the URL is used as a
|
||||
/// link target only when it is a clean `http(s)` URL with no markdown-breaking characters;
|
||||
/// otherwise the citation renders as plain text. This stops a hostile page title or URL from
|
||||
/// injecting a spoofed clickable link into the room.
|
||||
pub(super) fn render_with_citations(content: String, citations: &[WebSearchCitation]) -> String {
|
||||
if citations.is_empty() {
|
||||
return content;
|
||||
}
|
||||
|
||||
let mut text = rewrite_citation_superscripts(&content);
|
||||
|
||||
let mut sources = String::from("\n\nSources:");
|
||||
for (index, citation) in citations.iter().enumerate() {
|
||||
let n = index + 1;
|
||||
let title = escape_markdown_link_text(&citation.title);
|
||||
match sanitize_link_url(&citation.url) {
|
||||
// A citation that arrived with no title still renders as a usable link by showing the
|
||||
// URL as the link text, rather than an empty `[]( )` label.
|
||||
Some(url) if title.is_empty() => sources.push_str(&format!("\n[{n}] [{url}]({url})")),
|
||||
Some(url) => sources.push_str(&format!("\n[{n}] [{title}]({url})")),
|
||||
None if !title.is_empty() => sources.push_str(&format!("\n[{n}] {title}")),
|
||||
None => sources.push_str(&format!("\n[{n}] (source unavailable)")),
|
||||
}
|
||||
}
|
||||
|
||||
text.push_str(&sources);
|
||||
text
|
||||
}
|
||||
|
||||
/// Venice marks web-search citations with superscript runs in the reply text: a single `^1^`, a
|
||||
/// comma list `^1,2^`, or a caret-chained run `^2^3^10^` where consecutive citations share a
|
||||
/// caret. The whole run has to be matched at once: a per-citation pattern (string or regex)
|
||||
/// consumes the shared caret on the first match and orphans the rest (`^2^3^` would leave `3^`).
|
||||
/// So this matches each full run and expands it to one `[n]` per citation (`^2^3^` -> `[2][3]`).
|
||||
fn rewrite_citation_superscripts(content: &str) -> String {
|
||||
static RUN: OnceLock<Regex> = OnceLock::new();
|
||||
let run = RUN.get_or_init(|| {
|
||||
Regex::new(r"\^\d+(?:[,^]\d+)*\^").expect("citation superscript regex is valid")
|
||||
});
|
||||
|
||||
run.replace_all(content, |caps: ®ex::Captures| {
|
||||
caps[0]
|
||||
.split(['^', ','])
|
||||
.filter(|piece| !piece.is_empty())
|
||||
.map(|n| format!("[{n}]"))
|
||||
.collect::<String>()
|
||||
})
|
||||
.into_owned()
|
||||
}
|
||||
|
||||
/// Escapes the characters that would let citation title text break out of a markdown link label,
|
||||
/// and folds newlines to spaces so a multi-line title cannot inject extra markdown structure.
|
||||
fn escape_markdown_link_text(text: &str) -> String {
|
||||
text.replace('\\', "\\\\")
|
||||
.replace('[', "\\[")
|
||||
.replace(']', "\\]")
|
||||
.replace(['\r', '\n'], " ")
|
||||
}
|
||||
|
||||
/// Returns the URL as a markdown link target only when it is a clean `http(s)` URL with no
|
||||
/// characters that would break the `(...)` destination or smuggle a different scheme. Anything else
|
||||
/// returns `None`, so the caller renders the citation as plain text instead of a link.
|
||||
fn sanitize_link_url(url: &str) -> Option<String> {
|
||||
let url = url.trim();
|
||||
let is_http = url.starts_with("https://") || url.starts_with("http://");
|
||||
let is_clean = !url.contains(['(', ')', '<', '>', ' ', '\t', '\r', '\n']);
|
||||
|
||||
if is_http && is_clean {
|
||||
Some(url.to_owned())
|
||||
} else {
|
||||
None
|
||||
}
|
||||
}
|
||||
|
||||
@@ -61,6 +61,39 @@ pub struct TextGenerationConfig {
|
||||
#[serde(default)]
|
||||
pub max_context_tokens: u32,
|
||||
|
||||
/// Sampling and reasoning knobs that live at the top level of Venice's `/chat/completions`
|
||||
/// body, not inside the `venice_parameters` bag. Venice silently ignores a top-level knob
|
||||
/// placed in the bag, so these sit here as siblings and map straight to top-level wire fields
|
||||
/// in `chat.rs`. Each is omitted from the request when unset.
|
||||
#[serde(default)]
|
||||
pub top_p: Option<f32>,
|
||||
|
||||
#[serde(default)]
|
||||
pub frequency_penalty: Option<f32>,
|
||||
|
||||
#[serde(default)]
|
||||
pub presence_penalty: Option<f32>,
|
||||
|
||||
#[serde(default)]
|
||||
pub repetition_penalty: Option<f32>,
|
||||
|
||||
#[serde(default)]
|
||||
pub reasoning_effort: Option<String>,
|
||||
|
||||
/// Prompt-cache retention window (`default`, `extended`, or `24h`). This carries a named
|
||||
/// default rather than a bare `#[serde(default)]` (which would yield `None`), so a config that
|
||||
/// omits the key still ships `24h` and keeps caching on. Caching is the per-deployment cost
|
||||
/// lever, so the omitted-key case must not silently disable it. The value here must agree with
|
||||
/// the `Default` impl below.
|
||||
#[serde(default = "default_prompt_cache_retention")]
|
||||
pub prompt_cache_retention: Option<String>,
|
||||
|
||||
/// When set, the model's `reasoning_content` (its thinking) is appended to the reply. Off by
|
||||
/// default to match today's `strip_thinking_response: true` behavior, so existing deployments
|
||||
/// see no change.
|
||||
#[serde(default)]
|
||||
pub show_reasoning: bool,
|
||||
|
||||
/// Venice-specific request knobs, serialized 1:1 into the `venice_parameters` bag on the
|
||||
/// wire. Any unset field is omitted, so Venice applies its own server-side default.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
@@ -79,6 +112,16 @@ impl Default for TextGenerationConfig {
|
||||
// Matches Venice's own `availableContextTokens` (131072) and the non-OpenAI sibling
|
||||
// providers (ollama/localai/mistral all default to 128_000).
|
||||
max_context_tokens: 128_000,
|
||||
// Sampling knobs stay None so Venice applies its own server-side default. Caching is
|
||||
// the one exception: retention defaults to 24h here so a programmatic default caches
|
||||
// out of the box, agreeing with the `#[serde(default = ...)]` on the field.
|
||||
top_p: None,
|
||||
frequency_penalty: None,
|
||||
presence_penalty: None,
|
||||
repetition_penalty: None,
|
||||
reasoning_effort: None,
|
||||
prompt_cache_retention: default_prompt_cache_retention(),
|
||||
show_reasoning: false,
|
||||
// A usable starting point, not an everything-set dump: only these three are sent;
|
||||
// every other knob stays None so Venice applies its own default (omitting != false).
|
||||
venice_parameters: Some(VeniceParameters {
|
||||
@@ -95,6 +138,13 @@ fn default_text_model_id() -> String {
|
||||
"kimi-k2-5".to_owned()
|
||||
}
|
||||
|
||||
/// Defaults prompt-cache retention to 24h so caching is on unless a config explicitly opts out.
|
||||
/// A bare `#[serde(default)]` would deserialize an omitted key to `None`, which disables caching;
|
||||
/// this keeps the cost lever engaged for configs that never mention it.
|
||||
fn default_prompt_cache_retention() -> Option<String> {
|
||||
Some("24h".to_owned())
|
||||
}
|
||||
|
||||
/// The full `venice_parameters` knob set, mirroring Venice's `ChatCompletionRequest`
|
||||
/// schema field-for-field. Every field is optional with `skip_serializing_if`, so the
|
||||
/// request never carries a knob the user didn't set (the body is `additionalProperties: false`,
|
||||
@@ -133,6 +183,11 @@ pub struct VeniceParameters {
|
||||
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub disable_thinking: Option<bool>,
|
||||
|
||||
/// Response verbosity (`low`, `medium`, `high`). Venice accepts this both top-level and inside
|
||||
/// the bag; it lives here so the top-level config stays lean, and Venice reads it from the bag.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub verbosity: Option<String>,
|
||||
}
|
||||
|
||||
#[derive(Debug, Clone, Copy, Serialize, Deserialize)]
|
||||
|
||||
@@ -13,11 +13,16 @@ use crate::conversation::llm::{
|
||||
|
||||
use super::super::ControllerTrait;
|
||||
use super::config::Config;
|
||||
use super::recovery::UnsupportedFieldsCache;
|
||||
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct Controller {
|
||||
config: Config,
|
||||
http: reqwest::Client,
|
||||
// Per-model record of chat fields this Venice deployment has rejected as unsupported, learned at
|
||||
// runtime. `Arc`-backed inside, so the `Clone` derive shares one cache across all clones of an
|
||||
// agent's controller.
|
||||
unsupported_fields: UnsupportedFieldsCache,
|
||||
}
|
||||
|
||||
impl Controller {
|
||||
@@ -30,7 +35,11 @@ impl Controller {
|
||||
.build()
|
||||
.unwrap_or_else(|_| reqwest::Client::new());
|
||||
|
||||
Self { config, http }
|
||||
Self {
|
||||
config,
|
||||
http,
|
||||
unsupported_fields: UnsupportedFieldsCache::default(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -62,7 +71,14 @@ impl ControllerTrait for Controller {
|
||||
conversation: LLMConversation,
|
||||
params: TextGenerationParams,
|
||||
) -> anyhow::Result<TextGenerationResult> {
|
||||
super::chat::generate_text(&self.config, &self.http, conversation, params).await
|
||||
super::chat::generate_text(
|
||||
&self.config,
|
||||
&self.http,
|
||||
&self.unsupported_fields,
|
||||
conversation,
|
||||
params,
|
||||
)
|
||||
.await
|
||||
}
|
||||
|
||||
async fn speech_to_text(
|
||||
|
||||
@@ -185,8 +185,6 @@ fn image_format_to_mime_type(format: Option<&str>) -> mxlink::mime::Mime {
|
||||
"jpeg" | "jpg" => mxlink::mime::IMAGE_JPEG,
|
||||
"png" => mxlink::mime::IMAGE_PNG,
|
||||
// No mxlink::mime constant for webp; parse it, falling back to PNG on any surprise value.
|
||||
_ => "image/webp"
|
||||
.parse()
|
||||
.unwrap_or(mxlink::mime::IMAGE_PNG),
|
||||
_ => "image/webp".parse().unwrap_or(mxlink::mime::IMAGE_PNG),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -3,6 +3,7 @@ mod chat;
|
||||
mod config;
|
||||
mod controller;
|
||||
mod images;
|
||||
mod recovery;
|
||||
mod utils;
|
||||
mod wire;
|
||||
|
||||
|
||||
254
src/agent/provider/venice/recovery.rs
Normal file
254
src/agent/provider/venice/recovery.rs
Normal file
@@ -0,0 +1,254 @@
|
||||
//! Auto-recovery for Venice's strict request bodies.
|
||||
//!
|
||||
//! Venice's `/chat/completions` body is `additionalProperties: false`, so a model that does not
|
||||
//! support an optional knob rejects the whole request with a 400 instead of ignoring the field.
|
||||
//! Some knobs are documented as model-specific ("for supported models") and are pure
|
||||
//! optimization/tuning hints: dropping them changes nothing about the answer, only loses the
|
||||
//! optimization. When such a field is the reason for a 400, we strip it and retry, then remember
|
||||
//! the rejection per model so later requests skip the field (and the wasted round-trip) entirely.
|
||||
//!
|
||||
//! Universal sampling knobs (`temperature`, `top_p`, the penalties, `max_completion_tokens`) are
|
||||
//! deliberately NOT recoverable here: dropping one silently changes the model's output, so a model
|
||||
//! that rejects one is a real configuration problem the operator must see, not paper over.
|
||||
|
||||
use std::collections::{HashMap, HashSet};
|
||||
use std::sync::{Arc, OnceLock, RwLock};
|
||||
|
||||
use regex::Regex;
|
||||
|
||||
use super::wire::ChatCompletionRequest;
|
||||
|
||||
/// Top-level chat-completion fields baibot may drop to recover from a 400. Each is documented by
|
||||
/// Venice as model-specific or as a routing hint, and is meaning-preserving to omit (Venice falls
|
||||
/// back to its server-side default):
|
||||
/// - `prompt_cache_retention` — "extends retention ... for supported models" (cache TTL only)
|
||||
/// - `reasoning_effort` — "control reasoning effort level for supported models"
|
||||
/// - `prompt_cache_key` — cache-routing hint; dropping it only forfeits a cache-hit optimization
|
||||
///
|
||||
/// Recovery operates on TOP-LEVEL request fields only. Sub-fields inside the `venice_parameters`
|
||||
/// bag (`disable_thinking`, `enable_e2ee`, `character_slug`, …) are intentionally absent: a model
|
||||
/// that rejects one surfaces a clear 400 to the operator rather than being auto-stripped. Adding a
|
||||
/// new bag field does not extend recovery to it; only a name listed here is droppable.
|
||||
pub(super) const DROPPABLE_FIELDS: &[&str] = &[
|
||||
"prompt_cache_retention",
|
||||
"prompt_cache_key",
|
||||
"reasoning_effort",
|
||||
];
|
||||
|
||||
/// Per-model record of fields a Venice model has rejected as unsupported, learned at runtime from
|
||||
/// 400 responses. The `Arc` is shared across `Controller` clones, so a rejection learned once is
|
||||
/// seen by every clone of the same agent. The cache is process-lived only: a restart re-learns on
|
||||
/// the first request to each model, which costs one extra round-trip and nothing else, so there is
|
||||
/// no persistence to keep in sync with config changes.
|
||||
#[derive(Debug, Clone, Default)]
|
||||
pub(super) struct UnsupportedFieldsCache {
|
||||
inner: Arc<RwLock<HashMap<String, HashSet<String>>>>,
|
||||
}
|
||||
|
||||
impl UnsupportedFieldsCache {
|
||||
/// Fields already known unsupported for `model_id`. Returns an empty set on an unknown model or
|
||||
/// a poisoned lock, so a cache failure degrades to "strip nothing proactively" rather than
|
||||
/// breaking the request path.
|
||||
pub(super) fn known_for(&self, model_id: &str) -> HashSet<String> {
|
||||
self.inner
|
||||
.read()
|
||||
.ok()
|
||||
.and_then(|map| map.get(model_id).cloned())
|
||||
.unwrap_or_default()
|
||||
}
|
||||
|
||||
/// Records that `model_id` rejected `field`. A poisoned lock is ignored: failing to memoize only
|
||||
/// means the next request re-discovers the rejection, never a wrong result.
|
||||
pub(super) fn record(&self, model_id: &str, field: &str) {
|
||||
if let Ok(mut map) = self.inner.write() {
|
||||
map.entry(model_id.to_owned())
|
||||
.or_default()
|
||||
.insert(field.to_owned());
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Parses the offending field name out of a Venice 400 body. A field the schema does not allow is
|
||||
/// reported as `... field: 'prompt_cache_retention', value: '...'`, so this matches the `field: '..'`
|
||||
/// marker wherever it sits in the message. Returns `None` when the body carries no such marker (a
|
||||
/// different 400 class, e.g. a missing required field), so the caller surfaces that error instead.
|
||||
///
|
||||
/// Returns only the FIRST `field: '..'` match by design. If Venice ever names several rejected
|
||||
/// fields in one body, the retry loop strips this one, retries, and rediscovers the next on the
|
||||
/// following 400 — bounded and correct. Do not switch to `captures_iter` to "batch" them without
|
||||
/// re-checking the loop's per-field termination bound in `chat.rs`.
|
||||
pub(super) fn parse_rejected_field(body: &str) -> Option<String> {
|
||||
static RE: OnceLock<Regex> = OnceLock::new();
|
||||
let re =
|
||||
RE.get_or_init(|| Regex::new(r"field: '([^']+)'").expect("rejected-field regex is valid"));
|
||||
re.captures(body)
|
||||
.map(|caps| caps[1].to_owned())
|
||||
.filter(|field| !field.is_empty())
|
||||
}
|
||||
|
||||
/// Clears `field` from the request when it is one baibot may safely drop and it is currently set.
|
||||
/// Returns `true` only when a value was actually removed, which is what bounds the retry loop: once
|
||||
/// a field is `None`, a repeat rejection for the same name returns `false` and the caller stops
|
||||
/// instead of retrying forever. A field outside [`DROPPABLE_FIELDS`] always returns `false`, so a
|
||||
/// meaning-bearing knob is never silently dropped.
|
||||
pub(super) fn strip_droppable_field(request: &mut ChatCompletionRequest, field: &str) -> bool {
|
||||
if !DROPPABLE_FIELDS.contains(&field) {
|
||||
return false;
|
||||
}
|
||||
match field {
|
||||
"prompt_cache_retention" => request.prompt_cache_retention.take().is_some(),
|
||||
"prompt_cache_key" => request.prompt_cache_key.take().is_some(),
|
||||
"reasoning_effort" => request.reasoning_effort.take().is_some(),
|
||||
_ => false,
|
||||
}
|
||||
}
|
||||
|
||||
/// Pulls a human-readable message out of a Venice error body for surfacing in the room. Venice's
|
||||
/// usual envelope is `{"error": "..."}`; some OpenAI-compatible paths nest `{"error": {"message":
|
||||
/// "..."}}`. Falls back to the trimmed raw body (length-capped so a large body cannot flood the
|
||||
/// room) and finally to a fixed string for an empty body, so the caller always has something to
|
||||
/// show.
|
||||
pub(super) fn extract_error_message(body: &str) -> String {
|
||||
let trimmed = body.trim();
|
||||
if trimmed.is_empty() {
|
||||
return "no response body".to_owned();
|
||||
}
|
||||
|
||||
if let Ok(value) = serde_json::from_str::<serde_json::Value>(trimmed) {
|
||||
if let Some(msg) = value.get("error").and_then(|e| e.as_str()) {
|
||||
return msg.to_owned();
|
||||
}
|
||||
if let Some(msg) = value
|
||||
.get("error")
|
||||
.and_then(|e| e.get("message"))
|
||||
.and_then(|m| m.as_str())
|
||||
{
|
||||
return msg.to_owned();
|
||||
}
|
||||
}
|
||||
|
||||
const MAX: usize = 500;
|
||||
if trimmed.chars().count() > MAX {
|
||||
trimmed.chars().take(MAX).collect::<String>() + "…"
|
||||
} else {
|
||||
trimmed.to_owned()
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn full_request() -> ChatCompletionRequest {
|
||||
ChatCompletionRequest {
|
||||
model: "venice-uncensored".to_owned(),
|
||||
messages: vec![],
|
||||
temperature: Some(0.7),
|
||||
max_completion_tokens: Some(1024),
|
||||
top_p: None,
|
||||
frequency_penalty: None,
|
||||
presence_penalty: None,
|
||||
repetition_penalty: None,
|
||||
reasoning_effort: Some("high".to_owned()),
|
||||
prompt_cache_key: Some("cafef00d".to_owned()),
|
||||
prompt_cache_retention: Some("24h".to_owned()),
|
||||
venice_parameters: None,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parses_the_rejected_field_from_a_real_venice_body() {
|
||||
let body = r#"{"error":"Extra inputs are not permitted, field: 'prompt_cache_retention', value: 'default'","request_id":"qM_DmKSXKF07wRxmQJ-hc"}"#;
|
||||
assert_eq!(
|
||||
parse_rejected_field(body).as_deref(),
|
||||
Some("prompt_cache_retention")
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn returns_no_field_when_the_body_has_no_field_marker() {
|
||||
// A different 400 class (e.g. a genuinely malformed request) carries no `field: '..'`
|
||||
// marker, so there is nothing to strip and the caller must surface the error instead.
|
||||
assert_eq!(parse_rejected_field(r#"{"error":"Invalid request"}"#), None);
|
||||
assert_eq!(parse_rejected_field(""), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn strips_a_droppable_field_once_then_reports_no_progress() {
|
||||
let mut request = full_request();
|
||||
|
||||
// First strip clears the field and reports progress, so the caller retries.
|
||||
assert!(strip_droppable_field(
|
||||
&mut request,
|
||||
"prompt_cache_retention"
|
||||
));
|
||||
assert!(request.prompt_cache_retention.is_none());
|
||||
|
||||
// A repeat rejection for the same (now absent) field reports no progress: this is what
|
||||
// stops the retry loop instead of spinning forever.
|
||||
assert!(!strip_droppable_field(
|
||||
&mut request,
|
||||
"prompt_cache_retention"
|
||||
));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn refuses_to_strip_a_meaning_bearing_field() {
|
||||
let mut request = full_request();
|
||||
|
||||
// `temperature` is universal and changes the output; a rejection for it must surface, never
|
||||
// be silently dropped. The whole droppable set is the only thing strip will touch.
|
||||
assert!(!strip_droppable_field(&mut request, "temperature"));
|
||||
assert_eq!(request.temperature, Some(0.7));
|
||||
|
||||
assert!(!strip_droppable_field(
|
||||
&mut request,
|
||||
"max_completion_tokens"
|
||||
));
|
||||
assert_eq!(request.max_completion_tokens, Some(1024));
|
||||
|
||||
for field in DROPPABLE_FIELDS {
|
||||
assert!(
|
||||
strip_droppable_field(&mut full_request(), field),
|
||||
"every advertised droppable field must actually be strippable: {field}"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn cache_records_per_model_and_isolates_models() {
|
||||
let cache = UnsupportedFieldsCache::default();
|
||||
assert!(cache.known_for("venice-uncensored").is_empty());
|
||||
|
||||
cache.record("venice-uncensored", "prompt_cache_retention");
|
||||
cache.record("venice-uncensored", "reasoning_effort");
|
||||
|
||||
let known = cache.known_for("venice-uncensored");
|
||||
assert!(known.contains("prompt_cache_retention"));
|
||||
assert!(known.contains("reasoning_effort"));
|
||||
|
||||
// A rejection learned for one model must not leak to another.
|
||||
assert!(cache.known_for("kimi-k2-5").is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn extracts_a_human_message_from_error_envelopes() {
|
||||
assert_eq!(
|
||||
extract_error_message(r#"{"error":"Extra inputs are not permitted","request_id":"x"}"#),
|
||||
"Extra inputs are not permitted"
|
||||
);
|
||||
|
||||
// OpenAI-style nested envelope.
|
||||
assert_eq!(
|
||||
extract_error_message(r#"{"error":{"message":"context length exceeded"}}"#),
|
||||
"context length exceeded"
|
||||
);
|
||||
|
||||
// Unknown shape falls back to the raw body; empty falls back to a fixed string.
|
||||
assert_eq!(
|
||||
extract_error_message("plain text failure"),
|
||||
"plain text failure"
|
||||
);
|
||||
assert_eq!(extract_error_message(" "), "no response body");
|
||||
}
|
||||
}
|
||||
@@ -11,11 +11,13 @@ use crate::conversation::llm::{
|
||||
MessageContent as LLMMessageContent,
|
||||
};
|
||||
|
||||
use super::config::{Config, VeniceParameters, WebSearchMode};
|
||||
use super::chat::{append_reasoning, derive_prompt_cache_key, render_with_citations};
|
||||
use super::config::{Config, TextGenerationConfig, VeniceParameters, WebSearchMode};
|
||||
use super::controller::Controller;
|
||||
use super::utils::convert_llm_messages_to_venice;
|
||||
use super::wire::{
|
||||
ContentPart, EditImageRequest, GenerateImageRequest, MessageContent, SpeechRequest,
|
||||
ChatCompletionRequest, ContentPart, EditImageRequest, GenerateImageRequest, MessageContent,
|
||||
SpeechRequest, WebSearchCitation,
|
||||
};
|
||||
|
||||
#[test]
|
||||
@@ -55,11 +57,14 @@ speech_to_text:
|
||||
!json.contains("character_slug"),
|
||||
"an unset knob must be omitted entirely: {json}"
|
||||
);
|
||||
assert!(!json.contains("null"), "no nulls belong in the body: {json}");
|
||||
assert!(
|
||||
!json.contains("null"),
|
||||
"no nulls belong in the body: {json}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn converts_image_to_data_uri_and_skips_files() {
|
||||
fn converts_text_image_and_file_to_content_parts() {
|
||||
let messages = vec![
|
||||
LLMMessage {
|
||||
author: LLMAuthor::User,
|
||||
@@ -95,10 +100,10 @@ fn converts_image_to_data_uri_and_skips_files() {
|
||||
},
|
||||
];
|
||||
|
||||
let converted = convert_llm_messages_to_venice(messages);
|
||||
let converted = convert_llm_messages_to_venice(messages).expect("conversion should succeed");
|
||||
|
||||
// Text and image survive; the file is warn-skipped.
|
||||
assert_eq!(converted.len(), 2);
|
||||
// Text, image, AND file all survive now: the file is no longer warn-skipped.
|
||||
assert_eq!(converted.len(), 3);
|
||||
|
||||
match &converted[0].content {
|
||||
MessageContent::Text(text) => assert_eq!(text, "describe this"),
|
||||
@@ -112,9 +117,25 @@ fn converts_image_to_data_uri_and_skips_files() {
|
||||
"image should be inlined as a data URI: {}",
|
||||
image_url.url
|
||||
),
|
||||
other => panic!("expected an image part, got {other:?}"),
|
||||
},
|
||||
other => panic!("expected image parts, got {other:?}"),
|
||||
}
|
||||
|
||||
match &converted[2].content {
|
||||
MessageContent::Parts(parts) => match &parts[0] {
|
||||
ContentPart::File { file } => {
|
||||
assert!(
|
||||
file.file_data.starts_with("data:application/pdf;base64,"),
|
||||
"file should be inlined as a data URI: {}",
|
||||
file.file_data
|
||||
);
|
||||
assert_eq!(file.filename.as_deref(), Some("doc.pdf"));
|
||||
}
|
||||
other => panic!("expected a file part, got {other:?}"),
|
||||
},
|
||||
other => panic!("expected file parts, got {other:?}"),
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -183,7 +204,10 @@ fn speech_request_serializes_voice_and_omits_unset() {
|
||||
!json.contains("temperature"),
|
||||
"an unset knob must be omitted (not null): {json}"
|
||||
);
|
||||
assert!(!json.contains("null"), "no nulls belong in the body: {json}");
|
||||
assert!(
|
||||
!json.contains("null"),
|
||||
"no nulls belong in the body: {json}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -226,7 +250,10 @@ fn generate_image_request_pins_flags_and_omits_unset() {
|
||||
!json.contains("cfg_scale"),
|
||||
"an unset knob must be omitted: {json}"
|
||||
);
|
||||
assert!(!json.contains("null"), "no nulls belong in the body: {json}");
|
||||
assert!(
|
||||
!json.contains("null"),
|
||||
"no nulls belong in the body: {json}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -243,10 +270,7 @@ fn edit_image_request_carries_model_and_base64_image() {
|
||||
|
||||
let json = serde_json::to_string(&request).expect("serialize EditImageRequest");
|
||||
|
||||
assert!(
|
||||
json.contains("\"model\":\"firered-image-edit\""),
|
||||
"{json}"
|
||||
);
|
||||
assert!(json.contains("\"model\":\"firered-image-edit\""), "{json}");
|
||||
assert!(
|
||||
json.contains("\"image\":\"aGVsbG8=\""),
|
||||
"the base64 image string must be present: {json}"
|
||||
@@ -267,3 +291,259 @@ fn web_search_mode_off_deserializes_from_bare_yaml_off() {
|
||||
|
||||
assert!(matches!(params.enable_web_search, Some(WebSearchMode::Off)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn request_places_sampling_top_level_and_verbosity_in_the_bag() {
|
||||
// The whole config-shape decision in one assertion: top-level knobs serialize at the top
|
||||
// level, the dual-position `verbosity` serializes inside the bag. Venice silently ignores a
|
||||
// top-level knob misplaced into the bag, so this is the guard against a silent no-op.
|
||||
let request = ChatCompletionRequest {
|
||||
model: "kimi-k2-5".to_owned(),
|
||||
messages: vec![],
|
||||
temperature: Some(0.5),
|
||||
max_completion_tokens: Some(1024),
|
||||
top_p: Some(0.5),
|
||||
frequency_penalty: None,
|
||||
presence_penalty: None,
|
||||
repetition_penalty: None,
|
||||
reasoning_effort: Some("high".to_owned()),
|
||||
prompt_cache_key: Some("00000000cafef00d".to_owned()),
|
||||
prompt_cache_retention: Some("24h".to_owned()),
|
||||
venice_parameters: Some(VeniceParameters {
|
||||
verbosity: Some("high".to_owned()),
|
||||
..Default::default()
|
||||
}),
|
||||
};
|
||||
|
||||
let json = serde_json::to_value(&request).expect("serialize request");
|
||||
|
||||
assert_eq!(json["top_p"], 0.5);
|
||||
assert_eq!(json["reasoning_effort"], "high");
|
||||
assert_eq!(json["prompt_cache_retention"], "24h");
|
||||
assert_eq!(json["prompt_cache_key"], "00000000cafef00d");
|
||||
|
||||
assert!(
|
||||
json.get("verbosity").is_none(),
|
||||
"verbosity must not be a top-level field: {json}"
|
||||
);
|
||||
assert_eq!(json["venice_parameters"]["verbosity"], "high");
|
||||
assert!(
|
||||
json["venice_parameters"].get("top_p").is_none(),
|
||||
"top_p must not be inside the bag: {json}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn config_defaults_prompt_cache_retention_to_24h() {
|
||||
// The programmatic default.
|
||||
assert_eq!(
|
||||
TextGenerationConfig::default()
|
||||
.prompt_cache_retention
|
||||
.as_deref(),
|
||||
Some("24h")
|
||||
);
|
||||
|
||||
// A config that omits the key must ALSO default to 24h, via the named serde default. A bare
|
||||
// `#[serde(default)]` would yield None here and silently disable caching for such configs.
|
||||
let tg: TextGenerationConfig = serde_yaml_ng::from_str("model_id: kimi-k2-5\n")
|
||||
.expect("minimal config should deserialize");
|
||||
assert_eq!(
|
||||
tg.prompt_cache_retention.as_deref(),
|
||||
Some("24h"),
|
||||
"an omitted retention key must still default to 24h"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn cache_key_is_stable_for_same_inputs_and_varies_otherwise() {
|
||||
let key = derive_prompt_cache_key("system prompt", "2024-09-20 (Friday), 18:34:15 UTC");
|
||||
|
||||
// Identical inputs produce an identical key: this is what keeps turn 5 routing to the warm
|
||||
// server holding turns 1-4 (and what survives a process restart).
|
||||
assert_eq!(
|
||||
key,
|
||||
derive_prompt_cache_key("system prompt", "2024-09-20 (Friday), 18:34:15 UTC"),
|
||||
"identical inputs must produce an identical key"
|
||||
);
|
||||
assert_eq!(key.len(), 16, "the key is a 16-char hex string");
|
||||
assert!(key.chars().all(|c| c.is_ascii_hexdigit()));
|
||||
|
||||
// A different conversation start time or a different prompt must change the key.
|
||||
assert_ne!(
|
||||
key,
|
||||
derive_prompt_cache_key("system prompt", "2024-09-21 (Saturday), 09:00:00 UTC"),
|
||||
"a different start time must change the key"
|
||||
);
|
||||
assert_ne!(
|
||||
key,
|
||||
derive_prompt_cache_key("other prompt", "2024-09-20 (Friday), 18:34:15 UTC"),
|
||||
"a different prompt must change the key"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn citations_render_inline_refs_and_a_sources_block() {
|
||||
let citations = vec![WebSearchCitation {
|
||||
title: "Example Source".to_owned(),
|
||||
url: "https://example.com/a".to_owned(),
|
||||
}];
|
||||
|
||||
let rendered = render_with_citations("the sky is blue^1^".to_owned(), &citations);
|
||||
|
||||
assert!(
|
||||
rendered.contains("the sky is blue[1]"),
|
||||
"inline ^1^ becomes [1]: {rendered}"
|
||||
);
|
||||
assert!(
|
||||
rendered.contains("Sources:"),
|
||||
"a Sources block is appended: {rendered}"
|
||||
);
|
||||
assert!(
|
||||
rendered.contains("[1] [Example Source](https://example.com/a)"),
|
||||
"the source renders as a markdown link: {rendered}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn citations_absent_leaves_content_untouched() {
|
||||
let content = "plain answer, no web search".to_owned();
|
||||
assert_eq!(render_with_citations(content.clone(), &[]), content);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn citation_title_and_url_cannot_inject_markdown() {
|
||||
// A hostile page sets its title to break out of the link label and its URL to a non-http
|
||||
// scheme. Neither may produce a spoofed clickable link in the room.
|
||||
let citations = vec take".to_owned(),
|
||||
url: "javascript:alert(1)".to_owned(),
|
||||
}];
|
||||
|
||||
let rendered = render_with_citations("result^1^".to_owned(), &citations);
|
||||
|
||||
assert!(
|
||||
rendered.contains("evil\\](http://phish.example) take"),
|
||||
"the title's brackets must be escaped so it cannot close the link label: {rendered}"
|
||||
);
|
||||
assert!(
|
||||
!rendered.contains("(javascript:alert(1))"),
|
||||
"a non-http(s) URL must never become a markdown link target: {rendered}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn chained_and_comma_citation_runs_each_expand_to_separate_refs() {
|
||||
let citations = vec![
|
||||
WebSearchCitation {
|
||||
title: "One".to_owned(),
|
||||
url: "https://example.com/1".to_owned(),
|
||||
},
|
||||
WebSearchCitation {
|
||||
title: "Two".to_owned(),
|
||||
url: "https://example.com/2".to_owned(),
|
||||
},
|
||||
WebSearchCitation {
|
||||
title: "Three".to_owned(),
|
||||
url: "https://example.com/3".to_owned(),
|
||||
},
|
||||
];
|
||||
|
||||
// Caret-chained run: Venice shares the caret between consecutive citations (`^2^3^`). The whole
|
||||
// run must expand, not just the first, with no orphaned `3^` left behind.
|
||||
let chained = render_with_citations("alpha^2^3^ and beta^1^".to_owned(), &citations);
|
||||
assert!(
|
||||
chained.contains("alpha[2][3] and beta[1]"),
|
||||
"a chained ^2^3^ run must expand to [2][3] with no orphaned caret: {chained}"
|
||||
);
|
||||
|
||||
// Comma run.
|
||||
let comma = render_with_citations("gamma^1,3^".to_owned(), &citations);
|
||||
assert!(
|
||||
comma.contains("gamma[1][3]"),
|
||||
"a comma ^1,3^ run must expand to [1][3]: {comma}"
|
||||
);
|
||||
|
||||
// Multi-digit citation indices survive intact.
|
||||
let multidigit = render_with_citations("delta^2^10^".to_owned(), &citations);
|
||||
assert!(
|
||||
multidigit.contains("delta[2][10]"),
|
||||
"a multi-digit chained run must expand to [2][10]: {multidigit}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn malformed_citation_degrades_instead_of_failing() {
|
||||
// A citation arriving without a `url` must still deserialize (to an empty default) rather than
|
||||
// failing the whole response parse and losing an otherwise-good answer.
|
||||
let parsed: WebSearchCitation = serde_json::from_str(r#"{"title":"Only a title"}"#)
|
||||
.expect("a citation missing `url` should still deserialize");
|
||||
assert_eq!(parsed.url, "");
|
||||
|
||||
// Rendering citations with missing fields stays graceful: no empty `[]()` link, no panic.
|
||||
let citations = vec![
|
||||
WebSearchCitation {
|
||||
title: String::new(),
|
||||
url: "https://example.com/u".to_owned(),
|
||||
},
|
||||
WebSearchCitation {
|
||||
title: String::new(),
|
||||
url: String::new(),
|
||||
},
|
||||
];
|
||||
let rendered = render_with_citations("answer^1^2^".to_owned(), &citations);
|
||||
assert!(
|
||||
rendered.contains("[1] [https://example.com/u](https://example.com/u)"),
|
||||
"a citation with no title falls back to the URL as link text: {rendered}"
|
||||
);
|
||||
assert!(
|
||||
rendered.contains("[2] (source unavailable)"),
|
||||
"a citation with neither title nor URL renders a placeholder: {rendered}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn reasoning_is_appended_only_when_show_reasoning_is_set() {
|
||||
let base = "the answer".to_owned();
|
||||
|
||||
// Off (the default): thinking is dropped, never reaching the room.
|
||||
let off = append_reasoning(base.clone(), Some("secret thinking".to_owned()), false);
|
||||
assert_eq!(off, "the answer");
|
||||
|
||||
// On: thinking is appended below the answer in a collapsible <details> block (folded by
|
||||
// default, expandable in clients that support it).
|
||||
let on = append_reasoning(base.clone(), Some(" visible thinking ".to_owned()), true);
|
||||
assert!(on.starts_with("the answer"));
|
||||
assert!(on.contains("<details><summary>💭 Reasoning</summary>"));
|
||||
assert!(on.contains("</details>"));
|
||||
// The reasoning sits as its own markdown block (blank lines around it) and is trimmed.
|
||||
assert!(on.contains("\n\nvisible thinking\n\n"));
|
||||
|
||||
// On but empty or missing reasoning: nothing is appended.
|
||||
assert_eq!(
|
||||
append_reasoning(base.clone(), Some(" ".to_owned()), true),
|
||||
"the answer"
|
||||
);
|
||||
assert_eq!(append_reasoning(base.clone(), None, true), "the answer");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn oversized_file_is_rejected() {
|
||||
let messages = vec![LLMMessage {
|
||||
author: LLMAuthor::User,
|
||||
sender_id: None,
|
||||
timestamp: chrono::Utc::now(),
|
||||
content: LLMMessageContent::File(FileDetails::new(
|
||||
FileMessageEventContent::plain(
|
||||
"big.pdf".to_owned(),
|
||||
OwnedMxcUri::from("mxc://example.com/big"),
|
||||
),
|
||||
mime::APPLICATION_PDF,
|
||||
vec![0u8; 25 * 1024 * 1024 + 1],
|
||||
)),
|
||||
}];
|
||||
|
||||
assert!(
|
||||
convert_llm_messages_to_venice(messages).is_err(),
|
||||
"a file over the 25MB limit must be rejected"
|
||||
);
|
||||
}
|
||||
|
||||
@@ -3,21 +3,26 @@ use crate::conversation::llm::{
|
||||
};
|
||||
use crate::utils::base64::base64_encode;
|
||||
|
||||
use super::wire::{ChatMessage, ContentPart, ImageUrl, MessageContent};
|
||||
use super::wire::{ChatMessage, ContentPart, FilePart, ImageUrl, MessageContent};
|
||||
|
||||
pub fn convert_llm_messages_to_venice(messages: Vec<LLMMessage>) -> Vec<ChatMessage> {
|
||||
/// Venice's documented file-input ceiling is 25MB on the decoded bytes (swagger `file_data`).
|
||||
/// We check it here so an oversized file gets a clear message instead of an opaque 413 from the
|
||||
/// API; the 413 status branch in `chat.rs` is the backstop if a file slips past this guard.
|
||||
const MAX_FILE_BYTES: usize = 25 * 1024 * 1024;
|
||||
|
||||
pub fn convert_llm_messages_to_venice(
|
||||
messages: Vec<LLMMessage>,
|
||||
) -> anyhow::Result<Vec<ChatMessage>> {
|
||||
let mut venice_messages: Vec<ChatMessage> = Vec::with_capacity(messages.len());
|
||||
|
||||
for message in messages {
|
||||
if let Some(venice_message) = convert_llm_message_to_venice(message) {
|
||||
venice_messages.push(venice_message);
|
||||
}
|
||||
venice_messages.push(convert_llm_message_to_venice(message)?);
|
||||
}
|
||||
|
||||
venice_messages
|
||||
Ok(venice_messages)
|
||||
}
|
||||
|
||||
fn convert_llm_message_to_venice(message: LLMMessage) -> Option<ChatMessage> {
|
||||
fn convert_llm_message_to_venice(message: LLMMessage) -> anyhow::Result<ChatMessage> {
|
||||
let role = match message.author {
|
||||
LLMAuthor::Prompt => "system",
|
||||
LLMAuthor::Assistant => "assistant",
|
||||
@@ -25,7 +30,7 @@ fn convert_llm_message_to_venice(message: LLMMessage) -> Option<ChatMessage> {
|
||||
};
|
||||
|
||||
match message.content {
|
||||
LLMMessageContent::Text(text) => Some(ChatMessage {
|
||||
LLMMessageContent::Text(text) => Ok(ChatMessage {
|
||||
role: role.to_owned(),
|
||||
content: MessageContent::Text(text),
|
||||
}),
|
||||
@@ -38,18 +43,39 @@ fn convert_llm_message_to_venice(message: LLMMessage) -> Option<ChatMessage> {
|
||||
base64_encode(&image_details.data)
|
||||
);
|
||||
|
||||
Some(ChatMessage {
|
||||
Ok(ChatMessage {
|
||||
role: role.to_owned(),
|
||||
content: MessageContent::Parts(vec![ContentPart::ImageUrl {
|
||||
image_url: ImageUrl { url: data_uri },
|
||||
}]),
|
||||
})
|
||||
}
|
||||
LLMMessageContent::File(_file_details) => {
|
||||
tracing::warn!(
|
||||
"The Venice provider does not support file content. This file message will be skipped."
|
||||
LLMMessageContent::File(file_details) => {
|
||||
// Inline the file as a base64 data URI in a `file` content part. This is the input
|
||||
// type the openai_compat provider drops; baibot already extracts the bytes upstream.
|
||||
// The message reaches the room, so it carries no user-controlled filename: a crafted
|
||||
// name could otherwise inject markdown (a spoofed link) into the bot's reply.
|
||||
if file_details.data.len() > MAX_FILE_BYTES {
|
||||
return Err(anyhow::anyhow!(
|
||||
"The attached file is too large for Venice (the limit is 25MB)."
|
||||
));
|
||||
}
|
||||
|
||||
let data_uri = format!(
|
||||
"data:{};base64,{}",
|
||||
file_details.mime,
|
||||
base64_encode(&file_details.data)
|
||||
);
|
||||
None
|
||||
|
||||
Ok(ChatMessage {
|
||||
role: role.to_owned(),
|
||||
content: MessageContent::Parts(vec![ContentPart::File {
|
||||
file: FilePart {
|
||||
file_data: data_uri,
|
||||
filename: Some(file_details.filename()),
|
||||
},
|
||||
}]),
|
||||
})
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -25,6 +25,27 @@ pub struct ChatCompletionRequest {
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub max_completion_tokens: Option<u32>,
|
||||
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub top_p: Option<f32>,
|
||||
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub frequency_penalty: Option<f32>,
|
||||
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub presence_penalty: Option<f32>,
|
||||
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub repetition_penalty: Option<f32>,
|
||||
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub reasoning_effort: Option<String>,
|
||||
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub prompt_cache_key: Option<String>,
|
||||
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub prompt_cache_retention: Option<String>,
|
||||
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub venice_parameters: Option<VeniceParameters>,
|
||||
}
|
||||
@@ -37,8 +58,8 @@ pub struct ChatMessage {
|
||||
}
|
||||
|
||||
/// A message body is either a bare string or a list of content parts. Venice accepts both; we
|
||||
/// send the parts form only when a message carries an image (baibot keeps text and images in
|
||||
/// separate messages, so a parts list only ever holds images in v1).
|
||||
/// send the parts form when a message carries an image or a file (baibot keeps text, images, and
|
||||
/// files in separate messages, so a parts list holds a single image part or file part).
|
||||
#[derive(Debug, Serialize)]
|
||||
#[serde(untagged)]
|
||||
pub enum MessageContent {
|
||||
@@ -50,6 +71,7 @@ pub enum MessageContent {
|
||||
#[serde(tag = "type", rename_all = "snake_case")]
|
||||
pub enum ContentPart {
|
||||
ImageUrl { image_url: ImageUrl },
|
||||
File { file: FilePart },
|
||||
}
|
||||
|
||||
#[derive(Debug, Serialize)]
|
||||
@@ -58,12 +80,26 @@ pub struct ImageUrl {
|
||||
pub url: String,
|
||||
}
|
||||
|
||||
/// Standard OpenAI-shaped chat completion response. We only read `choices[0].message.content`;
|
||||
/// when web search is on, Venice inlines citations as `^n^` superscripts in that content and we
|
||||
/// pass it through untouched.
|
||||
#[derive(Debug, Serialize)]
|
||||
pub struct FilePart {
|
||||
/// A `data:<mime>;base64,<data>` URI carrying the file bytes inline.
|
||||
pub file_data: String,
|
||||
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub filename: Option<String>,
|
||||
}
|
||||
|
||||
/// Standard OpenAI-shaped chat completion response. We read `choices[0].message.content` and,
|
||||
/// when web search is on, the structured `venice_parameters.web_search_citations` (requested via
|
||||
/// `return_search_results_as_documents`) to rewrite the inline `^n^` superscripts into readable
|
||||
/// `[n]` references plus a `Sources:` block. `reasoning_content` carries the model's thinking when
|
||||
/// the model exposes it; it is appended only when `show_reasoning` is set.
|
||||
#[derive(Debug, Deserialize)]
|
||||
pub struct ChatCompletionResponse {
|
||||
pub choices: Vec<ChatChoice>,
|
||||
|
||||
#[serde(default)]
|
||||
pub venice_parameters: Option<ResponseVeniceParameters>,
|
||||
}
|
||||
|
||||
#[derive(Debug, Deserialize)]
|
||||
@@ -75,6 +111,31 @@ pub struct ChatChoice {
|
||||
pub struct ResponseMessage {
|
||||
#[serde(default)]
|
||||
pub content: Option<String>,
|
||||
|
||||
#[serde(default)]
|
||||
pub reasoning_content: Option<String>,
|
||||
}
|
||||
|
||||
/// The `venice_parameters` envelope on a chat-completion *response*, distinct from the request-side
|
||||
/// `VeniceParameters` bag. Only the citation list is read; other response-side fields are ignored.
|
||||
#[derive(Debug, Deserialize, Default)]
|
||||
pub struct ResponseVeniceParameters {
|
||||
#[serde(default)]
|
||||
pub web_search_citations: Vec<WebSearchCitation>,
|
||||
}
|
||||
|
||||
/// Only the `title` and `url` are read (for rendering the `Sources:` block). Venice also returns
|
||||
/// `content` and `date` per citation; serde drops them, the same way the response structs above
|
||||
/// ignore the response fields baibot does not use. Both fields default to empty so a single
|
||||
/// citation that arrives without one (schema drift on scraped results) degrades gracefully in the
|
||||
/// rendered list instead of failing the whole response deserialization.
|
||||
#[derive(Debug, Deserialize)]
|
||||
pub struct WebSearchCitation {
|
||||
#[serde(default)]
|
||||
pub title: String,
|
||||
|
||||
#[serde(default)]
|
||||
pub url: String,
|
||||
}
|
||||
|
||||
/// `/audio/transcriptions` response. We read `text`; the optional `duration`/`timestamps` the
|
||||
|
||||
Reference in New Issue
Block a user