Compare commits

..

79 Commits

Author SHA1 Message Date
Aine
9fc3998f8e Provider-neutral context management 2026-06-28 22:31:06 +01:00
Aine
0be08c0292 update venice.ai links 2026-06-28 10:41:19 +01:00
renovate[bot]
8904edb2a0 Update Rust crate quick_cache to 0.7.* 2026-06-28 07:16:51 +03:00
renovate[bot]
e18c083c4f Update docker.io/ollama/ollama Docker tag to v0.30.11 2026-06-27 06:22:52 +03:00
renovate[bot]
1948f58473 Update jdx/mise-action action to v4 2026-06-26 10:50:15 +03:00
Slavi Pantaleev
c8a90049ad CI: trim workflow comments
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 07:28:14 +03:00
Slavi Pantaleev
4f7778e62f CI: run the prek hook suite instead of hand-rolled cargo commands
CI previously re-implemented a subset of the prek hooks by hand (`cargo
test` + a bare `cargo clippy`), so `cargo fmt --check` and the stricter
`cargo clippy -- -D warnings` were enforced only by the local pre-commit
hook — easily bypassed with --no-verify (as PR #193 was). Run the same
prek suite CI-side so .pre-commit-config.yaml is the single source of
truth for what gets checked, on both commit and push.

Also drop the hard-coded `dtolnay/rust-toolchain@1.93.0` pin (which had
drifted from rust-toolchain.toml's 1.96.0) in favor of
actions-rust-lang/setup-rust-toolchain, which reads the toolchain version
and components from rust-toolchain.toml — so CI's rustfmt/clippy match
what developers run, and there is no second place to keep in sync. All
three actions are tag-pinned, so Renovate can manage them (unlike the
branch-pinned dtolnay ref).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 07:23:16 +03:00
Slavi Pantaleev
355b1b299c Run cargo fmt on PR #193 code to satisfy the fmt hook
The squashed thinking-notice PR was committed with --no-verify, so three
files were never run through `cargo fmt` under the project's pinned
toolchain (rust-toolchain.toml = 1.96.0). Format them so `cargo fmt
--all -- --check` passes — a prerequisite for wiring the prek suite
(which includes that check) into CI in the next commit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 07:20:20 +03:00
Aine
318a8fd91d Add thinking notice + fix Venice auto-recover caching (#193)
Opt-in 💭 thinking-notice for slow text generation, plus a fix making Venice unsupported-field auto-recovery survive the per-message controller rebuild. Prepares 1.24.0.

Co-authored-by: Aine <aine@etke.cc>
2026-06-26 07:07:50 +03:00
renovate[bot]
025accdeb0 Update Rust crate quick_cache to v0.6.24 2026-06-26 06:48:29 +03:00
renovate[bot]
edf5bc9fdb Update Rust crate anyhow to v1.0.103 2026-06-26 06:48:20 +03:00
Slavi Pantaleev
982ebc6657 CHANGELOG: set 1.23.1 release date to 2026-06-24
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 07:25:44 +03:00
Aine
4a75be742f Venice: add auto-recover for optional params and surface actual error message on 400s 2026-06-24 07:10:29 +03:00
renovate[bot]
aae3c4c08f Update ghcr.io/element-hq/element-web Docker tag to v1.12.22 2026-06-24 07:09:33 +03:00
Slavi Pantaleev
08cf50885d CHANGELOG: fix 1.23.0 date and add OpenAI-compatible CA-trust fix
Correct the 1.23.0 release date to 2026-06-23 and document the
system-CA-trust fix (2888cb9, via etke_openai_api_rust 0.1.10) that also
ships in this release.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 07:42:15 +03:00
Slavi Pantaleev
1eebb1a55b Merge pull request #189 from etkecc/venice-improvements
Venice improvements
2026-06-23 07:41:43 +03:00
Slavi Pantaleev
2888cb9450 Bump etke_openai_api_rust to 0.1.10 for system CA trust
0.1.10 enables ureq's `native-certs` feature, so the OpenAI-compatible
provider trusts the system CA store (honouring SSL_CERT_FILE) instead of
only the bundled webpki-roots. Without it, endpoints behind a private or
internal CA (FreeIPA, org PKI) fail the TLS handshake with "invalid peer
certificate: UnknownIssuer".

Fixes #188.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 07:16:58 +03:00
Aine
5de497482d Venice: add 💭 Reasoning spoiler block (html <details>) 2026-06-23 01:30:37 +01:00
Aine
9e5f6de965 Venice improvements 2026-06-22 17:18:08 +01:00
renovate[bot]
2ae641109f Update Rust crate reqwest to v0.13.4 2026-06-21 07:57:25 +03:00
Slavi Pantaleev
105ca7b506 Update reqwest to 0.13
Aligns baibot's direct reqwest dep with the 0.13 copy that
async-openai/matrix-sdk/mxlink already pull, instead of the lone 0.12
copy kept alive only by the anthropic fork. 0.13 renamed the
`rustls-tls` feature to `rustls`; update the feature list accordingly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 07:48:15 +03:00
Aine
de980f0165 Implement Venice.ai provider 2026-06-21 07:18:37 +03:00
renovate[bot]
cc3888a4cf Update forgejo.ellis.link/continuwuation/continuwuity Docker tag to v0.5.10 2026-06-20 21:49:16 +03:00
renovate[bot]
eb3285f104 Update actions/checkout action to v7 2026-06-19 11:02:19 +03:00
renovate[bot]
0a03dde523 Update Rust crate async-openai to v0.41.1 2026-06-19 11:02:07 +03:00
renovate[bot]
a5b575da68 Update docker.io/ollama/ollama Docker tag to v0.30.10 2026-06-18 09:55:57 +03:00
renovate[bot]
fe7afec920 Update ghcr.io/element-hq/synapse Docker tag to v1.155.0 2026-06-17 06:24:58 +03:00
renovate[bot]
b6641523da Update docker.io/ollama/ollama Docker tag to v0.30.9 2026-06-17 06:19:12 +03:00
renovate[bot]
c579977e59 Update dependency prek to v0.4.5 2026-06-15 16:13:50 +03:00
renovate[bot]
0a5f37c3eb Update docker.io/ollama/ollama Docker tag to v0.30.8 2026-06-14 07:09:07 +03:00
renovate[bot]
58246e4eb3 Update ghcr.io/element-hq/element-web Docker tag to v1.12.21 2026-06-09 22:04:39 +03:00
renovate[bot]
9b767a420e Update Rust crate regex to v1.12.4 2026-06-09 22:04:26 +03:00
renovate[bot]
54c27311c1 Update docker.io/ollama/ollama Docker tag to v0.30.7 2026-06-09 07:50:10 +03:00
renovate[bot]
3d15f11f76 Update docker.io/ollama/ollama Docker tag to v0.30.6 2026-06-05 12:33:56 +03:00
Slavi Pantaleev
ead70d71ff Prepare 1.21.1 release
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 10:07:16 +03:00
Slavi Pantaleev
49bb52a091 Update anthropic-rs fork to drop vulnerable rustls-webpki 0.101
The anthropic git dependency now uses reqwest 0.12 / rustls 0.23, pulling
rustls-webpki 0.103.13 instead of the 0.101.7 that was dragged in via the
old reqwest 0.11. This clears three RUSTSEC/Dependabot advisories:

- GHSA-82j2-j2ch-gfr8 (high): DoS via panic on malformed CRL BIT STRING
- GHSA-xgp8-3hg3-c2mh (low): name constraints accepted for wildcard names
- GHSA-965h-392x-2mh5 (low): name constraints for URI names incorrectly accepted

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 10:06:27 +03:00
Slavi Pantaleev
1fc2f0f65a Prepare 1.21.0 release
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 09:44:25 +03:00
Slavi Pantaleev
f7cb38b620 Fix clippy warnings in tests
- Avoid unwrap_or() on a statically-Some value by using the raw token count
- Use an array literal instead of vec! for the non-allocated test cases

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 09:41:00 +03:00
Slavi Pantaleev
04e7667d11 Default to gpt-image-2 for OpenAI image generation
Make gpt-image-2 the default image-generation model and add it to the
recognized model-id mapping, updating the sample provider configs to match.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 09:39:45 +03:00
Slavi Pantaleev
78cbfda481 Update Rust crate async-openai to 0.41.0
Adapt to upstream API changes:
- Handle the new ImageModel::GptImage2 variant in image-model match arms
- ImageSize dropped Copy (gained an Other(String) variant), so clone it

Supersedes #170.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 09:37:46 +03:00
renovate[bot]
36fd6a4dda Update Rust crate quick_cache to v0.6.23 2026-06-05 09:33:16 +03:00
renovate[bot]
1308196419 Update docker.io/ollama/ollama Docker tag to v0.30.5 2026-06-05 09:33:07 +03:00
renovate[bot]
b3546ebaf1 Update ghcr.io/element-hq/synapse Docker tag to v1.154.0 2026-06-04 18:50:24 +03:00
renovate[bot]
2499e6baca Update Rust crate chrono to v0.4.45 2026-06-04 18:50:01 +03:00
renovate[bot]
871b9d2f3b Update dependency prek to v0.4.4 2026-06-04 14:48:21 +03:00
renovate[bot]
4a36cf9446 Update docker.io/ollama/ollama Docker tag to v0.30.4 2026-06-04 07:13:27 +03:00
renovate[bot]
a2180452c9 Update docker.io/ollama/ollama Docker tag to v0.30.2 2026-06-03 07:39:57 +03:00
Slavi Pantaleev
0cb0fc18ce Prepare 1.20.0 release
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 23:05:29 +03:00
renovate[bot]
78bddac716 Update Rust crate tiktoken-rs to 0.12.* (#162)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-06-02 23:03:20 +03:00
Slavi Pantaleev
f7209221f6 Update matrix-sdk to v0.18.0 (and mxlink to v1.15.0)
matrix-sdk 0.18.0 introduces no breaking changes that affect baibot, but
it does require a matching mxlink release: mxlink 1.14.0 pins matrix-sdk
0.17.0, so without bumping mxlink the two matrix-sdk versions conflict.
mxlink 1.15.0 (released alongside this) tracks matrix-sdk 0.18.0, so we
bump the floor to >=1.15.0.

Verified locally: cargo build and the full test suite pass.

Supersedes Renovate PR #161.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 22:47:13 +03:00
renovate[bot]
5f2054e07f Update Rust crate async-openai to v0.40.3 2026-06-02 12:06:50 +03:00
renovate[bot]
8301c5a29f Update docker.io/ollama/ollama Docker tag to v0.30.0 2026-06-02 12:06:30 +03:00
Slavi Pantaleev
b325b6b1a5 Upgrade Rust (1.95.0 -> 1.96.0)
Bumps the base image in Dockerfile/Dockerfile.ci (via Renovate #158) and
keeps rust-toolchain.toml in sync, since rustup honors the toolchain file
inside the container build and would otherwise keep using 1.95.0.

Verified locally under 1.96.0: cargo check, clippy -D warnings, fmt check,
and cargo test --all-features all pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 09:02:05 +03:00
renovate[bot]
0600721663 Update ghcr.io/element-hq/element-web Docker tag to v1.12.20 2026-05-27 20:50:26 +03:00
renovate[bot]
dc17cf7e96 Update ghcr.io/element-hq/element-web Docker tag to v1.12.19 2026-05-27 13:50:35 +03:00
Slavi Pantaleev
d6a8f6ba0b Prepare 1.19.3 release
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 10:09:11 +03:00
renovate[bot]
c4c3d71195 Update dependency prek to v0.4.3 2026-05-27 09:56:15 +03:00
renovate[bot]
84ae29d034 Update dependency prek to v0.4.2 2026-05-26 15:14:23 +03:00
renovate[bot]
447f43df7b Update Rust crate async-openai to v0.40.2 2026-05-22 23:11:28 +03:00
renovate[bot]
dabac790ad Update Rust crate async-openai to v0.40.1 2026-05-22 10:19:16 +03:00
renovate[bot]
fa01f012a1 Update Rust crate serde_json to v1.0.150 2026-05-22 10:19:05 +03:00
Slavi Pantaleev
20cb33bc66 Prepare 1.19.2 release
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 12:23:20 +03:00
Slavi Pantaleev
2791bdb08b Merge pull request #150 from etkecc/chore/update-dependencies
Update dependencies
2026-05-21 09:39:11 +03:00
Slavi Pantaleev
5dd505202a Update dependencies
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 09:33:28 +03:00
Slavi Pantaleev
d61078002c Merge pull request #149 from etkecc/renovate/async-openai-0.x
Update Rust crate async-openai to 0.40.0
2026-05-21 09:26:46 +03:00
renovate[bot]
cf5b346558 Update Rust crate async-openai to 0.40.0 2026-05-21 05:55:20 +00:00
renovate[bot]
ebbb6658e1 Update dependency prek to v0.4.1 2026-05-20 09:13:57 +03:00
renovate[bot]
369dc0c1ba Update ghcr.io/element-hq/synapse Docker tag to v1.153.0 2026-05-19 21:46:36 +03:00
renovate[bot]
3185f44a93 Update Rust crate quick_cache to v0.6.22 2026-05-18 07:32:20 +03:00
renovate[bot]
aa8ccde0ed Update docker.io/postgres Docker tag to v18.4 2026-05-15 12:48:23 +03:00
renovate[bot]
d5df0d7416 Update docker.io/ollama/ollama Docker tag to v0.24.0 2026-05-15 12:48:07 +03:00
renovate[bot]
140ca9ed68 Update dependency prek to v0.4.0 2026-05-14 16:40:26 +03:00
renovate[bot]
89d77d52b6 Update Rust crate async-openai to v0.38.2 2026-05-14 07:32:06 +03:00
renovate[bot]
5092700275 Update docker.io/ollama/ollama Docker tag to v0.23.4 2026-05-14 07:31:22 +03:00
renovate[bot]
10c365124e Update docker.io/ollama/ollama Docker tag to v0.23.3 2026-05-13 07:28:16 +03:00
renovate[bot]
08bdf4f7a2 Update ghcr.io/element-hq/element-web Docker tag to v1.12.18 2026-05-12 20:38:24 +03:00
Slavi Pantaleev
af557a7e45 Fix multi-arch manifest publish: use buildx imagetools
docker/build-push-action now wraps single-platform images in an OCI
image index (to carry provenance attestations), so the per-arch
`*-amd64`/`*-arm64` tags are manifest lists. `docker manifest create`
refuses manifest-list sources ("X is a manifest list"). Switch to
`docker buildx imagetools create`, which flattens index sources
correctly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 09:22:50 +03:00
renovate[bot]
fa7eb11b1d Update Rust crate async-openai to v0.38.1 2026-05-12 02:39:22 +03:00
Slavi Pantaleev
ff1e128f0e Fix versioned image publish under workflow_run trigger
e978d3c switched the publish trigger from `push` to `workflow_run` and
correctly migrated the `type=raw,value=latest` rule to read the upstream
head_branch from `github.event.workflow_run.*` — but left the
`type=semver,pattern={{raw}}` rule unchanged. That rule still reads
`github.ref`, which under workflow_run dispatch is always
`refs/heads/main` (the default branch where the workflow file lives),
not the triggering tag ref. As a result, no semver tag was extracted,
metadata-action produced no tags, and `buildx` failed with
"tag is needed when pushing to registry". The `latest` tag kept
publishing because its rule was migrated; versioned tags (v1.19.0,
v1.19.1) silently stopped publishing.

Pass the upstream head_branch to the semver rule explicitly via `value`,
gated by `enable` so it only fires for v* tags. Mirrors the migration
the raw rule already received.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 17:28:54 +03:00
57 changed files with 4036 additions and 697 deletions

View File

@@ -14,13 +14,31 @@ concurrency:
group: ci-${{ github.event.pull_request.number || github.ref }} group: ci-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true cancel-in-progress: true
jobs: jobs:
test-and-clippy: prek:
name: Unit testing and linting name: Lint, format & test
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- uses: actions/checkout@v6 - uses: actions/checkout@v7
- uses: dtolnay/rust-toolchain@1.93.0
# Toolchain version + components come from rust-toolchain.toml. rustflags is
# cleared so plain builds don't fail on warnings; the clippy hook still does.
- uses: actions-rust-lang/setup-rust-toolchain@v1
with:
rustflags: ''
- name: Install SQLite3 - name: Install SQLite3
run: sudo apt-get update && sudo apt-get install -y libsqlite3-dev run: sudo apt-get update && sudo apt-get install -y libsqlite3-dev
- run: cargo test --all-features
- run: cargo clippy # just drives the prek recipes; mise provides the pinned prek (mise.toml).
- uses: taiki-e/install-action@v2
with:
tool: just
- uses: jdx/mise-action@v4
# Run the same prek hooks devs run locally; .pre-commit-config.yaml is the
# source of truth. Tests are a separate step for visible timing.
- name: Lint & format (prek hooks, excluding tests)
run: just prek-run-on-all --skip test-unit
- name: Unit tests
run: just test

View File

@@ -22,7 +22,7 @@ jobs:
json: ${{ steps.meta.outputs.json }} json: ${{ steps.meta.outputs.json }}
steps: steps:
- name: Checkout - name: Checkout
uses: actions/checkout@v6 uses: actions/checkout@v7
with: with:
ref: ${{ github.event.workflow_run.head_sha }} ref: ${{ github.event.workflow_run.head_sha }}
fetch-depth: 0 fetch-depth: 0
@@ -34,7 +34,7 @@ jobs:
ghcr.io/${{ github.repository }} ghcr.io/${{ github.repository }}
tags: | tags: |
type=raw,value=latest,enable=${{ github.event.workflow_run.head_branch == 'main' }} type=raw,value=latest,enable=${{ github.event.workflow_run.head_branch == 'main' }}
type=semver,pattern={{raw}} type=semver,pattern={{raw}},value=${{ github.event.workflow_run.head_branch }},enable=${{ startsWith(github.event.workflow_run.head_branch || '', 'v') }}
docker-build: docker-build:
if: | if: |
@@ -61,7 +61,7 @@ jobs:
steps: steps:
- name: Checkout - name: Checkout
uses: actions/checkout@v6 uses: actions/checkout@v7
with: with:
ref: ${{ github.event.workflow_run.head_sha }} ref: ${{ github.event.workflow_run.head_sha }}
fetch-depth: 0 fetch-depth: 0
@@ -77,7 +77,7 @@ jobs:
with: with:
tags: | tags: |
type=raw,value=latest,enable=${{ github.event.workflow_run.head_branch == 'main' }} type=raw,value=latest,enable=${{ github.event.workflow_run.head_branch == 'main' }}
type=semver,pattern={{raw}} type=semver,pattern={{raw}},value=${{ github.event.workflow_run.head_branch }},enable=${{ startsWith(github.event.workflow_run.head_branch || '', 'v') }}
flavor: | flavor: |
latest=auto latest=auto
suffix=-${{ matrix.arch }},onlatest=true suffix=-${{ matrix.arch }},onlatest=true
@@ -121,5 +121,4 @@ jobs:
- name: Create and push manifest - name: Create and push manifest
run: | run: |
docker manifest create ${{ matrix.image }} ${{ matrix.image }}-amd64 ${{ matrix.image }}-arm64 docker buildx imagetools create -t ${{ matrix.image }} ${{ matrix.image }}-amd64 ${{ matrix.image }}-arm64
docker manifest push ${{ matrix.image }}

View File

@@ -1,3 +1,83 @@
# (2026-06-28) Version 1.25.0
- (**Feature**) [♻️ Context management](./docs/configuration/text-generation.md#️-context-management) now works with every provider, not only [OpenAI](./docs/providers.md#openai). Token counting previously went through [tiktoken-rs](https://github.com/zurawiki/tiktoken-rs), which is accurate only for OpenAI models and silently mis-counted everything else (worst of all for non-English text). OpenAI agents keep using tiktoken-rs; every other provider, including the recommended [Venice](./docs/providers.md#venice), now uses a provider-neutral approximation that needs no per-model tokenizer (ASCII counted at about four characters per token, other scripts such as Cyrillic and CJK at about two), landing within roughly 10-20% of the real count. See the [context management docs](./docs/configuration/text-generation.md#️-context-management).
- (**Improvement**) Context management now trims a conversation on whole-turn boundaries for every provider, so an assistant reply is never kept without the user message it answered. This also adjusts how the OpenAI provider trims: a dangling assistant reply at the oldest edge of the kept history is now dropped along with its missing prompt, rather than left in place.
# (2026-06-26) Version 1.24.0
- (**Feature**) Add an opt-in 💭 **thinking notice** for text generation. When enabled, a slow response (for example, from a reasoning model that runs for minutes) posts a "thinking…" placeholder after a short delay, refreshes it periodically with varying flavor text, and then edits that same message into the final answer, so a long wait no longer looks like a stuck bot. The notice is **disabled by default** and configurable per-room or globally via `text-generation set-thinking-notice-enabled true`. Fast responses (under the delay threshold) never show a placeholder. See the [text-generation configuration docs](./docs/configuration/text-generation.md#-thinking-notice).
- (**Bugfix**) The [Venice](https://venice.ai) unsupported-field auto-recovery (added in 1.23.1) now actually remembers rejections across messages. The cache lived on the provider's controller, which is rebuilt on every message for room-local and global agents, so each turn started with an empty cache, re-sent the unsupported field, and logged the same `400 Bad Request` warning again. The cache is now process-global (keyed per Venice deployment), so a field a model rejects is dropped proactively on every later request instead of being re-discovered each turn.
# (2026-06-24) Version 1.23.1
- (**Bugfix**) The [Venice](https://venice.ai) provider now auto-recovers when a model rejects an optional knob it does not support. Venice's request body is strict (`additionalProperties: false`), so a model that lacks `prompt_cache_retention`, `reasoning_effort`, or `prompt_cache_key` rejected the whole request with a `400 Bad Request` — breaking agent creation and every reply. baibot now drops the unsupported field and retries, remembering the rejection per model so later requests skip it without a wasted round-trip. Only these meaning-preserving fields are dropped; sampling knobs that change the output (`temperature`, `top_p`, the penalties) are never silently removed and still surface as an error.
- (**Improvement**) When Venice rejects a request with a `400 Bad Request`, baibot now surfaces Venice's actual error message (e.g. `Extra inputs are not permitted, field: 'prompt_cache_retention'`) instead of a generic "configuration does not result in a working agent". This makes agent-creation failures self-explanatory. Other error statuses keep their bodies redacted, since those can carry account or rate-limit details.
# (2026-06-23) Version 1.23.0
- (**Feature**) The [Venice](https://venice.ai) provider now accepts file inputs (PDF, DOCX, and other documents, up to 25MB), the same way it already handled images. This makes Venice the second provider after OpenAI to accept files; the others (Anthropic and the OpenAI-compatible providers) skip them. See the [text-generation feature docs](./docs/features.md#-text-generation).
- (**Feature**) Add prompt caching to the Venice provider, on by default (`prompt_cache_retention: 24h`). baibot derives the cache key from the system prompt and the conversation start time (both fixed for the life of a conversation), so a long, stable system prompt stays cached across the day instead of being reprocessed and re-billed on every turn. See [Text Generation / Prompt Override](./docs/configuration/text-generation.md#️-prompt-override).
- (**Feature**) Wire up the rest of Venice's sampling and reasoning controls: top-level `top_p`, `frequency_penalty`, `presence_penalty`, `repetition_penalty`, and `reasoning_effort`; `verbosity` in the `venice_parameters` bag; and a `show_reasoning` toggle that appends the model's reasoning to the reply as a collapsible, folded-by-default `💭 Reasoning` block (off by default). See the [Venice configuration reference](./docs/providers.md#venice).
- (**Feature**) Render Venice web-search citations as readable `[n]` references with a `Sources:` list of links, instead of leaving Venice's raw `^n^` superscripts in the reply.
- (**Security**) Escape citation titles and validate citation URLs before rendering them, and drop user-supplied filenames from error messages, so a hostile web page or a crafted filename cannot inject a spoofed link into the bot's reply.
- (**Bugfix**) The [OpenAI-compatible](./docs/providers.md#openai-compatible) provider now trusts the system CA store (honoring `SSL_CERT_FILE`), so endpoints served behind a private/internal CA (FreeIPA, organization PKI) no longer fail the TLS handshake with `invalid peer certificate: UnknownIssuer`. Fixed upstream in `etke_openai_api_rust` 0.1.10. Thanks to [@shaba](https://github.com/shaba) for the report in [#188](https://github.com/etkecc/baibot/pull/188).
# (2026-06-21) Version 1.22.0
- (**Feature**) Add a native [Venice](https://venice.ai) provider with [🖌️ image-generation](./docs/features.md#️-image-creation) (incl. editing), [💬 text-generation](./docs/features.md#-text-generation) (incl. vision), [🗣️ text-to-speech](./docs/features.md#️-text-to-speech), [🦻 speech-to-text](./docs/features.md#-speech-to-text), and Venice's native web search via the full `venice_parameters` knob set. Unlike the [OpenAI-compatible](./docs/providers.md#openai-compatible) path (which drops images and can't reach Venice's audio or native image endpoints), it talks to Venice's API directly, using the knob-rich native `/image/generate` and `/image/edit` endpoints. See the [Venice provider docs](./docs/providers.md#venice).
# (2026-06-05) Version 1.21.1
- (**Security**) Update the [anthropic](https://github.com/etkecc/anthropic-rs) dependency to use [reqwest](https://crates.io/crates/reqwest) 0.12 / [rustls](https://crates.io/crates/rustls) 0.23, replacing the vulnerable `rustls-webpki` 0.101 line with 0.103.13. This resolves [`GHSA-82j2-j2ch-gfr8`](https://github.com/advisories/GHSA-82j2-j2ch-gfr8) (high — denial of service via panic on a malformed CRL), [`GHSA-xgp8-3hg3-c2mh`](https://github.com/advisories/GHSA-xgp8-3hg3-c2mh) and [`GHSA-965h-392x-2mh5`](https://github.com/advisories/GHSA-965h-392x-2mh5) (name-constraint validation issues).
# (2026-06-05) Version 1.21.0
- (**Improvement**) Default to OpenAI's `gpt-image-2` model for image generation (in newly-created OpenAI agents and the sample provider configs).
- (**Internal Improvement**) Update [async-openai](https://crates.io/crates/async-openai) from 0.40 to 0.41, which [resynchronizes with the upstream OpenAI API spec](https://github.com/64bit/async-openai/issues/557) after it had drifted out of sync — a mismatch that was already causing some breakage (hopefully now resolved). Adapts to the newly-added `gpt-image-2` image model and an `ImageSize` type change.
- (**Internal Improvement**) Dependency updates.
# (2026-06-02) Version 1.20.0
- (**Internal Improvement**) Update [matrix-sdk](https://crates.io/crates/matrix-sdk) from 0.17 to 0.18 and [mxlink](https://crates.io/crates/mxlink) to 1.15.0.
- (**Internal Improvement**) Update [tiktoken-rs](https://crates.io/crates/tiktoken-rs) to 0.12, backporting OpenAI [tiktoken](https://github.com/openai/tiktoken) 0.13.0 for better alignment with upstream tokenization behavior.
- (**Internal Improvement**) Bump the pinned Rust toolchain from 1.95.0 to 1.96.0 (in `rust-toolchain.toml` and the Docker build images).
- (**Internal Improvement**) Dependency updates.
# (2026-05-27) Version 1.19.3
- (**Internal Improvement**) Update [async-openai](https://crates.io/crates/async-openai) to 0.40.2, pulling in several upstream fixes (streaming HTTP error surfacing, default `ResponseTextParam.format` deserialization, etc.).
- (**Internal Improvement**) Dependency updates.
# (2026-05-21) Version 1.19.2
- (**Internal Improvement**) Update [async-openai](https://crates.io/crates/async-openai) to 0.40.0.
- (**Internal Improvement**) Dependency updates.
# (2026-05-09) Version 1.19.1 # (2026-05-09) Version 1.19.1
- (**Internal Improvement**) Update [async-openai](https://crates.io/crates/async-openai) to 0.38.0. - (**Internal Improvement**) Update [async-openai](https://crates.io/crates/async-openai) to 0.38.0.

881
Cargo.lock generated

File diff suppressed because it is too large Load Diff

View File

@@ -7,7 +7,7 @@ license = "AGPL-3.0-or-later"
readme = "README.md" readme = "README.md"
keywords = ["matrix", "chat", "bot", "AI", "LLM"] keywords = ["matrix", "chat", "bot", "AI", "LLM"]
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"] include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
version = "1.19.1" version = "1.25.0"
edition = "2024" edition = "2024"
[lib] [lib]
@@ -17,22 +17,25 @@ path = "src/lib.rs"
[dependencies] [dependencies]
anthropic = { git = "https://github.com/etkecc/anthropic-rs.git", branch = "fix-content-block-image" } anthropic = { git = "https://github.com/etkecc/anthropic-rs.git", branch = "fix-content-block-image" }
anyhow = "1.0.*" anyhow = "1.0.*"
async-openai = { version = "0.38.0", features = ["audio", "chat-completion", "image", "responses"] } async-openai = { version = "0.41.0", features = ["audio", "chat-completion", "image", "responses"] }
base64 = "0.22.*" base64 = "0.22.*"
chrono = { version = "0.4.*", default-features = false, features = ["std", "now"] } chrono = { version = "0.4.*", default-features = false, features = ["std", "now"] }
# We'd rather not depend on this, but we cannot use the ruma-events EventContent macro without it. # We'd rather not depend on this, but we cannot use the ruma-events EventContent macro without it.
matrix-sdk = { version = "0.17.0", default-features = false } matrix-sdk = { version = "0.18.0", default-features = false }
mime_guess = "2.0.*" mime_guess = "2.0.*"
mxidwc = "1.0.*" mxidwc = "1.0.*"
mxlink = ">=1.14.0" mxlink = ">=1.15.0"
etke_openai_api_rust = "0.1.*" etke_openai_api_rust = "0.1.*"
quick_cache = "0.6.*" quick_cache = "0.7.*"
regex = "1.12.*" regex = "1.12.*"
# HTTP client for the native `venice` provider. rustls only (no extra TLS stack), matching the
# reqwest copy async-openai/matrix-sdk/mxlink already use.
reqwest = { version = "0.13.*", default-features = false, features = ["json", "multipart", "rustls"] }
serde = { version = "1.0.*", features = ["derive"], default-features = false } serde = { version = "1.0.*", features = ["derive"], default-features = false }
serde_json = "1.0.*" serde_json = "1.0.*"
serde_yaml_ng = "0.10.*" serde_yaml_ng = "0.10.*"
tempfile = "3.27.*" tempfile = "3.27.*"
tiktoken-rs = { version = "0.11.*", default-features = false } tiktoken-rs = { version = "0.12.*", default-features = false }
tokio = { version = "1.52.*", features = ["rt", "rt-multi-thread", "macros"] } tokio = { version = "1.52.*", features = ["rt", "rt-multi-thread", "macros"] }
tracing = "0.1.*" tracing = "0.1.*"
tracing-subscriber = { version = "0.3.*", features = ["env-filter"] } tracing-subscriber = { version = "0.3.*", features = ["env-filter"] }

View File

@@ -4,7 +4,7 @@
# # # #
####################################### #######################################
FROM docker.io/rust:1.95.0-slim-trixie AS build FROM docker.io/rust:1.96.0-slim-trixie AS build
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev

View File

@@ -4,7 +4,7 @@
# # # #
####################################### #######################################
FROM docker.io/rust:1.95.0-slim-trixie AS build FROM docker.io/rust:1.96.0-slim-trixie AS build
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev

View File

@@ -13,7 +13,7 @@ It's influenced by [chaz](https://github.com/arcuru/chaz), but does **not** use
## 🌟 Features ## 🌟 Features
- 🎨 Encourages **[provider](./docs/providers.md) choice** ([Anthropic](./docs/providers.md#anthropic), [Groq](./docs/providers.md#groq), [LocalAI](./docs/providers.md#localai), [OpenAI](./docs/providers.md#openai) and [☁️ many more](./docs/providers.md#️-providers)) as well as **[mixing & matching models](./docs/features.md#-mixing--matching-models)**: - 🎨 Encourages **[provider](./docs/providers.md) choice** ([Anthropic](./docs/providers.md#anthropic), [Groq](./docs/providers.md#groq), [LocalAI](./docs/providers.md#localai), [OpenAI](./docs/providers.md#openai), [Venice](./docs/providers.md#venice) and [☁️ many more](./docs/providers.md#️-providers)) as well as **[mixing & matching models](./docs/features.md#-mixing--matching-models)**:
- Supports **different use purposes** (depending on the [☁️ provider](./docs/providers.md) & model): - Supports **different use purposes** (depending on the [☁️ provider](./docs/providers.md) & model):
@@ -30,7 +30,7 @@ It's influenced by [chaz](https://github.com/arcuru/chaz), but does **not** use
- 🔒 Supports [encryption](./docs/features.md#-encryption) for Matrix communication and Account-Data-stored configuration - 🔒 Supports [encryption](./docs/features.md#-encryption) for Matrix communication and Account-Data-stored configuration
- ♻️ Supports [context-management](./docs/configuration/text-generation.md#️-context-management) handling on some models (automatically adjusting the message history length, etc.) - ♻️ Supports [context-management](./docs/configuration/text-generation.md#️-context-management) for every [provider](./docs/providers.md) (automatically trimming older messages on whole-turn boundaries once a conversation outgrows the context window)
- 🛠️ Allows **customizing much of the bot's [configuration](./docs/configuration/README.md)** at runtime (using commands sent via chat) - 🛠️ Allows **customizing much of the bot's [configuration](./docs/configuration/README.md)** at runtime (using commands sent via chat)

View File

@@ -50,13 +50,22 @@ Example: `!bai config room text-generation set-auto-usage only_for_voice` (this
### ♻️ Context Management ### ♻️ Context Management
The bot also supports ♻️ **context management**, which automatically adjusts the message history length, etc. The bot also supports ♻️ **context management**, which automatically trims the oldest messages once a conversation grows past the context window. It drops whole turns at a time, so a reply is never separated from the message it answered.
This feature relies on [tokenization](https://en.wikipedia.org/wiki/Large_language_model#Tokenization) performed by the [tiktoken-rs](https://github.com/zurawiki/tiktoken-rs) library which is [poorly well-maintained](https://github.com/zurawiki/tiktoken-rs/issues/50) and only works well for [OpenAI](../providers.md#openai) models. Counting tokens precisely needs the model's own tokenizer. For [OpenAI](../providers.md#openai) models, the bot counts them with the [tiktoken-rs](https://github.com/zurawiki/tiktoken-rs) library. For every other provider, including the recommended [Venice](../providers.md#venice), the bot falls back to a provider-neutral **approximation** that needs no per-model tokenizer: it counts ASCII text at about four characters per token and other scripts (Cyrillic, CJK, and so on) at about two. Treat it as rough, within roughly 10-20% of the real count for typical text, which is plenty for keeping a long conversation inside the context window.
This setting is **disabled by default**, but can be enabled via `!bai config room text-generation set-context-management-enabled true` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings)). This setting is **disabled by default**, but can be enabled via `!bai config room text-generation set-context-management-enabled true` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings)).
### 💭 Thinking Notice
The bot can post a 💭 **"thinking…" notice** while text generation is running, useful for slow models (for example, reasoning models that may run for minutes) where the response would otherwise look stuck.
When enabled, a placeholder message appears only after a short delay (so fast responses get no notice), updates periodically with varying status text, and is then edited in place to become the final answer.
This setting is **disabled by default**, but can be enabled via `!bai config room text-generation set-thinking-notice-enabled true` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings)).
### 👤 Sender Context Mode ### 👤 Sender Context Mode
In multi-user rooms, it may be useful for the model to know which participant sent each message in the conversation context. In multi-user rooms, it may be useful for the model to know which participant sent each message in the conversation context.
@@ -101,6 +110,8 @@ Prompts may contain the following **placeholder variables** which will be replac
💡 `{{ baibot_now_utc }}` changes as time goes on, which prevents [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) from working. It's better to use `{{ baibot_conversation_start_time_utc }}` in prompts, as its value doesn't change yet still orients the bot to the current date/time. 💡 `{{ baibot_now_utc }}` changes as time goes on, which prevents [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) from working. It's better to use `{{ baibot_conversation_start_time_utc }}` in prompts, as its value doesn't change yet still orients the bot to the current date/time.
💡 On the [Venice provider](../providers.md#venice), baibot derives the prompt-cache key from the system prompt and the conversation start time, both stable for the life of a conversation, and ships `prompt_cache_retention: 24h` by default. A stable system prompt then stays cached across the whole conversation instead of being reprocessed (and re-billed) on every turn.
Here's a prompt that combines some of the above variables: Here's a prompt that combines some of the above variables:
> You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}." > You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."

View File

@@ -56,7 +56,7 @@ Example: `!bai config room text-to-speech set-speed-override 1.5` (this can also
### 👫 Voice override ### 👫 Voice override
The voice override setting lets you change the voice being used by the text-to-speech model configured at the [🤖 agent](../agents.md) level (usually `onyx` when using [OpenAI](../providers.md#openai)). The voice override setting lets you change the voice being used by the text-to-speech model configured at the [🤖 agent](../agents.md) level (e.g. `onyx` when using [OpenAI](../providers.md#openai), or `af_sky` when using [Venice](../providers.md#venice)).
Possible values (e.g. `onyx`) depend on the model you're using. For example, for [OpenAI](../providers.md#openai)'s Whisper model, [these voices](https://platform.openai.com/docs/guides/text-to-speech/voice-options) are available. Possible values (e.g. `onyx`) depend on the model you're using. For example, for [OpenAI](../providers.md#openai)'s Whisper model, [these voices](https://platform.openai.com/docs/guides/text-to-speech/voice-options) are available.

View File

@@ -26,7 +26,7 @@ Text Generation is the bot's ability to **respond to users' messages with text**
![Screenshot of Text Generation - a user sends a message and the bot replies in a new conversation thread](./screenshots/text-generation.webp) ![Screenshot of Text Generation - a user sends a message and the bot replies in a new conversation thread](./screenshots/text-generation.webp)
Some models also support vision and document understanding, so you may be able to mix text, images, and files (PDFs, text documents, etc.) in the same conversation. Note that certain providers may not support all file types or may have issues with specific files (e.g. scanned/image-based PDFs). If a file is rejected by the provider, the conversation thread may become unusable — start a new thread to work around this. Some models also support vision and document understanding, so you may be able to mix text, images, and files (PDFs, text documents, etc.) in the same conversation. File inputs (documents such as PDFs) are currently accepted only by the OpenAI and Venice providers; the others skip them. Note that certain providers may not support all file types or may have issues with specific files (e.g. scanned/image-based PDFs). If a file is rejected by the provider, the conversation thread may become unusable — start a new thread to work around this.
In multi-user (group) rooms, to avoid disturbing the normal conversation between people, the bot is auto-configured to only respond to messages starting with the command prefix (`!bai`) or direct mentions via the [💬 Text Generation / 🗟 Prefix Requirement Type](./configuration/text-generation.md#-prefix-requirement-type) setting. In multi-user (group) rooms, to avoid disturbing the normal conversation between people, the bot is auto-configured to only respond to messages starting with the command prefix (`!bai`) or direct mentions via the [💬 Text Generation / 🗟 Prefix Requirement Type](./configuration/text-generation.md#-prefix-requirement-type) setting.

View File

@@ -19,11 +19,12 @@ The list of supported providers is below.
- [OpenAI Compatible](#openai-compatible) - [OpenAI Compatible](#openai-compatible)
- [OpenRouter](#openrouter) - [OpenRouter](#openrouter)
- [Together AI](#together-ai) - [Together AI](#together-ai)
- [Venice](#venice)
### How to choose a provider ### How to choose a provider
If you're not sure which provider to start with, **we recommend [OpenAI](#openai)** as it's the most popular and has the **widest range of capabilities**: [💬 text-generation](./features.md#-text-generation) (incl. vision, incl. [🛠️ tools](./features.md#️-built-in-tools-openai-only)), [🖌️ image-generation](./features.md#️image-generation), [🦻 speech-to-text](./features.md#-speech-to-text), [🗣️ text-to-speech](./features.md#️-text-to-speech). If you're not sure which provider to start with, **we recommend [Venice](#venice)**: it's the most capable provider baibot supports (covering [💬 text-generation](./features.md#-text-generation) with vision, file inputs, prompt caching, and native web search, plus [🖌️ image-generation](./features.md#️-image-creation) incl. editing, [🦻 speech-to-text](./features.md#-speech-to-text), and [🗣️ text-to-speech](./features.md#️-text-to-speech)) and the only one that runs inference with no logging and no training on your data. If you'd rather start with the most widely-used option, [OpenAI](#openai) is a solid, well-supported choice too.
You don't need to choose just one though. The bot supports [mixing & matching models](./features.md#-mixing--matching-models), so you can use multiple providers at the same time. You don't need to choose just one though. The bot supports [mixing & matching models](./features.md#-mixing--matching-models), so you can use multiple providers at the same time.
@@ -171,3 +172,95 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- create a global agent: `!bai agent create-global together-ai my-together-ai-agent` - create a global agent: `!bai agent create-global together-ai my-together-ai-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/together-ai.yml). 💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/together-ai.yml).
### Venice
[Venice AI](https://venice.ai/chat?ref=kpXDe6) _(ref link with a $10 bonus for you)_ runs inference on Venice-controlled GPUs or zero-data-retention partner infrastructure and stores no prompts or responses, so your conversations don't linger anywhere. It serves both frontier proprietary models and the latest open-source ones.
- 🆔 Identifier: `venice`
- 🔗 Links: [🏠 Home page](https://venice.ai/chat?ref=kpXDe6), [👤 Sign up](https://venice.ai/chat?ref=kpXDe6), [📋 Models list](https://docs.venice.ai/models/overview)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation) (incl. editing, via the native knob-rich `/image/generate` and `/image/edit` endpoints), [💬 text-generation](./features.md#-text-generation) (incl. vision, file inputs like PDF and DOCX, and prompt caching; native web search via the `venice_parameters` config), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local venice my-venice-agent`
- create a global agent: `!bai agent create-global venice my-venice-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/venice.yml).
Unlike the [OpenAI Compatible](#openai-compatible) provider (which can talk to Venice but drops images and can't reach its audio or native image endpoints), this is a first-class Venice integration that exposes Venice's full parameter set. Image generation uses the native `/image/generate` endpoint rather than the OpenAI-compatible `/images/generations` shim, so every Venice-specific knob below is available.
#### Configuration reference
Every parameter below is optional unless marked otherwise. Omitting a knob lets Venice apply its own server-side default; this is **not** the same as setting it to `false`, which actively sends `false`.
**`text_generation`** (top-level knobs) — sampling, caching, and reasoning controls that sit directly on `text_generation`, next to `model_id`, `prompt`, `temperature`, `max_response_tokens`, and `max_context_tokens`. They map to top-level fields on Venice's request, separate from the `venice_parameters` bag below.
| Knob | What it does | Default |
|------|--------------|---------|
| `top_p` | Nucleus sampling, `0.0`–`1.0`. An alternative to `temperature`. | — |
| `frequency_penalty` | Penalize tokens by how often they have already appeared, `-2.0`–`2.0`. | — |
| `presence_penalty` | Penalize tokens that have appeared at all, `-2.0`–`2.0`. | — |
| `repetition_penalty` | Penalize repetition. Values above `1.0` discourage repeats. | — |
| `reasoning_effort` | Reasoning budget for models that support it: `low`, `medium`, `high`. | — |
| `prompt_cache_retention` | How long Venice keeps the prompt prefix cached: `default`, `extended`, or `24h`. `24h` is the lever that makes a long, stable system prompt cheap across a day of conversations. | `24h` |
| `show_reasoning` | Append the model's reasoning (its `reasoning_content`) below the answer, as a collapsible `💭 Reasoning` block that stays folded until clicked. Reads a field separate from the answer text, so it works regardless of `strip_thinking_response`. | `false` |
**`text_generation.venice_parameters`** — Venice-specific request knobs sent in the `venice_parameters` bag. Set any of them to override Venice's behavior. The `Default` column shows the value baibot's sample config ships; a `—` means the knob is left unset, so Venice's own default applies.
| Knob | What it does | Default |
|------|--------------|---------|
| `enable_web_search` | Web search mode: `auto` (model decides), `on` (always), or `off`. | `auto` |
| `enable_web_citations` | Append source citations to web-search answers. | — |
| `enable_web_scraping` | Allow the model to scrape page contents during web search. | — |
| `enable_x_search` | Include X (Twitter) in web search. | — |
| `include_search_results_in_stream` | Stream search results back as they arrive. | — |
| `return_search_results_as_documents` | Return search results as structured documents. | — |
| `include_venice_system_prompt` | Prepend Venice's own system prompt alongside yours. | — |
| `character_slug` | Use a public Venice character by its slug. | — |
| `strip_thinking_response` | Strip `<think></think>` blocks from reasoning models so the user sees only the answer. | `true` |
| `disable_thinking` | Disable the model's reasoning step entirely. | — |
| `enable_e2ee` | Run in end-to-end-encrypted mode rather than the default TEE-only mode. | `false` |
| `verbosity` | Response verbosity for models that support it: `low`, `medium`, `high`. | — |
**`text_to_speech`**:
| Knob | What it does | Default |
|------|--------------|---------|
| `model_id` | The Venice TTS model (e.g. `tts-kokoro`, `tts-qwen3-1-7b`, `tts-xai-v1`). | `tts-kokoro` |
| `voice` | The voice to synthesize with. Model-specific (Kokoro: `af_*`/`am_*`/`bf_*`/`bm_*`); a cloned-voice handle (`vv_<id>`) also works. | `af_sky` |
| `response_format` | Audio format: `mp3`, `opus`, `aac`, `flac`, `wav`, or `pcm`. | `mp3` |
| `speed` | Playback speed, `0.25`–`4.0`. | `1.0` |
| `prompt` | A style prompt steering emotion/delivery. Only Qwen 3 TTS honors it. | — |
| `temperature` | Sampling temperature, `0.0`–`2.0`. Only Qwen 3 / Orpheus / Chatterbox HD honor it. | — |
| `top_p` | Nucleus sampling, `0.0`–`1.0`. Only Qwen 3 TTS honors it. | — |
**`image_generation`**:
| Knob | What it does | Default |
|------|--------------|---------|
| `model_id` | The image-generation model. | `chroma` |
| `negative_prompt` | A description of what should **not** appear in the image. | — |
| `cfg_scale` | CFG scale, `0`–`20`. Higher values adhere more closely to the prompt. | — |
| `steps` | Number of inference steps. Model-specific; some models ignore it. | — |
| `style_preset` | A named style to apply (e.g. `3D Model`). | — |
| `seed` | Random seed, `-999999999`–`999999999`. Fix it for reproducible results. | random |
| `safe_mode` | Blur images classified as adult content. | `true` |
| `hide_watermark` | Hide the Venice watermark (may be ignored for some content). | `false` |
| `format` | Output format: `jpeg`, `png`, or `webp`. | `webp` |
| `width` / `height` | Image dimensions in pixels, each `1`–`1280`. | `1024` |
| `aspect_ratio` | Aspect ratio for models that support it (e.g. `1:1`, `16:9`). Alternative to `width`/`height`. | — |
| `resolution` | Resolution tier for models that support it (`1K`, `2K`, `4K`). | — |
| `quality` | Output quality for supported models: `low`, `medium`, `high`. Higher can cost more. | — |
| `lora_strength` | Lora strength, `0`–`100`. Only applies if the model uses additional Loras. | — |
| `embed_exif_metadata` | Embed the generation prompt into the image's EXIF metadata. | `false` |
| `enable_web_search` | Let the model pull the latest info from the web. Model-specific; costs extra credits. | — |
**`image_generation.edit`** — image editing reuses the `image_generation` block; only the model and a few output knobs differ:
| Knob | What it does | Default |
|------|--------------|---------|
| `model_id` | The image-edit model. | `firered-image-edit` |
| `output_format` | Output format: `jpeg`, `png`, or `webp`. When omitted, Venice infers it (PNG at 1K, JPEG at 2K/4K). | inferred |
| `aspect_ratio` | Aspect ratio of the result: `auto`, `1:1`, `3:2`, `16:9`, `21:9`, `9:16`, `2:3`, `3:4`, `4:5` (model-specific). | — |
| `resolution` | Resolution tier: `1K`, `2K`, `4K` (model-specific). | `1K` |
| `safe_mode` | Blur images classified as adult content. | `true` |

View File

@@ -21,7 +21,7 @@ text_to_speech:
speed: 1.0 speed: 1.0
response_format: opus response_format: opus
image_generation: image_generation:
model_id: gpt-image-1.5 model_id: gpt-image-2
style: null style: null
size: null size: null
quality: null quality: null

View File

@@ -0,0 +1,117 @@
base_url: https://api.venice.ai/api/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: kimi-k2-5
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 4096
max_context_tokens: 128000
# Prompt caching: how long Venice keeps the prompt prefix cached. "default", "extended", or "24h".
# "24h" (shipped by default) makes a long, stable system prompt cheap across a day of conversations.
prompt_cache_retention: 24h
# Top-level sampling and reasoning knobs (uncomment to override Venice's default):
# Nucleus sampling, 0.0-1.0 (an alternative to temperature).
# top_p: 0.9
# Penalize tokens by how often they have already appeared, -2.0-2.0.
# frequency_penalty: 0.0
# Penalize tokens that have appeared at all, -2.0-2.0.
# presence_penalty: 0.0
# Penalize repetition; values above 1.0 discourage repeats.
# repetition_penalty: 1.0
# Reasoning budget for models that support it: low, medium, high.
# reasoning_effort: medium
# Append the model's reasoning below the answer as a collapsible "💭 Reasoning" block (folded by
# default). Reads a field separate from the answer text, so it works alongside
# strip_thinking_response (which only strips <think> blocks from the answer).
# show_reasoning: true
# Venice-specific request parameters. Only the keys present below are sent to Venice; omit a
# key to fall back to Venice's own default. Omitting a knob is NOT the same as setting it to
# `false` — `false` actively sends `false`.
venice_parameters:
# Web search: "auto" (model decides), "on" (always), or "off".
enable_web_search: "auto"
# Strip <think></think> blocks from reasoning models so the user sees only the answer.
strip_thinking_response: true
# Run in TEE-only mode instead of end-to-end encryption (works across all models).
enable_e2ee: false
# Other available knobs — uncomment to override Venice's default:
# enable_web_citations: true
# enable_web_scraping: true
# include_venice_system_prompt: false
# include_search_results_in_stream: true
# return_search_results_as_documents: true
# enable_x_search: true
# disable_thinking: true
# Response verbosity for models that support it: low, medium, high.
# verbosity: medium
# character_slug: public-character-id
speech_to_text:
model_id: nvidia/parakeet-tdt-0.6b-v3
text_to_speech:
# The Venice TTS model. Others include tts-qwen3-1-7b, tts-xai-v1,
# tts-elevenlabs-turbo-v2-5, tts-minimax-speech-02-hd. See the models list endpoint.
model_id: tts-kokoro
# The voice to synthesize with. Voices are model-specific: Kokoro uses af_*/am_*/bf_*/bm_*
# (e.g. af_sky, am_adam), other models have their own sets. You can also pass a cloned-voice
# handle (vv_<id>) created via Venice's voice-cloning API. An incompatible voice returns an error.
voice: af_sky
# Output audio format: mp3, opus, aac, flac, wav, or pcm. mp3 is the broadest Matrix-client fit.
response_format: mp3
# Other available knobs — uncomment to override Venice's default:
# Playback speed, 0.25–4.0 (1.0 is normal).
# speed: 1.0
# A style prompt steering emotion/delivery (e.g. "Excited and energetic."). Only Qwen 3 TTS uses it.
# prompt: "Calm and warm."
# Sampling temperature, 0.0–2.0 (higher = more varied). Only Qwen 3 / Orpheus / Chatterbox HD use it.
# temperature: 0.9
# Nucleus sampling, 0.0–1.0. Only Qwen 3 TTS uses it.
# top_p: 1.0
image_generation:
# The image-generation model. See the models list endpoint for the full set.
model_id: chroma
# The image-edit model, used when editing an existing image rather than generating a new one.
# Editing shares this same image_generation config block; only the model differs.
edit:
model_id: firered-image-edit
# Other edit knobs — uncomment to override Venice's default:
# Output format: jpeg, png, or webp. When omitted, Venice infers it (PNG at 1K, JPEG at 2K/4K).
# output_format: png
# Aspect ratio of the result: auto, 1:1, 3:2, 16:9, 21:9, 9:16, 2:3, 3:4, 4:5 (model-specific).
# aspect_ratio: auto
# Resolution tier: 1K, 2K, 4K (model-specific). Defaults to 1K.
# resolution: 1K
# Blur images classified as adult content. Defaults to true.
# safe_mode: true
# Other generation knobs — uncomment to override Venice's default. Omitting a knob is NOT the same
# as setting it: an omitted knob lets Venice apply its own default, a set value is sent verbatim.
# A description of what should NOT appear in the image.
# negative_prompt: "blurry, watermark, text"
# CFG scale, 0–20. Higher values make the image adhere more closely to the prompt.
# cfg_scale: 7.5
# Number of inference steps. Model-specific; some models ignore it.
# steps: 8
# A named style to apply (e.g. "3D Model"). See Venice's image-styles reference.
# style_preset: "3D Model"
# Random seed, -999999999–999999999. Fix it for reproducible results; omit for a random seed.
# seed: 123456789
# Blur images classified as adult content. Defaults to true.
# safe_mode: true
# Hide the Venice watermark. Venice may ignore this for certain generated content. Defaults to false.
# hide_watermark: false
# Output format: jpeg, png, or webp. webp is smallest; png is highest-quality. Defaults to webp.
# format: webp
# Image dimensions in pixels, each 1–1280. Default 1024×1024.
# width: 1024
# height: 1024
# Aspect ratio (used by certain models, e.g. Nano Banana): "1:1", "16:9". An alternative to width/height.
# aspect_ratio: "1:1"
# Resolution tier (used by certain models): "1K", "2K", "4K".
# resolution: "1K"
# Output quality for supported models (e.g. GPT Image 2): low, medium, high. Higher can cost more.
# quality: high
# Lora strength, 0–100. Only applies if the model uses additional Loras.
# lora_strength: 50
# Embed the generation prompt into the image's EXIF metadata. Defaults to false.
# embed_exif_metadata: false
# Let the model pull the latest info from the web for the image. Model-specific; costs extra credits.
# enable_web_search: false

View File

@@ -111,7 +111,7 @@ agents:
# speed: 1.0 # speed: 1.0
# response_format: opus # response_format: opus
# image_generation: # image_generation:
# model_id: gpt-image-1.5 # model_id: gpt-image-2
# style: null # style: null
# size: null # size: null
# quality: null # quality: null

View File

@@ -1,6 +1,6 @@
services: services:
continuwuity: continuwuity:
image: forgejo.ellis.link/continuwuation/continuwuity:v0.5.9 image: forgejo.ellis.link/continuwuation/continuwuity:v0.5.10
user: "${UID}:${GID}" user: "${UID}:${GID}"
restart: unless-stopped restart: unless-stopped
cap_drop: cap_drop:

View File

@@ -1,6 +1,6 @@
services: services:
element-web: element-web:
image: ghcr.io/element-hq/element-web:v1.12.17 image: ghcr.io/element-hq/element-web:v1.12.22
user: "${UID}:${GID}" user: "${UID}:${GID}"
restart: unless-stopped restart: unless-stopped
environment: environment:

View File

@@ -1,6 +1,6 @@
services: services:
ollama: ollama:
image: docker.io/ollama/ollama:0.23.2 image: docker.io/ollama/ollama:0.30.11
restart: unless-stopped restart: unless-stopped
ports: ports:
- "${SERVICE_OLLAMA_BIND_PORT_HTTP}:11434" - "${SERVICE_OLLAMA_BIND_PORT_HTTP}:11434"

View File

@@ -1,6 +1,6 @@
services: services:
postgres: postgres:
image: docker.io/postgres:18.3-alpine image: docker.io/postgres:18.4-alpine
user: ${UID}:${GID} user: ${UID}:${GID}
restart: unless-stopped restart: unless-stopped
environment: environment:
@@ -14,7 +14,7 @@ services:
- /etc/passwd:/etc/passwd:ro - /etc/passwd:/etc/passwd:ro
synapse: synapse:
image: ghcr.io/element-hq/synapse:v1.152.1 image: ghcr.io/element-hq/synapse:v1.155.0
user: "${UID}:${GID}" user: "${UID}:${GID}"
restart: unless-stopped restart: unless-stopped
entrypoint: python entrypoint: python

View File

@@ -1,5 +1,5 @@
[tools] [tools]
prek = "0.3.13" prek = "0.4.5"
[settings] [settings]
# Disable automatic trust prompts - we trust this config # Disable automatic trust prompts - we trust this config

View File

@@ -1,4 +1,4 @@
[toolchain] [toolchain]
channel = "1.95.0" channel = "1.96.0"
components = ["rustfmt", "clippy"] components = ["rustfmt", "clippy"]
profile = "default" profile = "default"

View File

@@ -109,6 +109,9 @@ fn create_controller_from_provider_and_json_value_config(
AgentProvider::TogetherAI => { AgentProvider::TogetherAI => {
provider::openai_compat::create_controller_from_yaml_value_config(agent_id, config) provider::openai_compat::create_controller_from_yaml_value_config(agent_id, config)
} }
AgentProvider::Venice => {
provider::venice::create_controller_from_yaml_value_config(agent_id, config)
}
} }
} }
@@ -150,5 +153,9 @@ pub fn default_config_for_provider(provider: &AgentProvider) -> serde_yaml_ng::V
let config = super::provider::togetherai::default_config(); let config = super::provider::togetherai::default_config();
serde_yaml_ng::to_value(config).expect("Failed to serialize config") serde_yaml_ng::to_value(config).expect("Failed to serialize config")
} }
AgentProvider::Venice => {
let config = super::provider::venice::default_config();
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
} }
} }

View File

@@ -15,7 +15,7 @@ use crate::agent::provider::{
}; };
use crate::conversation::llm::{ use crate::conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage, Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
MessageContent as LLMMessageContent, shorten_messages_list_to_context_size, MessageContent as LLMMessageContent, TokenEstimate, shorten_messages_list_to_context_size,
}; };
use crate::strings; use crate::strings;
@@ -130,7 +130,7 @@ impl ControllerTrait for Controller {
tracing::trace!("Shortening messages list to context size"); tracing::trace!("Shortening messages list to context size");
conversation_messages = shorten_messages_list_to_context_size( conversation_messages = shorten_messages_list_to_context_size(
&text_generation_config.model_id, TokenEstimate::Approximate,
&prompt_message, &prompt_message,
conversation_messages, conversation_messages,
Some(text_generation_config.max_response_tokens), Some(text_generation_config.max_response_tokens),

View File

@@ -61,6 +61,7 @@ pub enum ControllerType {
OpenAI(Box<super::openai::Controller>), OpenAI(Box<super::openai::Controller>),
OpenAICompat(Box<super::openai_compat::Controller>), OpenAICompat(Box<super::openai_compat::Controller>),
Anthropic(Box<super::anthropic::Controller>), Anthropic(Box<super::anthropic::Controller>),
Venice(Box<super::venice::Controller>),
} }
impl ControllerTrait for ControllerType { impl ControllerTrait for ControllerType {
@@ -69,6 +70,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.supports_purpose(purpose), ControllerType::OpenAI(controller) => controller.supports_purpose(purpose),
ControllerType::OpenAICompat(controller) => controller.supports_purpose(purpose), ControllerType::OpenAICompat(controller) => controller.supports_purpose(purpose),
ControllerType::Anthropic(controller) => controller.supports_purpose(purpose), ControllerType::Anthropic(controller) => controller.supports_purpose(purpose),
ControllerType::Venice(controller) => controller.supports_purpose(purpose),
} }
} }
@@ -77,6 +79,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.text_generation_model_id(), ControllerType::OpenAI(controller) => controller.text_generation_model_id(),
ControllerType::OpenAICompat(controller) => controller.text_generation_model_id(), ControllerType::OpenAICompat(controller) => controller.text_generation_model_id(),
ControllerType::Anthropic(controller) => controller.text_generation_model_id(), ControllerType::Anthropic(controller) => controller.text_generation_model_id(),
ControllerType::Venice(controller) => controller.text_generation_model_id(),
} }
} }
@@ -85,6 +88,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.text_generation_prompt(), ControllerType::OpenAI(controller) => controller.text_generation_prompt(),
ControllerType::OpenAICompat(controller) => controller.text_generation_prompt(), ControllerType::OpenAICompat(controller) => controller.text_generation_prompt(),
ControllerType::Anthropic(controller) => controller.text_generation_prompt(), ControllerType::Anthropic(controller) => controller.text_generation_prompt(),
ControllerType::Venice(controller) => controller.text_generation_prompt(),
} }
} }
@@ -93,6 +97,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.text_to_speech_voice(), ControllerType::OpenAI(controller) => controller.text_to_speech_voice(),
ControllerType::OpenAICompat(controller) => controller.text_to_speech_voice(), ControllerType::OpenAICompat(controller) => controller.text_to_speech_voice(),
ControllerType::Anthropic(controller) => controller.text_to_speech_voice(), ControllerType::Anthropic(controller) => controller.text_to_speech_voice(),
ControllerType::Venice(controller) => controller.text_to_speech_voice(),
} }
} }
@@ -101,6 +106,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.text_to_speech_speed(), ControllerType::OpenAI(controller) => controller.text_to_speech_speed(),
ControllerType::OpenAICompat(controller) => controller.text_to_speech_speed(), ControllerType::OpenAICompat(controller) => controller.text_to_speech_speed(),
ControllerType::Anthropic(controller) => controller.text_to_speech_speed(), ControllerType::Anthropic(controller) => controller.text_to_speech_speed(),
ControllerType::Venice(controller) => controller.text_to_speech_speed(),
} }
} }
@@ -109,6 +115,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.text_generation_temperature(), ControllerType::OpenAI(controller) => controller.text_generation_temperature(),
ControllerType::OpenAICompat(controller) => controller.text_generation_temperature(), ControllerType::OpenAICompat(controller) => controller.text_generation_temperature(),
ControllerType::Anthropic(controller) => controller.text_generation_temperature(), ControllerType::Anthropic(controller) => controller.text_generation_temperature(),
ControllerType::Venice(controller) => controller.text_generation_temperature(),
} }
} }
@@ -117,6 +124,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.ping().await, ControllerType::OpenAI(controller) => controller.ping().await,
ControllerType::OpenAICompat(controller) => controller.ping().await, ControllerType::OpenAICompat(controller) => controller.ping().await,
ControllerType::Anthropic(controller) => controller.ping().await, ControllerType::Anthropic(controller) => controller.ping().await,
ControllerType::Venice(controller) => controller.ping().await,
} }
} }
@@ -135,6 +143,9 @@ impl ControllerTrait for ControllerType {
ControllerType::Anthropic(controller) => { ControllerType::Anthropic(controller) => {
controller.generate_text(conversation, params).await controller.generate_text(conversation, params).await
} }
ControllerType::Venice(controller) => {
controller.generate_text(conversation, params).await
}
} }
} }
@@ -154,6 +165,9 @@ impl ControllerTrait for ControllerType {
ControllerType::Anthropic(controller) => { ControllerType::Anthropic(controller) => {
controller.speech_to_text(mime_type, media, params).await controller.speech_to_text(mime_type, media, params).await
} }
ControllerType::Venice(controller) => {
controller.speech_to_text(mime_type, media, params).await
}
} }
} }
@@ -170,6 +184,7 @@ impl ControllerTrait for ControllerType {
ControllerType::Anthropic(controller) => { ControllerType::Anthropic(controller) => {
controller.generate_image(prompt, params).await controller.generate_image(prompt, params).await
} }
ControllerType::Venice(controller) => controller.generate_image(prompt, params).await,
} }
} }
@@ -189,6 +204,9 @@ impl ControllerTrait for ControllerType {
ControllerType::Anthropic(controller) => { ControllerType::Anthropic(controller) => {
controller.create_image_edit(prompt, images, params).await controller.create_image_edit(prompt, images, params).await
} }
ControllerType::Venice(controller) => {
controller.create_image_edit(prompt, images, params).await
}
} }
} }
@@ -203,6 +221,7 @@ impl ControllerTrait for ControllerType {
controller.text_to_speech(text, params).await controller.text_to_speech(text, params).await
} }
ControllerType::Anthropic(controller) => controller.text_to_speech(text, params).await, ControllerType::Anthropic(controller) => controller.text_to_speech(text, params).await,
ControllerType::Venice(controller) => controller.text_to_speech(text, params).await,
} }
} }
} }

View File

@@ -11,6 +11,7 @@ pub enum AgentProvider {
OpenAICompat, OpenAICompat,
OpenRouter, OpenRouter,
TogetherAI, TogetherAI,
Venice,
} }
impl AgentProvider { impl AgentProvider {
@@ -25,6 +26,7 @@ impl AgentProvider {
&Self::OpenAICompat, &Self::OpenAICompat,
&Self::OpenRouter, &Self::OpenRouter,
&Self::TogetherAI, &Self::TogetherAI,
&Self::Venice,
] ]
} }
@@ -39,6 +41,7 @@ impl AgentProvider {
Self::OpenAICompat => "openai-compatible", Self::OpenAICompat => "openai-compatible",
Self::OpenRouter => "openrouter", Self::OpenRouter => "openrouter",
Self::TogetherAI => "together-ai", Self::TogetherAI => "together-ai",
Self::Venice => "venice",
} }
} }
@@ -53,6 +56,7 @@ impl AgentProvider {
"openai-compatible" => Ok(Self::OpenAICompat), "openai-compatible" => Ok(Self::OpenAICompat),
"openrouter" => Ok(Self::OpenRouter), "openrouter" => Ok(Self::OpenRouter),
"together-ai" => Ok(Self::TogetherAI), "together-ai" => Ok(Self::TogetherAI),
"venice" => Ok(Self::Venice),
_ => Err("Unexpected string value"), _ => Err("Unexpected string value"),
} }
} }
@@ -181,6 +185,25 @@ impl AgentProvider {
text_generation_supports_vision: false, text_generation_supports_vision: false,
text_generation_supports_tools: false, text_generation_supports_tools: false,
}, },
Self::Venice => AgentProviderInfo {
id: Self::Venice.to_static_str(),
name: "Venice",
description: "Venice AI runs inference on Venice-controlled GPUs or zero-data-retention partner infrastructure and stores no prompts or responses. It serves frontier proprietary and open-source models with text-generation (including vision), speech-to-text, text-to-speech, native image generation and editing, and native web search.",
homepage_url: Some("https://venice.ai"),
wiki_url: None,
sign_up_url: Some("https://venice.ai"),
models_list_url: Some("https://api.venice.ai/api/v1/models"),
supported_purposes: vec![
AgentPurpose::ImageGeneration,
AgentPurpose::TextGeneration,
AgentPurpose::TextToSpeech,
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: true,
// Venice does native web search via `venice_parameters`, NOT baibot's built-in
// tools mechanism (the OpenAI web_search/code_interpreter block), so this is false.
text_generation_supports_tools: false,
},
} }
} }
} }

View File

@@ -1,6 +1,7 @@
use chrono::{DateTime, Utc}; use chrono::{DateTime, Utc};
use std::collections::HashMap; use std::collections::HashMap;
#[derive(Clone)]
pub struct TextGenerationPromptVariables { pub struct TextGenerationPromptVariables {
map: HashMap<String, String>, map: HashMap<String, String>,
} }

View File

@@ -10,6 +10,7 @@ pub mod openai;
pub mod openai_compat; pub mod openai_compat;
pub(super) mod openrouter; pub(super) mod openrouter;
pub(super) mod togetherai; pub(super) mod togetherai;
pub mod venice;
fn default_temperature() -> f32 { fn default_temperature() -> f32 {
1.0 1.0

View File

@@ -1,6 +1,6 @@
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
use super::OPENAI_IMAGE_MODEL_GPT_IMAGE_1_DOT_5; use super::OPENAI_IMAGE_MODEL_GPT_IMAGE_2;
use crate::agent::{default_prompt, provider::ConfigTrait}; use crate::agent::{default_prompt, provider::ConfigTrait};
#[derive(Debug, Clone, Serialize, Deserialize)] #[derive(Debug, Clone, Serialize, Deserialize)]
@@ -175,7 +175,7 @@ pub struct ImageGenerationConfig {
impl Default for ImageGenerationConfig { impl Default for ImageGenerationConfig {
fn default() -> Self { fn default() -> Self {
Self { Self {
model_id: OPENAI_IMAGE_MODEL_GPT_IMAGE_1_DOT_5.to_owned(), model_id: OPENAI_IMAGE_MODEL_GPT_IMAGE_2.to_owned(),
style: default_image_style(), style: default_image_style(),
size: default_image_size(), size: default_image_size(),
quality: default_image_quality(), quality: default_image_quality(),
@@ -193,6 +193,7 @@ impl ImageGenerationConfig {
"gpt-image-1" => Ok(async_openai::types::images::ImageModel::GptImage1), "gpt-image-1" => Ok(async_openai::types::images::ImageModel::GptImage1),
"gpt-image-1.5" => Ok(async_openai::types::images::ImageModel::GptImage1dot5), "gpt-image-1.5" => Ok(async_openai::types::images::ImageModel::GptImage1dot5),
"gpt-image-1-mini" => Ok(async_openai::types::images::ImageModel::GptImage1Mini), "gpt-image-1-mini" => Ok(async_openai::types::images::ImageModel::GptImage1Mini),
"gpt-image-2" => Ok(async_openai::types::images::ImageModel::GptImage2),
other => Ok(async_openai::types::images::ImageModel::Other( other => Ok(async_openai::types::images::ImageModel::Other(
other.to_owned(), other.to_owned(),
)), )),

View File

@@ -24,7 +24,7 @@ use crate::{
}, },
conversation::llm::{ conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage, Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
MessageContent as LLMMessageContent, shorten_messages_list_to_context_size, MessageContent as LLMMessageContent, TokenEstimate, shorten_messages_list_to_context_size,
}, },
utils::base64::base64_decode, utils::base64::base64_decode,
}; };
@@ -117,7 +117,7 @@ impl ControllerTrait for Controller {
tracing::trace!("Shortening messages list to context size"); tracing::trace!("Shortening messages list to context size");
conversation_messages = shorten_messages_list_to_context_size( conversation_messages = shorten_messages_list_to_context_size(
&text_generation_config.model_id, TokenEstimate::Tiktoken(&text_generation_config.model_id),
&prompt_message, &prompt_message,
conversation_messages, conversation_messages,
text_generation_config.max_response_tokens, text_generation_config.max_response_tokens,
@@ -271,6 +271,7 @@ impl ControllerTrait for Controller {
ImageModel::GptImage1 => ImageModel::GptImage1Mini, ImageModel::GptImage1 => ImageModel::GptImage1Mini,
ImageModel::GptImage1dot5 => ImageModel::GptImage1Mini, ImageModel::GptImage1dot5 => ImageModel::GptImage1Mini,
ImageModel::GptImage1Mini => ImageModel::GptImage1Mini, ImageModel::GptImage1Mini => ImageModel::GptImage1Mini,
ImageModel::GptImage2 => ImageModel::GptImage1Mini,
ImageModel::Other(_) => ImageModel::DallE2, ImageModel::Other(_) => ImageModel::DallE2,
} }
} else { } else {
@@ -310,7 +311,7 @@ impl ControllerTrait for Controller {
let size = if params.smallest_size_possible { let size = if params.smallest_size_possible {
Some(get_sticker_size(&model)) Some(get_sticker_size(&model))
} else { } else {
image_generation_config.size image_generation_config.size.clone()
}; };
let response_format = match model.clone() { let response_format = match model.clone() {
@@ -321,6 +322,7 @@ impl ControllerTrait for Controller {
ImageModel::GptImage1 => None, ImageModel::GptImage1 => None,
ImageModel::GptImage1Mini => None, ImageModel::GptImage1Mini => None,
ImageModel::GptImage1dot5 => None, ImageModel::GptImage1dot5 => None,
ImageModel::GptImage2 => None,
ImageModel::Other(_) => Some(ImageResponseFormat::B64Json), ImageModel::Other(_) => Some(ImageResponseFormat::B64Json),
}; };
@@ -430,6 +432,7 @@ impl ControllerTrait for Controller {
ImageModel::GptImage1 => None, ImageModel::GptImage1 => None,
ImageModel::GptImage1Mini => None, ImageModel::GptImage1Mini => None,
ImageModel::GptImage1dot5 => None, ImageModel::GptImage1dot5 => None,
ImageModel::GptImage2 => None,
ImageModel::Other(_) => Some(ImageResponseFormat::B64Json), ImageModel::Other(_) => Some(ImageResponseFormat::B64Json),
}; };
@@ -647,6 +650,7 @@ fn get_sticker_size(model: &ImageModel) -> async_openai::types::images::ImageSiz
ImageModel::GptImage1 => ImageSize::S1024x1024, ImageModel::GptImage1 => ImageSize::S1024x1024,
ImageModel::GptImage1Mini => ImageSize::S1024x1024, ImageModel::GptImage1Mini => ImageSize::S1024x1024,
ImageModel::GptImage1dot5 => ImageSize::S1024x1024, ImageModel::GptImage1dot5 => ImageSize::S1024x1024,
ImageModel::GptImage2 => ImageSize::S1024x1024,
ImageModel::Other(_) => ImageSize::S1024x1024, ImageModel::Other(_) => ImageSize::S1024x1024,
} }
} }

View File

@@ -16,7 +16,7 @@ use super::super::AgentInstantiationResult;
use super::ConfigTrait; use super::ConfigTrait;
use super::controller::ControllerType; use super::controller::ControllerType;
pub const OPENAI_IMAGE_MODEL_GPT_IMAGE_1_DOT_5: &str = "gpt-image-1.5"; pub const OPENAI_IMAGE_MODEL_GPT_IMAGE_2: &str = "gpt-image-2";
pub fn create_controller_from_yaml_value_config( pub fn create_controller_from_yaml_value_config(
agent_id: &str, agent_id: &str,

View File

@@ -15,7 +15,7 @@ use crate::{
}, },
conversation::llm::{ conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage, Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
MessageContent as LLMMessageContent, shorten_messages_list_to_context_size, MessageContent as LLMMessageContent, TokenEstimate, shorten_messages_list_to_context_size,
}, },
}; };
use crate::{ use crate::{
@@ -114,7 +114,7 @@ impl ControllerTrait for Controller {
tracing::trace!("Shortening messages list to context size"); tracing::trace!("Shortening messages list to context size");
conversation_messages = shorten_messages_list_to_context_size( conversation_messages = shorten_messages_list_to_context_size(
&text_generation_config.model_id, TokenEstimate::Approximate,
&prompt_message, &prompt_message,
conversation_messages, conversation_messages,
text_generation_config.max_response_tokens, text_generation_config.max_response_tokens,

View File

@@ -0,0 +1,156 @@
use crate::agent::AgentPurpose;
use crate::agent::provider::entity::{TextToSpeechParams, TextToSpeechResult};
use crate::agent::provider::{SpeechToTextParams, SpeechToTextResult};
use crate::strings;
use super::config::Config;
use super::wire::{SpeechRequest, TranscriptionResponse};
pub async fn speech_to_text(
config: &Config,
http: &reqwest::Client,
mime_type: &mxlink::mime::Mime,
media: Vec<u8>,
params: SpeechToTextParams,
) -> anyhow::Result<SpeechToTextResult> {
let Some(speech_to_text_config) = &config.speech_to_text else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::SpeechToText
),
));
};
// Unlike the openai_compat path (which writes the audio to a temp file because its library
// can't take bytes), reqwest's multipart takes the bytes directly.
let part = reqwest::multipart::Part::bytes(media)
.file_name("audio")
.mime_str(mime_type.as_ref())?;
let mut form = reqwest::multipart::Form::new()
.part("file", part)
.text("model", speech_to_text_config.model_id.clone())
.text("response_format", "json");
if let Some(language) = &params.language_override {
form = form.text("language", language.clone());
}
let url = format!(
"{}/audio/transcriptions",
config.base_url.trim_end_matches('/')
);
tracing::trace!(
model_id = speech_to_text_config.model_id,
language = ?params.language_override,
"Sending Venice audio transcription API request"
);
let response = http
.post(&url)
.bearer_auth(&config.api_key)
.multipart(form)
.send()
.await?;
let status = response.status();
if !status.is_success() {
// Body to the server log only, not into the returned error (which reaches the Matrix room).
let body = response.text().await.unwrap_or_default();
tracing::warn!(%status, body, "Venice audio transcription request failed");
return Err(anyhow::anyhow!(
"Venice audio transcription request failed with status {status}"
));
}
let response: TranscriptionResponse = response.json().await?;
Ok(SpeechToTextResult {
text: response.text,
})
}
pub async fn text_to_speech(
config: &Config,
http: &reqwest::Client,
input: &str,
params: TextToSpeechParams,
) -> anyhow::Result<TextToSpeechResult> {
let Some(text_to_speech_config) = &config.text_to_speech else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::TextToSpeech
),
));
};
// Per-call overrides win over the configured defaults.
let voice = params
.voice_override
.or_else(|| text_to_speech_config.voice.clone());
let speed = params.speed_override.or(text_to_speech_config.speed);
let response_format = text_to_speech_config.response_format.clone();
let mime_type = response_format_to_mime_type(response_format.as_deref());
let request = SpeechRequest {
model: text_to_speech_config.model_id.clone(),
input: input.to_owned(),
voice,
speed,
response_format,
prompt: text_to_speech_config.prompt.clone(),
temperature: text_to_speech_config.temperature,
top_p: text_to_speech_config.top_p,
};
let url = format!("{}/audio/speech", config.base_url.trim_end_matches('/'));
tracing::trace!(
model_id = text_to_speech_config.model_id,
voice = ?request.voice,
"Sending Venice text-to-speech API request"
);
let response = http
.post(&url)
.bearer_auth(&config.api_key)
.json(&request)
.send()
.await?;
let status = response.status();
if !status.is_success() {
// Body to the server log only, not into the returned error (which reaches the Matrix room).
let body = response.text().await.unwrap_or_default();
tracing::warn!(%status, body, "Venice text-to-speech request failed");
return Err(anyhow::anyhow!(
"Venice text-to-speech request failed with status {status}"
));
}
// The speech endpoint answers with raw binary audio; read the body directly.
let bytes = response.bytes().await?.to_vec();
Ok(TextToSpeechResult { bytes, mime_type })
}
/// Map a Venice TTS `response_format` to its MIME type. Defaults to `audio/mpeg` (the
/// IANA-registered MP3 type, RFC 3003) when the format is unset, matching Venice's own `mp3`
/// default. This deliberately uses `audio/mpeg` rather than the `audio/mp3` alias the openai
/// provider emits; baibot's downstream audio-filename mapping treats both as `.mp3`.
fn response_format_to_mime_type(response_format: Option<&str>) -> mxlink::mime::Mime {
let raw = match response_format.unwrap_or("mp3") {
"mp3" => "audio/mpeg",
"opus" => "audio/ogg",
"aac" => "audio/aac",
"flac" => "audio/flac",
"wav" => "audio/wav",
"pcm" => "audio/L8",
_ => "audio/mpeg",
};
raw.parse()
.unwrap_or(mxlink::mime::APPLICATION_OCTET_STREAM)
}

View File

@@ -0,0 +1,351 @@
use std::collections::hash_map::DefaultHasher;
use std::hash::{Hash, Hasher};
use std::sync::OnceLock;
use regex::Regex;
use crate::agent::AgentPurpose;
use crate::agent::provider::entity::{TextGenerationParams, TextGenerationResult};
use crate::conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
MessageContent as LLMMessageContent, TokenEstimate, shorten_messages_list_to_context_size,
};
use crate::strings;
use super::config::{Config, WebSearchMode};
use super::utils::convert_llm_messages_to_venice;
use super::wire::{ChatCompletionRequest, ChatCompletionResponse, WebSearchCitation};
pub async fn generate_text(
config: &Config,
http: &reqwest::Client,
unsupported: &super::recovery::UnsupportedFieldsCache,
conversation: LLMConversation,
params: TextGenerationParams,
) -> anyhow::Result<TextGenerationResult> {
let Some(text_generation_config) = &config.text_generation else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::TextGeneration
),
));
};
let prompt_text = params.prompt_variables.format(
params
.prompt_override
.unwrap_or(text_generation_config.prompt.clone().unwrap_or_default())
.trim(),
);
// Prompt-cache routing key. Hash ONLY conversation-stable inputs: the rendered system prompt
// and the conversation start time. Folding in anything per-turn (message content, the current
// time, the message count) would mint a fresh key every turn, miss the cache every lookup, and
// pay full price plus the hashing cost. The start time is rendered explicitly here so the key
// stays stable even when the user's prompt template never mentions the time variable; an
// unknown start time renders "unknown" and simply keys on the prompt alone.
let conversation_start_time = params
.prompt_variables
.format("{{ baibot_conversation_start_time_utc }}");
let prompt_cache_key = derive_prompt_cache_key(&prompt_text, &conversation_start_time);
let prompt_message = if prompt_text.is_empty() {
None
} else {
Some(LLMMessage {
author: LLMAuthor::Prompt,
sender_id: None,
content: LLMMessageContent::Text(prompt_text),
timestamp: chrono::Utc::now(),
})
};
let mut conversation_messages = conversation.messages;
if params.context_management_enabled {
conversation_messages = shorten_messages_list_to_context_size(
TokenEstimate::Approximate,
&prompt_message,
conversation_messages,
text_generation_config.max_response_tokens,
text_generation_config.max_context_tokens,
);
}
if let Some(prompt_message) = prompt_message {
conversation_messages.insert(0, prompt_message);
}
let messages = convert_llm_messages_to_venice(conversation_messages)?;
let temperature = params
.temperature_override
.unwrap_or(text_generation_config.temperature);
// When web search is active, ask Venice to return structured search results so we can render
// readable citations from them. Respect an explicit user choice and only fill the flag when
// the user left it unset.
let venice_parameters = text_generation_config
.venice_parameters
.clone()
.map(|mut vp| {
let web_search_active = matches!(
vp.enable_web_search,
Some(WebSearchMode::On | WebSearchMode::Auto)
);
if web_search_active && vp.return_search_results_as_documents.is_none() {
vp.return_search_results_as_documents = Some(true);
}
vp
});
let mut request = ChatCompletionRequest {
model: text_generation_config.model_id.clone(),
messages,
temperature: Some(temperature),
// Web search rides entirely inside `venice_parameters`; there is no `tools` array here.
// `max_tokens` is deprecated on Venice in favor of `max_completion_tokens`.
max_completion_tokens: text_generation_config.max_response_tokens,
top_p: text_generation_config.top_p,
frequency_penalty: text_generation_config.frequency_penalty,
presence_penalty: text_generation_config.presence_penalty,
repetition_penalty: text_generation_config.repetition_penalty,
reasoning_effort: text_generation_config.reasoning_effort.clone(),
prompt_cache_key: Some(prompt_cache_key),
prompt_cache_retention: text_generation_config.prompt_cache_retention.clone(),
venice_parameters,
};
let url = format!("{}/chat/completions", config.base_url.trim_end_matches('/'));
let model_id = text_generation_config.model_id.clone();
// Proactively drop fields this model has already rejected earlier in this process, so a known
// mismatch costs zero wasted round-trips after the first discovery. Venice's body is
// `additionalProperties: false`, so sending a known-unsupported field would 400 again.
for field in unsupported.known_for(&model_id) {
super::recovery::strip_droppable_field(&mut request, &field);
}
// Send with bounded auto-recovery. When Venice 400s because a model does not support an optional
// knob, it names the field (`field: '...'`); if that field is one we may safely drop, we strip
// it, remember the rejection for this model, and retry. The loop is bounded: each retry clears a
// distinct droppable field (strip returns false once it is gone), so after at most
// `DROPPABLE_FIELDS.len()` retries the request either succeeds or surfaces the error.
let response = loop {
tracing::trace!(
model = model_id,
messages_count = request.messages.len(),
"Sending Venice chat completion API request"
);
let response = http
.post(&url)
.bearer_auth(&config.api_key)
.json(&request)
.send()
.await?;
let status = response.status();
if status.is_success() {
break response;
}
// Always log the body server-side: Venice explains a rejected strict body there.
let body = response.text().await.unwrap_or_default();
tracing::warn!(%status, body, "Venice chat completion request failed");
// Recover from a strict-body 400 over an unsupported optional knob: strip the named field
// and retry. Only fields in `DROPPABLE_FIELDS` are eligible, so a meaning-bearing knob (a
// sampling parameter) is never silently dropped; that case falls through to surface below.
if status == reqwest::StatusCode::BAD_REQUEST
&& let Some(field) = super::recovery::parse_rejected_field(&body)
&& super::recovery::strip_droppable_field(&mut request, &field)
{
unsupported.record(&model_id, &field);
tracing::info!(
model = model_id,
field,
"Venice rejected an unsupported field; dropping it and retrying"
);
continue;
}
// A 413 almost always means an attached file pushed the request past Venice's size limit.
if status == reqwest::StatusCode::PAYLOAD_TOO_LARGE {
return Err(anyhow::anyhow!(
"The request was too large for Venice, most likely an attached file over the 25MB limit."
));
}
// A 400 is a complaint about the request baibot built, so the body is safe and useful to
// surface: it tells the operator (e.g. at agent-create time) exactly which field or value
// Venice rejected, instead of an opaque status. Other statuses keep the body OUT of the
// returned error, since it can carry account / rate-limit details that shouldn't reach the
// room.
if status == reqwest::StatusCode::BAD_REQUEST {
return Err(anyhow::anyhow!(
"Venice rejected the request (400 Bad Request): {}",
super::recovery::extract_error_message(&body)
));
}
return Err(anyhow::anyhow!(
"Venice chat completion request failed with status {status}"
));
};
let response: ChatCompletionResponse = response.json().await?;
let citations = response
.venice_parameters
.map(|vp| vp.web_search_citations)
.unwrap_or_default();
let Some(choice) = response.choices.into_iter().next() else {
return Err(anyhow::anyhow!(
"No choices were returned from the Venice chat completion API"
));
};
let Some(content) = choice.message.content else {
return Err(anyhow::anyhow!(
"No message content was returned from the Venice chat completion API"
));
};
let text = render_with_citations(content, &citations);
let text = append_reasoning(
text,
choice.message.reasoning_content,
text_generation_config.show_reasoning,
);
Ok(TextGenerationResult { text })
}
/// Builds the prompt-cache routing key from conversation-stable inputs. `DefaultHasher::new()` is a
/// fixed-seed SipHasher (keys 0,0), so it is deterministic across processes and restarts: identical
/// inputs always produce the same key, which is what lets a restarted bot keep hitting the warm
/// cache. The algorithm is not guaranteed stable across Rust std versions, so a rebuild on a new
/// toolchain can shift every key once, a one-time cache warm-up with no correctness effect.
pub(super) fn derive_prompt_cache_key(prompt_text: &str, conversation_start_time: &str) -> String {
let mut hasher = DefaultHasher::new();
prompt_text.hash(&mut hasher);
conversation_start_time.hash(&mut hasher);
format!("{:016x}", hasher.finish())
}
/// Appends the model's thinking to the reply only when the deployment opts in via `show_reasoning`.
/// `reasoning_content` is a field separate from the answer `content` (it is unaffected by
/// `strip_thinking_response`, which only strips inline `<think>` blocks from `content`), so reading
/// it here is independent of that knob. Default-off matches today's behavior: thinking never reaches
/// a room that did not ask for it.
///
/// The thinking renders as a Matrix-native collapsible `<details>` block: folded by default, one
/// click to expand, so it stays out of the way of the answer instead of dumping a wall of reasoning
/// inline. This survives the send path: the reply goes through markdown (`send_text_markdown`),
/// whose pulldown-cmark pass writes raw HTML verbatim rather than escaping it, and ruma's HTML
/// sanitizer allow-lists `<details>`/`<summary>`. Clients that do not render `<details>` degrade to
/// showing the summary and reasoning inline, so nothing is lost there either.
pub(super) fn append_reasoning(
text: String,
reasoning_content: Option<String>,
show_reasoning: bool,
) -> String {
if !show_reasoning {
return text;
}
match reasoning_content {
Some(reasoning) if !reasoning.trim().is_empty() => {
// The blank lines around the trimmed reasoning keep it a separate markdown block from
// the surrounding `<details>`/`</details>` HTML blocks, so the reasoning itself still
// renders as markdown (lists, code, emphasis) inside the collapsible.
let reasoning = reasoning.trim();
format!(
"{text}\n\n<details><summary>💭 Reasoning</summary>\n\n{reasoning}\n\n</details>"
)
}
_ => text,
}
}
/// Rewrites Venice's inline `^n^` citation superscripts into readable `[n]` references and appends
/// a `Sources:` list of markdown links, one per citation in order. Returns the content unchanged
/// when web search returned no citations, so non-search replies are never touched.
///
/// Citation `title` and `url` come from scraped web pages, so they are attacker-influenced. The
/// title is escaped so it cannot break out of the markdown link label, and the URL is used as a
/// link target only when it is a clean `http(s)` URL with no markdown-breaking characters;
/// otherwise the citation renders as plain text. This stops a hostile page title or URL from
/// injecting a spoofed clickable link into the room.
pub(super) fn render_with_citations(content: String, citations: &[WebSearchCitation]) -> String {
if citations.is_empty() {
return content;
}
let mut text = rewrite_citation_superscripts(&content);
let mut sources = String::from("\n\nSources:");
for (index, citation) in citations.iter().enumerate() {
let n = index + 1;
let title = escape_markdown_link_text(&citation.title);
match sanitize_link_url(&citation.url) {
// A citation that arrived with no title still renders as a usable link by showing the
// URL as the link text, rather than an empty `[]( )` label.
Some(url) if title.is_empty() => sources.push_str(&format!("\n[{n}] [{url}]({url})")),
Some(url) => sources.push_str(&format!("\n[{n}] [{title}]({url})")),
None if !title.is_empty() => sources.push_str(&format!("\n[{n}] {title}")),
None => sources.push_str(&format!("\n[{n}] (source unavailable)")),
}
}
text.push_str(&sources);
text
}
/// Venice marks web-search citations with superscript runs in the reply text: a single `^1^`, a
/// comma list `^1,2^`, or a caret-chained run `^2^3^10^` where consecutive citations share a
/// caret. The whole run has to be matched at once: a per-citation pattern (string or regex)
/// consumes the shared caret on the first match and orphans the rest (`^2^3^` would leave `3^`).
/// So this matches each full run and expands it to one `[n]` per citation (`^2^3^` -> `[2][3]`).
fn rewrite_citation_superscripts(content: &str) -> String {
static RUN: OnceLock<Regex> = OnceLock::new();
let run = RUN.get_or_init(|| {
Regex::new(r"\^\d+(?:[,^]\d+)*\^").expect("citation superscript regex is valid")
});
run.replace_all(content, |caps: &regex::Captures| {
caps[0]
.split(['^', ','])
.filter(|piece| !piece.is_empty())
.map(|n| format!("[{n}]"))
.collect::<String>()
})
.into_owned()
}
/// Escapes the characters that would let citation title text break out of a markdown link label,
/// and folds newlines to spaces so a multi-line title cannot inject extra markdown structure.
fn escape_markdown_link_text(text: &str) -> String {
text.replace('\\', "\\\\")
.replace('[', "\\[")
.replace(']', "\\]")
.replace(['\r', '\n'], " ")
}
/// Returns the URL as a markdown link target only when it is a clean `http(s)` URL with no
/// characters that would break the `(...)` destination or smuggle a different scheme. Anything else
/// returns `None`, so the caller renders the citation as plain text instead of a link.
fn sanitize_link_url(url: &str) -> Option<String> {
let url = url.trim();
let is_http = url.starts_with("https://") || url.starts_with("http://");
let is_clean = !url.contains(['(', ')', '<', '>', ' ', '\t', '\r', '\n']);
if is_http && is_clean {
Some(url.to_owned())
} else {
None
}
}

View File

@@ -0,0 +1,410 @@
use serde::{Deserialize, Serialize};
use crate::agent::{default_prompt, provider::ConfigTrait};
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct Config {
pub base_url: String,
pub api_key: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub text_generation: Option<TextGenerationConfig>,
#[serde(skip_serializing_if = "Option::is_none")]
pub speech_to_text: Option<SpeechToTextConfig>,
#[serde(skip_serializing_if = "Option::is_none")]
pub text_to_speech: Option<TextToSpeechConfig>,
#[serde(skip_serializing_if = "Option::is_none")]
pub image_generation: Option<ImageGenerationConfig>,
}
impl Default for Config {
fn default() -> Self {
Self {
base_url: "https://api.venice.ai/api/v1".to_owned(),
api_key: "YOUR_API_KEY_HERE".to_owned(),
text_generation: Some(TextGenerationConfig::default()),
speech_to_text: Some(SpeechToTextConfig::default()),
text_to_speech: Some(TextToSpeechConfig::default()),
image_generation: Some(ImageGenerationConfig::default()),
}
}
}
impl ConfigTrait for Config {
fn validate(&self) -> Result<(), String> {
if self.base_url.is_empty() {
return Err("The base URL must not be empty.".to_owned());
}
Ok(())
}
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct TextGenerationConfig {
#[serde(default = "default_text_model_id")]
pub model_id: String,
#[serde(default)]
pub prompt: Option<String>,
#[serde(default = "super::super::default_temperature")]
pub temperature: f32,
#[serde(default)]
pub max_response_tokens: Option<u32>,
#[serde(default)]
pub max_context_tokens: u32,
/// Sampling and reasoning knobs that live at the top level of Venice's `/chat/completions`
/// body, not inside the `venice_parameters` bag. Venice silently ignores a top-level knob
/// placed in the bag, so these sit here as siblings and map straight to top-level wire fields
/// in `chat.rs`. Each is omitted from the request when unset.
#[serde(default)]
pub top_p: Option<f32>,
#[serde(default)]
pub frequency_penalty: Option<f32>,
#[serde(default)]
pub presence_penalty: Option<f32>,
#[serde(default)]
pub repetition_penalty: Option<f32>,
#[serde(default)]
pub reasoning_effort: Option<String>,
/// Prompt-cache retention window (`default`, `extended`, or `24h`). This carries a named
/// default rather than a bare `#[serde(default)]` (which would yield `None`), so a config that
/// omits the key still ships `24h` and keeps caching on. Caching is the per-deployment cost
/// lever, so the omitted-key case must not silently disable it. The value here must agree with
/// the `Default` impl below.
#[serde(default = "default_prompt_cache_retention")]
pub prompt_cache_retention: Option<String>,
/// When set, the model's `reasoning_content` (its thinking) is appended to the reply. Off by
/// default to match today's `strip_thinking_response: true` behavior, so existing deployments
/// see no change.
#[serde(default)]
pub show_reasoning: bool,
/// Venice-specific request knobs, serialized 1:1 into the `venice_parameters` bag on the
/// wire. Any unset field is omitted, so Venice applies its own server-side default.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub venice_parameters: Option<VeniceParameters>,
}
impl Default for TextGenerationConfig {
fn default() -> Self {
Self {
model_id: default_text_model_id(),
prompt: Some(default_prompt().to_owned()),
temperature: super::super::default_temperature(),
// Reserved output budget: sent as the response cap AND subtracted from the context
// window when trimming history. Mirrors the openai_compat sibling's default.
max_response_tokens: Some(4096),
// Matches Venice's own `availableContextTokens` (131072) and the non-OpenAI sibling
// providers (ollama/localai/mistral all default to 128_000).
max_context_tokens: 128_000,
// Sampling knobs stay None so Venice applies its own server-side default. Caching is
// the one exception: retention defaults to 24h here so a programmatic default caches
// out of the box, agreeing with the `#[serde(default = ...)]` on the field.
top_p: None,
frequency_penalty: None,
presence_penalty: None,
repetition_penalty: None,
reasoning_effort: None,
prompt_cache_retention: default_prompt_cache_retention(),
show_reasoning: false,
// A usable starting point, not an everything-set dump: only these three are sent;
// every other knob stays None so Venice applies its own default (omitting != false).
venice_parameters: Some(VeniceParameters {
enable_web_search: Some(WebSearchMode::Auto),
strip_thinking_response: Some(true),
enable_e2ee: Some(false),
..Default::default()
}),
}
}
}
fn default_text_model_id() -> String {
"kimi-k2-5".to_owned()
}
/// Defaults prompt-cache retention to 24h so caching is on unless a config explicitly opts out.
/// A bare `#[serde(default)]` would deserialize an omitted key to `None`, which disables caching;
/// this keeps the cost lever engaged for configs that never mention it.
fn default_prompt_cache_retention() -> Option<String> {
Some("24h".to_owned())
}
/// The full `venice_parameters` knob set, mirroring Venice's `ChatCompletionRequest`
/// schema field-for-field. Every field is optional with `skip_serializing_if`, so the
/// request never carries a knob the user didn't set (the body is `additionalProperties: false`,
/// and an unset knob simply omits rather than sending `null`).
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
pub struct VeniceParameters {
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_web_search: Option<WebSearchMode>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_web_citations: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_web_scraping: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub include_venice_system_prompt: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub include_search_results_in_stream: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub return_search_results_as_documents: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_x_search: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_e2ee: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub character_slug: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub strip_thinking_response: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub disable_thinking: Option<bool>,
/// Response verbosity (`low`, `medium`, `high`). Venice accepts this both top-level and inside
/// the bag; it lives here so the top-level config stays lean, and Venice reads it from the bag.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub verbosity: Option<String>,
}
#[derive(Debug, Clone, Copy, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum WebSearchMode {
Auto,
On,
Off,
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct SpeechToTextConfig {
#[serde(default = "default_speech_to_text_model_id")]
pub model_id: String,
}
impl Default for SpeechToTextConfig {
fn default() -> Self {
Self {
model_id: default_speech_to_text_model_id(),
}
}
}
fn default_speech_to_text_model_id() -> String {
"nvidia/parakeet-tdt-0.6b-v3".to_owned()
}
/// `/audio/speech` (`CreateSpeechRequestSchema`) request knobs. Only `model_id` is required on
/// the wire; everything else is optional with `skip_serializing_if` so an unset knob is omitted
/// rather than sent as `null` (the body is `additionalProperties: false`). `voice` is a free
/// `Option<String>`, not a closed enum: Venice's voice set spans dozens of model-specific names
/// plus arbitrary cloned-voice handles (`vv_<id>`), so an enum would reject valid handles.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct TextToSpeechConfig {
#[serde(default = "default_text_to_speech_model_id")]
pub model_id: String,
#[serde(
default = "default_text_to_speech_voice",
skip_serializing_if = "Option::is_none"
)]
pub voice: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub speed: Option<f32>,
#[serde(
default = "default_text_to_speech_response_format",
skip_serializing_if = "Option::is_none"
)]
pub response_format: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub prompt: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub temperature: Option<f32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub top_p: Option<f32>,
}
impl Default for TextToSpeechConfig {
fn default() -> Self {
Self {
model_id: default_text_to_speech_model_id(),
voice: default_text_to_speech_voice(),
speed: None,
response_format: default_text_to_speech_response_format(),
prompt: None,
temperature: None,
top_p: None,
}
}
}
fn default_text_to_speech_model_id() -> String {
"tts-kokoro".to_owned()
}
fn default_text_to_speech_voice() -> Option<String> {
Some("af_sky".to_owned())
}
fn default_text_to_speech_response_format() -> Option<String> {
Some("mp3".to_owned())
}
/// `/image/generate` (`GenerateImageRequest`) request knobs, mirroring Venice's schema
/// field-for-field. Only `model_id` is required; every other knob is optional with
/// `skip_serializing_if` so unset knobs are omitted (the body is `additionalProperties: false`).
/// The full knob set is deliberate: the native `/image/generate` endpoint is the flagship's
/// reason to exist over the knob-dropping OpenAI-compat path, so the knobs ARE the feature.
/// The deprecated `inpaint` knob is intentionally absent.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct ImageGenerationConfig {
#[serde(default = "default_image_generation_model_id")]
pub model_id: String,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub negative_prompt: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub cfg_scale: Option<f32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub steps: Option<u32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub style_preset: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub seed: Option<i64>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub safe_mode: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub hide_watermark: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub format: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub width: Option<u32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub height: Option<u32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub aspect_ratio: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub resolution: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub quality: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub lora_strength: Option<u32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub embed_exif_metadata: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_web_search: Option<bool>,
/// Image-edit settings, nested here because baibot has a single `ImageGeneration` purpose
/// and edit shares its config gate. The gen and edit model sets are disjoint, so edit
/// carries its own model field.
#[serde(default)]
pub edit: ImageEditSettings,
}
impl Default for ImageGenerationConfig {
fn default() -> Self {
Self {
model_id: default_image_generation_model_id(),
negative_prompt: None,
cfg_scale: None,
steps: None,
style_preset: None,
seed: None,
safe_mode: None,
hide_watermark: None,
format: None,
width: None,
height: None,
aspect_ratio: None,
resolution: None,
quality: None,
lora_strength: None,
embed_exif_metadata: None,
enable_web_search: None,
edit: ImageEditSettings::default(),
}
}
}
fn default_image_generation_model_id() -> String {
"chroma".to_owned()
}
/// `/image/edit` (`EditImageRequest`) request knobs, mirroring Venice's schema. The source image
/// and prompt are supplied per-call (not config), so only the model and the output-shaping knobs
/// live here. Each knob is optional with `skip_serializing_if`.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct ImageEditSettings {
#[serde(default = "default_image_edit_model_id")]
pub model_id: String,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub output_format: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub aspect_ratio: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub resolution: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub safe_mode: Option<bool>,
}
impl Default for ImageEditSettings {
fn default() -> Self {
Self {
model_id: default_image_edit_model_id(),
output_format: None,
aspect_ratio: None,
resolution: None,
safe_mode: None,
}
}
}
fn default_image_edit_model_id() -> String {
"firered-image-edit".to_owned()
}

View File

@@ -0,0 +1,167 @@
use crate::agent::AgentPurpose;
use crate::agent::provider::entity::{
ImageEditResult, ImageGenerationResult, ImageSource, PingResult, TextGenerationParams,
TextGenerationResult, TextToSpeechParams, TextToSpeechResult,
};
use crate::agent::provider::{
ImageEditParams, ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
};
use crate::conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
MessageContent as LLMMessageContent,
};
use super::super::ControllerTrait;
use super::config::Config;
use super::recovery::UnsupportedFieldsCache;
#[derive(Debug, Clone)]
pub struct Controller {
config: Config,
http: reqwest::Client,
// Per-model record of chat fields this Venice deployment has rejected as unsupported, learned at
// runtime. `Arc`-backed inside, so the `Clone` derive shares one cache across all clones of an
// agent's controller.
unsupported_fields: UnsupportedFieldsCache,
}
impl Controller {
pub fn new(config: Config) -> Self {
// Image generation and text-to-speech can run long, so give the client a generous timeout
// instead of reqwest's default (none). `build` only fails on TLS/system init; fall back to
// the infallible `Client::new()` so this constructor stays infallible.
let http = reqwest::Client::builder()
.timeout(std::time::Duration::from_secs(120))
.build()
.unwrap_or_else(|_| reqwest::Client::new());
// Process-global per-deployment cache rather than a fresh one: dynamic agents rebuild their
// controller on every message, so an instance-owned cache would never retain a learned
// rejection and the same unsupported field would 400 (and warn) on every turn.
let unsupported_fields = UnsupportedFieldsCache::shared_for(&config.base_url);
Self {
config,
http,
unsupported_fields,
}
}
}
impl ControllerTrait for Controller {
async fn ping(&self) -> anyhow::Result<PingResult> {
if !self.supports_purpose(AgentPurpose::TextGeneration) {
return Ok(PingResult::Inconclusive);
}
// Mirror the openai/openai_compat ping: a real "Hello!" round-trip exercises the strict
// /chat/completions body and auth, so a successful ping proves text generation works.
let messages = vec![LLMMessage {
author: LLMAuthor::User,
sender_id: None,
content: LLMMessageContent::Text("Hello!".to_string()),
timestamp: chrono::Utc::now(),
}];
let conversation = LLMConversation { messages };
self.generate_text(conversation, TextGenerationParams::default())
.await?;
Ok(PingResult::Successful)
}
async fn generate_text(
&self,
conversation: LLMConversation,
params: TextGenerationParams,
) -> anyhow::Result<TextGenerationResult> {
super::chat::generate_text(
&self.config,
&self.http,
&self.unsupported_fields,
conversation,
params,
)
.await
}
async fn speech_to_text(
&self,
mime_type: &mxlink::mime::Mime,
media: Vec<u8>,
params: SpeechToTextParams,
) -> anyhow::Result<SpeechToTextResult> {
super::audio::speech_to_text(&self.config, &self.http, mime_type, media, params).await
}
async fn generate_image(
&self,
prompt: &str,
params: ImageGenerationParams,
) -> anyhow::Result<ImageGenerationResult> {
super::images::generate_image(&self.config, &self.http, prompt, params).await
}
async fn create_image_edit(
&self,
prompt: &str,
images: Vec<ImageSource>,
params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
super::images::create_image_edit(&self.config, &self.http, prompt, images, params).await
}
async fn text_to_speech(
&self,
input: &str,
params: TextToSpeechParams,
) -> anyhow::Result<TextToSpeechResult> {
super::audio::text_to_speech(&self.config, &self.http, input, params).await
}
fn supports_purpose(&self, purpose: AgentPurpose) -> bool {
match purpose {
AgentPurpose::TextGeneration => self.config.text_generation.is_some(),
AgentPurpose::SpeechToText => self.config.speech_to_text.is_some(),
AgentPurpose::TextToSpeech => self.config.text_to_speech.is_some(),
AgentPurpose::ImageGeneration => self.config.image_generation.is_some(),
AgentPurpose::CatchAll => true,
}
}
fn text_generation_model_id(&self) -> Option<String> {
self.config
.text_generation
.as_ref()
.map(|config| config.model_id.to_owned())
}
fn text_generation_prompt(&self) -> Option<String> {
self.config
.text_generation
.as_ref()
.and_then(|config| config.prompt.clone())
}
fn text_generation_temperature(&self) -> Option<f32> {
self.config
.text_generation
.as_ref()
.map(|config| config.temperature)
}
fn text_to_speech_voice(&self) -> Option<String> {
self.config
.text_to_speech
.as_ref()
.and_then(|config| config.voice.clone())
}
fn text_to_speech_speed(&self) -> Option<f32> {
self.config
.text_to_speech
.as_ref()
.and_then(|config| config.speed)
}
}

View File

@@ -0,0 +1,190 @@
use crate::agent::AgentPurpose;
use crate::agent::provider::entity::{ImageEditResult, ImageGenerationResult, ImageSource};
use crate::agent::provider::{ImageEditParams, ImageGenerationParams};
use crate::strings;
use crate::utils::base64::{base64_decode, base64_encode};
use super::config::Config;
use super::wire::{EditImageRequest, GenerateImageRequest, GenerateImageResponse};
/// Generate an image via Venice's native `/image/generate` endpoint.
///
/// This is the base64-in-JSON path: we pin `return_binary: false` so Venice answers with a JSON
/// envelope (`GenerateImageResponse`) carrying the image as a base64 string, which we decode. The
/// sibling `create_image_edit` is the *other* response shape (raw binary); the two must not be
/// crossed. `params` is advisory only; the Venice config drives the request.
pub async fn generate_image(
config: &Config,
http: &reqwest::Client,
prompt: &str,
_params: ImageGenerationParams,
) -> anyhow::Result<ImageGenerationResult> {
let Some(image_generation_config) = &config.image_generation else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::ImageGeneration
),
));
};
let request = GenerateImageRequest {
model: image_generation_config.model_id.clone(),
prompt: prompt.to_owned(),
// Pinned: baibot wants exactly one image, returned as base64-in-JSON so `GenerateImageResponse`
// can decode it. Flipping `return_binary` would make Venice answer with raw binary and break
// the JSON decode below, so neither knob is configurable.
return_binary: false,
variants: 1,
negative_prompt: image_generation_config.negative_prompt.clone(),
cfg_scale: image_generation_config.cfg_scale,
steps: image_generation_config.steps,
style_preset: image_generation_config.style_preset.clone(),
seed: image_generation_config.seed,
safe_mode: image_generation_config.safe_mode,
hide_watermark: image_generation_config.hide_watermark,
format: image_generation_config.format.clone(),
width: image_generation_config.width,
height: image_generation_config.height,
aspect_ratio: image_generation_config.aspect_ratio.clone(),
resolution: image_generation_config.resolution.clone(),
quality: image_generation_config.quality.clone(),
lora_strength: image_generation_config.lora_strength,
embed_exif_metadata: image_generation_config.embed_exif_metadata,
enable_web_search: image_generation_config.enable_web_search,
};
let url = format!("{}/image/generate", config.base_url.trim_end_matches('/'));
// The prompt is user content; keep it out of logs (mirrors the STT/TTS paths).
tracing::trace!(
model_id = image_generation_config.model_id,
"Sending Venice image generation API request"
);
let response = http
.post(&url)
.bearer_auth(&config.api_key)
.json(&request)
.send()
.await?;
let status = response.status();
if !status.is_success() {
// Body to the server log only, not into the returned error (which reaches the Matrix room).
let body = response.text().await.unwrap_or_default();
tracing::warn!(%status, body, "Venice image generation request failed");
return Err(anyhow::anyhow!(
"Venice image generation request failed with status {status}"
));
}
let response: GenerateImageResponse = response.json().await?;
tracing::trace!(request_id = ?response.id, "Venice image generation succeeded");
let Some(image_base64) = response.images.into_iter().next() else {
return Err(anyhow::anyhow!(
"The Venice image generation API returned no images"
));
};
// Swallow the decode error's detail (it can echo input bytes/offsets); the returned error
// reaches the Matrix room, so it stays generic while the real cause goes to the server log.
let bytes = base64_decode(&image_base64).map_err(|decode_err| {
tracing::warn!(%decode_err, "Venice image generation returned undecodable base64");
anyhow::anyhow!("Venice image generation returned invalid base64 image data")
})?;
Ok(ImageGenerationResult {
bytes,
mime_type: image_format_to_mime_type(image_generation_config.format.as_deref()),
revised_prompt: None,
})
}
/// Edit an image via Venice's native `/image/edit` endpoint.
///
/// This is the raw-binary path: the request is JSON carrying the source image as a base64 string
/// (Venice's `image` field is `anyOf` upload/base64/URL; we send base64, no multipart), and the
/// response body IS the edited image bytes (no JSON envelope). `params` is advisory only.
pub async fn create_image_edit(
config: &Config,
http: &reqwest::Client,
prompt: &str,
images: Vec<ImageSource>,
_params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
let Some(image_generation_config) = &config.image_generation else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::ImageGeneration
),
));
};
let edit_config = &image_generation_config.edit;
let Some(source) = images.into_iter().next() else {
return Err(anyhow::anyhow!("No image sources provided"));
};
let request = EditImageRequest {
model: edit_config.model_id.clone(),
prompt: prompt.to_owned(),
image: base64_encode(&source.bytes),
output_format: edit_config.output_format.clone(),
aspect_ratio: edit_config.aspect_ratio.clone(),
resolution: edit_config.resolution.clone(),
safe_mode: edit_config.safe_mode,
};
let url = format!("{}/image/edit", config.base_url.trim_end_matches('/'));
// The prompt is user content; keep it out of logs (mirrors the STT/TTS paths).
tracing::trace!(
model_id = edit_config.model_id,
"Sending Venice image edit API request"
);
let response = http
.post(&url)
.bearer_auth(&config.api_key)
.json(&request)
.send()
.await?;
let status = response.status();
if !status.is_success() {
// Body to the server log only, not into the returned error (which reaches the Matrix room).
let body = response.text().await.unwrap_or_default();
tracing::warn!(%status, body, "Venice image edit request failed");
return Err(anyhow::anyhow!(
"Venice image edit request failed with status {status}"
));
}
// The edit endpoint answers with raw binary image bytes, so read the body directly instead of
// parsing JSON. The actual format comes from the response Content-Type header; fall back to the
// configured `output_format` when the header is missing or unparseable.
let mime_type = response
.headers()
.get(reqwest::header::CONTENT_TYPE)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<mxlink::mime::Mime>().ok())
.unwrap_or_else(|| image_format_to_mime_type(edit_config.output_format.as_deref()));
let bytes = response.bytes().await?.to_vec();
Ok(ImageEditResult { bytes, mime_type })
}
/// Map a Venice image `format`/`output_format` value (`jpeg`/`png`/`webp`) to its MIME type.
/// Venice defaults to `webp` when the format is unset, so an absent value maps to `image/webp`.
fn image_format_to_mime_type(format: Option<&str>) -> mxlink::mime::Mime {
match format.unwrap_or("webp") {
"jpeg" | "jpg" => mxlink::mime::IMAGE_JPEG,
"png" => mxlink::mime::IMAGE_PNG,
// No mxlink::mime constant for webp; parse it, falling back to PNG on any surprise value.
_ => "image/webp".parse().unwrap_or(mxlink::mime::IMAGE_PNG),
}
}

View File

@@ -0,0 +1,48 @@
mod audio;
mod chat;
mod config;
mod controller;
mod images;
mod recovery;
mod utils;
mod wire;
#[cfg(test)]
mod tests;
pub use config::Config;
pub use controller::Controller;
use super::super::AgentInstantiationError;
use super::super::AgentInstantiationResult;
use super::ConfigTrait;
use super::controller::ControllerType;
pub fn create_controller_from_yaml_value_config(
agent_id: &str,
config: serde_yaml_ng::Value,
) -> AgentInstantiationResult<ControllerType> {
let config = match &config {
serde_yaml_ng::Value::Mapping(_) => {
let config: Config =
serde_yaml_ng::from_value(config).map_err(AgentInstantiationError::Yaml)?;
config
.validate()
.map_err(AgentInstantiationError::ConfigFailsValidation)?;
config
}
_ => {
return Err(AgentInstantiationError::ConfigForAgentIsNotAMapping(
agent_id.to_owned(),
));
}
};
Ok(ControllerType::Venice(Box::new(Controller::new(config))))
}
pub fn default_config() -> Config {
Config::default()
}

View File

@@ -0,0 +1,312 @@
//! Auto-recovery for Venice's strict request bodies.
//!
//! Venice's `/chat/completions` body is `additionalProperties: false`, so a model that does not
//! support an optional knob rejects the whole request with a 400 instead of ignoring the field.
//! Some knobs are documented as model-specific ("for supported models") and are pure
//! optimization/tuning hints: dropping them changes nothing about the answer, only loses the
//! optimization. When such a field is the reason for a 400, we strip it and retry, then remember
//! the rejection per model so later requests skip the field (and the wasted round-trip) entirely.
//!
//! Universal sampling knobs (`temperature`, `top_p`, the penalties, `max_completion_tokens`) are
//! deliberately NOT recoverable here: dropping one silently changes the model's output, so a model
//! that rejects one is a real configuration problem the operator must see, not paper over.
use std::collections::{HashMap, HashSet};
use std::sync::{Arc, OnceLock, RwLock};
use regex::Regex;
use super::wire::ChatCompletionRequest;
/// Top-level chat-completion fields baibot may drop to recover from a 400. Each is documented by
/// Venice as model-specific or as a routing hint, and is meaning-preserving to omit (Venice falls
/// back to its server-side default):
/// - `prompt_cache_retention` — "extends retention ... for supported models" (cache TTL only)
/// - `reasoning_effort` — "control reasoning effort level for supported models"
/// - `prompt_cache_key` — cache-routing hint; dropping it only forfeits a cache-hit optimization
///
/// Recovery operates on TOP-LEVEL request fields only. Sub-fields inside the `venice_parameters`
/// bag (`disable_thinking`, `enable_e2ee`, `character_slug`, …) are intentionally absent: a model
/// that rejects one surfaces a clear 400 to the operator rather than being auto-stripped. Adding a
/// new bag field does not extend recovery to it; only a name listed here is droppable.
pub(super) const DROPPABLE_FIELDS: &[&str] = &[
"prompt_cache_retention",
"prompt_cache_key",
"reasoning_effort",
];
/// Per-model record of fields a Venice model has rejected as unsupported, learned at runtime from
/// 400 responses. The `Arc` is shared across `Controller` clones, so a rejection learned once is
/// seen by every clone of the same agent. The cache is process-lived only: a restart re-learns on
/// the first request to each model, which costs one extra round-trip and nothing else, so there is
/// no persistence to keep in sync with config changes.
#[derive(Debug, Clone, Default)]
pub(super) struct UnsupportedFieldsCache {
inner: Arc<RwLock<HashMap<String, HashSet<String>>>>,
}
/// Process-global registry of per-deployment caches, keyed by Venice base URL. A `Controller` is
/// rebuilt from its config on every message for dynamic (global/room-local) agents (see
/// `agent::manager::available_room_agents_by_room_config_context`), so a cache stored on the
/// controller would be reset each message and the "process-lived" intent above would never hold.
/// The registry lets a rebuilt controller re-attach to the same cache. Keyed by base URL because
/// "which fields a model rejects" is a property of the deployment+model, not of the agent instance:
/// agents on the same Venice deployment share learnings, while a distinct deployment keeps its own,
/// so a model name that supports a field on one deployment is never proactively stripped against
/// another.
fn registry() -> &'static RwLock<HashMap<String, UnsupportedFieldsCache>> {
static REGISTRY: OnceLock<RwLock<HashMap<String, UnsupportedFieldsCache>>> = OnceLock::new();
REGISTRY.get_or_init(|| RwLock::new(HashMap::new()))
}
impl UnsupportedFieldsCache {
/// The process-global cache for `base_url`, created on first use and shared by every controller
/// for that deployment thereafter. This is what makes a learned rejection survive the
/// per-message controller rebuild. A poisoned registry lock degrades to an isolated cache:
/// correctness holds, only cross-controller sharing is lost for that call.
pub(super) fn shared_for(base_url: &str) -> Self {
if let Ok(map) = registry().read()
&& let Some(cache) = map.get(base_url)
{
return cache.clone();
}
match registry().write() {
Ok(mut map) => map.entry(base_url.to_owned()).or_default().clone(),
Err(_) => Self::default(),
}
}
/// Fields already known unsupported for `model_id`. Returns an empty set on an unknown model or
/// a poisoned lock, so a cache failure degrades to "strip nothing proactively" rather than
/// breaking the request path.
pub(super) fn known_for(&self, model_id: &str) -> HashSet<String> {
self.inner
.read()
.ok()
.and_then(|map| map.get(model_id).cloned())
.unwrap_or_default()
}
/// Records that `model_id` rejected `field`. A poisoned lock is ignored: failing to memoize only
/// means the next request re-discovers the rejection, never a wrong result.
pub(super) fn record(&self, model_id: &str, field: &str) {
if let Ok(mut map) = self.inner.write() {
map.entry(model_id.to_owned())
.or_default()
.insert(field.to_owned());
}
}
}
/// Parses the offending field name out of a Venice 400 body. A field the schema does not allow is
/// reported as `... field: 'prompt_cache_retention', value: '...'`, so this matches the `field: '..'`
/// marker wherever it sits in the message. Returns `None` when the body carries no such marker (a
/// different 400 class, e.g. a missing required field), so the caller surfaces that error instead.
///
/// Returns only the FIRST `field: '..'` match by design. If Venice ever names several rejected
/// fields in one body, the retry loop strips this one, retries, and rediscovers the next on the
/// following 400 — bounded and correct. Do not switch to `captures_iter` to "batch" them without
/// re-checking the loop's per-field termination bound in `chat.rs`.
pub(super) fn parse_rejected_field(body: &str) -> Option<String> {
static RE: OnceLock<Regex> = OnceLock::new();
let re =
RE.get_or_init(|| Regex::new(r"field: '([^']+)'").expect("rejected-field regex is valid"));
re.captures(body)
.map(|caps| caps[1].to_owned())
.filter(|field| !field.is_empty())
}
/// Clears `field` from the request when it is one baibot may safely drop and it is currently set.
/// Returns `true` only when a value was actually removed, which is what bounds the retry loop: once
/// a field is `None`, a repeat rejection for the same name returns `false` and the caller stops
/// instead of retrying forever. A field outside [`DROPPABLE_FIELDS`] always returns `false`, so a
/// meaning-bearing knob is never silently dropped.
pub(super) fn strip_droppable_field(request: &mut ChatCompletionRequest, field: &str) -> bool {
if !DROPPABLE_FIELDS.contains(&field) {
return false;
}
match field {
"prompt_cache_retention" => request.prompt_cache_retention.take().is_some(),
"prompt_cache_key" => request.prompt_cache_key.take().is_some(),
"reasoning_effort" => request.reasoning_effort.take().is_some(),
_ => false,
}
}
/// Pulls a human-readable message out of a Venice error body for surfacing in the room. Venice's
/// usual envelope is `{"error": "..."}`; some OpenAI-compatible paths nest `{"error": {"message":
/// "..."}}`. Falls back to the trimmed raw body (length-capped so a large body cannot flood the
/// room) and finally to a fixed string for an empty body, so the caller always has something to
/// show.
pub(super) fn extract_error_message(body: &str) -> String {
let trimmed = body.trim();
if trimmed.is_empty() {
return "no response body".to_owned();
}
if let Ok(value) = serde_json::from_str::<serde_json::Value>(trimmed) {
if let Some(msg) = value.get("error").and_then(|e| e.as_str()) {
return msg.to_owned();
}
if let Some(msg) = value
.get("error")
.and_then(|e| e.get("message"))
.and_then(|m| m.as_str())
{
return msg.to_owned();
}
}
const MAX: usize = 500;
if trimmed.chars().count() > MAX {
trimmed.chars().take(MAX).collect::<String>() + "…"
} else {
trimmed.to_owned()
}
}
#[cfg(test)]
mod tests {
use super::*;
fn full_request() -> ChatCompletionRequest {
ChatCompletionRequest {
model: "venice-uncensored".to_owned(),
messages: vec![],
temperature: Some(0.7),
max_completion_tokens: Some(1024),
top_p: None,
frequency_penalty: None,
presence_penalty: None,
repetition_penalty: None,
reasoning_effort: Some("high".to_owned()),
prompt_cache_key: Some("cafef00d".to_owned()),
prompt_cache_retention: Some("24h".to_owned()),
venice_parameters: None,
}
}
#[test]
fn parses_the_rejected_field_from_a_real_venice_body() {
let body = r#"{"error":"Extra inputs are not permitted, field: 'prompt_cache_retention', value: 'default'","request_id":"qM_DmKSXKF07wRxmQJ-hc"}"#;
assert_eq!(
parse_rejected_field(body).as_deref(),
Some("prompt_cache_retention")
);
}
#[test]
fn returns_no_field_when_the_body_has_no_field_marker() {
// A different 400 class (e.g. a genuinely malformed request) carries no `field: '..'`
// marker, so there is nothing to strip and the caller must surface the error instead.
assert_eq!(parse_rejected_field(r#"{"error":"Invalid request"}"#), None);
assert_eq!(parse_rejected_field(""), None);
}
#[test]
fn strips_a_droppable_field_once_then_reports_no_progress() {
let mut request = full_request();
// First strip clears the field and reports progress, so the caller retries.
assert!(strip_droppable_field(
&mut request,
"prompt_cache_retention"
));
assert!(request.prompt_cache_retention.is_none());
// A repeat rejection for the same (now absent) field reports no progress: this is what
// stops the retry loop instead of spinning forever.
assert!(!strip_droppable_field(
&mut request,
"prompt_cache_retention"
));
}
#[test]
fn refuses_to_strip_a_meaning_bearing_field() {
let mut request = full_request();
// `temperature` is universal and changes the output; a rejection for it must surface, never
// be silently dropped. The whole droppable set is the only thing strip will touch.
assert!(!strip_droppable_field(&mut request, "temperature"));
assert_eq!(request.temperature, Some(0.7));
assert!(!strip_droppable_field(
&mut request,
"max_completion_tokens"
));
assert_eq!(request.max_completion_tokens, Some(1024));
for field in DROPPABLE_FIELDS {
assert!(
strip_droppable_field(&mut full_request(), field),
"every advertised droppable field must actually be strippable: {field}"
);
}
}
#[test]
fn cache_records_per_model_and_isolates_models() {
let cache = UnsupportedFieldsCache::default();
assert!(cache.known_for("venice-uncensored").is_empty());
cache.record("venice-uncensored", "prompt_cache_retention");
cache.record("venice-uncensored", "reasoning_effort");
let known = cache.known_for("venice-uncensored");
assert!(known.contains("prompt_cache_retention"));
assert!(known.contains("reasoning_effort"));
// A rejection learned for one model must not leak to another.
assert!(cache.known_for("kimi-k2-5").is_empty());
}
#[test]
fn shared_for_survives_controller_rebuild_and_isolates_deployments() {
// Distinct, test-only base URLs so the process-global registry can't collide with another
// test running in parallel.
let url_a = "https://shared-for-test-a.invalid/api/v1";
let url_b = "https://shared-for-test-b.invalid/api/v1";
// A controller rebuilt for the same deployment (a fresh `shared_for` call, as happens per
// message for dynamic agents) re-attaches to the SAME cache, so an earlier learning holds.
let first = UnsupportedFieldsCache::shared_for(url_a);
first.record("model-x", "reasoning_effort");
let rebuilt = UnsupportedFieldsCache::shared_for(url_a);
assert!(
rebuilt.known_for("model-x").contains("reasoning_effort"),
"a rebuilt controller for the same deployment must see the earlier rejection"
);
// A different deployment keeps its own learnings: a model name that rejects a field on one
// Venice must not silence/strip it on another.
let other_deployment = UnsupportedFieldsCache::shared_for(url_b);
assert!(
other_deployment.known_for("model-x").is_empty(),
"rejections must not leak across deployments"
);
}
#[test]
fn extracts_a_human_message_from_error_envelopes() {
assert_eq!(
extract_error_message(r#"{"error":"Extra inputs are not permitted","request_id":"x"}"#),
"Extra inputs are not permitted"
);
// OpenAI-style nested envelope.
assert_eq!(
extract_error_message(r#"{"error":{"message":"context length exceeded"}}"#),
"context length exceeded"
);
// Unknown shape falls back to the raw body; empty falls back to a fixed string.
assert_eq!(
extract_error_message("plain text failure"),
"plain text failure"
);
assert_eq!(extract_error_message(" "), "no response body");
}
}

View File

@@ -0,0 +1,549 @@
use mxlink::matrix_sdk::ruma::OwnedMxcUri;
use mxlink::matrix_sdk::ruma::events::room::message::{
FileMessageEventContent, ImageMessageEventContent,
};
use mxlink::mime;
use super::super::ControllerTrait;
use crate::agent::AgentPurpose;
use crate::conversation::llm::{
Author as LLMAuthor, FileDetails, ImageDetails, Message as LLMMessage,
MessageContent as LLMMessageContent,
};
use super::chat::{append_reasoning, derive_prompt_cache_key, render_with_citations};
use super::config::{Config, TextGenerationConfig, VeniceParameters, WebSearchMode};
use super::controller::Controller;
use super::utils::convert_llm_messages_to_venice;
use super::wire::{
ChatCompletionRequest, ContentPart, EditImageRequest, GenerateImageRequest, MessageContent,
SpeechRequest, WebSearchCitation,
};
#[test]
fn config_round_trips_with_venice_parameters() {
let yaml = r#"
base_url: https://api.venice.ai/api/v1
api_key: test-key
text_generation:
model_id: kimi-k2-5
temperature: 0.7
max_response_tokens: 1024
max_context_tokens: 65536
venice_parameters:
enable_web_search: "auto"
enable_web_citations: true
speech_to_text:
model_id: nvidia/parakeet-tdt-0.6b-v3
"#;
let config: Config = serde_yaml_ng::from_str(yaml).expect("config should deserialize");
let tg = config.text_generation.expect("text_generation present");
let vp = tg.venice_parameters.expect("venice_parameters present");
assert!(matches!(vp.enable_web_search, Some(WebSearchMode::Auto)));
assert_eq!(vp.enable_web_citations, Some(true));
assert_eq!(vp.character_slug, None);
// The bag must serialize the enum to the exact wire string, and an unset knob must be ABSENT
// (not `null`) so the strict `additionalProperties: false` body is honored.
let json = serde_json::to_string(&vp).expect("serialize venice_parameters");
assert!(
json.contains("\"enable_web_search\":\"auto\""),
"web search should be the literal \"auto\": {json}"
);
assert!(
!json.contains("character_slug"),
"an unset knob must be omitted entirely: {json}"
);
assert!(
!json.contains("null"),
"no nulls belong in the body: {json}"
);
}
#[test]
fn converts_text_image_and_file_to_content_parts() {
let messages = vec![
LLMMessage {
author: LLMAuthor::User,
sender_id: None,
timestamp: chrono::Utc::now(),
content: LLMMessageContent::Text("describe this".to_owned()),
},
LLMMessage {
author: LLMAuthor::User,
sender_id: None,
timestamp: chrono::Utc::now(),
content: LLMMessageContent::Image(ImageDetails::new(
ImageMessageEventContent::plain(
"pic.png".to_owned(),
OwnedMxcUri::from("mxc://example.com/abc"),
),
mime::IMAGE_PNG,
vec![1, 2, 3],
)),
},
LLMMessage {
author: LLMAuthor::User,
sender_id: None,
timestamp: chrono::Utc::now(),
content: LLMMessageContent::File(FileDetails::new(
FileMessageEventContent::plain(
"doc.pdf".to_owned(),
OwnedMxcUri::from("mxc://example.com/def"),
),
mime::APPLICATION_PDF,
vec![4, 5, 6],
)),
},
];
let converted = convert_llm_messages_to_venice(messages).expect("conversion should succeed");
// Text, image, AND file all survive now: the file is no longer warn-skipped.
assert_eq!(converted.len(), 3);
match &converted[0].content {
MessageContent::Text(text) => assert_eq!(text, "describe this"),
other => panic!("expected bare text, got {other:?}"),
}
match &converted[1].content {
MessageContent::Parts(parts) => match &parts[0] {
ContentPart::ImageUrl { image_url } => assert!(
image_url.url.starts_with("data:image/png;base64,"),
"image should be inlined as a data URI: {}",
image_url.url
),
other => panic!("expected an image part, got {other:?}"),
},
other => panic!("expected image parts, got {other:?}"),
}
match &converted[2].content {
MessageContent::Parts(parts) => match &parts[0] {
ContentPart::File { file } => {
assert!(
file.file_data.starts_with("data:application/pdf;base64,"),
"file should be inlined as a data URI: {}",
file.file_data
);
assert_eq!(file.filename.as_deref(), Some("doc.pdf"));
}
other => panic!("expected a file part, got {other:?}"),
},
other => panic!("expected file parts, got {other:?}"),
}
}
#[test]
fn supports_purpose_truth_table() {
let config: Config = serde_yaml_ng::from_str(
r#"
base_url: https://api.venice.ai/api/v1
api_key: test-key
text_generation:
model_id: kimi-k2-5
speech_to_text:
model_id: nvidia/parakeet-tdt-0.6b-v3
"#,
)
.expect("config should deserialize");
let controller = Controller::new(config);
assert!(controller.supports_purpose(AgentPurpose::TextGeneration));
assert!(controller.supports_purpose(AgentPurpose::SpeechToText));
assert!(controller.supports_purpose(AgentPurpose::CatchAll));
assert!(!controller.supports_purpose(AgentPurpose::TextToSpeech));
assert!(!controller.supports_purpose(AgentPurpose::ImageGeneration));
}
#[test]
fn supports_purpose_true_when_image_and_tts_blocks_present() {
let config: Config = serde_yaml_ng::from_str(
r#"
base_url: https://api.venice.ai/api/v1
api_key: test-key
text_to_speech:
model_id: tts-kokoro
image_generation:
model_id: chroma
"#,
)
.expect("config should deserialize");
let controller = Controller::new(config);
assert!(controller.supports_purpose(AgentPurpose::TextToSpeech));
assert!(controller.supports_purpose(AgentPurpose::ImageGeneration));
}
#[test]
fn speech_request_serializes_voice_and_omits_unset() {
let request = SpeechRequest {
model: "tts-kokoro".to_owned(),
input: "hello".to_owned(),
voice: Some("af_sky".to_owned()),
speed: None,
response_format: Some("mp3".to_owned()),
prompt: None,
temperature: None,
top_p: None,
};
let json = serde_json::to_string(&request).expect("serialize SpeechRequest");
assert!(
json.contains("\"voice\":\"af_sky\""),
"voice should be present: {json}"
);
assert!(
!json.contains("temperature"),
"an unset knob must be omitted (not null): {json}"
);
assert!(
!json.contains("null"),
"no nulls belong in the body: {json}"
);
}
#[test]
fn generate_image_request_pins_flags_and_omits_unset() {
let request = GenerateImageRequest {
model: "chroma".to_owned(),
prompt: "a cat".to_owned(),
return_binary: false,
variants: 1,
negative_prompt: None,
cfg_scale: None,
steps: None,
style_preset: None,
seed: None,
safe_mode: None,
hide_watermark: None,
format: None,
width: None,
height: None,
aspect_ratio: None,
resolution: None,
quality: None,
lora_strength: None,
embed_exif_metadata: None,
enable_web_search: None,
};
let json = serde_json::to_string(&request).expect("serialize GenerateImageRequest");
assert!(json.contains("\"model\":\"chroma\""), "{json}");
assert!(
json.contains("\"return_binary\":false"),
"return_binary must be pinned false: {json}"
);
assert!(
json.contains("\"variants\":1"),
"variants must be pinned 1: {json}"
);
assert!(
!json.contains("cfg_scale"),
"an unset knob must be omitted: {json}"
);
assert!(
!json.contains("null"),
"no nulls belong in the body: {json}"
);
}
#[test]
fn edit_image_request_carries_model_and_base64_image() {
let request = EditImageRequest {
model: "firered-image-edit".to_owned(),
prompt: "make it a sunrise".to_owned(),
image: "aGVsbG8=".to_owned(),
output_format: None,
aspect_ratio: None,
resolution: None,
safe_mode: None,
};
let json = serde_json::to_string(&request).expect("serialize EditImageRequest");
assert!(json.contains("\"model\":\"firered-image-edit\""), "{json}");
assert!(
json.contains("\"image\":\"aGVsbG8=\""),
"the base64 image string must be present: {json}"
);
assert!(
!json.contains("output_format"),
"an unset knob must be omitted: {json}"
);
}
#[test]
fn web_search_mode_off_deserializes_from_bare_yaml_off() {
// `off` is a YAML-1.1 boolean but a plain string under serde_yaml_ng's YAML-1.2 core schema,
// so it deserializes straight into the lowercase `WebSearchMode::Off`. This pins that the
// sample config and docs can use the bare, unquoted `off` without it parsing as a boolean.
let params: VeniceParameters =
serde_yaml_ng::from_str("enable_web_search: off").expect("bare `off` should deserialize");
assert!(matches!(params.enable_web_search, Some(WebSearchMode::Off)));
}
#[test]
fn request_places_sampling_top_level_and_verbosity_in_the_bag() {
// The whole config-shape decision in one assertion: top-level knobs serialize at the top
// level, the dual-position `verbosity` serializes inside the bag. Venice silently ignores a
// top-level knob misplaced into the bag, so this is the guard against a silent no-op.
let request = ChatCompletionRequest {
model: "kimi-k2-5".to_owned(),
messages: vec![],
temperature: Some(0.5),
max_completion_tokens: Some(1024),
top_p: Some(0.5),
frequency_penalty: None,
presence_penalty: None,
repetition_penalty: None,
reasoning_effort: Some("high".to_owned()),
prompt_cache_key: Some("00000000cafef00d".to_owned()),
prompt_cache_retention: Some("24h".to_owned()),
venice_parameters: Some(VeniceParameters {
verbosity: Some("high".to_owned()),
..Default::default()
}),
};
let json = serde_json::to_value(&request).expect("serialize request");
assert_eq!(json["top_p"], 0.5);
assert_eq!(json["reasoning_effort"], "high");
assert_eq!(json["prompt_cache_retention"], "24h");
assert_eq!(json["prompt_cache_key"], "00000000cafef00d");
assert!(
json.get("verbosity").is_none(),
"verbosity must not be a top-level field: {json}"
);
assert_eq!(json["venice_parameters"]["verbosity"], "high");
assert!(
json["venice_parameters"].get("top_p").is_none(),
"top_p must not be inside the bag: {json}"
);
}
#[test]
fn config_defaults_prompt_cache_retention_to_24h() {
// The programmatic default.
assert_eq!(
TextGenerationConfig::default()
.prompt_cache_retention
.as_deref(),
Some("24h")
);
// A config that omits the key must ALSO default to 24h, via the named serde default. A bare
// `#[serde(default)]` would yield None here and silently disable caching for such configs.
let tg: TextGenerationConfig = serde_yaml_ng::from_str("model_id: kimi-k2-5\n")
.expect("minimal config should deserialize");
assert_eq!(
tg.prompt_cache_retention.as_deref(),
Some("24h"),
"an omitted retention key must still default to 24h"
);
}
#[test]
fn cache_key_is_stable_for_same_inputs_and_varies_otherwise() {
let key = derive_prompt_cache_key("system prompt", "2024-09-20 (Friday), 18:34:15 UTC");
// Identical inputs produce an identical key: this is what keeps turn 5 routing to the warm
// server holding turns 1-4 (and what survives a process restart).
assert_eq!(
key,
derive_prompt_cache_key("system prompt", "2024-09-20 (Friday), 18:34:15 UTC"),
"identical inputs must produce an identical key"
);
assert_eq!(key.len(), 16, "the key is a 16-char hex string");
assert!(key.chars().all(|c| c.is_ascii_hexdigit()));
// A different conversation start time or a different prompt must change the key.
assert_ne!(
key,
derive_prompt_cache_key("system prompt", "2024-09-21 (Saturday), 09:00:00 UTC"),
"a different start time must change the key"
);
assert_ne!(
key,
derive_prompt_cache_key("other prompt", "2024-09-20 (Friday), 18:34:15 UTC"),
"a different prompt must change the key"
);
}
#[test]
fn citations_render_inline_refs_and_a_sources_block() {
let citations = vec![WebSearchCitation {
title: "Example Source".to_owned(),
url: "https://example.com/a".to_owned(),
}];
let rendered = render_with_citations("the sky is blue^1^".to_owned(), &citations);
assert!(
rendered.contains("the sky is blue[1]"),
"inline ^1^ becomes [1]: {rendered}"
);
assert!(
rendered.contains("Sources:"),
"a Sources block is appended: {rendered}"
);
assert!(
rendered.contains("[1] [Example Source](https://example.com/a)"),
"the source renders as a markdown link: {rendered}"
);
}
#[test]
fn citations_absent_leaves_content_untouched() {
let content = "plain answer, no web search".to_owned();
assert_eq!(render_with_citations(content.clone(), &[]), content);
}
#[test]
fn citation_title_and_url_cannot_inject_markdown() {
// A hostile page sets its title to break out of the link label and its URL to a non-http
// scheme. Neither may produce a spoofed clickable link in the room.
let citations = vec![WebSearchCitation {
title: "evil](http://phish.example) take".to_owned(),
url: "javascript:alert(1)".to_owned(),
}];
let rendered = render_with_citations("result^1^".to_owned(), &citations);
assert!(
rendered.contains("evil\\](http://phish.example) take"),
"the title's brackets must be escaped so it cannot close the link label: {rendered}"
);
assert!(
!rendered.contains("(javascript:alert(1))"),
"a non-http(s) URL must never become a markdown link target: {rendered}"
);
}
#[test]
fn chained_and_comma_citation_runs_each_expand_to_separate_refs() {
let citations = vec![
WebSearchCitation {
title: "One".to_owned(),
url: "https://example.com/1".to_owned(),
},
WebSearchCitation {
title: "Two".to_owned(),
url: "https://example.com/2".to_owned(),
},
WebSearchCitation {
title: "Three".to_owned(),
url: "https://example.com/3".to_owned(),
},
];
// Caret-chained run: Venice shares the caret between consecutive citations (`^2^3^`). The whole
// run must expand, not just the first, with no orphaned `3^` left behind.
let chained = render_with_citations("alpha^2^3^ and beta^1^".to_owned(), &citations);
assert!(
chained.contains("alpha[2][3] and beta[1]"),
"a chained ^2^3^ run must expand to [2][3] with no orphaned caret: {chained}"
);
// Comma run.
let comma = render_with_citations("gamma^1,3^".to_owned(), &citations);
assert!(
comma.contains("gamma[1][3]"),
"a comma ^1,3^ run must expand to [1][3]: {comma}"
);
// Multi-digit citation indices survive intact.
let multidigit = render_with_citations("delta^2^10^".to_owned(), &citations);
assert!(
multidigit.contains("delta[2][10]"),
"a multi-digit chained run must expand to [2][10]: {multidigit}"
);
}
#[test]
fn malformed_citation_degrades_instead_of_failing() {
// A citation arriving without a `url` must still deserialize (to an empty default) rather than
// failing the whole response parse and losing an otherwise-good answer.
let parsed: WebSearchCitation = serde_json::from_str(r#"{"title":"Only a title"}"#)
.expect("a citation missing `url` should still deserialize");
assert_eq!(parsed.url, "");
// Rendering citations with missing fields stays graceful: no empty `[]()` link, no panic.
let citations = vec![
WebSearchCitation {
title: String::new(),
url: "https://example.com/u".to_owned(),
},
WebSearchCitation {
title: String::new(),
url: String::new(),
},
];
let rendered = render_with_citations("answer^1^2^".to_owned(), &citations);
assert!(
rendered.contains("[1] [https://example.com/u](https://example.com/u)"),
"a citation with no title falls back to the URL as link text: {rendered}"
);
assert!(
rendered.contains("[2] (source unavailable)"),
"a citation with neither title nor URL renders a placeholder: {rendered}"
);
}
#[test]
fn reasoning_is_appended_only_when_show_reasoning_is_set() {
let base = "the answer".to_owned();
// Off (the default): thinking is dropped, never reaching the room.
let off = append_reasoning(base.clone(), Some("secret thinking".to_owned()), false);
assert_eq!(off, "the answer");
// On: thinking is appended below the answer in a collapsible <details> block (folded by
// default, expandable in clients that support it).
let on = append_reasoning(base.clone(), Some(" visible thinking ".to_owned()), true);
assert!(on.starts_with("the answer"));
assert!(on.contains("<details><summary>💭 Reasoning</summary>"));
assert!(on.contains("</details>"));
// The reasoning sits as its own markdown block (blank lines around it) and is trimmed.
assert!(on.contains("\n\nvisible thinking\n\n"));
// On but empty or missing reasoning: nothing is appended.
assert_eq!(
append_reasoning(base.clone(), Some(" ".to_owned()), true),
"the answer"
);
assert_eq!(append_reasoning(base.clone(), None, true), "the answer");
}
#[test]
fn oversized_file_is_rejected() {
let messages = vec![LLMMessage {
author: LLMAuthor::User,
sender_id: None,
timestamp: chrono::Utc::now(),
content: LLMMessageContent::File(FileDetails::new(
FileMessageEventContent::plain(
"big.pdf".to_owned(),
OwnedMxcUri::from("mxc://example.com/big"),
),
mime::APPLICATION_PDF,
vec![0u8; 25 * 1024 * 1024 + 1],
)),
}];
assert!(
convert_llm_messages_to_venice(messages).is_err(),
"a file over the 25MB limit must be rejected"
);
}

View File

@@ -0,0 +1,81 @@
use crate::conversation::llm::{
Author as LLMAuthor, Message as LLMMessage, MessageContent as LLMMessageContent,
};
use crate::utils::base64::base64_encode;
use super::wire::{ChatMessage, ContentPart, FilePart, ImageUrl, MessageContent};
/// Venice's documented file-input ceiling is 25MB on the decoded bytes (swagger `file_data`).
/// We check it here so an oversized file gets a clear message instead of an opaque 413 from the
/// API; the 413 status branch in `chat.rs` is the backstop if a file slips past this guard.
const MAX_FILE_BYTES: usize = 25 * 1024 * 1024;
pub fn convert_llm_messages_to_venice(
messages: Vec<LLMMessage>,
) -> anyhow::Result<Vec<ChatMessage>> {
let mut venice_messages: Vec<ChatMessage> = Vec::with_capacity(messages.len());
for message in messages {
venice_messages.push(convert_llm_message_to_venice(message)?);
}
Ok(venice_messages)
}
fn convert_llm_message_to_venice(message: LLMMessage) -> anyhow::Result<ChatMessage> {
let role = match message.author {
LLMAuthor::Prompt => "system",
LLMAuthor::Assistant => "assistant",
LLMAuthor::User => "user",
};
match message.content {
LLMMessageContent::Text(text) => Ok(ChatMessage {
role: role.to_owned(),
content: MessageContent::Text(text),
}),
LLMMessageContent::Image(image_details) => {
// Inline the image as a base64 data URI, the same shape the OpenAI vision content
// part uses. This is the gap the openai_compat provider can't fill (it drops images).
let data_uri = format!(
"data:{};base64,{}",
image_details.mime,
base64_encode(&image_details.data)
);
Ok(ChatMessage {
role: role.to_owned(),
content: MessageContent::Parts(vec![ContentPart::ImageUrl {
image_url: ImageUrl { url: data_uri },
}]),
})
}
LLMMessageContent::File(file_details) => {
// Inline the file as a base64 data URI in a `file` content part. This is the input
// type the openai_compat provider drops; baibot already extracts the bytes upstream.
// The message reaches the room, so it carries no user-controlled filename: a crafted
// name could otherwise inject markdown (a spoofed link) into the bot's reply.
if file_details.data.len() > MAX_FILE_BYTES {
return Err(anyhow::anyhow!(
"The attached file is too large for Venice (the limit is 25MB)."
));
}
let data_uri = format!(
"data:{};base64,{}",
file_details.mime,
base64_encode(&file_details.data)
);
Ok(ChatMessage {
role: role.to_owned(),
content: MessageContent::Parts(vec![ContentPart::File {
file: FilePart {
file_data: data_uri,
filename: Some(file_details.filename()),
},
}]),
})
}
}
}

View File

@@ -0,0 +1,273 @@
//! Serde structs modeling Venice's `/chat/completions`, `/audio/transcriptions`,
//! `/audio/speech`, `/image/generate`, and `/image/edit` wire shapes. Request types are
//! `Serialize`-only (we build them, Venice never sends them back); response types are
//! `Deserialize`-only. Keeping the split means the untagged request content enum is never on a
//! deserialize path, so a surprise response shape can't fail to match it.
//!
//! Field names match Venice's schema 1:1 (so the config's `model_id` becomes `model` here). Every
//! request body is `additionalProperties: false`, so optional knobs carry `skip_serializing_if`
//! to omit rather than send `null`. `/audio/speech` and `/image/edit` return raw binary (no
//! response struct); only `/image/generate` returns JSON (`GenerateImageResponse`).
use serde::{Deserialize, Serialize};
use super::config::VeniceParameters;
#[derive(Debug, Serialize)]
pub struct ChatCompletionRequest {
pub model: String,
pub messages: Vec<ChatMessage>,
#[serde(skip_serializing_if = "Option::is_none")]
pub temperature: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub max_completion_tokens: Option<u32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub top_p: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub frequency_penalty: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub presence_penalty: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub repetition_penalty: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub reasoning_effort: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub prompt_cache_key: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub prompt_cache_retention: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub venice_parameters: Option<VeniceParameters>,
}
#[derive(Debug, Serialize)]
pub struct ChatMessage {
pub role: String,
pub content: MessageContent,
}
/// A message body is either a bare string or a list of content parts. Venice accepts both; we
/// send the parts form when a message carries an image or a file (baibot keeps text, images, and
/// files in separate messages, so a parts list holds a single image part or file part).
#[derive(Debug, Serialize)]
#[serde(untagged)]
pub enum MessageContent {
Text(String),
Parts(Vec<ContentPart>),
}
#[derive(Debug, Serialize)]
#[serde(tag = "type", rename_all = "snake_case")]
pub enum ContentPart {
ImageUrl { image_url: ImageUrl },
File { file: FilePart },
}
#[derive(Debug, Serialize)]
pub struct ImageUrl {
/// A `data:<mime>;base64,<data>` URI for inline images.
pub url: String,
}
#[derive(Debug, Serialize)]
pub struct FilePart {
/// A `data:<mime>;base64,<data>` URI carrying the file bytes inline.
pub file_data: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub filename: Option<String>,
}
/// Standard OpenAI-shaped chat completion response. We read `choices[0].message.content` and,
/// when web search is on, the structured `venice_parameters.web_search_citations` (requested via
/// `return_search_results_as_documents`) to rewrite the inline `^n^` superscripts into readable
/// `[n]` references plus a `Sources:` block. `reasoning_content` carries the model's thinking when
/// the model exposes it; it is appended only when `show_reasoning` is set.
#[derive(Debug, Deserialize)]
pub struct ChatCompletionResponse {
pub choices: Vec<ChatChoice>,
#[serde(default)]
pub venice_parameters: Option<ResponseVeniceParameters>,
}
#[derive(Debug, Deserialize)]
pub struct ChatChoice {
pub message: ResponseMessage,
}
#[derive(Debug, Deserialize)]
pub struct ResponseMessage {
#[serde(default)]
pub content: Option<String>,
#[serde(default)]
pub reasoning_content: Option<String>,
}
/// The `venice_parameters` envelope on a chat-completion *response*, distinct from the request-side
/// `VeniceParameters` bag. Only the citation list is read; other response-side fields are ignored.
#[derive(Debug, Deserialize, Default)]
pub struct ResponseVeniceParameters {
#[serde(default)]
pub web_search_citations: Vec<WebSearchCitation>,
}
/// Only the `title` and `url` are read (for rendering the `Sources:` block). Venice also returns
/// `content` and `date` per citation; serde drops them, the same way the response structs above
/// ignore the response fields baibot does not use. Both fields default to empty so a single
/// citation that arrives without one (schema drift on scraped results) degrades gracefully in the
/// rendered list instead of failing the whole response deserialization.
#[derive(Debug, Deserialize)]
pub struct WebSearchCitation {
#[serde(default)]
pub title: String,
#[serde(default)]
pub url: String,
}
/// `/audio/transcriptions` response. We read `text`; the optional `duration`/`timestamps` the
/// API can return are not used in v1.
#[derive(Debug, Deserialize)]
pub struct TranscriptionResponse {
pub text: String,
}
/// `/audio/speech` (`CreateSpeechRequestSchema`) request. `input` and `model` are always sent;
/// the rest are omitted when unset. The response is raw binary audio, so there is no response
/// struct.
#[derive(Debug, Serialize)]
pub struct SpeechRequest {
pub model: String,
pub input: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub voice: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub speed: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub response_format: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub prompt: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub temperature: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub top_p: Option<f32>,
}
/// `/image/generate` (`GenerateImageRequest`) request. `return_binary` is pinned `false` and
/// `variants` to `1` by the builder: baibot wants exactly one image returned as base64-in-JSON,
/// which `GenerateImageResponse` then decodes. Flipping `return_binary` would make Venice answer
/// with raw binary and break that JSON decode, so it is not configurable.
#[derive(Debug, Serialize)]
pub struct GenerateImageRequest {
pub model: String,
pub prompt: String,
pub return_binary: bool,
pub variants: u32,
#[serde(skip_serializing_if = "Option::is_none")]
pub negative_prompt: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub cfg_scale: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub steps: Option<u32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub style_preset: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub seed: Option<i64>,
#[serde(skip_serializing_if = "Option::is_none")]
pub safe_mode: Option<bool>,
#[serde(skip_serializing_if = "Option::is_none")]
pub hide_watermark: Option<bool>,
#[serde(skip_serializing_if = "Option::is_none")]
pub format: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub width: Option<u32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub height: Option<u32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub aspect_ratio: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub resolution: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub quality: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub lora_strength: Option<u32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub embed_exif_metadata: Option<bool>,
#[serde(skip_serializing_if = "Option::is_none")]
pub enable_web_search: Option<bool>,
}
/// `/image/generate` response when `return_binary` is false: a JSON envelope carrying the images
/// as base64 strings. We read `images[0]`; `request`/`timing` and other fields are ignored. `id`
/// is telemetry only (logged, never used for correctness), so it is optional: a response that
/// carries usable `images` must not fail to deserialize just because the telemetry field drifted.
#[derive(Debug, Deserialize)]
pub struct GenerateImageResponse {
#[serde(default)]
pub id: Option<String>,
pub images: Vec<String>,
}
/// `/image/edit` (`EditImageRequest`) request. The source `image` is a base64-encoded string
/// (Venice's `image` field is `anyOf` upload/base64/URL; we send base64-in-JSON, no multipart).
/// The response is raw binary, so there is no response struct.
#[derive(Debug, Serialize)]
pub struct EditImageRequest {
pub model: String,
pub prompt: String,
/// Base64-encoded source image bytes.
pub image: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub output_format: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub aspect_ratio: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub resolution: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub safe_mode: Option<bool>,
}

View File

@@ -1,8 +1,12 @@
use mxlink::matrix_sdk::{ use mxlink::matrix_sdk::{
Room, Room,
room::edit::EditedContent,
ruma::{ ruma::{
OwnedEventId, api::client::receipt::create_receipt::v3::ReceiptType, EventId, OwnedEventId,
events::room::message::OriginalSyncRoomMessageEvent, api::client::receipt::create_receipt::v3::ReceiptType,
events::room::message::{
OriginalSyncRoomMessageEvent, RoomMessageEventContentWithoutRelation,
},
}, },
}; };
@@ -52,6 +56,83 @@ impl Messaging {
} }
} }
/// Like `send_text_markdown_no_fail`, but logs send failures at `warn` instead of `error`.
/// For cosmetic, best-effort sends (the thinking-notice placeholder) where a failure is
/// acceptable-impact: the real response still ships, so this must NOT page the team.
pub async fn send_text_markdown_no_fail_quietly(
&self,
room: &Room,
message: String,
response_type: MessageResponseType,
) -> Option<mxlink::matrix_sdk::ruma::api::client::message::send_message_event::v3::Response>
{
let result = self
.bot
.matrix_link()
.messaging()
.send_text_markdown(room, message, response_type)
.await;
match result {
Ok(result) => Some(result),
Err(err) => {
tracing::warn!(
room_id = format!("{:?}", room.room_id()),
?err,
"Failed to send thinking-notice placeholder to room",
);
None
}
}
}
/// Edits an existing message's text in place (`m.replace`), best-effort.
///
/// Built and sent directly via matrix-sdk's `make_edit_event` + `room.send`,
/// deliberately bypassing the mxlink messaging layer: that layer overwrites
/// `relates_to` from a `MessageResponseType`, which would clobber the
/// `m.replace` relation and turn the edit into a brand-new message. The edit
/// stays in the original's thread by inheritance, so no `response_type` is needed.
/// Failures are warned and swallowed so a flickered notice never breaks the real response.
pub async fn edit_text_markdown_no_fail(
&self,
room: &Room,
event_id: &EventId,
markdown: String,
) -> Option<mxlink::matrix_sdk::ruma::api::client::message::send_message_event::v3::Response>
{
let new_content = RoomMessageEventContentWithoutRelation::text_markdown(markdown);
let edit_content = match room
.make_edit_event(event_id, EditedContent::RoomMessage(new_content))
.await
{
Ok(edit_content) => edit_content,
Err(err) => {
tracing::warn!(
room_id = format!("{:?}", room.room_id()),
?event_id,
?err,
"Failed to build edit event",
);
return None;
}
};
match room.send(edit_content).await {
Ok(result) => Some(result.response),
Err(err) => {
tracing::warn!(
room_id = format!("{:?}", room.room_id()),
?event_id,
?err,
"Failed to send edit to room",
);
None
}
}
}
pub async fn send_notice_markdown_no_fail( pub async fn send_notice_markdown_no_fail(
&self, &self,
room: &Room, room: &Room,

View File

@@ -9,7 +9,7 @@ fn agent_config_parsing_works() {
let sample_config = crate::agent::default_config_for_provider(&provider); let sample_config = crate::agent::default_config_for_provider(&provider);
let sample_config_pretty_yaml = serde_yaml_ng::to_string(&sample_config).unwrap(); let sample_config_pretty_yaml = serde_yaml_ng::to_string(&sample_config).unwrap();
let test_cases = vec![ let test_cases = [
// Invalid input // Invalid input
TestCase { TestCase {
input: r#"Hello"#.to_owned(), input: r#"Hello"#.to_owned(),

View File

@@ -38,6 +38,9 @@ pub enum ConfigTextGenerationSettingRelatedControllerType {
GetContextManagementEnabled, GetContextManagementEnabled,
SetContextManagementEnabled(Option<bool>), SetContextManagementEnabled(Option<bool>),
GetThinkingNoticeEnabled,
SetThinkingNoticeEnabled(Option<bool>),
GetPrefixRequirementType, GetPrefixRequirementType,
SetPrefixRequirementType(Option<TextGenerationPrefixRequirementType>), SetPrefixRequirementType(Option<TextGenerationPrefixRequirementType>),

View File

@@ -55,6 +55,44 @@ pub(super) fn determine(
); );
} }
if let Some(remaining_text) = text.strip_prefix("thinking-notice-enabled") {
let remaining_text = remaining_text.trim();
if !remaining_text.is_empty() {
return Err(ControllerType::Error(
strings::cfg::configuration_getter_used_with_extra_text(
"thinking-notice-enabled",
remaining_text,
)
.to_owned(),
));
}
return Ok(ConfigTextGenerationSettingRelatedControllerType::GetThinkingNoticeEnabled);
}
if let Some(value_string) = text.strip_prefix("set-thinking-notice-enabled") {
let value_string = value_string.trim().to_owned();
let value_opt = if value_string.is_empty() {
None
} else {
let value_string_lowercase = value_string.to_lowercase();
Some(match value_string_lowercase.as_str() {
"true" => true,
"false" => false,
_ => {
return Err(ControllerType::Error(
strings::cfg::configuration_value_unrecognized(&value_string).to_owned(),
));
}
})
};
return Ok(
ConfigTextGenerationSettingRelatedControllerType::SetThinkingNoticeEnabled(value_opt),
);
}
if let Some(remaining_text) = text.strip_prefix("prefix-requirement-type") { if let Some(remaining_text) = text.strip_prefix("prefix-requirement-type") {
let remaining_text = remaining_text.trim(); let remaining_text = remaining_text.trim();

View File

@@ -43,6 +43,27 @@ pub(super) async fn dispatch(
} }
} }
ConfigTextGenerationSettingRelatedControllerType::GetThinkingNoticeEnabled => {
let value = &room_settings.text_generation.thinking_notice_enabled;
setting_get::<bool>(bot, message_context, value).await
}
ConfigTextGenerationSettingRelatedControllerType::SetThinkingNoticeEnabled(value) => {
let value = value.to_owned();
let setter_callback = Box::new(move |room_settings: &mut RoomSettings| {
room_settings.text_generation.thinking_notice_enabled = value;
});
match config_type {
SettingsStorageSource::Room => {
room_setting_set::<bool>(bot, message_context, &value, setter_callback).await
}
SettingsStorageSource::Global => {
global_setting_set::<bool>(bot, message_context, &value, setter_callback).await
}
}
}
ConfigTextGenerationSettingRelatedControllerType::GetPrefixRequirementType => { ConfigTextGenerationSettingRelatedControllerType::GetPrefixRequirementType => {
let value = &room_settings.text_generation.prefix_requirement_type; let value = &room_settings.text_generation.prefix_requirement_type;
setting_get::<TextGenerationPrefixRequirementType>(bot, message_context, value).await setting_get::<TextGenerationPrefixRequirementType>(bot, message_context, value).await

View File

@@ -234,6 +234,44 @@ fn build_section_text_generation(command_prefix: &str, bot_username: &str) -> St
)); ));
message.push_str("\n\n"); message.push_str("\n\n");
// Thinking Notice
message.push_str(&format!(
"#### {}",
strings::help::cfg::text_generation_thinking_notice_heading()
));
message.push_str("\n\n");
message.push_str(&strings::help::cfg::text_generation_thinking_notice_intro());
message.push('\n');
message.push_str(
&strings::help::cfg::the_following_configuration_values_are_recognized(vec![true, false]),
);
message.push_str("\n\n");
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_show(
command_prefix,
"text-generation thinking-notice-enabled"
)
));
message.push('\n');
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_set(
command_prefix,
"text-generation set-thinking-notice-enabled VALUE"
)
));
message.push('\n');
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_unset(
command_prefix,
"text-generation set-thinking-notice-enabled"
)
));
message.push_str("\n\n");
// Sender Context // Sender Context
message.push_str(&format!( message.push_str(&format!(

View File

@@ -359,6 +359,33 @@ async fn generate_text_generation_section(
), ),
); );
// Thinking Notice
let effective_thinking_notice = room_config_context.text_generation_thinking_notice_enabled();
let room_config_thinking_notice = room_config_context
.room_config
.settings
.text_generation
.thinking_notice_enabled;
let global_config_thinking_notice = room_config_context
.global_config
.fallback_room_settings
.text_generation
.thinking_notice_enabled;
let thinking_notice_set_where = if room_config_thinking_notice.is_some() {
strings::cfg::status_badge_set_in_room_config()
} else if global_config_thinking_notice.is_some() {
strings::cfg::status_badge_set_in_global_config()
} else {
strings::cfg::status_badge_using_hardcoded_default()
};
message.push_str(&strings::cfg::status_text_generation_entry_thinking_notice(
effective_thinking_notice,
thinking_notice_set_where,
));
// Sender Context // Sender Context
let effective_sender_context = room_config_context.text_generation_sender_context_mode(); let effective_sender_context = room_config_context.text_generation_sender_context_mode();

View File

@@ -30,6 +30,13 @@ use crate::{
entity::MessageContext, entity::MessageContext,
}; };
/// How long text generation must run before the first "thinking…" placeholder appears.
/// Fast responses (under this threshold) never get a placeholder.
const THINKING_NOTICE_FIRST_DELAY: std::time::Duration = std::time::Duration::from_secs(3);
/// How often the "thinking…" placeholder is refreshed once it has appeared.
const THINKING_NOTICE_INTERVAL: std::time::Duration = std::time::Duration::from_secs(10);
#[derive(Debug, PartialEq)] #[derive(Debug, PartialEq)]
pub enum ChatCompletionControllerType { pub enum ChatCompletionControllerType {
// Invoked via a command prefix (e.g. `!bai Hello!`) // Invoked via a command prefix (e.g. `!bai Hello!`)
@@ -520,6 +527,16 @@ async fn handle_stage_text_generation(
conversation.start_time(), conversation.start_time(),
); );
// Cloned only when the thinking-notice is enabled; the original is moved into `params` below.
let notice_prompt_variables = if message_context
.room_config_context()
.text_generation_thinking_notice_enabled()
{
Some(prompt_variables.clone())
} else {
None
};
let params = TextGenerationParams { let params = TextGenerationParams {
context_management_enabled: message_context context_management_enabled: message_context
.room_config_context() .room_config_context()
@@ -536,10 +553,75 @@ async fn handle_stage_text_generation(
prompt_variables, prompt_variables,
}; };
let result = controller // When the thinking-notice is enabled, race generation against a timer that posts and then
.generate_text(conversation, params) // periodically edits a "thinking…" placeholder. `biased;` makes generation win a tie, and the
.instrument(span) // loop exits the instant generation resolves, so there is no detached task and no late edit can
.await; // ever clobber the real answer. `placeholder` is the event we must finalize in every exit path.
let (result, placeholder) = if let Some(notice_prompt_variables) = notice_prompt_variables {
let generation = controller
.generate_text(conversation, params)
.instrument(span);
tokio::pin!(generation);
let mut placeholder: Option<OwnedEventId> = None;
// Seed the flavor sequence per-generation so different turns don't all open on the same
// line; the monotonic increment then guarantees consecutive notices differ.
let mut notice_sequence: usize = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|since| since.subsec_nanos() as usize)
.unwrap_or(0);
let mut tick = tokio::time::interval_at(
tokio::time::Instant::now() + THINKING_NOTICE_FIRST_DELAY,
THINKING_NOTICE_INTERVAL,
);
// If an edit runs long, hold ~INTERVAL spacing rather than bursting the missed ticks.
tick.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Delay);
let result = loop {
tokio::select! {
biased;
generation_result = &mut generation => break generation_result,
_ = tick.tick() => {
let message = notice_prompt_variables.format(
strings::thinking::pick_message(start_time.elapsed(), notice_sequence),
);
notice_sequence = notice_sequence.wrapping_add(1);
match &placeholder {
None => {
placeholder = bot
.messaging()
.send_text_markdown_no_fail_quietly(
message_context.room(),
message,
response_type.clone(),
)
.await
.map(|response| response.event_id);
}
Some(event_id) => {
bot.messaging()
.edit_text_markdown_no_fail(
message_context.room(),
event_id,
message,
)
.await;
}
}
}
}
};
(result, placeholder)
} else {
let result = controller
.generate_text(conversation, params)
.instrument(span)
.await;
(result, None)
};
let duration = std::time::Instant::now().duration_since(start_time); let duration = std::time::Instant::now().duration_since(start_time);
@@ -560,17 +642,20 @@ async fn handle_stage_text_generation(
err, err,
); );
bot.messaging() let error_text = strings::agent::error_while_serving_purpose(
.send_error_markdown_no_fail( agent.identifier(),
message_context.room(), &AgentPurpose::TextGeneration,
&strings::agent::error_while_serving_purpose( &err,
agent.identifier(), );
&AgentPurpose::TextGeneration,
&err, finalize_thinking_notice_with_error(
), bot,
response_type, message_context,
) placeholder.as_ref(),
.await; &error_text,
response_type,
)
.await;
return None; return None;
} }
@@ -583,26 +668,70 @@ async fn handle_stage_text_generation(
"Agent returned empty text", "Agent returned empty text",
); );
bot.messaging() let empty_text = strings::agent::empty_response_returned(agent.identifier());
.send_error_markdown_no_fail(
message_context.room(), finalize_thinking_notice_with_error(
&strings::agent::empty_response_returned(agent.identifier()), bot,
response_type, message_context,
) placeholder.as_ref(),
.await; &empty_text,
response_type,
)
.await;
return None; return None;
} }
let send_message_response = bot // Finalize the answer into a single message. With a placeholder, edit it in place (so the
.messaging() // "thinking…" message becomes the answer); the TTS payload then points at that same event.
.send_text_markdown_no_fail(message_context.room(), text.clone(), response_type) // If the edit fails, fall back to a fresh send so the real answer is never lost.
.await?; let event_id = match &placeholder {
Some(event_id)
if bot
.messaging()
.edit_text_markdown_no_fail(message_context.room(), event_id, text.clone())
.await
.is_some() =>
{
event_id.clone()
}
_ => {
bot.messaging()
.send_text_markdown_no_fail(message_context.room(), text.clone(), response_type)
.await?
.event_id
}
};
Some(TextToSpeechEligiblePayload { Some(TextToSpeechEligiblePayload { text, event_id })
text, }
event_id: send_message_response.event_id,
}) /// Finalizes a thinking-notice placeholder (if one was posted) with error/notice text, so a
/// failed or empty generation never leaves an orphaned "thinking…" message behind. With no
/// placeholder, this is the original behavior: a fresh error notice.
async fn finalize_thinking_notice_with_error(
bot: &Bot,
message_context: &MessageContext,
placeholder: Option<&OwnedEventId>,
text: &str,
response_type: MessageResponseType,
) {
match placeholder {
Some(event_id) => {
bot.messaging()
.edit_text_markdown_no_fail(
message_context.room(),
event_id,
crate::utils::status::create_error_message_text(text),
)
.await;
}
None => {
bot.messaging()
.send_error_markdown_no_fail(message_context.room(), text, response_type)
.await;
}
}
} }
async fn handle_stage_speech_to_text_actual_transcribing( async fn handle_stage_speech_to_text_actual_transcribing(

View File

@@ -6,5 +6,5 @@ mod utils;
mod tests; mod tests;
pub use entity::*; pub use entity::*;
pub use tokenization::shorten_messages_list_to_context_size; pub use tokenization::{TokenEstimate, shorten_messages_list_to_context_size};
pub use utils::*; pub use utils::*;

View File

@@ -4,6 +4,22 @@ use tiktoken_rs::tokenizer;
use super::{Author, Message, MessageContent}; use super::{Author, Message, MessageContent};
/// How to count the tokens in a conversation when trimming it to fit the context window.
pub enum TokenEstimate<'a> {
/// Count via the [tiktoken-rs](https://github.com/zurawiki/tiktoken-rs) library.
/// Accurate for OpenAI models; every other model falls back to the gpt-4
/// tokenizer, which misreads it (badly so for non-English text). Use this only
/// for the OpenAI provider.
Tiktoken(&'a str),
/// Provider-neutral approximation that needs no per-model tokenizer. Expect it
/// to land within roughly 10-20% of the real count for typical text, leaning
/// slightly high: over-counting trims a little extra history, while
/// under-counting would overflow the model's real context window. Use this for
/// every non-OpenAI provider.
Approximate,
}
fn get_bpe_for_model(model: &str) -> &'static CoreBPE { fn get_bpe_for_model(model: &str) -> &'static CoreBPE {
let tokenizer = tokenizer::get_tokenizer(model) let tokenizer = tokenizer::get_tokenizer(model)
.or_else(|| tokenizer::get_tokenizer("gpt-4")) .or_else(|| tokenizer::get_tokenizer("gpt-4"))
@@ -13,21 +29,27 @@ fn get_bpe_for_model(model: &str) -> &'static CoreBPE {
} }
pub fn shorten_messages_list_to_context_size( pub fn shorten_messages_list_to_context_size(
model: &str, estimate: TokenEstimate<'_>,
prompt_message: &Option<Message>, prompt_message: &Option<Message>,
mut messages: Vec<Message>, mut messages: Vec<Message>,
max_response_tokens: Option<u32>, max_response_tokens: Option<u32>,
max_context_tokens: u32, max_context_tokens: u32,
) -> Vec<Message> { ) -> Vec<Message> {
// Loading the tokenization data is an expensive process, so // Loading the tiktoken data is expensive, so we resolve the counter once up
// se construct the BPE instance once and then use it for all messages. // front and reuse it for every message.
let bpe = get_bpe_for_model(model); let tiktoken = match estimate {
TokenEstimate::Tiktoken(model) => Some((get_bpe_for_model(model), model)),
TokenEstimate::Approximate => None,
};
let count = |message: &Message| match tiktoken {
Some((bpe, model)) => tiktoken_token_size_for_message(bpe, model, message),
None => approximate_token_size_for_message(message),
};
// We want to retain the prompt in all cases, so we always count it first. // We want to retain the prompt in all cases, so we always count it first.
// We also always reserve enough tokens for the maximum response we expect. // We also always reserve enough tokens for the maximum response we expect.
let mut current_context_length: u32 = if let Some(prompt_message) = prompt_message { let mut current_context_length: u32 = if let Some(prompt_message) = prompt_message {
calculate_token_size_for_message(bpe, model, prompt_message) count(prompt_message) + max_response_tokens.unwrap_or(0)
+ max_response_tokens.unwrap_or(0)
} else { } else {
0 0
}; };
@@ -37,7 +59,7 @@ pub fn shorten_messages_list_to_context_size(
let mut messages_to_keep: Vec<Message> = Vec::new(); let mut messages_to_keep: Vec<Message> = Vec::new();
for message in messages { for message in messages {
let tokens_for_message = calculate_token_size_for_message(bpe, model, &message); let tokens_for_message = count(&message);
if current_context_length + tokens_for_message > max_context_tokens { if current_context_length + tokens_for_message > max_context_tokens {
break; break;
@@ -48,14 +70,26 @@ pub fn shorten_messages_list_to_context_size(
messages_to_keep.push(message); messages_to_keep.push(message);
} }
// Cut on a turn boundary: the loop may stop right after an assistant reply
// whose triggering user message did not fit, which would leave the kept window
// starting on an orphaned reply. `messages_to_keep` is newest-first here, so
// the oldest kept messages are at the end; drop any trailing assistant messages
// until the window begins at the start of a turn (a user message).
while matches!(
messages_to_keep.last().map(|message| &message.author),
Some(Author::Assistant)
) {
messages_to_keep.pop();
}
messages_to_keep.reverse(); messages_to_keep.reverse();
messages_to_keep messages_to_keep
} }
/// Calculate the token size of a message for a given model, with a preloaded CoreBPE object. /// Token size of a message via tiktoken, for a preloaded CoreBPE object.
/// Related to `calculate_token_size_for_model_message`. /// Accurate only for OpenAI models (see [`TokenEstimate::Tiktoken`]).
fn calculate_token_size_for_message(bpe: &CoreBPE, model: &str, message: &Message) -> u32 { fn tiktoken_token_size_for_message(bpe: &CoreBPE, model: &str, message: &Message) -> u32 {
let (tokens_per_message, tokens_per_name) = if model.starts_with("gpt-3.5") { let (tokens_per_message, tokens_per_name) = if model.starts_with("gpt-3.5") {
( (
4, // every message follows <im_start>{role/name}\n{content}<im_end>\n 4, // every message follows <im_start>{role/name}\n{content}<im_end>\n
@@ -80,6 +114,52 @@ fn calculate_token_size_for_message(bpe: &CoreBPE, model: &str, message: &Messag
(text_length + role_length + tokens_per_message + tokens_per_name) as u32 (text_length + role_length + tokens_per_message + tokens_per_name) as u32
} }
/// ASCII text averages about four characters per token.
const ASCII_TOKENS_PER_CHAR: f32 = 0.25;
/// Non-ASCII scripts (Cyrillic, CJK, and others) pack more information per
/// character: real tokenizers land around two characters per token for them, so
/// each counts as half a token. CJK runs a touch denser than that, so its estimate
/// can read slightly low, still within the tolerance this approximation targets.
const WIDE_TOKENS_PER_CHAR: f32 = 0.5;
/// Structural per-message overhead (role marker plus message framing), mirroring
/// the small constant the tiktoken path adds.
const APPROX_TOKENS_PER_MESSAGE: u32 = 4;
/// Provider-neutral, tokenizer-free token size of a message
/// (see [`TokenEstimate::Approximate`]).
fn approximate_token_size_for_message(message: &Message) -> u32 {
let text_tokens = match &message.content {
MessageContent::Text(text) => approximate_token_size_for_text(text),
// Images and files are not counted as text, matching the tiktoken path.
MessageContent::Image(..) | MessageContent::File(..) => 0,
};
text_tokens + APPROX_TOKENS_PER_MESSAGE
}
/// Rough token estimate for a piece of text, with no tokenizer.
///
/// ASCII characters count as a quarter-token each (~4 chars/token); characters
/// outside ASCII count as half a token each (~2 chars/token), matching how real
/// tokenizers treat Cyrillic and CJK. Weighting non-ASCII up keeps the estimate
/// from badly under-counting non-English text, the case the tiktoken fallback gets
/// most wrong.
fn approximate_token_size_for_text(text: &str) -> u32 {
let mut estimate = 0.0_f32;
for character in text.chars() {
estimate += if character.is_ascii() {
ASCII_TOKENS_PER_CHAR
} else {
WIDE_TOKENS_PER_CHAR
};
}
estimate.ceil() as u32
}
pub mod test { pub mod test {
#[test] #[test]
fn message_size_counting_works() { fn message_size_counting_works() {
@@ -94,7 +174,7 @@ pub mod test {
timestamp: chrono::Utc::now(), timestamp: chrono::Utc::now(),
}; };
let tokens = super::calculate_token_size_for_message(bpe, model, &message); let tokens = super::tiktoken_token_size_for_message(bpe, model, &message);
assert_eq!(8, tokens); assert_eq!(8, tokens);
} }
@@ -105,7 +185,8 @@ pub mod test {
let bpe = super::get_bpe_for_model(model); let bpe = super::get_bpe_for_model(model);
let max_response_tokens: Option<u32> = Some(5); let max_response_tokens_value: u32 = 5;
let max_response_tokens: Option<u32> = Some(max_response_tokens_value);
let prompt = super::Message { let prompt = super::Message {
author: super::Author::Prompt, author: super::Author::Prompt,
@@ -117,7 +198,7 @@ pub mod test {
assert_eq!( assert_eq!(
prompt_length, prompt_length,
super::calculate_token_size_for_message(bpe, model, &prompt) super::tiktoken_token_size_for_message(bpe, model, &prompt)
); );
let mut conversation_messages = Vec::new(); let mut conversation_messages = Vec::new();
@@ -132,7 +213,7 @@ pub mod test {
assert_eq!( assert_eq!(
first_length, first_length,
super::calculate_token_size_for_message(bpe, model, &first) super::tiktoken_token_size_for_message(bpe, model, &first)
); );
conversation_messages.push(first); conversation_messages.push(first);
@@ -147,7 +228,7 @@ pub mod test {
assert_eq!( assert_eq!(
second_length, second_length,
super::calculate_token_size_for_message(bpe, model, &second) super::tiktoken_token_size_for_message(bpe, model, &second)
); );
conversation_messages.push(second); conversation_messages.push(second);
@@ -164,7 +245,7 @@ pub mod test {
assert_eq!( assert_eq!(
third_length, third_length,
super::calculate_token_size_for_message(bpe, model, &third) super::tiktoken_token_size_for_message(bpe, model, &third)
); );
conversation_messages.push(third.clone()); conversation_messages.push(third.clone());
@@ -181,7 +262,7 @@ pub mod test {
assert_eq!( assert_eq!(
forth_length, forth_length,
super::calculate_token_size_for_message(bpe, model, &forth) super::tiktoken_token_size_for_message(bpe, model, &forth)
); );
conversation_messages.push(forth.clone()); conversation_messages.push(forth.clone());
@@ -189,11 +270,11 @@ pub mod test {
assert_eq!(4, conversation_messages.len()); assert_eq!(4, conversation_messages.len());
let new_conversation_messages = super::shorten_messages_list_to_context_size( let new_conversation_messages = super::shorten_messages_list_to_context_size(
model, super::TokenEstimate::Tiktoken(model),
&Some(prompt), &Some(prompt),
conversation_messages, conversation_messages,
max_response_tokens, max_response_tokens,
prompt_length + max_response_tokens.unwrap_or(0) + forth_length + third_length, prompt_length + max_response_tokens_value + forth_length + third_length,
); );
assert_eq!(2, new_conversation_messages.len()); assert_eq!(2, new_conversation_messages.len());
@@ -215,7 +296,8 @@ pub mod test {
let bpe = super::get_bpe_for_model(model); let bpe = super::get_bpe_for_model(model);
let max_response_tokens: Option<u32> = Some(5); let max_response_tokens_value: u32 = 5;
let max_response_tokens: Option<u32> = Some(max_response_tokens_value);
let prompt = super::Message { let prompt = super::Message {
author: super::Author::User, author: super::Author::User,
@@ -227,7 +309,7 @@ pub mod test {
assert_eq!( assert_eq!(
prompt_length, prompt_length,
super::calculate_token_size_for_message(bpe, model, &prompt) super::tiktoken_token_size_for_message(bpe, model, &prompt)
); );
let mut conversation_messages = Vec::new(); let mut conversation_messages = Vec::new();
@@ -242,7 +324,7 @@ pub mod test {
assert_eq!( assert_eq!(
first_length, first_length,
super::calculate_token_size_for_message(bpe, model, &first) super::tiktoken_token_size_for_message(bpe, model, &first)
); );
conversation_messages.push(first); conversation_messages.push(first);
@@ -257,7 +339,7 @@ pub mod test {
assert_eq!( assert_eq!(
second_length, second_length,
super::calculate_token_size_for_message(bpe, model, &second) super::tiktoken_token_size_for_message(bpe, model, &second)
); );
conversation_messages.push(second); conversation_messages.push(second);
@@ -274,7 +356,7 @@ pub mod test {
assert_eq!( assert_eq!(
third_length, third_length,
super::calculate_token_size_for_message(bpe, model, &third) super::tiktoken_token_size_for_message(bpe, model, &third)
); );
conversation_messages.push(third.clone()); conversation_messages.push(third.clone());
@@ -291,7 +373,7 @@ pub mod test {
assert_eq!( assert_eq!(
forth_length, forth_length,
super::calculate_token_size_for_message(bpe, model, &forth) super::tiktoken_token_size_for_message(bpe, model, &forth)
); );
conversation_messages.push(forth.clone()); conversation_messages.push(forth.clone());
@@ -299,11 +381,11 @@ pub mod test {
assert_eq!(4, conversation_messages.len()); assert_eq!(4, conversation_messages.len());
let new_conversation_messages = super::shorten_messages_list_to_context_size( let new_conversation_messages = super::shorten_messages_list_to_context_size(
model, super::TokenEstimate::Tiktoken(model),
&Some(prompt), &Some(prompt),
conversation_messages, conversation_messages,
max_response_tokens, max_response_tokens,
prompt_length + max_response_tokens.unwrap_or(0) + forth_length + third_length, prompt_length + max_response_tokens_value + forth_length + third_length,
); );
assert_eq!(2, new_conversation_messages.len()); assert_eq!(2, new_conversation_messages.len());
@@ -318,4 +400,126 @@ pub mod test {
forth.content forth.content
); );
} }
#[test]
fn approximate_counting_weights_ascii_and_wide_scripts() {
// 12 ASCII characters at ~4 chars/token = 3 text tokens.
assert_eq!(3, super::approximate_token_size_for_text("Hello there!"));
// 5 CJK characters at ~0.5 token/char = 3 text tokens. The ASCII rate would
// have under-counted these to 2, the failure mode this path avoids.
assert_eq!(3, super::approximate_token_size_for_text("こんにちは"));
let message = super::Message {
author: super::Author::User,
sender_id: None,
content: super::MessageContent::Text("Hello there!".to_string()),
timestamp: chrono::Utc::now(),
};
// 3 text tokens plus the per-message overhead (4).
assert_eq!(7, super::approximate_token_size_for_message(&message));
}
#[test]
fn approximate_shortening_trims_to_budget() {
let prompt = super::Message {
author: super::Author::Prompt,
sender_id: None,
content: super::MessageContent::Text("You are a bot!".to_string()),
timestamp: chrono::Utc::now(),
};
let older = super::Message {
author: super::Author::User,
sender_id: None,
content: super::MessageContent::Text("This is the older message.".to_string()),
timestamp: chrono::Utc::now(),
};
let newer = super::Message {
// A user message, so it is a valid window start: keeping a lone
// assistant reply would be an orphan and get trimmed (see
// `shortening_cuts_on_a_turn_boundary`).
author: super::Author::User,
sender_id: None,
content: super::MessageContent::Text("This is the newer message.".to_string()),
timestamp: chrono::Utc::now(),
};
// Budget room for the prompt and only the newest message.
let max_context_tokens = super::approximate_token_size_for_message(&prompt)
+ super::approximate_token_size_for_message(&newer);
let new_conversation_messages = super::shorten_messages_list_to_context_size(
super::TokenEstimate::Approximate,
&Some(prompt),
vec![older, newer.clone()],
None,
max_context_tokens,
);
assert_eq!(1, new_conversation_messages.len());
assert_eq!(
new_conversation_messages.first().unwrap().content,
newer.content
);
}
#[test]
fn shortening_cuts_on_a_turn_boundary() {
// A four-message conversation of two full turns. All four messages are the
// same length, so they cost the same number of tokens.
let prompt = super::Message {
author: super::Author::Prompt,
sender_id: None,
content: super::MessageContent::Text("system".to_string()),
timestamp: chrono::Utc::now(),
};
let user_one = super::Message {
author: super::Author::User,
sender_id: None,
content: super::MessageContent::Text("user msg 1".to_string()),
timestamp: chrono::Utc::now(),
};
let asst_one = super::Message {
author: super::Author::Assistant,
sender_id: None,
content: super::MessageContent::Text("asst msg 1".to_string()),
timestamp: chrono::Utc::now(),
};
let user_two = super::Message {
author: super::Author::User,
sender_id: None,
content: super::MessageContent::Text("user msg 2".to_string()),
timestamp: chrono::Utc::now(),
};
let asst_two = super::Message {
author: super::Author::Assistant,
sender_id: None,
content: super::MessageContent::Text("asst msg 2".to_string()),
timestamp: chrono::Utc::now(),
};
let per_message = super::approximate_token_size_for_message(&user_one);
// Budget fits the prompt plus three messages. By raw token budget the loop
// would keep asst_two, user_two, and asst_one, but asst_one's own user
// message (user_one) does not fit, so it must be dropped too rather than
// left as an orphaned reply.
let max_context_tokens =
super::approximate_token_size_for_message(&prompt) + (per_message * 3);
let kept = super::shorten_messages_list_to_context_size(
super::TokenEstimate::Approximate,
&Some(prompt),
vec![user_one, asst_one, user_two.clone(), asst_two.clone()],
None,
max_context_tokens,
);
// Only the last whole turn survives; the orphaned asst_one is dropped.
assert_eq!(2, kept.len());
assert_eq!(kept.first().unwrap().content, user_two.content);
assert_eq!(kept.last().unwrap().content, asst_two.content);
}
} }

View File

@@ -136,6 +136,20 @@ impl RoomConfigContext {
.unwrap_or(false) .unwrap_or(false)
} }
pub fn text_generation_thinking_notice_enabled(&self) -> bool {
self.room_config
.settings
.text_generation
.thinking_notice_enabled
.or({
self.global_config
.fallback_room_settings
.text_generation
.thinking_notice_enabled
})
.unwrap_or(false)
}
pub fn text_generation_sender_context_mode(&self) -> TextGenerationSenderContextMode { pub fn text_generation_sender_context_mode(&self) -> TextGenerationSenderContextMode {
self.room_config self.room_config
.settings .settings

View File

@@ -14,6 +14,10 @@ pub struct RoomSettingsTextGeneration {
/// When enabled, the bot will automatically tokenize messages and try to shorten the message context intelligently. /// When enabled, the bot will automatically tokenize messages and try to shorten the message context intelligently.
pub context_management_enabled: Option<bool>, pub context_management_enabled: Option<bool>,
/// Controls whether a "thinking…" notice is posted while text generation runs longer than a threshold.
/// When enabled, a placeholder message appears for slow responses and is edited in place (with elapsed-tiered flavor text) until it becomes the final answer.
pub thinking_notice_enabled: Option<bool>,
/// Controls how each message in the conversation context is annotated with sender metadata. /// Controls how each message in the conversation context is annotated with sender metadata.
pub sender_context_mode: Option<TextGenerationSenderContextMode>, pub sender_context_mode: Option<TextGenerationSenderContextMode>,

View File

@@ -249,6 +249,10 @@ pub fn status_text_generation_entry_context_management(value: bool, set_where: &
format!("- ♻️ Context management: `{}` ({})\n", value, set_where) format!("- ♻️ Context management: `{}` ({})\n", value, set_where)
} }
pub fn status_text_generation_entry_thinking_notice(value: bool, set_where: &str) -> String {
format!("- 💭 Thinking notice: `{}` ({})\n", value, set_where)
}
pub fn status_text_generation_entry_sender_context( pub fn status_text_generation_entry_sender_context(
value: impl std::fmt::Display, value: impl std::fmt::Display,
set_where: &str, set_where: &str,

View File

@@ -128,7 +128,19 @@ pub fn text_generation_context_management_intro() -> String {
format!( format!(
"{}\n{}", "{}\n{}",
"Controls the bot's ability to **intelligently drop old messages from the conversation context** when it gets too large.", "Controls the bot's ability to **intelligently drop old messages from the conversation context** when it gets too large.",
"This feature relies on [tokenization](https://en.wikipedia.org/wiki/Large_language_model#Tokenization) performed by the [tiktoken-rs](https://github.com/zurawiki/tiktoken-rs) library which is [poorly well-maintained](https://github.com/zurawiki/tiktoken-rs/issues/50) and only works well for [OpenAI](./providers.md#openai) models.", "Counting tokens precisely needs the model's own tokenizer. For [OpenAI](./providers.md#openai) models the bot uses the [tiktoken-rs](https://github.com/zurawiki/tiktoken-rs) library; for every other provider (including the recommended [Venice](./providers.md#venice)) it falls back to a provider-neutral **approximation** (ASCII counted at ~4 characters per token, other scripts at ~2), within roughly 10-20% of the real count for typical text.",
)
}
pub fn text_generation_thinking_notice_heading() -> &'static str {
"💭 Thinking Notice"
}
pub fn text_generation_thinking_notice_intro() -> String {
format!(
"{}\n{}",
"Controls whether the bot posts a **\"thinking…\" notice** while text generation takes a while, so slow responses (for example, from reasoning models that run for minutes) don't look stuck.",
"When enabled, a placeholder message appears only after a short delay, updates periodically with varying status text, and is then edited in place to become the final answer. Disabled by default.",
) )
} }

View File

@@ -11,6 +11,7 @@ pub mod provider;
pub mod room_config; pub mod room_config;
pub mod speech_to_text; pub mod speech_to_text;
pub mod text_to_speech; pub mod text_to_speech;
pub mod thinking;
pub mod usage; pub mod usage;
pub const PROGRESS_INDICATOR_EMOJI: &str = "⏳"; pub const PROGRESS_INDICATOR_EMOJI: &str = "⏳";

146
src/strings/thinking.rs Normal file
View File

@@ -0,0 +1,146 @@
use std::time::Duration;
/// After this much elapsed generation time, the notice escalates to the "medium" pool.
const TIER_MEDIUM_AFTER: Duration = Duration::from_secs(30);
/// After this much elapsed generation time, the notice escalates to the "deep" pool.
const TIER_DEEP_AFTER: Duration = Duration::from_secs(90);
/// Early flavor: the response is just taking a moment longer than instant.
const MESSAGES_LIGHT: &[&str] = &[
"*{{ baibot_name }} is thinking…*",
"*{{ baibot_name }} is mulling that over…*",
"*Hmm, let me think about that…*",
"*{{ baibot_name }} is gathering some thoughts…*",
"*One moment, {{ baibot_name }} is working on it…*",
"*{{ baibot_name }} is putting the pieces together…*",
"*Give {{ baibot_name }} a second here…*",
"*Let me think this one through…*",
"*{{ baibot_name }} is warming up the gears…*",
"*{{ baibot_name }} is on it…*",
];
/// Mid flavor: this is a real question and the model is genuinely working.
const MESSAGES_MEDIUM: &[&str] = &[
"*Huh, good one. {{ baibot_name }} is really thinking now…*",
"*Still working on it, {{ baibot_name }} wants to get this right…*",
"*This one needs a bit more thought…*",
"*{{ baibot_name }} is digging into this…*",
"*Hang tight, {{ baibot_name }} is turning it over…*",
"*Not a quick one, this. {{ baibot_name }} is still at it…*",
"*{{ baibot_name }} is chewing on this properly now…*",
"*This deserves some real thought, bear with {{ baibot_name }}…*",
"*{{ baibot_name }} is working through the details…*",
"*Won't be long, {{ baibot_name }} is closing in on it…*",
];
/// Deep flavor: a long-running generation (e.g. a reasoning model going for minutes).
const MESSAGES_DEEP: &[&str] = &[
"*Okay, this is a hard one. {{ baibot_name }} is really deep in thought…*",
"*{{ baibot_name }} is in the weeds on this one, thanks for your patience…*",
"*Still here, still thinking. {{ baibot_name }} doesn't want to rush it…*",
"*A proper puzzle, this. {{ baibot_name }} is taking the time to do it justice…*",
"*{{ baibot_name }} is really wrestling with this one…*",
"*A meaty question. {{ baibot_name }} is still turning it over…*",
"*{{ baibot_name }} hasn't forgotten you, just thinking hard…*",
"*Almost there, {{ baibot_name }} is pulling it all together…*",
"*{{ baibot_name }} is going the extra mile on this one…*",
"*Deep thoughts in progress. {{ baibot_name }} appreciates your patience…*",
];
/// Returns the flavor pool matching how long generation has been running.
pub fn messages_for_elapsed(elapsed: Duration) -> &'static [&'static str] {
if elapsed >= TIER_DEEP_AFTER {
MESSAGES_DEEP
} else if elapsed >= TIER_MEDIUM_AFTER {
MESSAGES_MEDIUM
} else {
MESSAGES_LIGHT
}
}
/// Picks one (still-untemplated) flavor message from the tier active at `elapsed`,
/// indexed by a monotonic `sequence` the caller increments once per notice.
///
/// Using a monotonic counter (rather than a clock-derived value) guarantees two
/// things the gesture-novelty goal needs: consecutive notices never repeat a line
/// (the index advances by one each tick, so it differs whenever the tier has more
/// than one message), and the pool is fully walked before any line recurs. The
/// caller seeds `sequence` with a per-generation value so different turns don't all
/// open on the same line. A clock-derived index can't promise this: `interval_at`
/// ticks on a near-fixed schedule, so its sub-second component clusters and would
/// re-pick the same line.
pub fn pick_message(elapsed: Duration, sequence: usize) -> &'static str {
let pool = messages_for_elapsed(elapsed);
pool[sequence % pool.len()]
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn messages_for_elapsed_returns_the_right_tier() {
// Boundaries: light below 30s, medium [30s, 90s), deep at/after 90s.
// The pools have distinct content, so value comparison identifies the tier.
assert_eq!(messages_for_elapsed(Duration::from_secs(0)), MESSAGES_LIGHT);
assert_eq!(
messages_for_elapsed(Duration::from_secs(29)),
MESSAGES_LIGHT
);
assert_eq!(
messages_for_elapsed(Duration::from_secs(30)),
MESSAGES_MEDIUM
);
assert_eq!(
messages_for_elapsed(Duration::from_secs(89)),
MESSAGES_MEDIUM
);
assert_eq!(messages_for_elapsed(Duration::from_secs(90)), MESSAGES_DEEP);
assert_eq!(
messages_for_elapsed(Duration::from_secs(600)),
MESSAGES_DEEP
);
}
#[test]
fn every_tier_is_non_empty() {
for pool in [MESSAGES_LIGHT, MESSAGES_MEDIUM, MESSAGES_DEEP] {
assert!(!pool.is_empty());
for message in pool {
assert!(!message.trim().is_empty());
}
}
}
#[test]
fn every_tier_uses_the_bot_name_template() {
// Not every line names the bot (some are first-person flavor), but each
// tier exercises the template var so substitution is wired through.
for pool in [MESSAGES_LIGHT, MESSAGES_MEDIUM, MESSAGES_DEEP] {
assert!(pool.iter().any(|m| m.contains("{{ baibot_name }}")));
}
}
#[test]
fn pick_message_stays_within_the_active_tier() {
let deep = messages_for_elapsed(Duration::from_secs(120));
for sequence in 0..50 {
assert!(deep.contains(&pick_message(Duration::from_secs(120), sequence)));
}
}
#[test]
fn pick_message_never_repeats_on_consecutive_sequences() {
// The gesture-novelty guarantee: a monotonic sequence must not pick the same line twice
// in a row (and walks the whole tier before any line recurs).
let elapsed = Duration::from_secs(0);
let pool_len = messages_for_elapsed(elapsed).len();
for sequence in 0..(pool_len * 3) {
assert_ne!(
pick_message(elapsed, sequence),
pick_message(elapsed, sequence + 1),
);
}
}
}