Compare commits

...

263 Commits

Author SHA1 Message Date
Slavi Pantaleev
9f1d664319 Capitalize thinking-notice flavor messages and bump 1.24.0 release date
The "thinking…" presets started with a lowercase letter, which reads
oddly as a standalone message. Capitalize the first word of each
literal-led line; lines that open with {{ baibot_name }} are untouched.

Also bump the 1.24.0 CHANGELOG date to the actual release date.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 07:02:08 +03:00
Aine
0c78928dab add thinking notice 2026-06-25 19:41:15 +01:00
Slavi Pantaleev
982ebc6657 CHANGELOG: set 1.23.1 release date to 2026-06-24
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 07:25:44 +03:00
Aine
4a75be742f Venice: add auto-recover for optional params and surface actual error message on 400s 2026-06-24 07:10:29 +03:00
renovate[bot]
aae3c4c08f Update ghcr.io/element-hq/element-web Docker tag to v1.12.22 2026-06-24 07:09:33 +03:00
Slavi Pantaleev
08cf50885d CHANGELOG: fix 1.23.0 date and add OpenAI-compatible CA-trust fix
Correct the 1.23.0 release date to 2026-06-23 and document the
system-CA-trust fix (2888cb9, via etke_openai_api_rust 0.1.10) that also
ships in this release.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 07:42:15 +03:00
Slavi Pantaleev
1eebb1a55b Merge pull request #189 from etkecc/venice-improvements
Venice improvements
2026-06-23 07:41:43 +03:00
Slavi Pantaleev
2888cb9450 Bump etke_openai_api_rust to 0.1.10 for system CA trust
0.1.10 enables ureq's `native-certs` feature, so the OpenAI-compatible
provider trusts the system CA store (honouring SSL_CERT_FILE) instead of
only the bundled webpki-roots. Without it, endpoints behind a private or
internal CA (FreeIPA, org PKI) fail the TLS handshake with "invalid peer
certificate: UnknownIssuer".

Fixes #188.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 07:16:58 +03:00
Aine
5de497482d Venice: add 💭 Reasoning spoiler block (html <details>) 2026-06-23 01:30:37 +01:00
Aine
9e5f6de965 Venice improvements 2026-06-22 17:18:08 +01:00
renovate[bot]
2ae641109f Update Rust crate reqwest to v0.13.4 2026-06-21 07:57:25 +03:00
Slavi Pantaleev
105ca7b506 Update reqwest to 0.13
Aligns baibot's direct reqwest dep with the 0.13 copy that
async-openai/matrix-sdk/mxlink already pull, instead of the lone 0.12
copy kept alive only by the anthropic fork. 0.13 renamed the
`rustls-tls` feature to `rustls`; update the feature list accordingly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 07:48:15 +03:00
Aine
de980f0165 Implement Venice.ai provider 2026-06-21 07:18:37 +03:00
renovate[bot]
cc3888a4cf Update forgejo.ellis.link/continuwuation/continuwuity Docker tag to v0.5.10 2026-06-20 21:49:16 +03:00
renovate[bot]
eb3285f104 Update actions/checkout action to v7 2026-06-19 11:02:19 +03:00
renovate[bot]
0a03dde523 Update Rust crate async-openai to v0.41.1 2026-06-19 11:02:07 +03:00
renovate[bot]
a5b575da68 Update docker.io/ollama/ollama Docker tag to v0.30.10 2026-06-18 09:55:57 +03:00
renovate[bot]
fe7afec920 Update ghcr.io/element-hq/synapse Docker tag to v1.155.0 2026-06-17 06:24:58 +03:00
renovate[bot]
b6641523da Update docker.io/ollama/ollama Docker tag to v0.30.9 2026-06-17 06:19:12 +03:00
renovate[bot]
c579977e59 Update dependency prek to v0.4.5 2026-06-15 16:13:50 +03:00
renovate[bot]
0a5f37c3eb Update docker.io/ollama/ollama Docker tag to v0.30.8 2026-06-14 07:09:07 +03:00
renovate[bot]
58246e4eb3 Update ghcr.io/element-hq/element-web Docker tag to v1.12.21 2026-06-09 22:04:39 +03:00
renovate[bot]
9b767a420e Update Rust crate regex to v1.12.4 2026-06-09 22:04:26 +03:00
renovate[bot]
54c27311c1 Update docker.io/ollama/ollama Docker tag to v0.30.7 2026-06-09 07:50:10 +03:00
renovate[bot]
3d15f11f76 Update docker.io/ollama/ollama Docker tag to v0.30.6 2026-06-05 12:33:56 +03:00
Slavi Pantaleev
ead70d71ff Prepare 1.21.1 release
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 10:07:16 +03:00
Slavi Pantaleev
49bb52a091 Update anthropic-rs fork to drop vulnerable rustls-webpki 0.101
The anthropic git dependency now uses reqwest 0.12 / rustls 0.23, pulling
rustls-webpki 0.103.13 instead of the 0.101.7 that was dragged in via the
old reqwest 0.11. This clears three RUSTSEC/Dependabot advisories:

- GHSA-82j2-j2ch-gfr8 (high): DoS via panic on malformed CRL BIT STRING
- GHSA-xgp8-3hg3-c2mh (low): name constraints accepted for wildcard names
- GHSA-965h-392x-2mh5 (low): name constraints for URI names incorrectly accepted

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 10:06:27 +03:00
Slavi Pantaleev
1fc2f0f65a Prepare 1.21.0 release
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 09:44:25 +03:00
Slavi Pantaleev
f7cb38b620 Fix clippy warnings in tests
- Avoid unwrap_or() on a statically-Some value by using the raw token count
- Use an array literal instead of vec! for the non-allocated test cases

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 09:41:00 +03:00
Slavi Pantaleev
04e7667d11 Default to gpt-image-2 for OpenAI image generation
Make gpt-image-2 the default image-generation model and add it to the
recognized model-id mapping, updating the sample provider configs to match.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 09:39:45 +03:00
Slavi Pantaleev
78cbfda481 Update Rust crate async-openai to 0.41.0
Adapt to upstream API changes:
- Handle the new ImageModel::GptImage2 variant in image-model match arms
- ImageSize dropped Copy (gained an Other(String) variant), so clone it

Supersedes #170.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 09:37:46 +03:00
renovate[bot]
36fd6a4dda Update Rust crate quick_cache to v0.6.23 2026-06-05 09:33:16 +03:00
renovate[bot]
1308196419 Update docker.io/ollama/ollama Docker tag to v0.30.5 2026-06-05 09:33:07 +03:00
renovate[bot]
b3546ebaf1 Update ghcr.io/element-hq/synapse Docker tag to v1.154.0 2026-06-04 18:50:24 +03:00
renovate[bot]
2499e6baca Update Rust crate chrono to v0.4.45 2026-06-04 18:50:01 +03:00
renovate[bot]
871b9d2f3b Update dependency prek to v0.4.4 2026-06-04 14:48:21 +03:00
renovate[bot]
4a36cf9446 Update docker.io/ollama/ollama Docker tag to v0.30.4 2026-06-04 07:13:27 +03:00
renovate[bot]
a2180452c9 Update docker.io/ollama/ollama Docker tag to v0.30.2 2026-06-03 07:39:57 +03:00
Slavi Pantaleev
0cb0fc18ce Prepare 1.20.0 release
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 23:05:29 +03:00
renovate[bot]
78bddac716 Update Rust crate tiktoken-rs to 0.12.* (#162)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-06-02 23:03:20 +03:00
Slavi Pantaleev
f7209221f6 Update matrix-sdk to v0.18.0 (and mxlink to v1.15.0)
matrix-sdk 0.18.0 introduces no breaking changes that affect baibot, but
it does require a matching mxlink release: mxlink 1.14.0 pins matrix-sdk
0.17.0, so without bumping mxlink the two matrix-sdk versions conflict.
mxlink 1.15.0 (released alongside this) tracks matrix-sdk 0.18.0, so we
bump the floor to >=1.15.0.

Verified locally: cargo build and the full test suite pass.

Supersedes Renovate PR #161.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 22:47:13 +03:00
renovate[bot]
5f2054e07f Update Rust crate async-openai to v0.40.3 2026-06-02 12:06:50 +03:00
renovate[bot]
8301c5a29f Update docker.io/ollama/ollama Docker tag to v0.30.0 2026-06-02 12:06:30 +03:00
Slavi Pantaleev
b325b6b1a5 Upgrade Rust (1.95.0 -> 1.96.0)
Bumps the base image in Dockerfile/Dockerfile.ci (via Renovate #158) and
keeps rust-toolchain.toml in sync, since rustup honors the toolchain file
inside the container build and would otherwise keep using 1.95.0.

Verified locally under 1.96.0: cargo check, clippy -D warnings, fmt check,
and cargo test --all-features all pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 09:02:05 +03:00
renovate[bot]
0600721663 Update ghcr.io/element-hq/element-web Docker tag to v1.12.20 2026-05-27 20:50:26 +03:00
renovate[bot]
dc17cf7e96 Update ghcr.io/element-hq/element-web Docker tag to v1.12.19 2026-05-27 13:50:35 +03:00
Slavi Pantaleev
d6a8f6ba0b Prepare 1.19.3 release
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 10:09:11 +03:00
renovate[bot]
c4c3d71195 Update dependency prek to v0.4.3 2026-05-27 09:56:15 +03:00
renovate[bot]
84ae29d034 Update dependency prek to v0.4.2 2026-05-26 15:14:23 +03:00
renovate[bot]
447f43df7b Update Rust crate async-openai to v0.40.2 2026-05-22 23:11:28 +03:00
renovate[bot]
dabac790ad Update Rust crate async-openai to v0.40.1 2026-05-22 10:19:16 +03:00
renovate[bot]
fa01f012a1 Update Rust crate serde_json to v1.0.150 2026-05-22 10:19:05 +03:00
Slavi Pantaleev
20cb33bc66 Prepare 1.19.2 release
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 12:23:20 +03:00
Slavi Pantaleev
2791bdb08b Merge pull request #150 from etkecc/chore/update-dependencies
Update dependencies
2026-05-21 09:39:11 +03:00
Slavi Pantaleev
5dd505202a Update dependencies
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 09:33:28 +03:00
Slavi Pantaleev
d61078002c Merge pull request #149 from etkecc/renovate/async-openai-0.x
Update Rust crate async-openai to 0.40.0
2026-05-21 09:26:46 +03:00
renovate[bot]
cf5b346558 Update Rust crate async-openai to 0.40.0 2026-05-21 05:55:20 +00:00
renovate[bot]
ebbb6658e1 Update dependency prek to v0.4.1 2026-05-20 09:13:57 +03:00
renovate[bot]
369dc0c1ba Update ghcr.io/element-hq/synapse Docker tag to v1.153.0 2026-05-19 21:46:36 +03:00
renovate[bot]
3185f44a93 Update Rust crate quick_cache to v0.6.22 2026-05-18 07:32:20 +03:00
renovate[bot]
aa8ccde0ed Update docker.io/postgres Docker tag to v18.4 2026-05-15 12:48:23 +03:00
renovate[bot]
d5df0d7416 Update docker.io/ollama/ollama Docker tag to v0.24.0 2026-05-15 12:48:07 +03:00
renovate[bot]
140ca9ed68 Update dependency prek to v0.4.0 2026-05-14 16:40:26 +03:00
renovate[bot]
89d77d52b6 Update Rust crate async-openai to v0.38.2 2026-05-14 07:32:06 +03:00
renovate[bot]
5092700275 Update docker.io/ollama/ollama Docker tag to v0.23.4 2026-05-14 07:31:22 +03:00
renovate[bot]
10c365124e Update docker.io/ollama/ollama Docker tag to v0.23.3 2026-05-13 07:28:16 +03:00
renovate[bot]
08bdf4f7a2 Update ghcr.io/element-hq/element-web Docker tag to v1.12.18 2026-05-12 20:38:24 +03:00
Slavi Pantaleev
af557a7e45 Fix multi-arch manifest publish: use buildx imagetools
docker/build-push-action now wraps single-platform images in an OCI
image index (to carry provenance attestations), so the per-arch
`*-amd64`/`*-arm64` tags are manifest lists. `docker manifest create`
refuses manifest-list sources ("X is a manifest list"). Switch to
`docker buildx imagetools create`, which flattens index sources
correctly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 09:22:50 +03:00
renovate[bot]
fa7eb11b1d Update Rust crate async-openai to v0.38.1 2026-05-12 02:39:22 +03:00
Slavi Pantaleev
ff1e128f0e Fix versioned image publish under workflow_run trigger
e978d3c switched the publish trigger from `push` to `workflow_run` and
correctly migrated the `type=raw,value=latest` rule to read the upstream
head_branch from `github.event.workflow_run.*` — but left the
`type=semver,pattern={{raw}}` rule unchanged. That rule still reads
`github.ref`, which under workflow_run dispatch is always
`refs/heads/main` (the default branch where the workflow file lives),
not the triggering tag ref. As a result, no semver tag was extracted,
metadata-action produced no tags, and `buildx` failed with
"tag is needed when pushing to registry". The `latest` tag kept
publishing because its rule was migrated; versioned tags (v1.19.0,
v1.19.1) silently stopped publishing.

Pass the upstream head_branch to the semver rule explicitly via `value`,
gated by `enable` so it only fires for v* tags. Mirrors the migration
the raw rule already received.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 17:28:54 +03:00
Slavi Pantaleev
c9da927c66 Prepare 1.19.1 release
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 16:40:24 +03:00
Slavi Pantaleev
7b16c5c1c3 Update async-openai to v0.38.0
Closes #137. The 0.36 -> 0.38 bump (skipping 0.37) introduces Tower-based
middleware support and a fix to the ReasoningItem `type` field, but
neither affects baibot's call sites — no code adaptation was needed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 16:40:10 +03:00
Slavi Pantaleev
f809bf8d7c Prepare 1.19.0 release
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 16:34:49 +03:00
Slavi Pantaleev
44aea427fc Update matrix-sdk to v0.17.0 and Rust toolchain to 1.95.0
matrix-sdk 0.17.0 brings a few breaking API changes that we adapt to:

- Drop the `native-tls` feature on matrix-sdk; it was removed upstream
  in 2026-04 (matrix-sdk now hardwires rustls). The previous workaround
  comment referencing rust-mxlink issue #1 is no longer relevant.
- `Relation::Reply` is now a tuple variant wrapping a `Reply` struct;
  pattern matches updated to bind through `reply.in_reply_to.event_id`.

Also bumps the Rust toolchain from 1.93.0 to 1.95.0 (closes #85, which
proposed the Dockerfile bump in isolation). rustc 1.94+ trips a
query-depth overflow when computing async layouts in the matrix-sdk
timeline future graph; matrix-rust-sdk PR #6489 raises the limit, but
\`recursion_limit\` is per-crate, so we repeat \`#![recursion_limit = "256"]\`
on this crate root.

Bumps the mxlink floor to 1.14.0, which carries the matching adapter.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 16:29:41 +03:00
renovate[bot]
8f45a57ced Update docker.io/ollama/ollama Docker tag to v0.23.2 2026-05-09 16:06:25 +03:00
renovate[bot]
6a9100309c Update ghcr.io/element-hq/synapse Docker tag to v1.152.1 2026-05-09 16:06:17 +03:00
renovate[bot]
1953878e6b Update Rust crate tokio to v1.52.3 2026-05-09 08:12:03 +03:00
renovate[bot]
35761a5bf7 Update ghcr.io/element-hq/element-web Docker tag to v1.12.17 2026-05-09 08:11:38 +03:00
renovate[bot]
c7fa0cc0ee Update forgejo.ellis.link/continuwuation/continuwuity Docker tag to v0.5.9 2026-05-09 08:11:25 +03:00
renovate[bot]
87fcf9d019 Update dependency prek to v0.3.13 2026-05-09 08:11:04 +03:00
renovate[bot]
4b52bc906c Update forgejo.ellis.link/continuwuation/continuwuity Docker tag to v0.5.8 2026-04-25 22:16:54 +03:00
renovate[bot]
517a6c5e33 Update docker.io/ollama/ollama Docker tag to v0.21.2 2026-04-24 08:17:02 +03:00
renovate[bot]
636c8a35eb Update Rust crate async-openai to 0.36.0 2026-04-24 06:54:26 +03:00
renovate[bot]
a50de600da Update docker.io/ollama/ollama Docker tag to v0.21.1 2026-04-22 07:15:30 +03:00
Slavi Pantaleev
e978d3cb2f Publish only after successful CI
Run the publish workflow from the workflow_run event so Docker publishing happens only after the CI workflow completes successfully for push events on main or v* tags.

Check out the exact SHA validated by CI and derive Docker metadata from the upstream CI ref, so publishing follows the tested revision instead of the default branch tip.
2026-04-20 22:09:37 +03:00
Slavi Pantaleev
2b1bdbd3d2 Split CI and publish workflows
This supersedes 7d183b9 ("Run CI for pull requests"), which mixed validation and publishing in one workflow and regressed docker-manifest by dropping the package-write permission it needs to publish the manifest.

Split the workflows so CI handles pull requests, branch pushes, tags, and manual runs, while publishing stays focused on Docker delivery with the manifest permission fixed explicitly at the job level.
2026-04-20 22:08:07 +03:00
Slavi Pantaleev
7d183b91d1 Run CI for pull requests 2026-04-20 21:40:03 +03:00
renovate[bot]
420f380417 Update Rust crate async-openai to 0.35.0 (#125)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-04-20 21:39:09 +03:00
renovate[bot]
6f541e2361 Update forgejo.ellis.link/continuwuation/continuwuity Docker tag to v0.5.7 2026-04-17 21:52:46 +03:00
renovate[bot]
0737c2761e Update docker.io/ollama/ollama Docker tag to v0.21.0 2026-04-17 10:41:04 +03:00
renovate[bot]
45852a0d53 Update Rust crate tokio to v1.52.1 2026-04-17 10:38:22 +03:00
renovate[bot]
0fcf8d8703 Update Rust crate tokio to 1.52.* 2026-04-15 09:43:04 +03:00
renovate[bot]
41905b006a Update docker.io/ollama/ollama Docker tag to v0.20.7 2026-04-14 07:33:22 +03:00
renovate[bot]
ee3d27701b Update docker.io/ollama/ollama Docker tag to v0.20.6 2026-04-13 08:58:29 +03:00
Slavi Pantaleev
1b5c2fbdc1 Prepare 1.18.0 release
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 10:55:55 +03:00
Slavi Pantaleev
a865a26093 Update tiktoken-rs to 0.11, adding support for newer GPT models
Supersedes https://github.com/etkecc/baibot/pull/116

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 10:52:22 +03:00
Slavi Pantaleev
1563e83672 Update dependencies 2026-04-11 10:40:09 +03:00
renovate[bot]
cf114b37b3 Update docker.io/ollama/ollama Docker tag to v0.20.5 2026-04-10 09:21:42 +03:00
renovate[bot]
b8ef2b978b Update Rust crate tokio to v1.51.1 2026-04-08 19:25:58 +03:00
renovate[bot]
7e37aee1b9 Update ghcr.io/element-hq/synapse Docker tag to v1.151.0 2026-04-08 11:57:52 +03:00
renovate[bot]
53836e556a Update ghcr.io/element-hq/element-web Docker tag to v1.12.15 2026-04-08 11:57:29 +03:00
renovate[bot]
89052bbfd9 Update ghcr.io/element-hq/element-web Docker tag to v1.12.14 2026-04-07 19:53:11 +03:00
renovate[bot]
07eb12d406 Update docker.io/ollama/ollama Docker tag to v0.20.3 2026-04-07 08:45:11 +03:00
renovate[bot]
3272cd6fb2 Update Rust crate tokio to 1.51.* 2026-04-04 15:52:04 +03:00
renovate[bot]
7ab97a39ca Update docker.io/ollama/ollama Docker tag to v0.20.2 2026-04-04 10:35:45 +03:00
renovate[bot]
11c5a9942e Update docker.io/ollama/ollama Docker tag to v0.20.1 2026-04-04 09:17:02 +03:00
renovate[bot]
251dec454c Update docker.io/ollama/ollama Docker tag to v0.20.0 2026-04-03 07:40:18 +03:00
renovate[bot]
1031dbf672 Update docker.io/ollama/ollama Docker tag to v0.19.0 2026-03-30 08:31:02 +03:00
renovate[bot]
581f00b9fb Update docker.io/ollama/ollama Docker tag to v0.18.3 2026-03-26 08:18:33 +02:00
Slavi Pantaleev
e57778d2bd Prepare 1.17.0 release 2026-03-25 20:18:56 +02:00
kschwank
2d659964a7 Add sender context mode for text generation (#104)
Add a per-room/global `sender_context_mode` setting that optionally
prefixes conversation messages with sender metadata before sending them
to the model provider.

This helps models distinguish between participants in multi-user rooms.
                                                                                                                                                                                                                                                             
Three modes are supported:
- `disabled` (default, no change)
- `matrix_user_id` (prefixes with `[sender=@user:server]`)
- `matrix_user_id_and_timestamp` (adds `send_at`; Example: `[sender=@user:server sent_at=<ISO 8601>]`)
                                                                                                                                                                                                                                                             
Sender context is applied to user and assistant text messages only,
skipping system prompts and non-text content.

Mixed-sender merged turns (something we intentionally do for Anthropic)
have their `sender_id` cleared to avoid misattribution.
2026-03-25 20:14:34 +02:00
renovate[bot]
661e7263fb Update ghcr.io/element-hq/synapse Docker tag to v1.150.0 2026-03-24 17:08:21 +02:00
renovate[bot]
1705c16762 Update ghcr.io/element-hq/element-web Docker tag to v1.12.13 2026-03-24 14:28:10 +02:00
Slavi Pantaleev
2455117e41 Prepare 1.16.1 release 2026-03-24 14:05:35 +02:00
Slavi Pantaleev
3e9c110afc Update dependencies 2026-03-24 14:04:40 +02:00
Slavi Pantaleev
d9b5524c97 Fix OpenAI response input for async-openai 0.34
async-openai 0.34 adds a required phase field to EasyInputMessage,
which broke our Responses API request construction.

Set phase to None because baibot does not currently model assistant commentary vs final-answer turns,
so omitting phase preserves the previous behavior while keeping the request compatible with the new crate.
2026-03-24 12:51:56 +02:00
renovate[bot]
9b169a7d28 Update Rust crate async-openai to 0.34.0 2026-03-24 12:40:25 +02:00
Slavi Pantaleev
cb29419d75 Prepare 1.16.0 release
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 12:25:52 +02:00
Slavi Pantaleev
51ca8c9948 Update dependencies 2026-03-20 12:23:51 +02:00
Slavi Pantaleev
12b938d2d1 Skip files with application/octet-stream MIME type
Files with unrecognized MIME types (application/octet-stream) are
not supported by any LLM provider and would cause errors that
permanently break the conversation thread. Instead, represent them
as a text message describing the attachment so the LLM is still
aware a file was sent without the thread becoming unusable.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 12:22:38 +02:00
Slavi Pantaleev
f3d1b32ad7 Use mime_guess for file extension MIME type detection
Replace the hand-maintained extension-to-MIME mapping with the
mime_guess crate, which was already in the dependency tree.
This covers hundreds of file extensions out of the box.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 12:22:38 +02:00
Slavi Pantaleev
3d3bd3c9f9 Add support for file attachments (m.file) in conversations
Files sent as m.file Matrix messages are now downloaded, MIME-detected,
and forwarded to LLM providers alongside the conversation context,
similar to how m.image is already handled.

- OpenAI provider: sends files inline as base64 data URLs
- Anthropic provider: skips files with a warning (library limitation)
- OpenAI-compat provider: skips files with a warning (library limitation)

Controller routing respects the existing prefix requirement setting.
MIME detection expanded to cover PDF, text, code, and document formats.
Docs updated to reflect file support and known limitations.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 12:22:38 +02:00
renovate[bot]
a25845e89e Update Rust crate quick_cache to v0.6.21 2026-03-19 19:22:00 +02:00
renovate[bot]
479f54b93d Update docker.io/ollama/ollama Docker tag to v0.18.2 2026-03-19 08:15:14 +02:00
renovate[bot]
90ab6807ac Update Rust crate quick_cache to v0.6.20 2026-03-18 09:15:53 +02:00
renovate[bot]
5022f79bf5 Update docker.io/ollama/ollama Docker tag to v0.18.1 2026-03-17 07:25:56 +02:00
renovate[bot]
2ba3b5a437 Update docker.io/ollama/ollama Docker tag to v0.18.0 2026-03-14 22:34:33 +02:00
renovate[bot]
2819011b9d Update Rust crate tracing-subscriber to v0.3.23 2026-03-13 20:44:43 +02:00
renovate[bot]
ede9065f77 Update Rust crate async-openai to v0.33.1 2026-03-13 11:54:12 +02:00
renovate[bot]
5e61f5b3a3 Update ghcr.io/element-hq/synapse Docker tag to v1.149.1 2026-03-11 20:05:48 +02:00
renovate[bot]
8633e82f62 Update Rust crate quick_cache to v0.6.19 2026-03-11 20:05:32 +02:00
Slavi Pantaleev
6aa3d70d57 Bump default OpenAI text-generation model (gpt-5.2 -> gpt-5.4) 2026-03-11 09:53:02 +02:00
renovate[bot]
75e29ee3fb Update Rust crate tempfile to 3.27.* 2026-03-11 08:50:06 +02:00
renovate[bot]
eab7978b9f Update ghcr.io/element-hq/element-web Docker tag to v1.12.12 2026-03-10 22:06:39 +02:00
renovate[bot]
765e7c17b2 Update ghcr.io/element-hq/synapse Docker tag to v1.149.0 2026-03-10 20:19:06 +02:00
Slavi Pantaleev
748d2b7fd4 Prepare 1.15.0 release
Add the 1.15.0 changelog entry covering access-token authentication support, Rust 1.93.0 toolchain pinning, docs updates, and dependency updates. Bump crate version to 1.15.0 in Cargo.toml and Cargo.lock.
2026-03-07 11:13:42 +02:00
Slavi Pantaleev
527759dd02 Document Matrix authentication modes
Add a dedicated configuration/authentication doc covering password and access-token setup,
including environment-variable mappings and token-generation example.

Link to it from configuration docs and align the sample config wording with Matrix Authentication Service/OIDC terminology.
2026-03-07 11:08:50 +02:00
Slavi Pantaleev
3290255bad Pin local Rust toolchain to 1.93.0
Add rust-toolchain.toml so local development and ad-hoc cargo commands use the same known-good compiler version as CI.
This avoids the matrix-sdk query-depth overflow seen on newer stable toolchains.

Related to 9b987395b3

Ref: https://github.com/etkecc/baibot/pull/83#issuecomment-4008463792
2026-03-07 10:40:56 +02:00
Slavi Pantaleev
9b987395b3 Pin GitHub CI Rust toolchain to 1.93.0
Using the floating stable toolchain currently breaks this project. Repro via just build-debug:

```
Compiling matrix-sdk v0.16.0
error: queries overflow the depth limit!
help: consider increasing the recursion limit by adding #![recursion_limit = "256"] to your crate (matrix_sdk)
note: query depth increased by 130 when computing layout of matrix-sdk Client::sync async body
error: could not compile matrix-sdk (lib) due to 1 previous error
```

Pinning CI to 1.93.0 keeps CI on the known-good toolchain until upstream/toolchain compatibility is addressed.
We should have had that to begin with (we're pinning as much as we can anyway), but..

Ref: https://github.com/etkecc/baibot/pull/83#issuecomment-4008463792
2026-03-07 10:38:35 +02:00
Taylor Southwick
4852d1fe92 Add support for access tokens using MAS (#83)
* Add support for access tokens using MAS

* use 1.13.0

* Update dependencies

* Harden auth credential selection in matrix link init

Use the same non-empty access-token criterion for auth mode selection and bind the token directly from the branch condition.
Return explicit configuration errors for missing or empty `device_id`/`password` instead of panicking, so invalid auth config fails gracefully.

* Centralize and harden user auth config handling

Move authentication-mode resolution into typed config parsing with ConfigUserAuth,
so downstream login setup consumes validated credentials instead of re-checking raw optional fields.

Enforce explicit password-vs-token selection, validate token/device/user-id requirements in one place,
and normalize empty auth env overrides to unset values for consistent behavior across YAML and environment input.

* Add auth config unit tests

Move auth_config tests into a dedicated cfg test module file to keep production config code compact while preserving behavior coverage. The tests cover password/token mode selection, missing/both auth method rejection, missing device_id, and empty-value handling.

* Use conventional mxlink version requirement

Replace the unconventional wildcard lower-bound expression with a standard semver lower bound for readability and tooling consistency.

---------

Co-authored-by: Slavi Pantaleev <slavi@devture.com>
2026-03-07 10:26:40 +02:00
renovate[bot]
8bd313f0d4 Update docker/build-push-action action to v7 2026-03-06 10:15:37 +02:00
renovate[bot]
711e1099d6 Update docker.io/ollama/ollama Docker tag to v0.17.7 2026-03-06 08:16:12 +02:00
renovate[bot]
91c8dd8f7d Update docker/metadata-action action to v6 2026-03-05 22:22:11 +02:00
renovate[bot]
afc5572d6a Update docker/login-action action to v4 2026-03-04 17:13:28 +02:00
renovate[bot]
73e13dcf2f Update docker.io/ollama/ollama Docker tag to v0.17.6 2026-03-04 07:36:04 +02:00
renovate[bot]
2bebd109b1 Update forgejo.ellis.link/continuwuation/continuwuity Docker tag to v0.5.6 2026-03-04 07:35:22 +02:00
renovate[bot]
7f7c58be1f Update Rust crate tokio to 1.50.* 2026-03-03 16:27:53 +02:00
renovate[bot]
47e5a464a0 Update docker.io/ollama/ollama Docker tag to v0.17.5 2026-03-01 08:09:47 +02:00
renovate[bot]
85f751e514 Update docker.io/ollama/ollama Docker tag to v0.17.4 2026-02-27 07:08:48 +02:00
renovate[bot]
95acad3558 Update docker.io/postgres Docker tag to v18.3 2026-02-27 06:36:26 +02:00
renovate[bot]
304056c59a Update docker.io/ollama/ollama Docker tag to v0.17.2 2026-02-27 06:36:19 +02:00
renovate[bot]
bedc0335f1 Update docker.io/ollama/ollama Docker tag to v0.17.1 2026-02-26 13:33:16 +02:00
renovate[bot]
a8be8c3c1e Update ghcr.io/element-hq/element-web Docker tag to v1.12.11 2026-02-24 16:54:27 +02:00
renovate[bot]
826fa728a9 Update ghcr.io/element-hq/synapse Docker tag to v1.148.0 2026-02-24 16:53:10 +02:00
renovate[bot]
5aef8e8b2f Update Rust crate tempfile to 3.26.* 2026-02-24 08:21:45 +02:00
renovate[bot]
f70f20181e Update Rust crate chrono to v0.4.44 2026-02-24 08:16:54 +02:00
renovate[bot]
fcdd4f39ee Update docker.io/ollama/ollama Docker tag to v0.17.0 2026-02-24 08:16:31 +02:00
renovate[bot]
891adfec49 Update Rust crate anyhow to v1.0.102 2026-02-20 08:49:48 +02:00
renovate[bot]
35ab79844b Update docker.io/ollama/ollama Docker tag to v0.16.3 2026-02-20 08:49:37 +02:00
Slavi Pantaleev
bbc122fbb1 Release 1.14.3
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 06:37:11 +02:00
renovate[bot]
2413c8b88b Update actions/checkout action to v6 2026-02-18 06:31:39 +02:00
renovate[bot]
10c3c64469 Update Rust crate async-openai to 0.33.0 2026-02-18 06:25:58 +02:00
renovate[bot]
b3307b404b Update docker.io/rust Docker tag to v1.93.1 2026-02-18 06:25:48 +02:00
Slavi Pantaleev
7a0d1e830d Add Renovate configuration for automated dependency updates
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 06:06:02 +02:00
Slavi Pantaleev
b3bd241823 Release 1.14.2
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 05:53:46 +02:00
Slavi Pantaleev
de3d8b054f Update dependencies 2026-02-18 05:44:41 +02:00
Slavi Pantaleev
0a55e276a2 Update dependencies 2026-02-18 05:40:25 +02:00
Slavi Pantaleev
1f2c65d2e6 Refactor dev services to support homeserver choice (Continuwuity or Synapse)
The dev environment previously hardcoded Synapse (bundled with Postgres
and Element Web) in a monolithic etc/services/core/ directory.

With Continuwuity now available as a lighter alternative (no external DB),
this refactors the service layout so developers choose their homeserver
once and everything derives from that choice. Continuwuity is the new
default for its smaller footprint.

Key changes:
- Break etc/services/core/ into etc/services/synapse/ and
  etc/services/element-web/, each with their own compose.yml
- Add `homeserver` variable in justfile (reads var/homeserver,
  defaults to continuwuity)
- Add `homeserver-init` recipe to persist the choice
- Use placeholders (__HOMESERVER_SERVER_NAME__, __HOMESERVER_URL__,
  __HOMESERVER_CLIENT_URL__) in config templates, resolved at
  prepare time based on the chosen homeserver
- Make services-start/stop/prepare/tail-logs delegate to the chosen
  homeserver's recipes + element-web
- Make users-prepare delegate to {homeserver}-users-prepare
- Update docs/development.md for the new homeserver choice flow

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 05:28:32 +02:00
Slavi Pantaleev
3b5e4745f2 Add optional Continuwuity homeserver service for development/testing
Adds Continuwuity as an alternative to Synapse for local development,
useful for testing baibot compatibility with different homeserver
implementations. Follows the same optional service pattern as localai/ollama.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 04:47:22 +02:00
Slavi Pantaleev
407bb022d9 Release 1.14.1 2026-02-10 14:42:20 +02:00
Slavi Pantaleev
faf92cac09 Switch from deprecated serde_yaml to serde_yaml_ng
serde_yaml is deprecated and unmaintained. serde_yaml_ng is the community
fork with a compatible API, so this is a straightforward rename across the
codebase.
2026-02-10 14:34:57 +02:00
Slavi Pantaleev
a82e9a1d1f Add prek pre-commit hooks via mise, fix formatting and clippy warnings
- Add mise.toml (prek 0.3.2) and .pre-commit-config.yaml with hooks for
  trailing whitespace, end-of-file, YAML check, merge conflicts, large files,
  cargo fmt, cargo clippy (-D warnings), and unit tests
- Add prek/mise recipes to justfile
- Run cargo fmt to fix formatting issues
- Fix all clippy warnings: collapse nested if statements, derive Default for Avatar
2026-02-10 14:33:19 +02:00
Slavi Pantaleev
8f87f05a08 Update dependencies to fix security vulnerabilities
- Bump mxlink (>=1.11.0 -> >=1.12.0): pulls in fixes for time and bytes CVEs
- Bump async-openai (0.32.3 -> 0.32.4)
- Bump tempfile (3.24.* -> 3.25.*)
- Run cargo update to bump transitive dependencies, notably:
  - time (0.3.46 -> 0.3.47): fix stack exhaustion DoS
2026-02-10 14:28:38 +02:00
Slavi Pantaleev
61d18b2e13 Release 1.14.0 2026-02-04 03:19:30 +02:00
Slavi Pantaleev
d831c08306 Add support for OpenAI built-in tools (web_search, code_interpreter)
Document the new tools feature in README, features.md, and providers.md.
Update the `!bai providers` command output to show vision/tools support
consistently for all providers.

Based on #62 by @yeslayla which migrated the OpenAI provider to the
Responses API.

See: https://github.com/etkecc/baibot/pull/62
2026-02-04 03:17:19 +02:00
Slavi Pantaleev
c70387b0c3 Fix sticker generation for newer GPT image models
Sticker generation was failing when using newer GPT image models
(gpt-image-1, gpt-image-1-mini, gpt-image-1.5). The issue occurred
because stickers requested 256x256 size, but these models only support
1024x1024, 1536x1024, 1024x1536, and auto.

To reproduce, send `!bai sticker Something` to an agent configured
with a GPT image model. The error was:

  invalid_request_error: Invalid value: '256x256'. Supported values
  are: '1024x1024', '1024x1536', '1536x1024', and 'auto'. (param: size)
  (code: invalid_value)

The fix replaces the hardcoded 256x256 size override with a
`smallest_size_possible` flag, letting each provider determine the
appropriate sticker size based on the model being used.

The `openai_compat` provider still defaults to requesting 256x256 in all cases
(regardless of model name).
2026-02-04 02:39:51 +02:00
Slavi Pantaleev
38516f2e17 Update services 2026-02-04 02:28:12 +02:00
Slavi Pantaleev
5481b5a763 Update dependencies 2026-02-04 01:54:53 +02:00
Slavi Pantaleev
26bc437678 Add tools config to OpenAI example in config.yml.dist 2026-02-04 01:41:27 +02:00
Layla
ec93f1ee2a Implement OpenAI's response API and add support for built-in tools (web search & code interpreter). 2026-02-04 01:41:27 +02:00
Slavi Pantaleev
b920b6e556 Release 1.13.0 2026-01-23 00:33:29 +02:00
Slavi Pantaleev
7136d34843 Update dependencies 2026-01-23 00:23:25 +02:00
Slavi Pantaleev
e0b4a40dd8 Configure switching to cheaper models (gpt-image-1-mini) for gpt-image-1 & gpt-image-1.5 2026-01-22 22:31:51 +02:00
Slavi Pantaleev
691aeeb1c7 Upgrade Rust (1.92.0 -> 1.93.0) 2026-01-22 22:23:18 +02:00
Slavi Pantaleev
257ffae9e7 Release 1.12.0 2025-12-21 23:27:08 +02:00
Slavi Pantaleev
f7bf3d7b60 Upgrade async-openai (0.32.1 -> 0.32.2)
This brings in an important bugfix related to image generation.
Ref:
- https://github.com/64bit/async-openai/issues/507
- https://github.com/64bit/async-openai/pull/508
2025-12-21 23:21:03 +02:00
Slavi Pantaleev
3a88b0d656 Make gpt-image-1.5 the default image model for OpenAI 2025-12-21 12:53:44 +02:00
Slavi Pantaleev
08c689a889 Upgrade async-openai (0.31.1 -> 0.32.1) and adapt, adding support for gpt-image-1.5 2025-12-21 12:23:25 +02:00
Slavi Pantaleev
ae8e878817 Update services 2025-12-21 11:33:08 +02:00
Slavi Pantaleev
edbd72ece6 Release 1.11.0 2025-12-15 10:03:47 +02:00
Slavi Pantaleev
22906aa2d3 Upgrade Rust (1.91.1 -> 1.92.0) 2025-12-15 09:31:34 +02:00
Slavi Pantaleev
99bde53ef6 Upgrade services 2025-12-15 09:23:59 +02:00
Slavi Pantaleev
b3fd8e548f Minor documentation updates 2025-12-15 09:22:54 +02:00
Slavi Pantaleev
062fbbb8ef Add support for custom avatars (via file path) and for not touching the already-set avatar
This is based on the work done in https://github.com/etkecc/baibot/pull/60 by https://github.com/Fmstrat (Ben Curtis),
with various changes on top to make the code more idiomatic and flexible.

This commit squashes the following patches (newest first):

- Improve handling of `user.avatar` configuration (null & empty string being the same now) and add support for a special `keep` value
- Minor import reordering
- Simplify avatar configuration (`user.avatar.source` -> `user.avatar`)
- Combine `logo_bytes` and `mime_type` determination logic and do not fall back to default avatar if reading the custom avatar file fails
- Switch from deprecated `mime_guess::guess_mime_type(avatar_path)` to `mime_guess::from_path(avatar_path).first_or_octet_stream()`
- Relax `mime_guess` constraint and order alphabetically
- Use `mime` from `mxlink`
- Add mime-type support and switch to user.avatar.source
- (Original work by Fmstrat) Add support for custom avatars

Co-authored-by: Fmstrat <nospam@nowsci.com>
2025-12-15 08:07:24 +02:00
Slavi Pantaleev
2801c78ad9 Bump default OpenAI text-generation model (gpt-5.1 -> gpt-5.2) 2025-12-12 16:05:16 +02:00
Slavi Pantaleev
f4c698ad33 Release 1.10.0 2025-12-06 07:01:39 +02:00
Slavi Pantaleev
bd39001417 Update dependencies and matrix-sdk (0.14.0 -> 0.16.0) 2025-12-06 07:00:24 +02:00
Slavi Pantaleev
2692d0322e Release 1.9.0 2025-11-30 12:29:58 +02:00
Slavi Pantaleev
8eb70f0f2c Upgrade async-openai from our own etkecc fork to upstream's 0.31.1
Switches `async-openai` from our own etkecc fork (0.28.1-patched) to the
official crates.io version 0.31.1.

We adapt to async-openai's types reorganization and making use of crate
features to only enable what we need.
2025-11-30 11:48:56 +02:00
Slavi Pantaleev
5c0a7be7a2 Add changelog entry for v1.8.3 2025-11-28 14:42:48 +02:00
Slavi Pantaleev
0a8f9fc3e5 Release 1.8.3 2025-11-28 14:40:38 +02:00
Slavi Pantaleev
1ac3b2e060 Update dependencies 2025-11-28 14:23:07 +02:00
Slavi Pantaleev
a3ef9fd1bf Upgrade services 2025-11-28 14:19:59 +02:00
Slavi Pantaleev
df507eb201 Make use of the BAIBOT_USER_ENCRYPTION_RECOVERY_RESET_ALLOWED environment variable for configuring config.user.encryption.recovery_reset_allowed 2025-11-28 14:18:35 +02:00
Slavi Pantaleev
4dcd9eff40 Make use of the BAIBOT_PERSISTENCE_SESSION_ENCRYPTION_KEY environment variable for configuring config.persistence.session_encryption_key
Seems like we already had a constant defined, but weren't making use of it.
2025-11-28 14:17:23 +02:00
Slavi Pantaleev
ea760ce755 Release 1.8.2 2025-11-20 06:00:52 +02:00
Slavi Pantaleev
1528df6a55 Upgrade Rust (1.90.0 -> 1.91.1) 2025-11-20 05:51:45 +02:00
Slavi Pantaleev
b0fa024297 Update services 2025-11-20 05:50:36 +02:00
Slavi Pantaleev
3ec203128a Update dependencies 2025-11-20 05:49:15 +02:00
Slavi Pantaleev
da97361e1b Bump default OpenAI text-generation model (gpt-5 -> gpt-5.1) 2025-11-20 05:25:20 +02:00
Slavi Pantaleev
b430fe0189 Update sample OpenAI config (for gpt-5) misleading users into using max_response_tokens & remove openai-o1.yml sample config
No need to have both sample configs now.

Fixes https://github.com/etkecc/baibot/issues/57
2025-11-20 05:25:20 +02:00
Slavi Pantaleev
f03126a9e1 Upgrade Rust (1.89.0 -> 1.90.0) 2025-10-26 09:00:01 +02:00
Slavi Pantaleev
7d46b926c1 Update services 2025-10-26 08:30:42 +02:00
Slavi Pantaleev
6f3c048195 Release 1.8.1 2025-09-12 16:54:11 +03:00
Slavi Pantaleev
b47cf598b5 Update dependencies 2025-09-12 16:53:35 +03:00
Slavi Pantaleev
265ad7e1cb Release 1.8.0 2025-09-08 15:19:06 +03:00
Slavi Pantaleev
a159f67e45 Upgrade Rust (1.88.0 -> 1.89.0) and Debian base (12/bookworm -> 13/trixie) in Dockerfiles 2025-09-08 15:01:25 +03:00
Slavi Pantaleev
624b9de35b Upgrade mxlink (1.9.0 -> 1.10.0) and matrix-sdk (0.13.0 -> 0.14.0) 2025-09-08 14:26:37 +03:00
Slavi Pantaleev
941bf7ca42 Update sample configs for OpenAI (gpt-5) to specify max_completion_tokens, not max_response_tokens
Fixup for b43f61f5ff
2025-09-08 10:37:00 +03:00
Slavi Pantaleev
ef0f1671da Update services 2025-09-08 10:04:48 +03:00
Slavi Pantaleev
b43f61f5ff Change default OpenAI model (gpt-4.1 -> gpt-5)
Ref: https://openai.com/index/introducing-gpt-5/
2025-08-08 07:22:19 +03:00
Slavi Pantaleev
1967d2b34c Release 1.7.6 2025-07-11 16:43:34 +03:00
Slavi Pantaleev
bb3734ad24 Upgrade mxlink (1.8.1 -> 1.9.0) and matrix-sdk (0.12.0 -> 0.13.0) 2025-07-11 16:41:46 +03:00
Slavi Pantaleev
eb6db34177 Update dependencies 2025-07-11 16:23:58 +03:00
Slavi Pantaleev
6e845caa2e Release 1.7.5 2025-06-27 16:40:08 +03:00
Slavi Pantaleev
1004966785 Update Cargo.lock-pinned dependencies 2025-06-27 16:15:53 +03:00
Slavi Pantaleev
7ae1864c2e Upgrade Rust (1.86.0 -> 1.88.0)
Ref: https://releases.rs/docs/1.88.0/
2025-06-27 16:09:12 +03:00
Slavi Pantaleev
68a2fb161f Release 1.7.4 2025-06-10 16:38:03 +03:00
Slavi Pantaleev
3a3eb58d7b Update Cargo.lock-pinned dependencies 2025-06-10 16:37:13 +03:00
Slavi Pantaleev
74d988e650 Update mxlink (1.8.0 -> 1.8.1) 2025-06-10 16:33:44 +03:00
Slavi Pantaleev
ed8bedcd7e Release 1.7.3 2025-06-10 15:28:56 +03:00
Slavi Pantaleev
2842632969 Disable async-openai feature of tiktoken-rs crate
We don't make use of this feature, so it's not necessary.

Ref: af52be12fa/tiktoken-rs/Cargo.toml (L21)
2025-06-10 15:26:25 +03:00
Slavi Pantaleev
10a5bd2abb Update OpenAI model in sample config (gpt-4o -> gpt-4.1) 2025-06-10 15:20:00 +03:00
Slavi Pantaleev
5308b75f52 Update dependencies 2025-06-10 15:18:58 +03:00
Slavi Pantaleev
dad61e1270 Adjust docker run command on installation instructions to ensure /tmp is writable
Ref: https://github.com/etkecc/baibot/issues/43
2025-06-09 10:54:11 +03:00
Slavi Pantaleev
91986a129c Release 1.7.2 2025-05-11 23:20:58 +03:00
Slavi Pantaleev
264f683d6a Allow image_generation.size to be null for OpenAI and default it to that
The API spec for image creation and image editing says "string or null",
so we're allowing `null` now to trigger automatic selection.
2025-05-11 23:20:07 +03:00
Slavi Pantaleev
62f0f4fa0d Release 1.7.1 2025-05-11 22:20:26 +03:00
Slavi Pantaleev
69627abd74 Add image-editing feature documentation to the !bai usage command and adjust texts a bit 2025-05-11 22:19:15 +03:00
Slavi Pantaleev
d2660be33c Update sample config for OpenAI to use gpt-image-1, not dall-e-3 2025-05-10 12:30:16 +03:00
Slavi Pantaleev
ce81fe69bd Release 1.7.0 2025-05-10 12:22:58 +03:00
Slavi Pantaleev
1162636b88 Upgrade Rust (1.85.1 -> 1.86.0) 2025-05-10 12:22:46 +03:00
Slavi Pantaleev
8c90e13a79 Upgrade services 2025-05-10 12:11:08 +03:00
Slavi Pantaleev
274b614d25 Update dependencies 2025-05-10 11:58:06 +03:00
Slavi Pantaleev
7bd46821dc Update README to mention the images editing feature 2025-05-10 11:49:43 +03:00
Slavi Pantaleev
a84135ff32 fmt 2025-05-10 11:47:50 +03:00
Slavi Pantaleev
231528a0d8 Document which providers support vision 2025-05-10 11:47:00 +03:00
Slavi Pantaleev
d8e47b0578 Document vision support for text-generation 2025-05-10 11:39:33 +03:00
Slavi Pantaleev
96c1542f4a Add an image-editing screenshot that demos multiple image support 2025-05-10 11:36:14 +03:00
Slavi Pantaleev
2f9c3dfce0 Use patched anthropic-rs library to add Vision support to text conversations
Related to: https://github.com/AbdelStark/anthropic-rs/pull/11
2025-05-10 10:25:15 +03:00
Slavi Pantaleev
de958208b2 Use patched async-openai library to work around a few upstream issues
Related to:

- CreateImageEditRequest forces an application/octet-stream content type
  for images (https://github.com/64bit/async-openai/issues/364)

- CreateImageEditRequest only deals with a single image
  (https://github.com/64bit/async-openai/issues/363)

This is a continuation of 8f86289373
and fixes the Image Editing feature for OpenAI.
2025-05-10 10:00:36 +03:00
Slavi Pantaleev
ac4f2080ce Make some improvements as suggested by clippy 2025-05-10 09:33:44 +03:00
Slavi Pantaleev
3ffa50b7b9 Use ImageInput|AudioInput::from_vec_u8 helper 2025-05-10 09:29:06 +03:00
Slavi Pantaleev
8f86289373 Initial work on Vision support in text conversations and Image Editing
This is a huge patch which does some major refactoring like:

- renaming "Image Generation" to "Image Creation" in most places,
  to better match its new command (`!bai image create`)

- relocating image creation command (`!bai image` -> `!bai image create`),
  so it wouldn't conflict with the new image editing command (`!bai image edit`)

- introducing a new image editing command (`!bai image edit`), which
  is meant to work only with the OpenAI provider, but doesn't fully work yet
  due to https://github.com/64bit/async-openai/issues/364, though a next patch will fix it

- adding support for reading images off of Matrix conversations and forwarding them to
  text conversations. Works for OpenAI, but not for Anthropic yet
  (requires custom patches) and not for OpenAI-Compat (no support for
  images there)

- relocating some utils around (base64, mime)
2025-05-10 09:18:01 +03:00
Slavi Pantaleev
e0dcc39a72 Default OpenAI image-generation model to gpt-image-1 (previously dall-e-3) 2025-05-03 09:42:46 +03:00
Slavi Pantaleev
c94376109c Avoid passing response_format to OpenAI's image generation API for the gpt-image-1 model
The API reference for `response_format` says:

> This parameter isn't supported for gpt-image-1 which will always return base64-encoded images.

Related to https://github.com/etkecc/baibot/issues/40
2025-05-03 09:42:44 +03:00
Slavi Pantaleev
256ed05662 Make style, quality and internal response_format image generation parameters optional
Some OpenAI models (like `gpt-image-1`) either don't support these or
only support specific other values.

Related to https://github.com/etkecc/baibot/issues/40
2025-05-03 09:42:39 +03:00
Slavi Pantaleev
8222681e27 Default OpenAI text-generation model to gpt-4.1 (previously gpt-4o) 2025-05-03 09:42:37 +03:00
Slavi Pantaleev
f304b93c68 Improve in-room error reporting details when image generation fails
What previously was a generic error message like:

> ⚠️ Error: An error occurred while processing your message. Please try again.

.. now becomes a much more helpful error message like:

> ⚠️ Error: There was a problem performing image-generation via the room-local/my-openai-agent agent:
>
> invalid_request_error: Invalid value: 'standard'. Supported values are: 'low', 'medium', 'high', and 'auto'. (param: quality) (code: invalid_value)

Related to https://github.com/etkecc/baibot/issues/40
2025-05-03 09:42:28 +03:00
Slavi Pantaleev
889d8a1d04 Upgrade Synapse (v1.127.1 -> v1.128.0) and Element Web (v1.11.96 -> v1.11.99) 2025-05-03 08:50:00 +03:00
Slavi Pantaleev
6082bfaf56 Release 1.6.0 2025-04-12 08:04:23 +03:00
Slavi Pantaleev
1d629e0859 Update dependencies 2025-04-12 07:57:12 +03:00
Slavi Pantaleev
49471c1df0 Upgrade mxlink (1.6.0 -> 1.7.0) and matrix-sdk (0.10.0 -> 0.11.0) 2025-04-12 07:49:26 +03:00
138 changed files with 8805 additions and 2292 deletions

26
.github/workflows/ci.yml vendored Normal file
View File

@@ -0,0 +1,26 @@
name: CI
on:
workflow_dispatch:
pull_request:
branches: [ "main" ]
push:
branches:
- "**"
tags: [ "v*" ]
permissions:
contents: read
pull-requests: read
concurrency:
group: ci-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
jobs:
test-and-clippy:
name: Unit testing and linting
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: dtolnay/rust-toolchain@1.93.0
- name: Install SQLite3
run: sudo apt-get update && sudo apt-get install -y libsqlite3-dev
- run: cargo test --all-features
- run: cargo clippy

124
.github/workflows/publish.yml vendored Normal file
View File

@@ -0,0 +1,124 @@
name: Publish
on:
workflow_run:
workflows: [ "CI" ]
types: [ "completed" ]
permissions:
contents: read
concurrency:
group: publish-${{ github.event.workflow_run.id || github.ref }}
cancel-in-progress: false
jobs:
docker-clean-metadata:
if: |
github.event.workflow_run.conclusion == 'success' &&
github.event.workflow_run.event == 'push' &&
(
github.event.workflow_run.head_branch == 'main' ||
startsWith(github.event.workflow_run.head_branch || '', 'v')
)
runs-on: ubuntu-latest
outputs:
json: ${{ steps.meta.outputs.json }}
steps:
- name: Checkout
uses: actions/checkout@v7
with:
ref: ${{ github.event.workflow_run.head_sha }}
fetch-depth: 0
- name: Extract metadata (tags, labels) for Docker
id: meta
uses: docker/metadata-action@v6
with:
images: |
ghcr.io/${{ github.repository }}
tags: |
type=raw,value=latest,enable=${{ github.event.workflow_run.head_branch == 'main' }}
type=semver,pattern={{raw}},value=${{ github.event.workflow_run.head_branch }},enable=${{ startsWith(github.event.workflow_run.head_branch || '', 'v') }}
docker-build:
if: |
github.event.workflow_run.conclusion == 'success' &&
github.event.workflow_run.event == 'push' &&
(
github.event.workflow_run.head_branch == 'main' ||
startsWith(github.event.workflow_run.head_branch || '', 'v')
)
permissions:
contents: read
packages: write
attestations: write
id-token: write
strategy:
matrix:
include:
- os: self-hosted
arch: amd64
- os: ubuntu-24.04-arm
arch: arm64
runs-on: ${{ matrix.os }}
steps:
- name: Checkout
uses: actions/checkout@v7
with:
ref: ${{ github.event.workflow_run.head_sha }}
fetch-depth: 0
- name: Log in to the GitHub Container registry
uses: docker/login-action@v4
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Extract metadata (tags, labels) for Docker
id: meta
uses: docker/metadata-action@v6
with:
tags: |
type=raw,value=latest,enable=${{ github.event.workflow_run.head_branch == 'main' }}
type=semver,pattern={{raw}},value=${{ github.event.workflow_run.head_branch }},enable=${{ startsWith(github.event.workflow_run.head_branch || '', 'v') }}
flavor: |
latest=auto
suffix=-${{ matrix.arch }},onlatest=true
images: |
ghcr.io/${{ github.repository }}
- name: Build and push Docker images
uses: docker/build-push-action@v7
with:
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
docker-manifest:
if: |
github.event.workflow_run.conclusion == 'success' &&
github.event.workflow_run.event == 'push' &&
(
github.event.workflow_run.head_branch == 'main' ||
startsWith(github.event.workflow_run.head_branch || '', 'v')
)
permissions:
contents: read
packages: write
needs:
- docker-build
- docker-clean-metadata
runs-on: ubuntu-latest
strategy:
matrix:
image: ${{ fromJson(needs.docker-clean-metadata.outputs.json).tags }}
steps:
- name: Log in to the GitHub Container registry
uses: docker/login-action@v4
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Create and push manifest
run: |
docker buildx imagetools create -t ${{ matrix.image }} ${{ matrix.image }}-amd64 ${{ matrix.image }}-arm64

View File

@@ -1,107 +0,0 @@
name: CI (main and tags)
on:
push:
branches: [ "main" ]
tags: [ "v*" ]
permissions:
checks: write
contents: write
packages: write
pull-requests: read
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: false
jobs:
test-and-clippy:
name: Unit testing and linting
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- name: Install SQLite3
run: sudo apt-get update && sudo apt-get install -y libsqlite3-dev
- run: cargo test --all-features
- run: cargo clippy
docker-clean-metadata:
runs-on: ubuntu-latest
outputs:
json: ${{ steps.meta.outputs.json }}
steps:
- name: Extract metadata (tags, labels) for Docker
id: meta
uses: docker/metadata-action@v5
with:
images: |
ghcr.io/${{ github.repository }}
tags: |
type=raw,value=latest,enable=${{ github.ref_name == 'main' }}
type=semver,pattern={{raw}}
docker-build:
permissions:
contents: read
packages: write
attestations: write
id-token: write
strategy:
matrix:
include:
- os: self-hosted
arch: amd64
- os: ubuntu-24.04-arm
arch: arm64
runs-on: ${{ matrix.os }}
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Log in to the GitHub Container registry
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Extract metadata (tags, labels) for Docker
id: meta
uses: docker/metadata-action@v5
with:
tags: |
type=raw,value=latest,enable=${{ github.ref_name == 'main' }}
type=semver,pattern={{raw}}
flavor: |
latest=auto
suffix=-${{ matrix.arch }},onlatest=true
images: |
ghcr.io/${{ github.repository }}
- name: Build and push Docker images
uses: docker/build-push-action@v6
with:
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
docker-manifest:
needs:
- docker-build
- docker-clean-metadata
runs-on: ubuntu-latest
strategy:
matrix:
image: ${{ fromJson(needs.docker-clean-metadata.outputs.json).tags }}
steps:
- name: Log in to the GitHub Container registry
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Create and push manifest
run: |
docker manifest create ${{ matrix.image }} ${{ matrix.image }}-amd64 ${{ matrix.image }}-arm64
docker manifest push ${{ matrix.image }}

36
.pre-commit-config.yaml Normal file
View File

@@ -0,0 +1,36 @@
repos:
# Fast built-in hooks (Rust-native, no dependencies)
- repo: builtin
hooks:
- id: trailing-whitespace
- id: end-of-file-fixer
- id: check-yaml
- id: check-merge-conflict
- id: check-added-large-files
args: ['--maxkb=1024']
# Local hooks that run project-specific tools
- repo: local
hooks:
- id: cargo-fmt-check
name: Cargo Format Check
entry: cargo fmt --all -- --check
language: system
files: '\.rs$'
pass_filenames: false
- id: cargo-clippy
name: Cargo Clippy
entry: cargo clippy -- -D warnings
language: system
files: '\.rs$'
pass_filenames: false
priority: 100
- id: test-unit
name: Unit Tests
entry: just test
language: system
files: '\.rs$'
pass_filenames: false
priority: 100

View File

@@ -1,3 +1,270 @@
# (2026-06-26) Version 1.24.0
- (**Feature**) Add an opt-in 💭 **thinking notice** for text generation. When enabled, a slow response (for example, from a reasoning model that runs for minutes) posts a "thinking…" placeholder after a short delay, refreshes it periodically with varying flavor text, and then edits that same message into the final answer, so a long wait no longer looks like a stuck bot. The notice is **disabled by default** and configurable per-room or globally via `text-generation set-thinking-notice-enabled true`. Fast responses (under the delay threshold) never show a placeholder. See the [text-generation configuration docs](./docs/configuration/text-generation.md#-thinking-notice).
- (**Bugfix**) The [Venice](https://venice.ai) unsupported-field auto-recovery (added in 1.23.1) now actually remembers rejections across messages. The cache lived on the provider's controller, which is rebuilt on every message for room-local and global agents, so each turn started with an empty cache, re-sent the unsupported field, and logged the same `400 Bad Request` warning again. The cache is now process-global (keyed per Venice deployment), so a field a model rejects is dropped proactively on every later request instead of being re-discovered each turn.
# (2026-06-24) Version 1.23.1
- (**Bugfix**) The [Venice](https://venice.ai) provider now auto-recovers when a model rejects an optional knob it does not support. Venice's request body is strict (`additionalProperties: false`), so a model that lacks `prompt_cache_retention`, `reasoning_effort`, or `prompt_cache_key` rejected the whole request with a `400 Bad Request` — breaking agent creation and every reply. baibot now drops the unsupported field and retries, remembering the rejection per model so later requests skip it without a wasted round-trip. Only these meaning-preserving fields are dropped; sampling knobs that change the output (`temperature`, `top_p`, the penalties) are never silently removed and still surface as an error.
- (**Improvement**) When Venice rejects a request with a `400 Bad Request`, baibot now surfaces Venice's actual error message (e.g. `Extra inputs are not permitted, field: 'prompt_cache_retention'`) instead of a generic "configuration does not result in a working agent". This makes agent-creation failures self-explanatory. Other error statuses keep their bodies redacted, since those can carry account or rate-limit details.
# (2026-06-23) Version 1.23.0
- (**Feature**) The [Venice](https://venice.ai) provider now accepts file inputs (PDF, DOCX, and other documents, up to 25MB), the same way it already handled images. This makes Venice the second provider after OpenAI to accept files; the others (Anthropic and the OpenAI-compatible providers) skip them. See the [text-generation feature docs](./docs/features.md#-text-generation).
- (**Feature**) Add prompt caching to the Venice provider, on by default (`prompt_cache_retention: 24h`). baibot derives the cache key from the system prompt and the conversation start time (both fixed for the life of a conversation), so a long, stable system prompt stays cached across the day instead of being reprocessed and re-billed on every turn. See [Text Generation / Prompt Override](./docs/configuration/text-generation.md#️-prompt-override).
- (**Feature**) Wire up the rest of Venice's sampling and reasoning controls: top-level `top_p`, `frequency_penalty`, `presence_penalty`, `repetition_penalty`, and `reasoning_effort`; `verbosity` in the `venice_parameters` bag; and a `show_reasoning` toggle that appends the model's reasoning to the reply as a collapsible, folded-by-default `💭 Reasoning` block (off by default). See the [Venice configuration reference](./docs/providers.md#venice).
- (**Feature**) Render Venice web-search citations as readable `[n]` references with a `Sources:` list of links, instead of leaving Venice's raw `^n^` superscripts in the reply.
- (**Security**) Escape citation titles and validate citation URLs before rendering them, and drop user-supplied filenames from error messages, so a hostile web page or a crafted filename cannot inject a spoofed link into the bot's reply.
- (**Bugfix**) The [OpenAI-compatible](./docs/providers.md#openai-compatible) provider now trusts the system CA store (honoring `SSL_CERT_FILE`), so endpoints served behind a private/internal CA (FreeIPA, organization PKI) no longer fail the TLS handshake with `invalid peer certificate: UnknownIssuer`. Fixed upstream in `etke_openai_api_rust` 0.1.10. Thanks to [@shaba](https://github.com/shaba) for the report in [#188](https://github.com/etkecc/baibot/pull/188).
# (2026-06-21) Version 1.22.0
- (**Feature**) Add a native [Venice](https://venice.ai) provider with [🖌️ image-generation](./docs/features.md#️-image-creation) (incl. editing), [💬 text-generation](./docs/features.md#-text-generation) (incl. vision), [🗣️ text-to-speech](./docs/features.md#️-text-to-speech), [🦻 speech-to-text](./docs/features.md#-speech-to-text), and Venice's native web search via the full `venice_parameters` knob set. Unlike the [OpenAI-compatible](./docs/providers.md#openai-compatible) path (which drops images and can't reach Venice's audio or native image endpoints), it talks to Venice's API directly, using the knob-rich native `/image/generate` and `/image/edit` endpoints. See the [Venice provider docs](./docs/providers.md#venice).
# (2026-06-05) Version 1.21.1
- (**Security**) Update the [anthropic](https://github.com/etkecc/anthropic-rs) dependency to use [reqwest](https://crates.io/crates/reqwest) 0.12 / [rustls](https://crates.io/crates/rustls) 0.23, replacing the vulnerable `rustls-webpki` 0.101 line with 0.103.13. This resolves [`GHSA-82j2-j2ch-gfr8`](https://github.com/advisories/GHSA-82j2-j2ch-gfr8) (high — denial of service via panic on a malformed CRL), [`GHSA-xgp8-3hg3-c2mh`](https://github.com/advisories/GHSA-xgp8-3hg3-c2mh) and [`GHSA-965h-392x-2mh5`](https://github.com/advisories/GHSA-965h-392x-2mh5) (name-constraint validation issues).
# (2026-06-05) Version 1.21.0
- (**Improvement**) Default to OpenAI's `gpt-image-2` model for image generation (in newly-created OpenAI agents and the sample provider configs).
- (**Internal Improvement**) Update [async-openai](https://crates.io/crates/async-openai) from 0.40 to 0.41, which [resynchronizes with the upstream OpenAI API spec](https://github.com/64bit/async-openai/issues/557) after it had drifted out of sync — a mismatch that was already causing some breakage (hopefully now resolved). Adapts to the newly-added `gpt-image-2` image model and an `ImageSize` type change.
- (**Internal Improvement**) Dependency updates.
# (2026-06-02) Version 1.20.0
- (**Internal Improvement**) Update [matrix-sdk](https://crates.io/crates/matrix-sdk) from 0.17 to 0.18 and [mxlink](https://crates.io/crates/mxlink) to 1.15.0.
- (**Internal Improvement**) Update [tiktoken-rs](https://crates.io/crates/tiktoken-rs) to 0.12, backporting OpenAI [tiktoken](https://github.com/openai/tiktoken) 0.13.0 for better alignment with upstream tokenization behavior.
- (**Internal Improvement**) Bump the pinned Rust toolchain from 1.95.0 to 1.96.0 (in `rust-toolchain.toml` and the Docker build images).
- (**Internal Improvement**) Dependency updates.
# (2026-05-27) Version 1.19.3
- (**Internal Improvement**) Update [async-openai](https://crates.io/crates/async-openai) to 0.40.2, pulling in several upstream fixes (streaming HTTP error surfacing, default `ResponseTextParam.format` deserialization, etc.).
- (**Internal Improvement**) Dependency updates.
# (2026-05-21) Version 1.19.2
- (**Internal Improvement**) Update [async-openai](https://crates.io/crates/async-openai) to 0.40.0.
- (**Internal Improvement**) Dependency updates.
# (2026-05-09) Version 1.19.1
- (**Internal Improvement**) Update [async-openai](https://crates.io/crates/async-openai) to 0.38.0.
- (**Internal Improvement**) Dependency updates.
# (2026-05-09) Version 1.19.0
- (**Internal Improvement**) Update [matrix-sdk](https://crates.io/crates/matrix-sdk) from 0.16 to 0.17 and [mxlink](https://crates.io/crates/mxlink) to 1.14.0. matrix-sdk 0.17 dropped its `native-tls` feature and now uses [rustls](https://github.com/rustls/rustls) exclusively as its TLS backend.
- (**Internal Improvement**) Bump the pinned Rust toolchain from 1.93.0 to 1.95.0 (in `rust-toolchain.toml` and the Docker build images).
- (**Internal Improvement**) Dependency updates.
# (2026-04-11) Version 1.18.0
- (**Bugfix**) Fix the bot not sending a welcome message when joining a room on homeservers (like [Continuwuity](https://continuwuity.org/)) that place the join membership event in the sync response's `state` block rather than the `timeline` block, via [mxlink](https://crates.io/crates/mxlink) 1.13.1
- (**Improvement**) Update [tiktoken-rs](https://crates.io/crates/tiktoken-rs) to 0.11, adding tokenization support for newer GPT models (gpt-5.x, codex, etc.) and fixing context sizes for o1-mini/chatgpt-4o/gpt-4.5
- (**Internal Improvement**) Dependency updates
# (2026-03-25) Version 1.17.0
- (**Feature**) Add `text-generation sender-context-mode` for attaching sender metadata to conversation messages. See the [💬 Text Generation](./docs/configuration/text-generation.md#-sender-context-mode) documentation for details. Thanks to [kschwank](https://github.com/kschwank) for the contribution in [#104](https://github.com/etkecc/baibot/pull/104)!
# (2026-03-24) Version 1.16.1
- (**Bugfix**) Fix compatibility with [async-openai](https://crates.io/crates/async-openai) 0.34.0 by populating the new `phase` field required for OpenAI Responses API message inputs. baibot does not currently distinguish between assistant `commentary` and `final_answer` turns, so using `None` preserves the previous behavior while remaining compatible with the updated crate.
- (**Internal Improvement**) Dependency updates.
# (2026-03-20) Version 1.16.0
- (**Feature**) Add support for file attachments (`m.file` Matrix messages) in conversations. Files like PDFs, text documents, spreadsheets, code files, etc. are now downloaded and forwarded to the LLM alongside the conversation context, similar to how images (`m.image`) are already handled. See the [💬 Text Generation](./docs/features.md#-text-generation) documentation for details and known limitations.
- (**Improvement**) Use the [mime_guess](https://crates.io/crates/mime_guess) crate for MIME type detection from file extensions, replacing a hand-maintained mapping. This covers hundreds of file extensions out of the box.
# (2026-03-07) Version 1.15.0
- (**Feature**) Add support for authentication via access tokens (for [Matrix Authentication Service](https://github.com/element-hq/matrix-authentication-service)/OIDC-enabled homeservers) as an alternative to password authentication. See [🔐 Authentication](./docs/configuration/authentication.md) for setup details. Thanks to [Taylor Southwick](https://github.com/twsouthwick) for the contribution in [#83](https://github.com/etkecc/baibot/pull/83)!
- (**Internal Improvement**) Pin the Rust toolchain to `1.93.0` in both CI and local development to avoid `matrix-sdk` build failures on newer stable toolchains.
- (**Internal Improvement**) Documentation updates.
- (**Internal Improvement**) Dependency updates.
# (2026-02-18) Version 1.14.3
- (**Internal Improvement**) Add [Renovate](https://docs.renovatebot.com/) configuration for automated dependency updates
- (**Internal Improvement**) Dependency updates
# (2026-02-18) Version 1.14.2
- (**Internal Improvement**) Dependency updates
- (**Internal Improvement**) Reorganize the development environment to support [Continuwuity](https://continuwuity.org/) as a homeserver choice (in addition to [Synapse](https://github.com/element-hq/synapse)). Continuwuity is now the default for its lighter footprint (no external database required). See [development docs](./docs/development.md) for details.
# (2026-02-10) Version 1.14.1
- (**Security**) Dependency updates to fix security vulnerabilities ([time](https://crates.io/crates/time) stack exhaustion DoS, [bytes](https://crates.io/crates/bytes) integer overflow), via [mxlink](https://crates.io/crates/mxlink) 1.12.0
- (**Internal Improvement**) Switch from deprecated [serde_yaml](https://crates.io/crates/serde_yaml) to its maintained fork [serde_yaml_ng](https://crates.io/crates/serde_yaml_ng)
- (**Internal Improvement**) Add [prek](https://github.com/nicholasgasior/prek) pre-commit hooks via [mise](https://mise.jdx.dev/) for automated code quality checks (formatting, clippy, tests)
- (**Internal Improvement**) Fix clippy warnings and formatting issues
# (2026-02-04) Version 1.14.0
- (**Feature**) The `openai` provider now uses OpenAI's [Responses API](https://platform.openai.com/docs/api-reference/responses) (instead of the older Chat Completions API), adding support for [🛠️ built-in tools](./docs/features.md#️-built-in-tools-openai-only) (`web_search` and `code_interpreter`). These tools are **disabled by default** and can be enabled via the `text_generation.tools` configuration (see the [sample configuration](https://github.com/etkecc/baibot/blob/c70387b0c38d8d0f30bba2179a2a21a3710dbeaf/docs/sample-provider-configs/openai.yml#L12-L15)). To enable tools on an existing agent, you need to [update the agent](./docs/agents.md#updating-agents) to re-create it with the `text_generation.tools` section added and enable the tools you need. Thanks to [Layla Manley](https://github.com/yeslayla) for the contribution in [#62](https://github.com/etkecc/baibot/pull/62)!
- (**Bugfix**) Fix sticker generation for newer GPT image models (`gpt-image-1`, `gpt-image-1-mini`, `gpt-image-1.5`) which don't support the previously hardcoded `256x256` size (minimum is `1024x1024`)
- (**Internal Improvement**) Dependency updates
# (2026-01-23) Version 1.13.0
- (**Improvement**) Extend auto-switching to support cheaper models (`gpt-image-1-mini`) for `gpt-image-1` and `gpt-image-1.5` when generating stickers ([e0b4a40](https://github.com/etkecc/baibot/commit/e0b4a40))
- (**Internal Improvement**) Upgrade Rust compiler (1.92.0 -> 1.93.0) ([691aeeb](https://github.com/etkecc/baibot/commit/691aeeb))
- (**Internal Improvement**) Dependency updates
# (2025-12-21) Version 1.12.0
- (**Improvement**) Upgrade [async-openai](https://crates.io/crates/async-openai) (0.31.1 -> 0.32.2) and add support for OpenAI's `gpt-image-1.5` model ([08c689a](https://github.com/etkecc/baibot/commit/08c689a), [f7bf3d7](https://github.com/etkecc/baibot/commit/f7bf3d7))
- (**Internal Improvement**) Dependency updates
# (2025-12-15) Version 1.11.0
- (**Feature**) Add support for custom avatars via file path and for keeping the already-set avatar (for those who wish to manage it by themselves via other means). See the [sample config](./etc/app/config.yml.dist) for details. ([062fbbb](https://github.com/etkecc/baibot/commit/062fbbb8ef9ad600db483a431c5c782402191023))
- (**Internal Improvement**) Dependency updates ([99bde53](https://github.com/etkecc/baibot/commit/99bde53ef648a5a9086a96778fde4a9dbc1ede58))
- (**Internal Improvement**) Documentation updates ([b3fd8e5](https://github.com/etkecc/baibot/commit/b3fd8e548f83fe46398ced4760d7e2bb7588c24d))
- (**Internal Improvement**) Upgrade Rust compiler (1.91.1 -> 1.92.0) ([22906aa](https://github.com/etkecc/baibot/commit/22906aa2d3cae51815fad2560a545eaa69c247b6))
# (2025-12-06) Version 1.10.0
- (**Internal Improvement**) Dependency updates. This version is based on [mxlink](https://crates.io/crates/mxlink)@1.11.0 (which is based on the newly released [matrix-sdk](https://crates.io/crates/matrix-sdk)@[0.16.0](https://github.com/matrix-org/matrix-rust-sdk/releases/tag/matrix-sdk-0.16.0).
# (2025-11-30) Version 1.9.0
- (**Internal Improvement**) Upgrade [async-openai](https://crates.io/crates/async-openai) from our own etkecc fork (0.28.1-patched) to the official upstream version 0.31.1. This upgrade required some code adaptations to the new module structure, etc. While tested, regressions are possible.
# (2025-11-28) Version 1.8.3
- (**Improvement**) Add support for the `BAIBOT_PERSISTENCE_SESSION_ENCRYPTION_KEY` environment variable for configuring `persistence.session_encryption_key`
- (**Improvement**) Add support for the `BAIBOT_USER_ENCRYPTION_RECOVERY_RESET_ALLOWED` environment variable for configuring `user.encryption.recovery_reset_allowed`
- (**Internal Improvement**) Dependency updates.
# (2025-11-20) Version 1.8.2
- (**Internal Improvement**) Dependency and compiler updates (Rust 1.89.0 -> 1.91.1).
# (2025-09-12) Version 1.8.1
- (**Internal Improvement**) Dependency updates.
# (2025-09-08) Version 1.8.0
- (**Internal Improvement**) Upgrade [mxlink](https://crates.io/crates/mxlink) (1.9.0 -> 1.10.0) and [matrix-sdk](https://crates.io/crates/matrix-sdk) (0.13.0 -> 0.14.0)
- (**Internal Improvement**) Upgrade [Rust](https://www.rust-lang.org/) (1.88.0 -> 1.89.0)
- (**Internal Improvement**) Upgrade Debian base for container images (12/bookworm -> 13/trixie)
# (2025-07-11) Version 1.7.6
- (**Internal Improvement**) Dependency updates. This version is based on [mxlink](https://crates.io/crates/mxlink)@1.9.0 (which is based on the newly released [matrix-sdk](https://crates.io/crates/matrix-sdk)@[0.13.0](https://github.com/matrix-org/matrix-rust-sdk/releases/tag/matrix-sdk-0.13.0), which contains fixes for some security vulnerabilities)
# (2025-06-10) Version 1.7.5
- (**Internal Improvement**) Dependency and compiler updates (Rust 1.86 -> 1.86).
# (2025-06-10) Version 1.7.4
- (**Internal Improvement**) Dependency updates.
# (2025-06-10) Version 1.7.3
- (**Internal Improvement**) Dependency updates. This version is based on [mxlink](https://crates.io/crates/mxlink)@1.8.0 (which is based on the newly released [matrix-sdk](https://crates.io/crates/matrix-sdk)@[0.12.0](https://github.com/matrix-org/matrix-rust-sdk/releases/tag/matrix-sdk-0.12.0), which contains fixes for important security vulnerabilities)
# (2025-05-11) Version 1.7.2
- (**Bugfix**) Allow `image_generation.size` configuration value for OpenAI to be `null` to allow the model to choose the size automatically and default to that
# (2025-05-11) Version 1.7.1
- (**Bugfix**) Fix lack of documentation for the new [image-editing](./docs/features.md#-image-editing) feature in the `!bai usage` command's output
# (2025-05-10) Version 1.7.0
- (**Feature**) Add vision support to the OpenAI and Anthropic providers. You can now mix text and images in your conversations - fixes [issue #5](https://github.com/etkecc/baibot/issues/5)
- (**Feature**) Add [image-editing](./docs/features.md#-image-editing) support to the OpenAI provider
- (**Improvement**) Add compatibility with OpenAI's `gpt-image-1` model - fixes [issue #40](https://github.com/etkecc/baibot/issues/40)
- (**Change**) Rework [image-creation](./docs/features.md#-image-creation) to avoid command conflicts with [image-editing](./docs/features.md#-image-editing). The image-creation command syntax is now `!bai image create <prompt>` (previously: `!bai image <prompt>`).
- (**Internal Improvement**) Dependency and compiler updates
> [!WARNING]
> Unlike other releases, this release is not published to [crates.io](https://crates.io), because it relies on multiple library forks (`async-openai` and `anthropic-rs`) sourced from Github.
# (2025-04-12) Version 1.6.0
- (**Internal Improvement**) Dependency updates. This version is based on [mxlink](https://crates.io/crates/mxlink)@1.7.0 (which is based on the newly released [matrix-sdk](https://crates.io/crates/matrix-sdk)@[0.11.0](https://github.com/matrix-org/matrix-rust-sdk/releases/tag/matrix-sdk-0.11.0))
# (2025-03-31) Version 1.5.1
- (**Internal Improvement**) Dependency updates

3174
Cargo.lock generated

File diff suppressed because it is too large Load Diff

View File

@@ -7,7 +7,7 @@ license = "AGPL-3.0-or-later"
readme = "README.md"
keywords = ["matrix", "chat", "bot", "AI", "LLM"]
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
version = "1.5.1"
version = "1.24.0"
edition = "2024"
[lib]
@@ -15,25 +15,28 @@ name = "baibot"
path = "src/lib.rs"
[dependencies]
anthropic = "=0.0.8"
anthropic = { git = "https://github.com/etkecc/anthropic-rs.git", branch = "fix-content-block-image" }
anyhow = "1.0.*"
async-openai = "0.28.*"
async-openai = { version = "0.41.0", features = ["audio", "chat-completion", "image", "responses"] }
base64 = "0.22.*"
chrono = { version = "0.4.*", default-features = false, features = ["std", "now"] }
# We'd rather not depend on this, but we cannot use the ruma-events EventContent macro without it.
# We add the `native-tls` feature, because of https://github.com/etkecc/rust-mxlink/issues/1
matrix-sdk = { version = "0.10.0", default-features = false, features = ["native-tls"] }
matrix-sdk = { version = "0.18.0", default-features = false }
mime_guess = "2.0.*"
mxidwc = "1.0.*"
mxlink = ">=1.6.0"
mxlink = ">=1.15.0"
etke_openai_api_rust = "0.1.*"
quick_cache = "0.6.*"
regex = "1.11.*"
regex = "1.12.*"
# HTTP client for the native `venice` provider. rustls only (no extra TLS stack), matching the
# reqwest copy async-openai/matrix-sdk/mxlink already use.
reqwest = { version = "0.13.*", default-features = false, features = ["json", "multipart", "rustls"] }
serde = { version = "1.0.*", features = ["derive"], default-features = false }
serde_json = "1.0.*"
serde_yaml = "0.9.*"
tempfile = "3.19.*"
tiktoken-rs = { version = "0.6.*", features = ["async-openai"] }
tokio = { version = "1.44.*", features = ["rt", "rt-multi-thread", "macros"] }
serde_yaml_ng = "0.10.*"
tempfile = "3.27.*"
tiktoken-rs = { version = "0.12.*", default-features = false }
tokio = { version = "1.52.*", features = ["rt", "rt-multi-thread", "macros"] }
tracing = "0.1.*"
tracing-subscriber = { version = "0.3.*", features = ["env-filter"] }
url = "2.5.*"

View File

@@ -4,7 +4,7 @@
# #
#######################################
FROM docker.io/rust:1.85.1-slim-bookworm AS build
FROM docker.io/rust:1.96.0-slim-trixie AS build
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev
@@ -39,7 +39,7 @@ RUN --mount=type=cache,target=/target,sharing=locked \
# #
#######################################
FROM docker.io/debian:bookworm-slim
FROM docker.io/debian:trixie-slim
RUN apt-get update && apt-get install -y ca-certificates sqlite3 && \
apt-get clean && \

View File

@@ -4,7 +4,7 @@
# #
#######################################
FROM docker.io/rust:1.85.1-slim-bookworm AS build
FROM docker.io/rust:1.96.0-slim-trixie AS build
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev
@@ -20,7 +20,7 @@ RUN cargo build --release
# #
#######################################
FROM docker.io/debian:bookworm-slim
FROM docker.io/debian:trixie-slim
RUN apt-get update && apt-get install -y ca-certificates sqlite3 && \
apt-get clean && \

View File

@@ -13,14 +13,14 @@ It's influenced by [chaz](https://github.com/arcuru/chaz), but does **not** use
## 🌟 Features
- 🎨 Encourages **[provider](./docs/providers.md) choice** ([Anthropic](./docs/providers.md#anthropic), [Groq](./docs/providers.md#groq), [LocalAI](./docs/providers.md#localai), [OpenAI](./docs/providers.md#openai) and [☁️ many more](./docs/providers.md#️-providers)) as well as **[mixing & matching models](./docs/features.md#-mixing--matching-models)**:
- 🎨 Encourages **[provider](./docs/providers.md) choice** ([Anthropic](./docs/providers.md#anthropic), [Groq](./docs/providers.md#groq), [LocalAI](./docs/providers.md#localai), [OpenAI](./docs/providers.md#openai), [Venice](./docs/providers.md#venice) and [☁️ many more](./docs/providers.md#️-providers)) as well as **[mixing & matching models](./docs/features.md#-mixing--matching-models)**:
- Supports **different use purposes** (depending on the [☁️ provider](./docs/providers.md) & model):
- [💬 text-generation](./docs/features.md#-text-generation): communicating with you via text
- [💬 text-generation](./docs/features.md#-text-generation): communicating with you via text (though certain models may "see" images as well). The [OpenAI provider](./docs/providers.md#openai) also supports [🛠️ built-in tools](./docs/features.md#️-built-in-tools-openai-only) (web search, code interpreter)
- [🦻 speech-to-text](./docs/features.md#-speech-to-text): turning your voice messages into text
- [🗣️ text-to-speech](./docs/features.md#%EF%B8%8F-text-to-speech): turning bot or users text messages into voice messages
- [🖌️ image-generation](./docs/features.md#%EF%B8%8F-image-generation): generating images based on instructions
- [🖌️ image-generation](./docs/features.md#image-generation): creating and editing images based on instructions
- 🪄 Supports [seamless voice interaction](./docs/features.md#seamless-voice-interaction) (turning user voice messages into text, answering in text, then turning that text back into voice)

View File

@@ -43,7 +43,8 @@ Administrators cannot be changed without adjusting the bot's configuration on th
Room-local agent managers are users privileged to **create their own [agents](./agents.md)** (see `!bai agent`) in rooms.
**⚠️ WARNING**: Letting regular users create agents which contact arbitrary network services **may be a security issue**.
> [!WARNING]
> Letting regular users create agents which contact arbitrary network services **may be a security issue**.
The following commands are available:
- **Show** the currently allowed users: `!bai access room-local-agent-managers`

View File

@@ -35,7 +35,7 @@ Depending on where the agent is defined (within a room, globally, or [statically
When creating an agent, you will be given some sample [YAML](https://en.wikipedia.org/wiki/YAML) configuration which you can use to customize the agent's behavior.
This configuration varies depending on the [☁️ provider](./providers.md) used and the capabilities of the agent. Based on the configuration keys you pass, certain features will be enabled or disabled. For example, if you skip the `image_generation` key for an [OpenAI](./providers.md#openai) agent, it won't be able to generate images (see [🖌️ Image Generation](./features.md#-image-generation)).
This configuration varies depending on the [☁️ provider](./providers.md) used and the capabilities of the agent. Based on the configuration keys you pass, certain features will be enabled or disabled. For example, if you skip the `image_generation` key for an [OpenAI](./providers.md#openai) agent, it won't be able to generate images (see [🖌️ Image Creation](./features.md#-image-creation), [🎨 Image Editing](./features.md#-image-editing), [🫵 Sticker Creation](./features.md#-sticker-creation)).
After making your modifications to the sample YAML, you submit it back to the bot and the new agent will be created.

View File

@@ -12,12 +12,17 @@ This file is created from the template found in [etc/app/config.yml.dist](../../
Certain keys can be left unset, in which case [📝 hardcoded defaults](../../src/entity/cfg/defaults.rs) would be used.
Each configuration key found in the YAML configuration can be overridden by setting an environment variable (dots should be replaced with `_`). Example:
Some configuration keys found in the YAML configuration can be overridden by setting an environment variable (dots should be replaced with `_`). Example:
- to override `command_prefix`, set an environment variable `BAIBOT_COMMAND_PREFIX`
- to override `homeserver.server_name`, set an environment variable `BAIBOT_HOMESERVER_SERVER_NAME`
The static configuration contains an `initial_global_config` key, which is used to populate the bot's global configuration (stored as [dynamic configuration](#dynamic-configuration)) the first time the bot starts. Modifying this subsequently will not have any effect. After initial global configuration creation, it's expected to be managed dynamically via chat commands.
You can see the list of supported environment variables in the [🦀 src/entity/cfg/env.rs](../../src/entity/cfg/env.rs) file.
> [!WARNING]
> The static configuration contains an `initial_global_config` key, which is used to populate the bot's global configuration (stored as [dynamic configuration](#dynamic-configuration)) the first time the bot starts. Modifying this subsequently will not have any effect. After initial global configuration creation, it's expected to be managed dynamically via chat commands.
For Matrix-account authentication setup, see [🔐 Authentication](./authentication.md).
### Dynamic configuration
@@ -40,7 +45,7 @@ You can adjust the following settings per room and/or globally:
- [💬 Text Generation](text-generation.md)
- [🦻 Speech-to-Text](speech-to-text.md)
- [🗣️ Text-to-Speech](text-to-speech.md)
- [🖌️ Image Generation](image-generation.md)
- [🖌️ Image Creation](image-generation.md)
- [🤝 Handlers](handlers.md)
Refer to the bot's help messages (as a response to a `!bai config` help command) for the most up-to-date information on what Room Settings can be configured.

View File

@@ -0,0 +1,23 @@
## 🔐 Authentication
baibot supports 2 authentication modes for the Matrix account (`user.*` keys in config).
Set **exactly one** mode. If both are set (or neither is set), startup validation fails.
### Password authentication
- Config key: `user.password`
- Environment variable: `BAIBOT_USER_PASSWORD`
### Access token authentication
- Config keys: `user.access_token` + `user.device_id`
- Environment variables: `BAIBOT_USER_ACCESS_TOKEN` + `BAIBOT_USER_DEVICE_ID`
Access-token authentication is useful for OIDC-enabled homeservers (e.g. those using [Matrix Authentication Service](https://github.com/element-hq/matrix-authentication-service)).
Example token-generation command:
```sh
mas-cli manage issue-compatibility-token <username> [device_id]
```

View File

@@ -8,10 +8,10 @@ You can also use **different models within the same room** (e.g. [💬 text-gene
The bot supports the following use-purposes:
- [💬 text-generation](../features.md#-text-generation): communicating with you via text
- [💬 text-generation](../features.md#-text-generation): communicating with you via text (though certain models may also process images and files)
- [🦻 speech-to-text](../features.md#-speech-to-text): turning your voice messages into text
- [🗣️ text-to-speech](../features.md#️-text-to-speech): turning bot or users text messages into voice messages
- [🖌️ image-generation](../features.md#-image-generation): generating images based on instructions
- [🖌️ image-generation](../features.md#image-generation): generating images based on instructions
In a given room, each different purpose can be served by a different [provider](../providers.md) and model. This combination of provider and model configuration is called an [🤖 agent](../agents.md). Each purpose can be served by a different **handler** agent.

View File

@@ -1,9 +1,11 @@
## 🖌️ Image Generation
## Image Generation
The Image Generation feature is not configurable at this moment.
The Image Creation and Image Editing features are not configurable at this moment.
You may also wish to see:
- [🌟 Features / 🖌️ Image Generation](../features.md#-image-generation) for a higher-level introduction to the Image Generation features
- [📖 Usage / 🖌️ Image Generation](../usage.md#-image-generation) section for more details on how to use the bot for Image Generation in a room
- [🌟 Features / Image Generation / 🖌️ Image Creation](../features.md#-image-creation) for a higher-level introduction to the Image Creation features
- [🌟 Features / Image Generation / 🎨 Image Editing](../features.md#-image-editing) for a higher-level introduction to the Image Editing features
- [📖 Usage / Image Generation / 🖌️ Creating Images](../usage.md#-creating-images) section for more details on how to use the bot for Image Creation in a room
- [📖 Usage / Image Generation / 🎨 Editing images](../usage.md#-editing-images) section for more details on how to use the bot for Image Editing in a room

View File

@@ -57,6 +57,34 @@ This feature relies on [tokenization](https://en.wikipedia.org/wiki/Large_langua
This setting is **disabled by default**, but can be enabled via `!bai config room text-generation set-context-management-enabled true` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings)).
### 💭 Thinking Notice
The bot can post a 💭 **"thinking…" notice** while text generation is running, useful for slow models (for example, reasoning models that may run for minutes) where the response would otherwise look stuck.
When enabled, a placeholder message appears only after a short delay (so fast responses get no notice), updates periodically with varying status text, and is then edited in place to become the final answer.
This setting is **disabled by default**, but can be enabled via `!bai config room text-generation set-thinking-notice-enabled true` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings)).
### 👤 Sender Context Mode
In multi-user rooms, it may be useful for the model to know which participant sent each message in the conversation context.
To support this, the bot has a `text-generation sender-context-mode` setting, which can be set to:
- (default) `disabled`: do not attach sender metadata to messages before sending them to the model
- `matrix_user_id`: prefix text messages with the sender's Matrix user ID, for example: `[sender=@alice:example.com] Hello bot`
- `matrix_user_id_and_timestamp`: prefix text messages with the sender's Matrix user ID and the message timestamp, for example: `[sender=@alice:example.com sent_at=2026-03-23T14:30:00Z] Hello bot`
This sender metadata is attached to conversation messages before they are sent to the model provider. It applies to user and assistant text messages, but not to system prompts or non-text content.
⚠️ Enabling this sends Matrix user IDs, and optionally timestamps, to the model provider.
Example: `!bai config room text-generation set-sender-context-mode matrix_user_id` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
### ⌨️ Prompt Override
You can override the [system prompt](https://huggingface.co/docs/transformers/en/tasks/prompting) configured at the [🤖 agent](../agents.md) level.
@@ -82,6 +110,8 @@ Prompts may contain the following **placeholder variables** which will be replac
💡 `{{ baibot_now_utc }}` changes as time goes on, which prevents [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) from working. It's better to use `{{ baibot_conversation_start_time_utc }}` in prompts, as its value doesn't change yet still orients the bot to the current date/time.
💡 On the [Venice provider](../providers.md#venice), baibot derives the prompt-cache key from the system prompt and the conversation start time, both stable for the life of a conversation, and ships `prompt_cache_retention: 24h` by default. A stable system prompt then stays cached across the whole conversation instead of being reprocessed (and re-billed) on every turn.
Here's a prompt that combines some of the above variables:
> You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."

View File

@@ -56,7 +56,7 @@ Example: `!bai config room text-to-speech set-speed-override 1.5` (this can also
### 👫 Voice override
The voice override setting lets you change the voice being used by the text-to-speech model configured at the [🤖 agent](../agents.md) level (usually `onyx` when using [OpenAI](../providers.md#openai)).
The voice override setting lets you change the voice being used by the text-to-speech model configured at the [🤖 agent](../agents.md) level (e.g. `onyx` when using [OpenAI](../providers.md#openai), or `af_sky` when using [Venice](../providers.md#venice)).
Possible values (e.g. `onyx`) depend on the model you're using. For example, for [OpenAI](../providers.md#openai)'s Whisper model, [these voices](https://platform.openai.com/docs/guides/text-to-speech/voice-options) are available.

View File

@@ -18,6 +18,27 @@ For local development, we run all dependency services in [🐋 Docker](https://w
- (Optional) an API key for some Large Language Model [☁️ provider](./providers.md) (e.g. [OpenAI](./providers.md#openai)), though we recommend using [LocalAI](#localai) or [Ollama](#ollama) for local development
### Choosing a homeserver
The development environment supports two homeserver implementations:
- **[Continuwuity](https://continuwuity.org/)** (default) — lightweight, no external database required. Good for most development needs.
- **[Synapse](https://github.com/element-hq/synapse)** — the reference implementation, bundled with Postgres. Use this if you need Synapse-specific behavior.
To choose a homeserver (optional — defaults to Continuwuity if skipped):
```sh
just homeserver-init continuwuity # or: just homeserver-init synapse
```
The choice is stored in `var/homeserver` and affects all subsequent commands.
> **Note:** If you switch homeservers after initial setup, you will need to:
> - Delete `var/app/local/` and/or `var/app/container/` (app config and data)
> - Delete `var/services/element-web/` (to regenerate its config)
> - Re-run the prepare and user registration steps
### Getting started guide
Developing [locally](#running-locally) is possible, but requires a [Rust](https://www.rust-lang.org/) toolchain.
@@ -28,11 +49,12 @@ In any case, you will need [🐋 Docker](https://www.docker.com/) as [dependency
#### Running locally
1. Start the core dependency services (Postgres, Synapse, Element Web): `just services-start`
2. (Only the first time around) Prepare initial app configuration in `var/app/local/config.yml`: `just app-local-prepare`
3. (Only the first time around) [Prepare your configuration file](#prepare-your-configuration-file)
4. (Only the first time around) Prepare initial default Matrix user accounts (`admin` and `baibot`): `just users-prepare`
5. (Optional) Start additional services depending on which [agent provider you've chosen](#choosing-an-agent-provider):
1. (Optional) Choose a homeserver: `just homeserver-init continuwuity` (or `synapse`). Default is `continuwuity`.
2. Start the homeserver and Element Web: `just services-start`
3. (Only the first time around) Prepare initial app configuration in `var/app/local/config.yml`: `just app-local-prepare`
4. (Only the first time around) [Prepare your configuration file](#prepare-your-configuration-file)
5. (Only the first time around) Prepare initial default Matrix user accounts (`admin` and `baibot`): `just users-prepare`
6. (Optional) Start additional services depending on which [agent provider you've chosen](#choosing-an-agent-provider):
- for [LocalAI](#localai):
- Start services: `just localai-start`
- Wait a while for LocalAI to start up. It has a lot of models to download. Monitor progress using `just localai-tail-logs`
@@ -40,12 +62,12 @@ In any case, you will need [🐋 Docker](https://www.docker.com/) as [dependency
- for [Ollama](#ollama):
- Start services: `just ollama-start`
- (Only the first time around) Pull the model configured in `agents.static_definitions` in the configuration file: `just ollama-pull-model gemma2:2b`
6. Start the bot: `just run-locally`
7. Go to http://element.127.0.0.1.nip.io:42025/ and login with `admin` / `admin`
8. Create a new room and invite `@baibot:synapse.127.0.0.1.nip.io`
9. When done, stop the bot (`Ctrl` + `C`)
10. Stop the core dependency services: `just services-stop`
11. (Optional) Stop additional services:
7. Start the bot: `just run-locally`
8. Go to http://element.127.0.0.1.nip.io:42025/ and login with `admin` / `admin`
9. Create a new room and invite `@baibot:continuwuity.127.0.0.1.nip.io` (or `@baibot:synapse.127.0.0.1.nip.io` if using Synapse)
10. When done, stop the bot (`Ctrl` + `C`)
11. Stop the services: `just services-stop`
12. (Optional) Stop additional services:
- for [LocalAI](#localai): `just localai-stop`
- for [Ollama](#ollama): `just ollama-stop`
@@ -54,11 +76,12 @@ In any case, you will need [🐋 Docker](https://www.docker.com/) as [dependency
You can avoid having a [Rust](https://www.rust-lang.org/) toolchain installed locally and build/run this in a container.
1. Start the core dependency services (Postgres, Synapse, Element Web): `just services-start`
2. (Only the first time around) Prepare initial app configuration in `var/app/container/config.yml`: `just app-container-prepare`
3. (Only the first time around) [Prepare your configuration file](#prepare-your-configuration-file)
4. (Only the first time around) Prepare initial default Matrix user accounts (`admin` and `baibot`): `just users-prepare`
5. (Optional) Start additional services depending on which [agent provider you've chosen](#choosing-an-agent-provider):
1. (Optional) Choose a homeserver: `just homeserver-init continuwuity` (or `synapse`). Default is `continuwuity`.
2. Start the homeserver and Element Web: `just services-start`
3. (Only the first time around) Prepare initial app configuration in `var/app/container/config.yml`: `just app-container-prepare`
4. (Only the first time around) [Prepare your configuration file](#prepare-your-configuration-file)
5. (Only the first time around) Prepare initial default Matrix user accounts (`admin` and `baibot`): `just users-prepare`
6. (Optional) Start additional services depending on which [agent provider you've chosen](#choosing-an-agent-provider):
- for [LocalAI](#localai):
- Start services: `just localai-start`
- Wait a while for LocalAI to start up. It has a lot of models to download. Monitor progress using `just localai-tail-logs`
@@ -66,12 +89,12 @@ You can avoid having a [Rust](https://www.rust-lang.org/) toolchain installed lo
- for [Ollama](#ollama):
- Start services: `just ollama-start`
- (Only the first time around) Pull the model configured in `agents.static_definitions` in the configuration file: `just ollama-pull-model gemma2:2b`
6. Start the bot: `just run-in-container`
7. Go to http://element.127.0.0.1.nip.io:42025/ and login with `admin` / `admin`
8. Create a new room and invite `@baibot:synapse.127.0.0.1.nip.io`
9. When done, stop the bot (`Ctrl` + `C`)
10. Stop the dependency services: `just services-stop`
11. (Optional) Stop additional services:
7. Start the bot: `just run-in-container`
8. Go to http://element.127.0.0.1.nip.io:42025/ and login with `admin` / `admin`
9. Create a new room and invite `@baibot:continuwuity.127.0.0.1.nip.io` (or `@baibot:synapse.127.0.0.1.nip.io` if using Synapse)
10. When done, stop the bot (`Ctrl` + `C`)
11. Stop the services: `just services-stop`
12. (Optional) Stop additional services:
- for [LocalAI](#localai): `just localai-stop`
- for [Ollama](#ollama): `just ollama-stop`
@@ -93,7 +116,7 @@ For getting started most quickly (and locally), we recommend using [LocalAI](#lo
**Ollama is most lightweight** (~2GB for the container image + ~1.6GB for the model), but supports only [💬 text-generation](./features.md#-text-generation).
**LocalAI requires 4x more disk space** (~6GB for the container image + ~12GB for the models), but supports [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text) and [🖼️ image-generation](./features.md#️-image-generation).
**LocalAI requires 4x more disk space** (~6GB for the container image + ~12GB for the models), but supports [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text) and [🖼️ image-generation](./features.md#️-image-creation).
**OpenAI supports all of these capabilities** as well and does not require powerful hardware or lots of disk space. However, it requires signup and an API key.

View File

@@ -8,7 +8,7 @@ You can also use **different models within the same room** (e.g. [💬 text-gene
The bot supports the following use-purposes:
- [💬 text-generation](#-text-generation): communicating with you via text
- [💬 text-generation](#-text-generation): communicating with you via text (though certain models may also process images and files)
- [🦻 speech-to-text](#-speech-to-text): turning your voice messages into text
- [🗣️ text-to-speech](#%EF%B8%8F-text-to-speech): turning bot or users text messages into voice messages
- [🖌️ image-generation](#%EF%B8%8F-image-generation): generating images based on instructions
@@ -22,14 +22,18 @@ For more information about configuring handlers, see the [🤝 Handlers / Config
### 💬 Text Generation
Text Generation is the bot's ability to **respond to users' text messages with text**.
Text Generation is the bot's ability to **respond to users' messages with text**.
![Screenshot of Text Generation - a user sends a message and the bot replies in a new conversation thread](./screenshots/text-generation.webp)
Some models also support vision and document understanding, so you may be able to mix text, images, and files (PDFs, text documents, etc.) in the same conversation. File inputs (documents such as PDFs) are currently accepted only by the OpenAI and Venice providers; the others skip them. Note that certain providers may not support all file types or may have issues with specific files (e.g. scanned/image-based PDFs). If a file is rejected by the provider, the conversation thread may become unusable — start a new thread to work around this.
In multi-user (group) rooms, to avoid disturbing the normal conversation between people, the bot is auto-configured to only respond to messages starting with the command prefix (`!bai`) or direct mentions via the [💬 Text Generation / 🗟 Prefix Requirement Type](./configuration/text-generation.md#-prefix-requirement-type) setting.
Normally, the bot only responds to allowed [👥 Users](./access.md#-users). In certain cases, it's useful for an allowed user to provoke the bot to respond even in foreign threads or reply chains. You can learn more about this feature in the [On-demand involvement](./features.md#on-demand-involvement) section below.
If needed, the bot can also attach sender metadata to conversation messages before sending them to the model, which can help the model distinguish between participants in multi-user rooms. See [🛠️ Configuration / 💬 Text Generation / 👤 Sender Context Mode](./configuration/text-generation.md#-sender-context-mode).
A few other features (like [🗣️ Text-to-Speech](#️-text-to-speech) and [🦻 Speech-to-Text](#-speech-to-text)) combine well with Text Generation, so you **don't necessarily need to communicate with the bot via text** (with [Seamless voice interaction](#seamless-voice-interaction), you can communicate only with voice).
You may also wish to see:
@@ -38,6 +42,23 @@ You may also wish to see:
- [📖 Usage / 💬 Text Generation](./usage.md#-text-generation) section for more details on how to use the bot for Text Generation in a room
#### 🛠️ Built-in Tools (OpenAI only)
The [OpenAI provider](./providers.md#openai) supports built-in tools that extend the model's capabilities:
- [🔍 Web Search](https://platform.openai.com/docs/guides/tools-web-search) (`web_search`): allows the model to search the web for up-to-date information. [🖼️ Screenshot](./screenshots/text-generation-tools-web-search.webp)
- [💻 Code Interpreter](https://platform.openai.com/docs/guides/tools-code-interpreter) (`code_interpreter`): allows the model to write and execute Python code in a sandbox
These tools are **disabled by default** and need to be explicitly enabled in the agent's `text_generation.tools` configuration. See the [OpenAI sample configuration](https://github.com/etkecc/baibot/blob/c70387b0c38d8d0f30bba2179a2a21a3710dbeaf/docs/sample-provider-configs/openai.yml#L12-L15) for reference.
To enable tools on an existing dynamically-created agent, you need to [update the agent](./agents.md#updating-agents) to re-create it with the `text_generation.tools` section added and enable the tools you need
💡 **Note**: These tools run on OpenAI's infrastructure and may incur additional costs. Web search results include citations that are incorporated into the response.
#### On-demand involvement
In the following 2 cases, it's useful to involve the bot in conversations on-demand:
@@ -139,26 +160,42 @@ To operate in this mode, you can:
- optionally adjust [🦻 Speech-to-Text / 🪄 Message Type for non-threaded only-transcribed messages](./configuration/speech-to-text.md#-message-type-for-non-threaded-only-transcribed-messages), if you'd like to bot to send messages of type `notice` (for better compatibility with other bots in the room) instead of sending regular `text` messages (default)
### 🖌️ Image Generation
### Image Generation
Image generation is the bot's ability to **generate images** based on text prompts.
#### 🖌️ Image Creation
See a [🖼️ Screenshot of the Image Generation feature](./screenshots/image-generation.webp).
Image creation is the bot's ability to **create images** based on text prompts.
See a [🖼️ Screenshot of the Image Creation feature](./screenshots/image-creation.webp).
You may also wish to see:
- [🛠️ Configuration / 🖌️ Image Generation](./configuration/image-generation.md) for configuration options related to Image Generation
- [📖 Usage / 🖌️ Image Generation](./usage.md#-image-generation) section for more details on how to use the bot for Image Generation in a room
- [🫵 Sticker Generation](#-sticker-generation) - a special case of Image Generation
- [📖 Usage / Image Generation / 🖌️ Creating Images](./usage.md#-creating-images) section for more details on how to use the bot for Image Creation in a room
- [🖌️ Image Editing](#️-image-editing) - another image generation feature
- [🫵 Sticker Creation](#-sticker-creation) - a special case of Image Creation
### 🫵 Sticker Generation
#### 🎨 Image Editing
Sticker generation is the bot's ability to **generate sticker** images based on text prompts. It's a special case of [🖌️ Image Generation](#️-image-generation).
Image editing is the bot's ability to **edit images** based on a prompt and one or more existing images.
See a [🖼️ Screenshot of the Sticker Generation feature](./screenshots/sticker-generation.webp).
See a [🖼️ Screenshot of the Image Editing feature (manipulating a single image)](./screenshots/image-editing-single-image.webp) and a [🖼️ Screenshot of the Image Editing feature (manipulating multiple images)](./screenshots/image-editing-multiple-images.webp).
See [📖 Usage / 🖌️ Image Generation / Generating Stickers](./usage.md#generating-stickers) for details.
You may also wish to see:
- [🛠️ Configuration / 🖌️ Image Generation](./configuration/image-generation.md) for configuration options related to Image Generation
- [📖 Usage / Image Generation / 🎨 Editing images](./usage.md#-editing-images) section for more details on how to use the bot for Image Editing in a room
- [🖌️ Image Creation](#️-image-creation) - another image generation feature
#### 🫵 Sticker Creation
Sticker generation is the bot's ability to **generate sticker** images based on text prompts. It's a special case of [🖌️ Image Creation](#️-image-creation).
See a [🖼️ Screenshot of the Sticker Creation feature](./screenshots/sticker-generation.webp).
See [📖 Usage / Image Generation / 🫵 Creating Stickers](./usage.md#-creating-stickers) for details.
### 🔒 Encryption

View File

@@ -53,6 +53,7 @@ CONTAINER_IMAGE_NAME=ghcr.io/etkecc/baibot:v1.0.0
--env BAIBOT_PERSISTENCE_DATA_DIR_PATH=/data \
--mount type=bind,src=/path/to/config.yml,dst=/app/config.yml,ro \
--mount type=bind,src=/path/to/data,dst=/data \
--tmpfs=/tmp:rw,noexec,nosuid,size=1024m \
$CONTAINER_IMAGE_NAME
```

View File

@@ -19,11 +19,12 @@ The list of supported providers is below.
- [OpenAI Compatible](#openai-compatible)
- [OpenRouter](#openrouter)
- [Together AI](#together-ai)
- [Venice](#venice)
### How to choose a provider
If you're not sure which provider to start with, **we recommend [OpenAI](#openai)** as it's the most popular and has the **widest range of capabilities**: [💬 text-generation](./features.md#-text-generation), [🖌️ image-generation](./features.md#️-image-generation), [🦻 speech-to-text](./features.md#-speech-to-text), [🗣️ text-to-speech](./features.md#️-text-to-speech).
If you're not sure which provider to start with, **we recommend [OpenAI](#openai)** as it's the most popular and has the **widest range of capabilities**: [💬 text-generation](./features.md#-text-generation) (incl. vision, incl. [🛠️ tools](./features.md#️-built-in-tools-openai-only)), [🖌️ image-generation](./features.md#️image-generation), [🦻 speech-to-text](./features.md#-speech-to-text), [🗣️ text-to-speech](./features.md#️-text-to-speech).
You don't need to choose just one though. The bot supports [mixing & matching models](./features.md#-mixing--matching-models), so you can use multiple providers at the same time.
@@ -47,7 +48,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `anthropic`
- 🔗 Links: [🏠 Home page](https://www.anthropic.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Anthropic), [👤 Sign up](https://console.anthropic.com/), [📋 Models list](https://docs.anthropic.com/en/docs/about-claude/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (incl. vision, no tools)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local anthropic my-anthropic-agent`
- create a global agent: `!bai agent create-global anthropic my-anthropic-agent`
@@ -61,7 +62,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `groq`
- 🔗 Links: [🏠 Home page](https://groq.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Groq), [👤 Sign up](https://console.groq.com/login), [📋 Models list](https://console.groq.com/docs/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local groq my-groq-agent`
- create a global agent: `!bai agent create-global groq my-groq-agent`
@@ -75,7 +76,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `localai`
- 🔗 Links: [🏠 Home page](https://localai.io/), [📋 Models list](https://localai.io/gallery.html)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local localai my-localai-agent`
- create a global agent: `!bai agent create-global localai my-localai-agent`
@@ -89,7 +90,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `mistral`
- 🔗 Links: [🏠 Home page](https://mistral.ai/), [🌐 Wiki](https://en.wikipedia.org/wiki/Mistral_AI), [👤 Sign up](https://auth.mistral.ai/ui/registration), [📋 Models list](https://docs.mistral.ai/getting-started/models/)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local mistral my-mistral-agent`
- create a global agent: `!bai agent create-global mistral my-mistral-agent`
@@ -103,7 +104,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `ollama`
- 🔗 Links: [🏠 Home page](https://ollama.com/), [📋 Models list](https://ollama.com/library)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local ollama my-ollama-agent`
- create a global agent: `!bai agent create-global ollama my-ollama-agent`
@@ -120,15 +121,12 @@ For services which are not fully compatible with the OpenAI API, consider using
- 🆔 Identifier: `openai`
- 🔗 Links: [🏠 Home page](https://openai.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/OpenAI), [👤 Sign up](https://platform.openai.com/signup), [📋 Models list](https://platform.openai.com/docs/models)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-generation), [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (incl. vision, incl. [🛠️ tools](./features.md#️-built-in-tools-openai-only)), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local openai my-openai-agent`
- create a global agent: `!bai agent create-global openai my-openai-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which:
- in the general case looks [like this](./sample-provider-configs/openai.yml)
- for the [o1](https://platform.openai.com/docs/models/o1) models needs to look [like this](./sample-provider-configs/openai-o1.yml)
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/openai.yml).
### OpenAI Compatible
@@ -140,7 +138,7 @@ Some of these popular services already have **shortcut** providers (leading to t
This provider is just as featureful as the [OpenAI](#openai) provider, but is more compatible with services which do not fully adhere to the [OpenAI API spec](https://github.com/openai/openai-openapi/).
- 🆔 Identifier: `openai-compatible`
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-generation), [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (no vision, no tools), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local openai-compatible my-openai-compatible-agent`
- create a global agent: `!bai agent create-global openai-compatible my-openai-compatible-agent`
@@ -154,7 +152,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- 🆔 Identifier: `openrouter`
- 🔗 Links: [🏠 Home page](https://openrouter.ai/), [👤 Sign up](https://openrouter.ai/), [📋 Models list](https://openrouter.ai/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local openrouter my-openrouter-agent`
- create a global agent: `!bai agent create-global openrouter my-openrouter-agent`
@@ -168,9 +166,101 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- 🆔 Identifier: `together-ai`
- 🔗 Links: [🏠 Home page](https://www.together.ai/), [👤 Sign up](https://api.together.ai/signup), [📋 Models list](https://api.together.xyz/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local together-ai my-together-ai-agent`
- create a global agent: `!bai agent create-global together-ai my-together-ai-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/together-ai.yml).
### Venice
[Venice AI](https://venice.ai) runs inference on Venice-controlled GPUs or zero-data-retention partner infrastructure and stores no prompts or responses, so your conversations don't linger anywhere. It serves both frontier proprietary models and the latest open-source ones.
- 🆔 Identifier: `venice`
- 🔗 Links: [🏠 Home page](https://venice.ai), [👤 Sign up](https://venice.ai), [📋 Models list](https://api.venice.ai/api/v1/models)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation) (incl. editing, via the native knob-rich `/image/generate` and `/image/edit` endpoints), [💬 text-generation](./features.md#-text-generation) (incl. vision, file inputs like PDF and DOCX, and prompt caching; native web search via the `venice_parameters` config), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local venice my-venice-agent`
- create a global agent: `!bai agent create-global venice my-venice-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/venice.yml).
Unlike the [OpenAI Compatible](#openai-compatible) provider (which can talk to Venice but drops images and can't reach its audio or native image endpoints), this is a first-class Venice integration that exposes Venice's full parameter set. Image generation uses the native `/image/generate` endpoint rather than the OpenAI-compatible `/images/generations` shim, so every Venice-specific knob below is available.
#### Configuration reference
Every parameter below is optional unless marked otherwise. Omitting a knob lets Venice apply its own server-side default; this is **not** the same as setting it to `false`, which actively sends `false`.
**`text_generation`** (top-level knobs) — sampling, caching, and reasoning controls that sit directly on `text_generation`, next to `model_id`, `prompt`, `temperature`, `max_response_tokens`, and `max_context_tokens`. They map to top-level fields on Venice's request, separate from the `venice_parameters` bag below.
| Knob | What it does | Default |
|------|--------------|---------|
| `top_p` | Nucleus sampling, `0.0`–`1.0`. An alternative to `temperature`. | — |
| `frequency_penalty` | Penalize tokens by how often they have already appeared, `-2.0`–`2.0`. | — |
| `presence_penalty` | Penalize tokens that have appeared at all, `-2.0`–`2.0`. | — |
| `repetition_penalty` | Penalize repetition. Values above `1.0` discourage repeats. | — |
| `reasoning_effort` | Reasoning budget for models that support it: `low`, `medium`, `high`. | — |
| `prompt_cache_retention` | How long Venice keeps the prompt prefix cached: `default`, `extended`, or `24h`. `24h` is the lever that makes a long, stable system prompt cheap across a day of conversations. | `24h` |
| `show_reasoning` | Append the model's reasoning (its `reasoning_content`) below the answer, as a collapsible `💭 Reasoning` block that stays folded until clicked. Reads a field separate from the answer text, so it works regardless of `strip_thinking_response`. | `false` |
**`text_generation.venice_parameters`** — Venice-specific request knobs sent in the `venice_parameters` bag. Set any of them to override Venice's behavior. The `Default` column shows the value baibot's sample config ships; a `—` means the knob is left unset, so Venice's own default applies.
| Knob | What it does | Default |
|------|--------------|---------|
| `enable_web_search` | Web search mode: `auto` (model decides), `on` (always), or `off`. | `auto` |
| `enable_web_citations` | Append source citations to web-search answers. | — |
| `enable_web_scraping` | Allow the model to scrape page contents during web search. | — |
| `enable_x_search` | Include X (Twitter) in web search. | — |
| `include_search_results_in_stream` | Stream search results back as they arrive. | — |
| `return_search_results_as_documents` | Return search results as structured documents. | — |
| `include_venice_system_prompt` | Prepend Venice's own system prompt alongside yours. | — |
| `character_slug` | Use a public Venice character by its slug. | — |
| `strip_thinking_response` | Strip `<think></think>` blocks from reasoning models so the user sees only the answer. | `true` |
| `disable_thinking` | Disable the model's reasoning step entirely. | — |
| `enable_e2ee` | Run in end-to-end-encrypted mode rather than the default TEE-only mode. | `false` |
| `verbosity` | Response verbosity for models that support it: `low`, `medium`, `high`. | — |
**`text_to_speech`**:
| Knob | What it does | Default |
|------|--------------|---------|
| `model_id` | The Venice TTS model (e.g. `tts-kokoro`, `tts-qwen3-1-7b`, `tts-xai-v1`). | `tts-kokoro` |
| `voice` | The voice to synthesize with. Model-specific (Kokoro: `af_*`/`am_*`/`bf_*`/`bm_*`); a cloned-voice handle (`vv_<id>`) also works. | `af_sky` |
| `response_format` | Audio format: `mp3`, `opus`, `aac`, `flac`, `wav`, or `pcm`. | `mp3` |
| `speed` | Playback speed, `0.25`–`4.0`. | `1.0` |
| `prompt` | A style prompt steering emotion/delivery. Only Qwen 3 TTS honors it. | — |
| `temperature` | Sampling temperature, `0.0`–`2.0`. Only Qwen 3 / Orpheus / Chatterbox HD honor it. | — |
| `top_p` | Nucleus sampling, `0.0`–`1.0`. Only Qwen 3 TTS honors it. | — |
**`image_generation`**:
| Knob | What it does | Default |
|------|--------------|---------|
| `model_id` | The image-generation model. | `chroma` |
| `negative_prompt` | A description of what should **not** appear in the image. | — |
| `cfg_scale` | CFG scale, `0`–`20`. Higher values adhere more closely to the prompt. | — |
| `steps` | Number of inference steps. Model-specific; some models ignore it. | — |
| `style_preset` | A named style to apply (e.g. `3D Model`). | — |
| `seed` | Random seed, `-999999999`–`999999999`. Fix it for reproducible results. | random |
| `safe_mode` | Blur images classified as adult content. | `true` |
| `hide_watermark` | Hide the Venice watermark (may be ignored for some content). | `false` |
| `format` | Output format: `jpeg`, `png`, or `webp`. | `webp` |
| `width` / `height` | Image dimensions in pixels, each `1`–`1280`. | `1024` |
| `aspect_ratio` | Aspect ratio for models that support it (e.g. `1:1`, `16:9`). Alternative to `width`/`height`. | — |
| `resolution` | Resolution tier for models that support it (`1K`, `2K`, `4K`). | — |
| `quality` | Output quality for supported models: `low`, `medium`, `high`. Higher can cost more. | — |
| `lora_strength` | Lora strength, `0`–`100`. Only applies if the model uses additional Loras. | — |
| `embed_exif_metadata` | Embed the generation prompt into the image's EXIF metadata. | `false` |
| `enable_web_search` | Let the model pull the latest info from the web. Model-specific; costs extra credits. | — |
**`image_generation.edit`** — image editing reuses the `image_generation` block; only the model and a few output knobs differ:
| Knob | What it does | Default |
|------|--------------|---------|
| `model_id` | The image-edit model. | `firered-image-edit` |
| `output_format` | Output format: `jpeg`, `png`, or `webp`. When omitted, Venice infers it (PNG at 1K, JPEG at 2K/4K). | inferred |
| `aspect_ratio` | Aspect ratio of the result: `auto`, `1:1`, `3:2`, `16:9`, `21:9`, `9:16`, `2:3`, `3:4`, `4:5` (model-specific). | — |
| `resolution` | Resolution tier: `1K`, `2K`, `4K` (model-specific). | `1K` |
| `safe_mode` | Blur images classified as adult content. | `true` |

View File

@@ -1,24 +0,0 @@
base_url: https://api.openai.com/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: o1-mini
# o1 models do not support a system prompt
prompt: null
temperature: 1.0
# o1 models do not support max_response_tokens.
# They use `max_completion_tokens` as an alternative
max_response_tokens: null
max_completion_tokens: 16384
max_context_tokens: 128000
speech_to_text:
model_id: whisper-1
text_to_speech:
model_id: tts-1-hd
voice: onyx
speed: 1.0
response_format: opus
image_generation:
model_id: dall-e-3
style: vivid
size: 1024x1024
quality: standard

View File

@@ -1,11 +1,18 @@
base_url: https://api.openai.com/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: gpt-4o
model_id: gpt-5.4
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 16384
max_context_tokens: 128000
# Reasoning models need to use `max_completion_tokens` instead of `max_response_tokens`.
# If you're dealing with a non-reasoning model, specify `max_response_tokens` and unset `max_completion_tokens`.
max_response_tokens: null
max_completion_tokens: 128000
max_context_tokens: 400000
# Built-in tools
tools:
web_search: false
code_interpreter: false
speech_to_text:
model_id: whisper-1
text_to_speech:
@@ -14,7 +21,7 @@ text_to_speech:
speed: 1.0
response_format: opus
image_generation:
model_id: dall-e-3
style: vivid
size: 1024x1024
quality: standard
model_id: gpt-image-2
style: null
size: null
quality: null

View File

@@ -0,0 +1,117 @@
base_url: https://api.venice.ai/api/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: kimi-k2-5
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 4096
max_context_tokens: 128000
# Prompt caching: how long Venice keeps the prompt prefix cached. "default", "extended", or "24h".
# "24h" (shipped by default) makes a long, stable system prompt cheap across a day of conversations.
prompt_cache_retention: 24h
# Top-level sampling and reasoning knobs (uncomment to override Venice's default):
# Nucleus sampling, 0.0-1.0 (an alternative to temperature).
# top_p: 0.9
# Penalize tokens by how often they have already appeared, -2.0-2.0.
# frequency_penalty: 0.0
# Penalize tokens that have appeared at all, -2.0-2.0.
# presence_penalty: 0.0
# Penalize repetition; values above 1.0 discourage repeats.
# repetition_penalty: 1.0
# Reasoning budget for models that support it: low, medium, high.
# reasoning_effort: medium
# Append the model's reasoning below the answer as a collapsible "💭 Reasoning" block (folded by
# default). Reads a field separate from the answer text, so it works alongside
# strip_thinking_response (which only strips <think> blocks from the answer).
# show_reasoning: true
# Venice-specific request parameters. Only the keys present below are sent to Venice; omit a
# key to fall back to Venice's own default. Omitting a knob is NOT the same as setting it to
# `false` — `false` actively sends `false`.
venice_parameters:
# Web search: "auto" (model decides), "on" (always), or "off".
enable_web_search: "auto"
# Strip <think></think> blocks from reasoning models so the user sees only the answer.
strip_thinking_response: true
# Run in TEE-only mode instead of end-to-end encryption (works across all models).
enable_e2ee: false
# Other available knobs — uncomment to override Venice's default:
# enable_web_citations: true
# enable_web_scraping: true
# include_venice_system_prompt: false
# include_search_results_in_stream: true
# return_search_results_as_documents: true
# enable_x_search: true
# disable_thinking: true
# Response verbosity for models that support it: low, medium, high.
# verbosity: medium
# character_slug: public-character-id
speech_to_text:
model_id: nvidia/parakeet-tdt-0.6b-v3
text_to_speech:
# The Venice TTS model. Others include tts-qwen3-1-7b, tts-xai-v1,
# tts-elevenlabs-turbo-v2-5, tts-minimax-speech-02-hd. See the models list endpoint.
model_id: tts-kokoro
# The voice to synthesize with. Voices are model-specific: Kokoro uses af_*/am_*/bf_*/bm_*
# (e.g. af_sky, am_adam), other models have their own sets. You can also pass a cloned-voice
# handle (vv_<id>) created via Venice's voice-cloning API. An incompatible voice returns an error.
voice: af_sky
# Output audio format: mp3, opus, aac, flac, wav, or pcm. mp3 is the broadest Matrix-client fit.
response_format: mp3
# Other available knobs — uncomment to override Venice's default:
# Playback speed, 0.25–4.0 (1.0 is normal).
# speed: 1.0
# A style prompt steering emotion/delivery (e.g. "Excited and energetic."). Only Qwen 3 TTS uses it.
# prompt: "Calm and warm."
# Sampling temperature, 0.0–2.0 (higher = more varied). Only Qwen 3 / Orpheus / Chatterbox HD use it.
# temperature: 0.9
# Nucleus sampling, 0.0–1.0. Only Qwen 3 TTS uses it.
# top_p: 1.0
image_generation:
# The image-generation model. See the models list endpoint for the full set.
model_id: chroma
# The image-edit model, used when editing an existing image rather than generating a new one.
# Editing shares this same image_generation config block; only the model differs.
edit:
model_id: firered-image-edit
# Other edit knobs — uncomment to override Venice's default:
# Output format: jpeg, png, or webp. When omitted, Venice infers it (PNG at 1K, JPEG at 2K/4K).
# output_format: png
# Aspect ratio of the result: auto, 1:1, 3:2, 16:9, 21:9, 9:16, 2:3, 3:4, 4:5 (model-specific).
# aspect_ratio: auto
# Resolution tier: 1K, 2K, 4K (model-specific). Defaults to 1K.
# resolution: 1K
# Blur images classified as adult content. Defaults to true.
# safe_mode: true
# Other generation knobs — uncomment to override Venice's default. Omitting a knob is NOT the same
# as setting it: an omitted knob lets Venice apply its own default, a set value is sent verbatim.
# A description of what should NOT appear in the image.
# negative_prompt: "blurry, watermark, text"
# CFG scale, 0–20. Higher values make the image adhere more closely to the prompt.
# cfg_scale: 7.5
# Number of inference steps. Model-specific; some models ignore it.
# steps: 8
# A named style to apply (e.g. "3D Model"). See Venice's image-styles reference.
# style_preset: "3D Model"
# Random seed, -999999999–999999999. Fix it for reproducible results; omit for a random seed.
# seed: 123456789
# Blur images classified as adult content. Defaults to true.
# safe_mode: true
# Hide the Venice watermark. Venice may ignore this for certain generated content. Defaults to false.
# hide_watermark: false
# Output format: jpeg, png, or webp. webp is smallest; png is highest-quality. Defaults to webp.
# format: webp
# Image dimensions in pixels, each 1–1280. Default 1024×1024.
# width: 1024
# height: 1024
# Aspect ratio (used by certain models, e.g. Nano Banana): "1:1", "16:9". An alternative to width/height.
# aspect_ratio: "1:1"
# Resolution tier (used by certain models): "1K", "2K", "4K".
# resolution: "1K"
# Output quality for supported models (e.g. GPT Image 2): low, medium, high. Higher can cost more.
# quality: high
# Lora strength, 0–100. Only applies if the model uses additional Loras.
# lora_strength: 50
# Embed the generation prompt into the image's EXIF metadata. Defaults to false.
# embed_exif_metadata: false
# Let the model pull the latest info from the web for the image. Model-specific; costs extra credits.
# enable_web_search: false

Binary file not shown.

After

Width:  |  Height:  |  Size: 298 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 339 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 285 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 684 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 66 KiB

View File

@@ -11,6 +11,8 @@ This is related to the [💬 Text Generation](./features.md#-text-generation) fe
If there's a text-generation handler agent configured, the bot **may** respond to messages sent in the room.
Some models also support vision and document understanding, so you may be able to mix text, images, and files (PDFs, text documents, etc.) in the same conversation.
See screenshots of:
- 🖼️ [the default Text Generation flow](./screenshots/text-generation.webp) in 1:1 rooms
@@ -64,34 +66,48 @@ The speech-to-text feature triggers automatically by default, but can be adjuste
If all your messages are in the same language, you can improve accuracy & latency by configuring the language (see [🦻 Speech-to-Text / 🔤 Language](./configuration/speech-to-text.md#-language)).
### 🖌️ Image Generation
This is related to the [🖌️ Image Generation](./features.md#️-image-generation) feature.
### Image Generation
This feature is not configurable at the moment. The configuration (size, quality, style) specified at the [🤖 agent](./agents.md) level will be used.
Capabilities depend on the [☁️ provider](./providers.md) and model used.
#### Generating images
Simply send a command like `!bai image A beautiful sunset over the ocean` and the bot will start a threaded conversation and post an image based on your prompt.
#### 🖌️ Creating images
See a [🖼️ Screenshot of the Image Generation feature](./screenshots/image-generation.webp).
Simply send a command like `!bai image create A beautiful sunset over the ocean` and the bot will start a threaded conversation and post an image based on your prompt.
You can then, respond in the same message thread with:
See a [🖼️ Screenshot of the Image Creation feature](./screenshots/image-creation.webp).
You can then respond in the same message thread with:
- more messages, to add more criteria to your prompt.
- a message saying `again`, to generate one more image with the current prompt.
#### Generating stickers
#### 🎨 Editing images
A variation of [generating images](#generating-images) is to generate "sticker images".
Simply send a command like `!bai image edit Turn the following image into an anime-style drawing` and the bot will start a threaded conversation asking for more details.
See a [🖼️ Screenshot of the Sticker Generation feature](./screenshots/sticker-generation.webp).
See a [🖼️ Screenshot of the Image Editing feature (manipulating a single image)](./screenshots/image-editing-single-image.webp) and a [🖼️ Screenshot of the Image Editing feature (manipulating multiple images)](./screenshots/image-editing-multiple-images.webp).
To generate a sticker, send a command like `!bai sticker A huge ramen bowl with lots of chashu and a mountain of beansprouts on top`.
You can then respond in the same message thread with:
The difference from [generating images](#generating-images) is that the bot will:
- more messages, to add more criteria to your prompt.
- one or more images, to provide the images that the bot will operate on.
- a message saying `go`, to start the image generation process.
- a message saying `again`, to prompt the bot to generate one more image edit with the current prompt.
#### 🫵 Creating stickers
A variation of [creating images](#creating-images) is creating "sticker images".
See a [🖼️ Screenshot of the Sticker Creation feature](./screenshots/sticker-generation.webp).
To create a sticker, send a command like `!bai sticker A huge ramen bowl with lots of chashu and a mountain of beansprouts on top`.
The difference from [creating images](#creating-images) is that the bot will:
- generate a smaller-resolution image (currently hardcoded to `256x256`) - smaller/quicker, but still good enough for a sticker
- potentially switch to a different (cheaper or otherwise more suitable) model, if available

View File

@@ -1,16 +1,31 @@
homeserver:
# The canonical homeserver domain name
server_name: synapse.127.0.0.1.nip.io
url: http://synapse.127.0.0.1.nip.io:42020
server_name: __HOMESERVER_SERVER_NAME__
url: __HOMESERVER_URL__
user:
mxid_localpart: baibot
# Authentication: set EITHER password OR access_token + device_id.
#
# Password-based login (traditional homeservers):
password: baibot
# Access token login (for Matrix Authentication Service/OIDC-enabled homeservers):
# Generate a token via: mas-cli manage issue-compatibility-token <username> [device_id]
# access_token: null
# device_id: null
# The name the bot uses as a display name and when it refers to itself.
# Leave empty to use the default (baibot).
name: baibot
# An optional path to an image file to be used as a custom avatar image.
# - null or empty string: use the default avatar
# - "keep": don't touch the avatar, keep whatever is already set
# - any other value: path to a custom avatar image file
avatar: null
encryption:
# An optional passphrase to use for backing up and recovering the bot's encryption keys.
# You can use any string here.
@@ -39,7 +54,7 @@ room:
access:
# Space-separated list of MXID patterns which specify who is an admin.
admin_patterns:
- "@admin:synapse.127.0.0.1.nip.io"
- "@admin:__HOMESERVER_SERVER_NAME__"
persistence:
# This is unset here, because we expect the configuration to come from an environment variable (BAIBOT_PERSISTENCE_DATA_DIR_PATH).
@@ -76,13 +91,18 @@ agents:
# base_url: https://api.openai.com/v1
# api_key: ""
# text_generation:
# model_id: gpt-4o
# model_id: gpt-5.4
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
# temperature: 1.0
# max_response_tokens: 16384
# # Reasoning models need to use `max_completion_tokens` instead of `max_response_tokens`.
# max_completion_tokens: ~
# max_context_tokens: 128000
# # If you're dealing with a non-reasoning model, specify `max_response_tokens` and unset `max_completion_tokens`.
# max_response_tokens: null
# max_completion_tokens: 128000
# max_context_tokens: 400000
# # Built-in tools
# tools:
# web_search: false
# code_interpreter: false
# speech_to_text:
# model_id: whisper-1
# text_to_speech:
@@ -91,10 +111,10 @@ agents:
# speed: 1.0
# response_format: opus
# image_generation:
# model_id: dall-e-3
# style: vivid
# size: 1024x1024
# quality: standard
# model_id: gpt-image-2
# style: null
# size: null
# quality: null
#
# - id: localai
# provider: localai
@@ -146,7 +166,7 @@ initial_global_config:
# Space-separated list of MXID patterns which specify who can use the bot.
# By default, we let anyone on the homeserver use the bot.
user_patterns:
- "@*:synapse.127.0.0.1.nip.io"
- "@*:__HOMESERVER_SERVER_NAME__"
# Controls logging.
#

View File

@@ -0,0 +1,23 @@
services:
continuwuity:
image: forgejo.ellis.link/continuwuation/continuwuity:v0.5.10
user: "${UID}:${GID}"
restart: unless-stopped
cap_drop:
- ALL
read_only: true
environment:
CONDUWUIT_CONFIG: /etc/continuwuity/continuwuity.toml
CONDUWUIT_DATABASE_PATH: /var/lib/continuwuity
ports:
- "${SERVICE_CONTINUWUITY_BIND_PORT_CLIENT_API}:6167"
volumes:
- ../../etc/services/continuwuity/config:/etc/continuwuity:ro
- ./continuwuity/data:/var/lib/continuwuity
tmpfs:
- /tmp:rw,noexec,nosuid,size=500m
networks:
default:
name: ${NETWORK_NAME}
external: true

View File

@@ -0,0 +1,19 @@
[global]
server_name = "continuwuity.127.0.0.1.nip.io"
address = "0.0.0.0"
port = 6167
database_path = "/var/lib/continuwuity"
allow_registration = true
yes_i_am_very_very_sure_i_want_an_open_registration_server_prone_to_abuse = true
new_user_displayname_suffix = ""
max_request_size = 20_000_000
allow_federation = false
trusted_servers = ["matrix.org"]
log = "info,state_res=warn,rocket=off,_=off,sled=off"

View File

@@ -0,0 +1,48 @@
#!/bin/sh
set -eu
if [ $# -ne 3 ]; then
echo "Usage: $0 <env-file> <username> <password>"
exit 1
fi
ENV_FILE="$1"
USERNAME="$2"
PASSWORD="$3"
SERVER="http://$(grep '^SERVICE_CONTINUWUITY_BIND_PORT_CLIENT_API=' "${ENV_FILE}" | cut -d= -f2)"
REGISTER_URL="${SERVER}/_matrix/client/v3/register"
echo "Registering user '${USERNAME}' on ${SERVER}..."
SESSION_RESPONSE=$(curl -s -X POST "${REGISTER_URL}" \
-H 'Content-Type: application/json' \
-d "{\"username\": \"${USERNAME}\", \"password\": \"${PASSWORD}\"}")
SESSION_ID=$(echo "${SESSION_RESPONSE}" | grep -o '"session":"[^"]*"' | head -1 | cut -d'"' -f4)
if [ -z "${SESSION_ID}" ]; then
echo "Error: Could not get session ID. Response: ${SESSION_RESPONSE}"
exit 1
fi
# Determine the required auth flow from the server response.
# The first user requires m.login.registration_token (bootstrap token from logs).
# Subsequent users use m.login.dummy (open registration).
if echo "${SESSION_RESPONSE}" | grep -q 'm.login.registration_token'; then
CONTAINER_ID=$(docker ps -q --filter name=baibot-continuwuity-continuwuity)
REG_TOKEN=$(docker logs "${CONTAINER_ID}" 2>&1 | sed 's/\x1b\[[0-9;]*m//g' | grep 'using the registration token' | grep -oP 'registration token \K[A-Za-z0-9]+' | head -1)
AUTH_BODY="{\"type\": \"m.login.registration_token\", \"token\": \"${REG_TOKEN}\", \"session\": \"${SESSION_ID}\"}"
else
AUTH_BODY="{\"type\": \"m.login.dummy\", \"session\": \"${SESSION_ID}\"}"
fi
RESULT=$(curl -s -X POST "${REGISTER_URL}" \
-H 'Content-Type: application/json' \
-d "{\"username\": \"${USERNAME}\", \"password\": \"${PASSWORD}\", \"auth\": ${AUTH_BODY}}")
if echo "${RESULT}" | grep -q '"user_id"'; then
echo "Successfully registered user: $(echo "${RESULT}" | grep -o '"user_id":"[^"]*"' | cut -d'"' -f4)"
else
echo "Registration failed. Response: ${RESULT}"
exit 1
fi

View File

@@ -0,0 +1,21 @@
services:
element-web:
image: ghcr.io/element-hq/element-web:v1.12.22
user: "${UID}:${GID}"
restart: unless-stopped
environment:
ELEMENT_WEB_PORT: 8080
ports:
- "${SERVICE_ELEMENT_WEB_BIND_PORT_HTTP}:8080"
volumes:
- ./element-web/config.json:/app/config.json:ro
tmpfs:
- /var/cache/nginx:rw,mode=777
- /var/run:rw,mode=777
- /tmp/element-web-config:rw,mode=777
- /etc/nginx/conf.d:rw,mode=777
networks:
default:
name: ${NETWORK_NAME}
external: true

View File

@@ -1,5 +1,5 @@
{
"default_hs_url": "http://synapse.127.0.0.1.nip.io:42020",
"default_hs_url": "__HOMESERVER_CLIENT_URL__",
"default_is_url": "https://vector.im",
"integrations_ui_url": "https://scalar.vector.im/",
"integrations_rest_url": "https://scalar.vector.im/api",

View File

@@ -3,6 +3,8 @@ SERVICE_SYNAPSE_BIND_PORT_FEDERATION_API=127.0.0.1:42028
SERVICE_ELEMENT_WEB_BIND_PORT_HTTP=127.0.0.1:42025
SERVICE_CONTINUWUITY_BIND_PORT_CLIENT_API=127.0.0.1:42030
SERVICE_OLLAMA_BIND_PORT_HTTP=127.0.0.1:42026
# See https://localai.io/basics/container/#all-in-one-images for the list of available images

View File

@@ -1,6 +1,6 @@
services:
ollama:
image: docker.io/ollama/ollama:0.6.3
image: docker.io/ollama/ollama:0.30.10
restart: unless-stopped
ports:
- "${SERVICE_OLLAMA_BIND_PORT_HTTP}:11434"

View File

@@ -1,6 +1,6 @@
services:
postgres:
image: docker.io/postgres:16.8-alpine
image: docker.io/postgres:18.4-alpine
user: ${UID}:${GID}
restart: unless-stopped
environment:
@@ -8,12 +8,13 @@ services:
POSTGRES_PASSWORD: synapse-password
POSTGRES_DB: homeserver
POSTGRES_INITDB_ARGS: --lc-collate C --lc-ctype C --encoding UTF8
PGDATA: /data
volumes:
- ./postgres:/var/lib/postgresql/data
- ./postgres:/data
- /etc/passwd:/etc/passwd:ro
synapse:
image: ghcr.io/element-hq/synapse:v1.127.1
image: ghcr.io/element-hq/synapse:v1.155.0
user: "${UID}:${GID}"
restart: unless-stopped
entrypoint: python
@@ -22,25 +23,9 @@ services:
- "${SERVICE_SYNAPSE_BIND_PORT_CLIENT_API}:8008"
- "${SERVICE_SYNAPSE_BIND_PORT_FEDERATION_API}:8008"
volumes:
- ../../etc/services/core/synapse/config:/config:ro
- ../../etc/services/synapse/config:/config:ro
- ./synapse/media-store:/media-store
element-web:
image: ghcr.io/element-hq/element-web:v1.11.96
user: "${UID}:${GID}"
restart: unless-stopped
environment:
ELEMENT_WEB_PORT: 8080
ports:
- "${SERVICE_ELEMENT_WEB_BIND_PORT_HTTP}:8080"
volumes:
- ../../etc/services/core/element-web/config.json:/app/config.json:ro
tmpfs:
- /var/cache/nginx:rw,mode=777
- /var/run:rw,mode=777
- /tmp/element-web-config:rw,mode=777
- /etc/nginx/conf.d:rw,mode=777
networks:
default:
name: ${NETWORK_NAME}

214
justfile
View File

@@ -2,10 +2,33 @@ project_name := "baibot"
container_image_name := "localhost/baibot"
project_container_network := "baibot"
admin_username := "admin"
admin_password := "admin"
bot_username := "baibot"
bot_password := "baibot"
homeserver := `cat var/homeserver 2>/dev/null || echo continuwuity`
mise_data_dir := env("MISE_DATA_DIR", justfile_directory() / "var/mise")
mise_trusted_config_paths := justfile_directory() / "mise.toml"
# Show help by default
default:
@just --list --justfile {{ justfile() }}
# Selects which homeserver implementation to use (continuwuity or synapse)
homeserver-init value:
#!/bin/sh
mkdir -p {{ justfile_directory() }}/var
echo {{ value }} > {{ justfile_directory() }}/var/homeserver
echo ""
echo "⚠️ If you had already prepared your app configuration (var/app/local/config.yml or var/app/container/config.yml),"
echo " you will need to update it manually or delete it and re-run the prepare step."
echo " You should also delete var/app/local/data and/or var/app/container/data,"
echo " as old application state is not compatible across homeserver implementations."
echo ""
echo "⚠️ If Element Web was already prepared, delete var/services/element-web/ to regenerate its config."
# Builds and runs a development binary
run-locally *extra_args: app-local-prepare
RUST_BACKTRACE=1 \
@@ -65,9 +88,13 @@ docker-compose services_type *extra_args:
-p {{ project_name }}-{{ services_type }} \
{{ extra_args }}
# Runs a docker-compose command against the core services
docker-compose-core *extra_args:
just docker-compose core {{ extra_args }}
# Runs a docker-compose command against the synapse services
docker-compose-synapse *extra_args:
just docker-compose synapse {{ extra_args }}
# Runs a docker-compose command against the element-web services
docker-compose-element-web *extra_args:
just docker-compose element-web {{ extra_args }}
# Runs a docker-compose command against the localai services
docker-compose-localai *extra_args:
@@ -77,17 +104,52 @@ docker-compose-localai *extra_args:
docker-compose-ollama *extra_args:
just docker-compose ollama {{ extra_args }}
# Runs all core dependency components (in the background)
services-start: services-prepare (docker-compose-core "up" "-d")
# Runs a docker-compose command against the continuwuity services
docker-compose-continuwuity *extra_args:
just docker-compose continuwuity {{ extra_args }}
# Stops all core dependency components
services-stop: (docker-compose-core "down")
# Runs the homeserver and Element Web (in the background)
services-start: services-prepare
just -f {{ justfile_directory() }}/justfile {{ homeserver }}-start
just -f {{ justfile_directory() }}/justfile element-web-start
# Tails the logs for all running core services
services-tail-logs: (docker-compose-core "logs" "-f")
# Stops Element Web and the homeserver
services-stop:
just -f {{ justfile_directory() }}/justfile element-web-stop
just -f {{ justfile_directory() }}/justfile {{ homeserver }}-stop
# Prepares the core services for running
services-prepare: _prepare-var-services-env _prepare-var-services-postgres _prepare-var-services-synapse _prepare-container-network
# Tails the logs for the homeserver and Element Web
services-tail-logs:
just -f {{ justfile_directory() }}/justfile {{ homeserver }}-tail-logs
# Prepares the homeserver and Element Web for running
services-prepare:
just -f {{ justfile_directory() }}/justfile {{ homeserver }}-prepare
just -f {{ justfile_directory() }}/justfile element-web-prepare
# Runs Synapse (in the background)
synapse-start: synapse-prepare (docker-compose-synapse "up" "-d")
# Stops Synapse
synapse-stop: (docker-compose-synapse "down")
# Tails the logs for Synapse
synapse-tail-logs: (docker-compose-synapse "logs" "-f")
# Prepares Synapse for running
synapse-prepare: _prepare-var-services-env _prepare-var-services-postgres _prepare-var-services-synapse _prepare-container-network
# Runs Element Web (in the background)
element-web-start: element-web-prepare (docker-compose-element-web "up" "-d")
# Stops Element Web
element-web-stop: (docker-compose-element-web "down")
# Tails the logs for Element Web
element-web-tail-logs: (docker-compose-element-web "logs" "-f")
# Prepares Element Web for running
element-web-prepare: _prepare-var-services-env _prepare-var-services-element-web _prepare-container-network
# Runs LocalAI (in the background)
localai-start: localai-prepare (docker-compose-localai "up" "-d")
@@ -113,6 +175,27 @@ ollama-tail-logs: (docker-compose-ollama "logs" "-f")
# Prepares Ollama for running
ollama-prepare: _prepare-var-services-env _prepare-var-services-ollama _prepare-container-network
# Runs Continuwuity (in the background)
continuwuity-start: continuwuity-prepare (docker-compose-continuwuity "up" "-d")
# Stops Continuwuity
continuwuity-stop: (docker-compose-continuwuity "down")
# Tails the logs for Continuwuity
continuwuity-tail-logs: (docker-compose-continuwuity "logs" "-f")
# Prepares Continuwuity for running
continuwuity-prepare: _prepare-var-services-env _prepare-var-services-continuwuity _prepare-container-network
# Registers a user on Continuwuity via the Matrix Client-Server API
continuwuity-register-user username password:
{{ justfile_directory() }}/etc/services/continuwuity/register-user.sh {{ justfile_directory() }}/var/services/env {{ username }} {{ password }}
# Prepares the Continuwuity user accounts
continuwuity-users-prepare: continuwuity-prepare
just -f {{ justfile_directory() }}/justfile continuwuity-register-user "{{ admin_username }}" "{{ admin_password }}"
just -f {{ justfile_directory() }}/justfile continuwuity-register-user "{{ bot_username }}" "{{ bot_password }}"
# Pulls an Ollama model
ollama-pull-model model_id:
just -f {{ justfile_directory() }}/justfile docker-compose-ollama \
@@ -126,16 +209,20 @@ app-local-prepare: _prepare-var-app-local-config_yml _prepare-var-app-local-data
app-container-prepare: _prepare-var-app-container-config_yml _prepare-var-app-container-data
# Prepares the user accounts
users-prepare: services-prepare
just -f {{ justfile_directory() }}/justfile synapse-register-admin-user "admin" "admin"
just -f {{ justfile_directory() }}/justfile synapse-register-regular-user "baibot" "baibot"
users-prepare:
just -f {{ justfile_directory() }}/justfile {{ homeserver }}-users-prepare
# Prepares the Synapse user accounts
synapse-users-prepare: synapse-prepare
just -f {{ justfile_directory() }}/justfile synapse-register-admin-user "{{ admin_username }}" "{{ admin_password }}"
just -f {{ justfile_directory() }}/justfile synapse-register-regular-user "{{ bot_username }}" "{{ bot_password }}"
# Starts a Postgres CLI (psql)
postgres-cli: services-prepare (docker-compose-core "exec" "postgres" "/bin/sh" "-c" "'PGUSER=synapse PGPASSWORD=synapse-password PGDATABASE=homeserver psql -h postgres'")
postgres-cli: synapse-prepare (docker-compose-synapse "exec" "postgres" "/bin/sh" "-c" "'PGUSER=synapse PGPASSWORD=synapse-password PGDATABASE=homeserver psql -h postgres'")
# Creates an administrator user
synapse-register-admin-user username password: services-prepare
just -f {{ justfile_directory() }}/justfile docker-compose-core \
# Creates an administrator user on Synapse
synapse-register-admin-user username password: synapse-prepare
just -f {{ justfile_directory() }}/justfile docker-compose-synapse \
exec synapse \
register_new_matrix_user \
--admin \
@@ -144,9 +231,9 @@ synapse-register-admin-user username password: services-prepare
-c /config/homeserver.yaml \
http://localhost:8008
# Create a regular user
synapse-register-regular-user username password: services-prepare
just -f {{ justfile_directory() }}/justfile docker-compose-core \
# Creates a regular user on Synapse
synapse-register-regular-user username password: synapse-prepare
just -f {{ justfile_directory() }}/justfile docker-compose-synapse \
exec synapse \
register_new_matrix_user \
--no-admin \
@@ -159,6 +246,44 @@ synapse-register-regular-user username password: services-prepare
clippy *extra_args:
cargo clippy {{ extra_args }}
# Checks that the code compiles without building
check:
cargo check
# Invokes mise with the project-local data directory
mise *args: _ensure_mise_data_directory
#!/bin/sh
export MISE_DATA_DIR="{{ mise_data_dir }}"
export MISE_TRUSTED_CONFIG_PATHS="{{ mise_trusted_config_paths }}"
mise {{ args }}
# Runs prek (pre-commit hooks manager) with the given arguments
prek *args: _ensure_mise_tools_installed
@just --justfile {{ justfile() }} mise exec -- prek {{ args }}
# Runs pre-commit hooks on staged files
prek-run-on-staged *args: _ensure_mise_tools_installed
@just --justfile {{ justfile() }} mise exec -- prek run {{ args }}
# Runs pre-commit hooks on all files
prek-run-on-all *args: _ensure_mise_tools_installed
@just --justfile {{ justfile() }} mise exec -- prek run --all-files {{ args }}
# Installs the git pre-commit hook (runs prek automatically before each commit)
prek-install-git-pre-commit-hook: _ensure_mise_tools_installed
@just --justfile {{ justfile() }} mise exec -- prek install
# Internal - ensures var/mise directory exists
_ensure_mise_data_directory:
#!/bin/sh
if [ ! -d "{{ mise_data_dir }}" ]; then
mkdir -p "{{ mise_data_dir }}"
fi
# Internal - ensures mise tools are installed
_ensure_mise_tools_installed: _ensure_mise_data_directory
@just --justfile {{ justfile() }} mise install --quiet
_prepare-var-services-env:
#!/bin/sh
cd {{ justfile_directory() }};
@@ -188,6 +313,22 @@ _prepare-var-services-synapse:
mkdir -p var/services/synapse/media-store
fi
_prepare-var-services-element-web:
#!/bin/sh
cd {{ justfile_directory() }};
if [ ! -f var/services/element-web/config.json ]; then
mkdir -p var/services/element-web
cp {{ justfile_directory() }}/etc/services/element-web/config.json.dist var/services/element-web/config.json
homeserver="{{ homeserver }}"
if [ "$homeserver" = "continuwuity" ]; then
sed --in-place 's|__HOMESERVER_CLIENT_URL__|http://continuwuity.127.0.0.1.nip.io:42030|g' var/services/element-web/config.json
elif [ "$homeserver" = "synapse" ]; then
sed --in-place 's|__HOMESERVER_CLIENT_URL__|http://synapse.127.0.0.1.nip.io:42020|g' var/services/element-web/config.json
fi
fi
_prepare-var-services-ollama:
#!/bin/sh
cd {{ justfile_directory() }};
@@ -196,6 +337,14 @@ _prepare-var-services-ollama:
mkdir -p var/services/ollama
fi
_prepare-var-services-continuwuity:
#!/bin/sh
cd {{ justfile_directory() }};
if [ ! -f var/services/continuwuity ]; then
mkdir -p var/services/continuwuity/data
fi
_prepare-var-services-localai:
#!/bin/sh
cd {{ justfile_directory() }};
@@ -219,6 +368,15 @@ _prepare-var-app-local-config_yml:
if [ ! -f var/app/local/config.yml ]; then
mkdir -p var/app/local
cp {{ justfile_directory() }}/etc/app/config.yml.dist var/app/local/config.yml
homeserver="{{ homeserver }}"
if [ "$homeserver" = "continuwuity" ]; then
sed --in-place 's/__HOMESERVER_SERVER_NAME__/continuwuity.127.0.0.1.nip.io/g' var/app/local/config.yml
sed --in-place 's|__HOMESERVER_URL__|http://continuwuity.127.0.0.1.nip.io:42030|g' var/app/local/config.yml
elif [ "$homeserver" = "synapse" ]; then
sed --in-place 's/__HOMESERVER_SERVER_NAME__/synapse.127.0.0.1.nip.io/g' var/app/local/config.yml
sed --in-place 's|__HOMESERVER_URL__|http://synapse.127.0.0.1.nip.io:42020|g' var/app/local/config.yml
fi
fi
_prepare-var-app-local-data:
@@ -236,7 +394,18 @@ _prepare-var-app-container-config_yml:
if [ ! -f var/app/container/config.yml ]; then
mkdir -p var/app/container
cp {{ justfile_directory() }}/etc/app/config.yml.dist var/app/container/config.yml
sed --in-place 's/synapse.127.0.0.1.nip.io:42020/synapse:8008/g' var/app/container/config.yml
homeserver="{{ homeserver }}"
if [ "$homeserver" = "continuwuity" ]; then
sed --in-place 's/__HOMESERVER_SERVER_NAME__/continuwuity.127.0.0.1.nip.io/g' var/app/container/config.yml
sed --in-place 's|__HOMESERVER_URL__|http://continuwuity.127.0.0.1.nip.io:42030|g' var/app/container/config.yml
sed --in-place 's/continuwuity.127.0.0.1.nip.io:42030/continuwuity:6167/g' var/app/container/config.yml
elif [ "$homeserver" = "synapse" ]; then
sed --in-place 's/__HOMESERVER_SERVER_NAME__/synapse.127.0.0.1.nip.io/g' var/app/container/config.yml
sed --in-place 's|__HOMESERVER_URL__|http://synapse.127.0.0.1.nip.io:42020|g' var/app/container/config.yml
sed --in-place 's/synapse.127.0.0.1.nip.io:42020/synapse:8008/g' var/app/container/config.yml
fi
sed --in-place 's/127.0.0.1:42026/ollama:11434/g' var/app/container/config.yml
sed --in-place 's/127.0.0.1:42027/localai:8080/g' var/app/container/config.yml
fi
@@ -248,4 +417,3 @@ _prepare-var-app-container-data:
if [ ! -f var/app/container/data ]; then
mkdir -p var/app/container/data
fi

6
mise.toml Normal file
View File

@@ -0,0 +1,6 @@
[tools]
prek = "0.4.5"
[settings]
# Disable automatic trust prompts - we trust this config
yes = true

9
renovate.json Normal file
View File

@@ -0,0 +1,9 @@
{
"$schema": "https://docs.renovatebot.com/renovate-schema.json",
"extends": [
"config:recommended"
],
"labels": [
"dependencies"
]
}

4
rust-toolchain.toml Normal file
View File

@@ -0,0 +1,4 @@
[toolchain]
channel = "1.96.0"
components = ["rustfmt", "clippy"]
profile = "default"

View File

@@ -33,11 +33,11 @@ pub struct AgentDefinition {
)]
pub provider: AgentProvider,
pub config: serde_yaml::Value,
pub config: serde_yaml_ng::Value,
}
impl AgentDefinition {
pub fn new(id: String, provider: AgentProvider, config: serde_yaml::Value) -> Self {
pub fn new(id: String, provider: AgentProvider, config: serde_yaml_ng::Value) -> Self {
Self {
id,
provider,

View File

@@ -15,7 +15,7 @@ pub enum Error {
// Contains the error from the constructor function
ConstructionFailed(anyhow::Error),
// Contains the error from the YAML deserialization function
Yaml(serde_yaml::Error),
Yaml(serde_yaml_ng::Error),
}
pub type Result<T> = std::result::Result<T, Error>;
@@ -69,7 +69,7 @@ pub(super) fn create(
pub fn create_from_provider_and_yaml_value_config(
provider: &AgentProvider,
identifier: &PublicIdentifier,
config: serde_yaml::Value,
config: serde_yaml_ng::Value,
) -> Result<AgentInstance> {
let definition = AgentDefinition::new(identifier.prefixless(), provider.to_owned(), config);
@@ -79,7 +79,7 @@ pub fn create_from_provider_and_yaml_value_config(
fn create_controller_from_provider_and_json_value_config(
agent_id: &str,
provider: &AgentProvider,
config: serde_yaml::Value,
config: serde_yaml_ng::Value,
) -> Result<ControllerType> {
match provider {
AgentProvider::Anthropic => {
@@ -109,46 +109,53 @@ fn create_controller_from_provider_and_json_value_config(
AgentProvider::TogetherAI => {
provider::openai_compat::create_controller_from_yaml_value_config(agent_id, config)
}
AgentProvider::Venice => {
provider::venice::create_controller_from_yaml_value_config(agent_id, config)
}
}
}
pub fn default_config_for_provider(provider: &AgentProvider) -> serde_yaml::Value {
pub fn default_config_for_provider(provider: &AgentProvider) -> serde_yaml_ng::Value {
match provider {
AgentProvider::Anthropic => {
let config = super::provider::anthropic::default_config();
serde_yaml::to_value(config).expect("Failed to serialize config")
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
AgentProvider::Groq => {
let config = super::provider::groq::default_config();
serde_yaml::to_value(config).expect("Failed to serialize config")
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
AgentProvider::LocalAI => {
let config = super::provider::localai::default_config();
serde_yaml::to_value(config).expect("Failed to serialize config")
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
AgentProvider::Mistral => {
let config = super::provider::mistral::default_config();
serde_yaml::to_value(config).expect("Failed to serialize config")
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
AgentProvider::Ollama => {
let config = super::provider::ollama::default_config();
serde_yaml::to_value(config).expect("Failed to serialize config")
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
AgentProvider::OpenAI => {
let config = super::provider::openai::default_config();
serde_yaml::to_value(config).expect("Failed to serialize config")
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
AgentProvider::OpenAICompat => {
let config = super::provider::openai_compat::default_config();
serde_yaml::to_value(config).expect("Failed to serialize config")
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
AgentProvider::OpenRouter => {
let config = super::provider::openrouter::default_config();
serde_yaml::to_value(config).expect("Failed to serialize config")
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
AgentProvider::TogetherAI => {
let config = super::provider::togetherai::default_config();
serde_yaml::to_value(config).expect("Failed to serialize config")
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
AgentProvider::Venice => {
let config = super::provider::venice::default_config();
serde_yaml_ng::to_value(config).expect("Failed to serialize config")
}
}
}

View File

@@ -7,13 +7,15 @@ use anthropic::types::ContentBlock;
use super::super::ControllerTrait;
use crate::agent::AgentPurpose;
use crate::agent::provider::entity::{
ImageGenerationResult, PingResult, TextGenerationParams, TextGenerationResult,
TextToSpeechParams, TextToSpeechResult,
ImageEditResult, ImageGenerationResult, ImageSource, PingResult, TextGenerationParams,
TextGenerationResult, TextToSpeechParams, TextToSpeechResult,
};
use crate::agent::provider::{
ImageEditParams, ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
};
use crate::agent::provider::{ImageGenerationParams, SpeechToTextParams, SpeechToTextResult};
use crate::conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
shorten_messages_list_to_context_size,
MessageContent as LLMMessageContent, shorten_messages_list_to_context_size,
};
use crate::strings;
@@ -69,7 +71,8 @@ impl ControllerTrait for Controller {
let messages = vec![LLMMessage {
author: LLMAuthor::User,
message_text: "Hello!".to_string(),
sender_id: None,
content: LLMMessageContent::Text("Hello!".to_string()),
timestamp: chrono::Utc::now(),
}];
@@ -106,7 +109,8 @@ impl ControllerTrait for Controller {
} else {
Some(LLMMessage {
author: LLMAuthor::Prompt,
message_text: prompt_text,
sender_id: None,
content: LLMMessageContent::Text(prompt_text),
timestamp: chrono::Utc::now(),
})
};
@@ -144,8 +148,10 @@ impl ControllerTrait for Controller {
.temperature_override
.unwrap_or(text_generation_config.temperature);
if let Some(prompt_message) = prompt_message {
request.system = prompt_message.message_text;
if let Some(prompt_message) = prompt_message
&& let LLMMessageContent::Text(text) = &prompt_message.content
{
request.system = text.clone();
}
request.model = text_generation_config.model_id.clone();
@@ -172,15 +178,8 @@ impl ControllerTrait for Controller {
ContentBlock::Text { text } => {
text_parts.push(text);
}
ContentBlock::Image {
source,
media_type,
data: _,
} => {
text_parts.push(format!(
"The model responded with an image of type {}: {}",
media_type, source
));
ContentBlock::Image { .. } => {
text_parts.push("The model responded with an image".to_string());
}
}
}
@@ -213,6 +212,15 @@ impl ControllerTrait for Controller {
Err(anyhow::anyhow!("Image generation not supported"))
}
async fn create_image_edit(
&self,
_prompt: &str,
_images: Vec<ImageSource>,
_params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
Err(anyhow::anyhow!("Image editing is not supported"))
}
async fn text_to_speech(
&self,
_input: &str,

View File

@@ -12,12 +12,12 @@ use super::controller::ControllerType;
pub fn create_controller_from_yaml_value_config(
agent_id: &str,
config: serde_yaml::Value,
config: serde_yaml_ng::Value,
) -> AgentInstantiationResult<ControllerType> {
let config = match &config {
serde_yaml::Value::Mapping(_) => {
serde_yaml_ng::Value::Mapping(_) => {
let config: Config =
serde_yaml::from_value(config).map_err(AgentInstantiationError::Yaml)?;
serde_yaml_ng::from_value(config).map_err(AgentInstantiationError::Yaml)?;
config
.validate()

View File

@@ -1,6 +1,10 @@
use anthropic::types::{ContentBlock, Message, MessagesRequest, MessagesRequestBuilder, Role};
use anthropic::types::{
ContentBlock, ImageSource, Message, MessagesRequest, MessagesRequestBuilder, Role,
};
use crate::conversation::llm::{Author as LLMAuthor, Message as LLMMessage};
use crate::conversation::llm::{
Author as LLMAuthor, Message as LLMMessage, MessageContent as LLMMessageContent,
};
pub(super) fn create_anthropic_message_request(llm_messages: Vec<LLMMessage>) -> MessagesRequest {
let mut messages = vec![];
@@ -14,9 +18,24 @@ pub(super) fn create_anthropic_message_request(llm_messages: Vec<LLMMessage>) ->
}
};
let content = vec![ContentBlock::Text {
text: message.message_text,
}];
let content = match &message.content {
LLMMessageContent::Text(text) => vec![ContentBlock::Text { text: text.clone() }],
LLMMessageContent::Image(image_details) => {
vec![ContentBlock::Image {
source: ImageSource::Base64 {
media_type: image_details.mime.to_string(),
data: crate::utils::base64::base64_encode(&image_details.data),
},
}]
}
LLMMessageContent::File(file_details) => {
tracing::warn!(
"The Anthropic provider's library does not support file/document content. This file message ({}) will be skipped.",
file_details.filename(),
);
continue;
}
};
let message = Message { role, content };

View File

@@ -1,10 +1,10 @@
use crate::{agent::AgentPurpose, conversation::llm::Conversation};
use super::{
ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
ImageEditParams, ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
entity::{
ImageGenerationResult, PingResult, TextGenerationParams, TextGenerationResult,
TextToSpeechParams, TextToSpeechResult,
ImageEditResult, ImageGenerationResult, ImageSource, PingResult, TextGenerationParams,
TextGenerationResult, TextToSpeechParams, TextToSpeechResult,
},
};
@@ -42,6 +42,13 @@ pub trait ControllerTrait {
params: ImageGenerationParams,
) -> impl std::future::Future<Output = anyhow::Result<ImageGenerationResult>> + Send;
fn create_image_edit(
&self,
prompt: &str,
images: Vec<ImageSource>,
params: ImageEditParams,
) -> impl std::future::Future<Output = anyhow::Result<ImageEditResult>> + Send;
fn text_to_speech(
&self,
text: &str,
@@ -54,6 +61,7 @@ pub enum ControllerType {
OpenAI(Box<super::openai::Controller>),
OpenAICompat(Box<super::openai_compat::Controller>),
Anthropic(Box<super::anthropic::Controller>),
Venice(Box<super::venice::Controller>),
}
impl ControllerTrait for ControllerType {
@@ -62,6 +70,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.supports_purpose(purpose),
ControllerType::OpenAICompat(controller) => controller.supports_purpose(purpose),
ControllerType::Anthropic(controller) => controller.supports_purpose(purpose),
ControllerType::Venice(controller) => controller.supports_purpose(purpose),
}
}
@@ -70,6 +79,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.text_generation_model_id(),
ControllerType::OpenAICompat(controller) => controller.text_generation_model_id(),
ControllerType::Anthropic(controller) => controller.text_generation_model_id(),
ControllerType::Venice(controller) => controller.text_generation_model_id(),
}
}
@@ -78,6 +88,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.text_generation_prompt(),
ControllerType::OpenAICompat(controller) => controller.text_generation_prompt(),
ControllerType::Anthropic(controller) => controller.text_generation_prompt(),
ControllerType::Venice(controller) => controller.text_generation_prompt(),
}
}
@@ -86,6 +97,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.text_to_speech_voice(),
ControllerType::OpenAICompat(controller) => controller.text_to_speech_voice(),
ControllerType::Anthropic(controller) => controller.text_to_speech_voice(),
ControllerType::Venice(controller) => controller.text_to_speech_voice(),
}
}
@@ -94,6 +106,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.text_to_speech_speed(),
ControllerType::OpenAICompat(controller) => controller.text_to_speech_speed(),
ControllerType::Anthropic(controller) => controller.text_to_speech_speed(),
ControllerType::Venice(controller) => controller.text_to_speech_speed(),
}
}
@@ -102,6 +115,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.text_generation_temperature(),
ControllerType::OpenAICompat(controller) => controller.text_generation_temperature(),
ControllerType::Anthropic(controller) => controller.text_generation_temperature(),
ControllerType::Venice(controller) => controller.text_generation_temperature(),
}
}
@@ -110,6 +124,7 @@ impl ControllerTrait for ControllerType {
ControllerType::OpenAI(controller) => controller.ping().await,
ControllerType::OpenAICompat(controller) => controller.ping().await,
ControllerType::Anthropic(controller) => controller.ping().await,
ControllerType::Venice(controller) => controller.ping().await,
}
}
@@ -128,6 +143,9 @@ impl ControllerTrait for ControllerType {
ControllerType::Anthropic(controller) => {
controller.generate_text(conversation, params).await
}
ControllerType::Venice(controller) => {
controller.generate_text(conversation, params).await
}
}
}
@@ -147,6 +165,9 @@ impl ControllerTrait for ControllerType {
ControllerType::Anthropic(controller) => {
controller.speech_to_text(mime_type, media, params).await
}
ControllerType::Venice(controller) => {
controller.speech_to_text(mime_type, media, params).await
}
}
}
@@ -163,6 +184,29 @@ impl ControllerTrait for ControllerType {
ControllerType::Anthropic(controller) => {
controller.generate_image(prompt, params).await
}
ControllerType::Venice(controller) => controller.generate_image(prompt, params).await,
}
}
async fn create_image_edit(
&self,
prompt: &str,
images: Vec<ImageSource>,
params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
match &self {
ControllerType::OpenAI(controller) => {
controller.create_image_edit(prompt, images, params).await
}
ControllerType::OpenAICompat(controller) => {
controller.create_image_edit(prompt, images, params).await
}
ControllerType::Anthropic(controller) => {
controller.create_image_edit(prompt, images, params).await
}
ControllerType::Venice(controller) => {
controller.create_image_edit(prompt, images, params).await
}
}
}
@@ -177,6 +221,7 @@ impl ControllerTrait for ControllerType {
controller.text_to_speech(text, params).await
}
ControllerType::Anthropic(controller) => controller.text_to_speech(text, params).await,
ControllerType::Venice(controller) => controller.text_to_speech(text, params).await,
}
}
}

View File

@@ -11,6 +11,7 @@ pub enum AgentProvider {
OpenAICompat,
OpenRouter,
TogetherAI,
Venice,
}
impl AgentProvider {
@@ -25,6 +26,7 @@ impl AgentProvider {
&Self::OpenAICompat,
&Self::OpenRouter,
&Self::TogetherAI,
&Self::Venice,
]
}
@@ -39,6 +41,7 @@ impl AgentProvider {
Self::OpenAICompat => "openai-compatible",
Self::OpenRouter => "openrouter",
Self::TogetherAI => "together-ai",
Self::Venice => "venice",
}
}
@@ -53,6 +56,7 @@ impl AgentProvider {
"openai-compatible" => Ok(Self::OpenAICompat),
"openrouter" => Ok(Self::OpenRouter),
"together-ai" => Ok(Self::TogetherAI),
"venice" => Ok(Self::Venice),
_ => Err("Unexpected string value"),
}
}
@@ -68,6 +72,8 @@ impl AgentProvider {
sign_up_url: Some("https://console.anthropic.com/"),
models_list_url: Some("https://docs.anthropic.com/en/docs/about-claude/models"),
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: true,
text_generation_supports_tools: false,
},
Self::Groq => AgentProviderInfo {
id: Self::Groq.to_static_str(),
@@ -78,11 +84,13 @@ impl AgentProvider {
sign_up_url: Some("https://console.groq.com/login"),
models_list_url: Some("https://console.groq.com/docs/models"),
supported_purposes: vec![AgentPurpose::TextGeneration, AgentPurpose::SpeechToText],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::LocalAI => AgentProviderInfo {
id: Self::LocalAI.to_static_str(),
name: "LocalAI",
description: "LocalAI is the free, Open Source OpenAI alternative. LocalAI act as a drop-in replacement REST API that’s compatible with OpenAI API specifications for local inferencing. It allows you to run LLMs, generate images, audio (and not only) locally or on-prem with consumer grade hardware, supporting multiple model families and architectures.",
description: "LocalAI is the free, Open Source OpenAI alternative. LocalAI act as a drop-in replacement REST API that's compatible with OpenAI API specifications for local inferencing. It allows you to run LLMs, generate images, audio (and not only) locally or on-prem with consumer grade hardware, supporting multiple model families and architectures.",
homepage_url: Some("https://localai.io/"),
wiki_url: None,
sign_up_url: None,
@@ -92,6 +100,8 @@ impl AgentProvider {
AgentPurpose::TextToSpeech,
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::Mistral => AgentProviderInfo {
id: Self::Mistral.to_static_str(),
@@ -102,6 +112,8 @@ impl AgentProvider {
sign_up_url: Some("https://auth.mistral.ai/ui/registration"),
models_list_url: Some("https://docs.mistral.ai/getting-started/models/"),
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::Ollama => AgentProviderInfo {
id: Self::Ollama.to_static_str(),
@@ -112,6 +124,8 @@ impl AgentProvider {
sign_up_url: None,
models_list_url: Some("https://ollama.com/library"),
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::OpenAI => AgentProviderInfo {
id: Self::OpenAI.to_static_str(),
@@ -127,6 +141,8 @@ impl AgentProvider {
AgentPurpose::TextToSpeech,
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: true,
text_generation_supports_tools: true,
},
Self::OpenAICompat => AgentProviderInfo {
id: Self::OpenAICompat.to_static_str(),
@@ -142,6 +158,8 @@ impl AgentProvider {
AgentPurpose::TextToSpeech,
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::OpenRouter => AgentProviderInfo {
id: Self::OpenRouter.to_static_str(),
@@ -152,6 +170,8 @@ impl AgentProvider {
sign_up_url: Some("https://openrouter.ai/"),
models_list_url: Some("https://openrouter.ai/models"),
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::TogetherAI => AgentProviderInfo {
id: Self::TogetherAI.to_static_str(),
@@ -162,6 +182,27 @@ impl AgentProvider {
sign_up_url: Some("https://api.together.ai/signup"),
models_list_url: Some("https://api.together.xyz/models"),
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::Venice => AgentProviderInfo {
id: Self::Venice.to_static_str(),
name: "Venice",
description: "Venice AI runs inference on Venice-controlled GPUs or zero-data-retention partner infrastructure and stores no prompts or responses. It serves frontier proprietary and open-source models with text-generation (including vision), speech-to-text, text-to-speech, native image generation and editing, and native web search.",
homepage_url: Some("https://venice.ai"),
wiki_url: None,
sign_up_url: Some("https://venice.ai"),
models_list_url: Some("https://api.venice.ai/api/v1/models"),
supported_purposes: vec![
AgentPurpose::ImageGeneration,
AgentPurpose::TextGeneration,
AgentPurpose::TextToSpeech,
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: true,
// Venice does native web search via `venice_parameters`, NOT baibot's built-in
// tools mechanism (the OpenAI web_search/code_interpreter block), so this is false.
text_generation_supports_tools: false,
},
}
}
@@ -182,4 +223,6 @@ pub struct AgentProviderInfo {
pub sign_up_url: Option<&'static str>,
pub models_list_url: Option<&'static str>,
pub supported_purposes: Vec<AgentPurpose>,
pub text_generation_supports_vision: bool,
pub text_generation_supports_tools: bool,
}

View File

@@ -0,0 +1,63 @@
use mxlink::mime;
#[derive(Default)]
pub struct ImageGenerationParams {
pub smallest_size_possible: bool,
pub cheaper_model_switching_allowed: bool,
pub cheaper_quality_switching_allowed: bool,
}
impl ImageGenerationParams {
pub fn with_smallest_size_possible(mut self, value: bool) -> Self {
self.smallest_size_possible = value;
self
}
pub fn with_cheaper_model_switching_allowed(mut self, value: bool) -> Self {
self.cheaper_model_switching_allowed = value;
self
}
pub fn with_cheaper_quality_switching_allowed(mut self, value: bool) -> Self {
self.cheaper_quality_switching_allowed = value;
self
}
}
pub struct ImageGenerationResult {
pub bytes: Vec<u8>,
pub mime_type: mime::Mime,
pub revised_prompt: Option<String>,
}
#[derive(Default)]
pub struct ImageEditParams {}
pub struct ImageEditResult {
pub bytes: Vec<u8>,
pub mime_type: mime::Mime,
}
pub struct ImageSource {
pub filename: String,
pub bytes: Vec<u8>,
pub mime_type: mime::Mime,
}
impl ImageSource {
pub fn new(filename: String, bytes: Vec<u8>, mime_type: mime::Mime) -> Self {
Self {
filename,
bytes,
mime_type,
}
}
}
impl From<ImageSource> for async_openai::types::images::ImageInput {
fn from(value: ImageSource) -> Self {
async_openai::types::images::ImageInput::from_vec_u8(value.filename, value.bytes)
}
}

View File

@@ -1,31 +0,0 @@
#[derive(Default)]
pub struct ImageGenerationParams {
pub size_override: Option<String>,
pub cheaper_model_switching_allowed: bool,
pub cheaper_quality_switching_allowed: bool,
}
impl ImageGenerationParams {
pub fn with_size_override(mut self, value: Option<String>) -> Self {
self.size_override = value;
self
}
pub fn with_cheaper_model_switching_allowed(mut self, value: bool) -> Self {
self.cheaper_model_switching_allowed = value;
self
}
pub fn with_cheaper_quality_switching_allowed(mut self, value: bool) -> Self {
self.cheaper_quality_switching_allowed = value;
self
}
}
pub struct ImageGenerationResult {
pub bytes: Vec<u8>,
pub mime_type: mxlink::mime::Mime,
pub revised_prompt: Option<String>,
}

View File

@@ -1,12 +1,14 @@
mod agent_provider;
mod image_generation;
mod image;
mod ping;
mod speech_to_text;
mod text_generation;
mod text_to_speech;
pub use agent_provider::{AgentProvider, AgentProviderInfo};
pub use image_generation::{ImageGenerationParams, ImageGenerationResult};
pub use image::{
ImageEditParams, ImageEditResult, ImageGenerationParams, ImageGenerationResult, ImageSource,
};
pub use ping::PingResult;
pub use speech_to_text::{SpeechToTextParams, SpeechToTextResult};
pub use text_generation::{

View File

@@ -1,6 +1,7 @@
use chrono::{DateTime, Utc};
use std::collections::HashMap;
#[derive(Clone)]
pub struct TextGenerationPromptVariables {
map: HashMap<String, String>,
}

View File

@@ -10,6 +10,7 @@ pub mod openai;
pub mod openai_compat;
pub(super) mod openrouter;
pub(super) mod togetherai;
pub mod venice;
fn default_temperature() -> f32 {
1.0
@@ -20,6 +21,7 @@ pub use controller::{ControllerTrait, ControllerType};
pub use config::ConfigTrait;
pub use entity::{
AgentProvider, AgentProviderInfo, ImageGenerationParams, PingResult, SpeechToTextParams,
SpeechToTextResult, TextGenerationParams, TextGenerationPromptVariables, TextToSpeechParams,
AgentProvider, AgentProviderInfo, ImageEditParams, ImageGenerationParams, ImageSource,
PingResult, SpeechToTextParams, SpeechToTextResult, TextGenerationParams,
TextGenerationPromptVariables, TextToSpeechParams,
};

View File

@@ -1,5 +1,6 @@
use serde::{Deserialize, Serialize};
use super::OPENAI_IMAGE_MODEL_GPT_IMAGE_2;
use crate::agent::{default_prompt, provider::ConfigTrait};
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -63,6 +64,9 @@ pub struct TextGenerationConfig {
#[serde(default)]
pub max_context_tokens: u32,
#[serde(default)]
pub tools: ToolsConfig,
}
impl Default for TextGenerationConfig {
@@ -71,15 +75,25 @@ impl Default for TextGenerationConfig {
model_id: default_text_model_id(),
prompt: Some(default_prompt().to_owned()),
temperature: super::super::default_temperature(),
max_response_tokens: Some(16_384),
max_completion_tokens: None,
max_context_tokens: 128_000,
max_response_tokens: None,
max_completion_tokens: Some(128_000),
max_context_tokens: 400_000,
tools: ToolsConfig::default(),
}
}
}
fn default_text_model_id() -> String {
"gpt-4o".to_owned()
"gpt-5.4".to_owned()
}
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
pub struct ToolsConfig {
#[serde(default)]
pub web_search: bool,
#[serde(default)]
pub code_interpreter: bool,
}
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -103,16 +117,16 @@ fn default_speech_to_text_model_id() -> String {
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct TextToSpeechConfig {
#[serde(default = "default_text_to_speech_model_id")]
pub model_id: async_openai::types::SpeechModel,
pub model_id: async_openai::types::audio::SpeechModel,
#[serde(default = "default_text_to_speech_voice")]
pub voice: async_openai::types::Voice,
pub voice: async_openai::types::audio::Voice,
#[serde(default = "default_text_to_speech_speed")]
pub speed: f32,
#[serde(default = "default_text_to_speech_response_format")]
pub response_format: async_openai::types::SpeechResponseFormat,
pub response_format: async_openai::types::audio::SpeechResponseFormat,
}
impl Default for TextToSpeechConfig {
@@ -126,22 +140,22 @@ impl Default for TextToSpeechConfig {
}
}
fn default_text_to_speech_model_id() -> async_openai::types::SpeechModel {
async_openai::types::SpeechModel::Tts1Hd
fn default_text_to_speech_model_id() -> async_openai::types::audio::SpeechModel {
async_openai::types::audio::SpeechModel::Tts1Hd
}
fn default_text_to_speech_voice() -> async_openai::types::Voice {
async_openai::types::Voice::Onyx
fn default_text_to_speech_voice() -> async_openai::types::audio::Voice {
async_openai::types::audio::Voice::Onyx
}
fn default_text_to_speech_speed() -> f32 {
1.0
}
fn default_text_to_speech_response_format() -> async_openai::types::SpeechResponseFormat {
fn default_text_to_speech_response_format() -> async_openai::types::audio::SpeechResponseFormat {
// The API defaults to mp3, but we prefer Opus because it's smaller.
// Our clients should all have support for it.
async_openai::types::SpeechResponseFormat::Opus
async_openai::types::audio::SpeechResponseFormat::Opus
}
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -149,19 +163,19 @@ pub struct ImageGenerationConfig {
pub model_id: String,
#[serde(default = "default_image_style")]
pub style: async_openai::types::ImageStyle,
pub style: Option<async_openai::types::images::ImageStyle>,
#[serde(default = "default_image_size")]
pub size: async_openai::types::ImageSize,
pub size: Option<async_openai::types::images::ImageSize>,
#[serde(default = "default_image_quality")]
pub quality: async_openai::types::ImageQuality,
pub quality: Option<async_openai::types::images::ImageQuality>,
}
impl Default for ImageGenerationConfig {
fn default() -> Self {
Self {
model_id: "dall-e-3".to_owned(),
model_id: OPENAI_IMAGE_MODEL_GPT_IMAGE_2.to_owned(),
style: default_image_style(),
size: default_image_size(),
quality: default_image_quality(),
@@ -172,23 +186,29 @@ impl Default for ImageGenerationConfig {
impl ImageGenerationConfig {
pub fn model_id_as_openai_image_model(
&self,
) -> Result<async_openai::types::ImageModel, String> {
) -> Result<async_openai::types::images::ImageModel, String> {
match self.model_id.as_str() {
"dall-e-2" => Ok(async_openai::types::ImageModel::DallE2),
"dall-e-3" => Ok(async_openai::types::ImageModel::DallE3),
other => Ok(async_openai::types::ImageModel::Other(other.to_owned())),
"dall-e-2" => Ok(async_openai::types::images::ImageModel::DallE2),
"dall-e-3" => Ok(async_openai::types::images::ImageModel::DallE3),
"gpt-image-1" => Ok(async_openai::types::images::ImageModel::GptImage1),
"gpt-image-1.5" => Ok(async_openai::types::images::ImageModel::GptImage1dot5),
"gpt-image-1-mini" => Ok(async_openai::types::images::ImageModel::GptImage1Mini),
"gpt-image-2" => Ok(async_openai::types::images::ImageModel::GptImage2),
other => Ok(async_openai::types::images::ImageModel::Other(
other.to_owned(),
)),
}
}
}
fn default_image_style() -> async_openai::types::ImageStyle {
async_openai::types::ImageStyle::Vivid
fn default_image_style() -> Option<async_openai::types::images::ImageStyle> {
None
}
fn default_image_size() -> async_openai::types::ImageSize {
async_openai::types::ImageSize::S1024x1024
fn default_image_size() -> Option<async_openai::types::images::ImageSize> {
None
}
fn default_image_quality() -> async_openai::types::ImageQuality {
async_openai::types::ImageQuality::Standard
fn default_image_quality() -> Option<async_openai::types::images::ImageQuality> {
None
}

View File

@@ -4,34 +4,39 @@ use async_openai::{
Client as OpenAIClient,
config::OpenAIConfig,
types::{
ChatCompletionRequestMessage, CreateChatCompletionRequestArgs, CreateImageRequestArgs,
CreateSpeechRequestArgs, CreateTranscriptionRequestArgs,
audio::{AudioInput, CreateSpeechRequestArgs, CreateTranscriptionRequestArgs},
images::{
CreateImageEditRequestArgs, CreateImageRequestArgs, Image, ImageInput, ImageModel,
ImageResponseFormat,
},
responses::{
CodeInterpreterContainerAuto, CodeInterpreterTool, CodeInterpreterToolContainer,
CreateResponseArgs, OutputItem, OutputMessageContent, Tool, WebSearchTool,
},
},
};
use super::super::ControllerTrait;
use crate::{
agent::{
AgentPurpose,
provider::{
entity::{ImageGenerationResult, PingResult, TextToSpeechParams, TextToSpeechResult},
openai::utils::convert_string_to_enum,
},
},
strings,
};
use crate::{
agent::{
provider::{
ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
entity::{TextGenerationParams, TextGenerationResult},
},
utils::base64_decode,
agent::provider::{
ImageEditParams, ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
entity::{TextGenerationParams, TextGenerationResult},
},
conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
shorten_messages_list_to_context_size,
MessageContent as LLMMessageContent, shorten_messages_list_to_context_size,
},
utils::base64::base64_decode,
};
use crate::{
agent::{
AgentPurpose,
provider::entity::{
ImageEditResult, ImageGenerationResult, ImageSource, PingResult, TextToSpeechParams,
TextToSpeechResult,
},
},
strings,
};
use super::config::Config;
@@ -62,7 +67,8 @@ impl ControllerTrait for Controller {
let messages = vec![LLMMessage {
author: LLMAuthor::User,
message_text: "Hello!".to_string(),
sender_id: None,
content: LLMMessageContent::Text("Hello!".to_string()),
timestamp: chrono::Utc::now(),
}];
@@ -99,7 +105,8 @@ impl ControllerTrait for Controller {
} else {
Some(LLMMessage {
author: LLMAuthor::Prompt,
message_text: prompt_text,
sender_id: None,
content: LLMMessageContent::Text(prompt_text),
timestamp: chrono::Utc::now(),
})
};
@@ -124,28 +131,45 @@ impl ControllerTrait for Controller {
conversation_messages.insert(0, prompt_message);
}
let openai_conversation_messages: Vec<ChatCompletionRequestMessage> =
super::utils::convert_llm_messages_to_openai_messages(conversation_messages);
let input =
super::utils::convert_llm_messages_to_openai_response_input(conversation_messages);
let messages_count = openai_conversation_messages.len();
let messages_count = match &input {
async_openai::types::responses::InputParam::Items(items) => items.len(),
_ => 1,
};
let temperature = params
.temperature_override
.unwrap_or(text_generation_config.temperature);
let mut request_builder = CreateChatCompletionRequestArgs::default();
let mut request_builder = CreateResponseArgs::default();
request_builder
.model(&text_generation_config.model_id)
.temperature(temperature)
.messages(openai_conversation_messages);
.input(input);
if let Some(max_response_tokens) = text_generation_config.max_response_tokens {
request_builder.max_tokens(max_response_tokens);
let mut tools = Vec::new();
if text_generation_config.tools.web_search {
tools.push(Tool::WebSearch(WebSearchTool::default()));
}
if text_generation_config.tools.code_interpreter {
tools.push(Tool::CodeInterpreter(CodeInterpreterTool {
container: CodeInterpreterToolContainer::Auto(
CodeInterpreterContainerAuto::default(),
),
}));
}
if let Some(max_completion_tokens) = text_generation_config.max_completion_tokens {
request_builder.max_completion_tokens(max_completion_tokens);
if !tools.is_empty() {
request_builder.tools(tools);
}
if let Some(max_response_tokens) = text_generation_config.max_response_tokens {
request_builder.max_output_tokens(max_response_tokens);
} else if let Some(max_completion_tokens) = text_generation_config.max_completion_tokens {
request_builder.max_output_tokens(max_completion_tokens);
}
let request = request_builder.build()?;
@@ -155,33 +179,28 @@ impl ControllerTrait for Controller {
model = format!("{:?}", request.model),
?messages_count,
request = request_as_json,
"Sending OpenAI chat completion API request"
"Sending OpenAI response API request"
);
}
let response = self.client.chat().create(request).await?;
let response = self.client.responses().create(request).await?;
tracing::trace!(
?response,
"Got response from the OpenAI chat completion API"
);
tracing::trace!(?response, "Got response from the OpenAI response API");
// We only request 1 result, so there should only be 1 choice.
if let Some(choice) = response.choices.into_iter().next() {
match choice.message.content {
Some(text) => {
return Ok(TextGenerationResult { text });
}
None => {
return Err(anyhow::anyhow!(
"No content was found in the response choice from the OpenAI chat completion API"
));
for item in response.output {
if let OutputItem::Message(message) = item {
for content in message.content {
if let OutputMessageContent::OutputText(text_content) = content {
return Ok(TextGenerationResult {
text: text_content.text,
});
}
}
}
}
Err(anyhow::anyhow!(
"No response messages choices were returned from the OpenAI chat completion API"
"No response messages choices were returned from the OpenAI response API"
))
}
@@ -205,12 +224,7 @@ impl ControllerTrait for Controller {
let request = CreateTranscriptionRequestArgs::default()
.model(&speech_to_text_config.model_id)
.file(async_openai::types::AudioInput {
source: async_openai::types::InputSource::VecU8 {
filename,
vec: media,
},
})
.file(AudioInput::from_vec_u8(filename, media))
.language(language.clone())
.build()?;
@@ -220,7 +234,7 @@ impl ControllerTrait for Controller {
"Sending OpenAI speech-to-text API request"
);
let response = self.client.audio().transcribe(request).await?;
let response = self.client.audio().transcription().create(request).await?;
tracing::trace!(
?response,
@@ -252,11 +266,13 @@ impl ControllerTrait for Controller {
let model = if params.cheaper_model_switching_allowed {
// Switch to a cheaper model
match original_model {
async_openai::types::ImageModel::DallE2 => async_openai::types::ImageModel::DallE2,
async_openai::types::ImageModel::DallE3 => async_openai::types::ImageModel::DallE2,
async_openai::types::ImageModel::Other(_) => {
async_openai::types::ImageModel::DallE2
}
ImageModel::DallE2 => ImageModel::DallE2,
ImageModel::DallE3 => ImageModel::DallE2,
ImageModel::GptImage1 => ImageModel::GptImage1Mini,
ImageModel::GptImage1dot5 => ImageModel::GptImage1Mini,
ImageModel::GptImage1Mini => ImageModel::GptImage1Mini,
ImageModel::GptImage2 => ImageModel::GptImage1Mini,
ImageModel::Other(_) => ImageModel::DallE2,
}
} else {
original_model
@@ -265,33 +281,72 @@ impl ControllerTrait for Controller {
let quality = if params.cheaper_quality_switching_allowed {
// Switch to a cheaper quality
match &image_generation_config.quality {
async_openai::types::ImageQuality::Standard => {
async_openai::types::ImageQuality::Standard
}
async_openai::types::ImageQuality::HD => {
async_openai::types::ImageQuality::Standard
}
Some(quality) => match quality {
async_openai::types::images::ImageQuality::Standard => {
Some(async_openai::types::images::ImageQuality::Standard)
}
async_openai::types::images::ImageQuality::HD => {
Some(async_openai::types::images::ImageQuality::Standard)
}
// New quality levels - keep as-is or downgrade to Standard
async_openai::types::images::ImageQuality::High => {
Some(async_openai::types::images::ImageQuality::Standard)
}
async_openai::types::images::ImageQuality::Medium => {
Some(async_openai::types::images::ImageQuality::Medium)
}
async_openai::types::images::ImageQuality::Low => {
Some(async_openai::types::images::ImageQuality::Low)
}
async_openai::types::images::ImageQuality::Auto => {
Some(async_openai::types::images::ImageQuality::Auto)
}
},
None => None,
}
} else {
image_generation_config.quality.clone()
};
let size = params
.size_override
.map(|s| {
convert_string_to_enum::<async_openai::types::ImageSize>(&s)
.unwrap_or(image_generation_config.size)
})
.unwrap_or(image_generation_config.size);
let size = if params.smallest_size_possible {
Some(get_sticker_size(&model))
} else {
image_generation_config.size.clone()
};
let request = CreateImageRequestArgs::default()
.model(model)
.prompt(prompt.to_owned())
.response_format(async_openai::types::ImageResponseFormat::B64Json)
.size(size)
.style(image_generation_config.style.clone())
.quality(quality)
.build()?;
let response_format = match model.clone() {
ImageModel::DallE2 => Some(ImageResponseFormat::B64Json),
ImageModel::DallE3 => Some(ImageResponseFormat::B64Json),
// gpt-image-1 only outputs base64 and we don't need to specify the response format.
// In fact, specifying the response format results in an error.
ImageModel::GptImage1 => None,
ImageModel::GptImage1Mini => None,
ImageModel::GptImage1dot5 => None,
ImageModel::GptImage2 => None,
ImageModel::Other(_) => Some(ImageResponseFormat::B64Json),
};
let mut request_builder = CreateImageRequestArgs::default();
request_builder.model(model).prompt(prompt.to_owned());
if let Some(response_format) = response_format {
request_builder.response_format(response_format);
}
if let Some(style) = &image_generation_config.style {
request_builder.style(style.clone());
}
if let Some(quality) = quality {
request_builder.quality(quality.clone());
}
if let Some(size) = size {
request_builder.size(size);
}
let request = request_builder.build()?;
tracing::trace!(
?prompt,
@@ -302,15 +357,15 @@ impl ControllerTrait for Controller {
"Sending OpenAI image generation API request"
);
let response = self.client.images().create(request).await?;
let response = self.client.images().generate(request).await?;
if let Some(image) = response.data.into_iter().next() {
match image.deref() {
async_openai::types::Image::B64Json {
Image::B64Json {
b64_json,
revised_prompt,
} => {
let bytes = base64_decode(b64_json)?;
let bytes = base64_decode(b64_json.as_ref())?;
return Ok(ImageGenerationResult {
bytes,
@@ -329,6 +384,109 @@ impl ControllerTrait for Controller {
))
}
async fn create_image_edit(
&self,
prompt: &str,
images: Vec<ImageSource>,
_params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
let Some(image_generation_config) = &self.config.image_generation else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::ImageGeneration
),
));
};
if images.is_empty() {
return Err(anyhow::anyhow!("No image sources provided"));
}
let mut image_inputs: Vec<ImageInput> = Vec::new();
for image in images {
image_inputs.push(image.into());
}
let dalle2_size = match image_generation_config.size {
Some(async_openai::types::images::ImageSize::S256x256) => {
Some(async_openai::types::images::ImageSize::S256x256)
}
Some(async_openai::types::images::ImageSize::S512x512) => {
Some(async_openai::types::images::ImageSize::S512x512)
}
Some(async_openai::types::images::ImageSize::S1024x1024) => {
Some(async_openai::types::images::ImageSize::S1024x1024)
}
_ => None,
};
let model = image_generation_config
.model_id_as_openai_image_model()
.map_err(|err| anyhow::anyhow!(err))?;
let response_format = match model.clone() {
ImageModel::DallE2 => Some(ImageResponseFormat::B64Json),
ImageModel::DallE3 => Some(ImageResponseFormat::B64Json),
// gpt-image-1 only outputs base64 and we don't need to specify the response format.
// In fact, specifying the response format results in an error.
ImageModel::GptImage1 => None,
ImageModel::GptImage1Mini => None,
ImageModel::GptImage1dot5 => None,
ImageModel::GptImage2 => None,
ImageModel::Other(_) => Some(ImageResponseFormat::B64Json),
};
let mut request_builder = CreateImageEditRequestArgs::default();
request_builder
.image(image_inputs)
.prompt(prompt.to_owned())
.model(model);
if let Some(size) = dalle2_size {
request_builder.size(size);
}
if let Some(response_format) = response_format {
request_builder.response_format(response_format);
}
let request = request_builder
.build()
.map_err(|e| anyhow::anyhow!("Failed to build CreateImageEditRequest: {}", e))?;
tracing::trace!(
model = format!("{:?}", request.model),
size = format!("{:?}", request.size),
response_format = format!("{:?}", request.response_format),
"Sending OpenAI image edit API request"
);
let response = self.client.images().edit(request).await?;
if let Some(image_data) = response.data.into_iter().next() {
match image_data.deref() {
Image::B64Json { b64_json, .. } => {
let bytes = base64_decode(b64_json.as_ref())?;
return Ok(ImageEditResult {
bytes,
mime_type: mxlink::mime::IMAGE_PNG,
});
}
Image::Url { url, .. } => {
tracing::warn!(?url, "Received URL instead of B64Json for image edit");
return Err(anyhow::anyhow!(
"Unexpected image type (URL) when B64Json was requested"
));
}
}
}
Err(anyhow::anyhow!(
"The OpenAI image edit API returned no images"
))
}
async fn text_to_speech(
&self,
input: &str,
@@ -346,7 +504,7 @@ impl ControllerTrait for Controller {
let voice = if let Some(voice_string) = params.voice_override {
// This is a hacky way to construct a Voice enum from the string we have.
let voice: serde_json::Result<async_openai::types::Voice> =
let voice: serde_json::Result<async_openai::types::audio::Voice> =
serde_json::from_str(&format!("\"{}\"", voice_string));
match voice {
Ok(voice) => voice,
@@ -386,7 +544,7 @@ impl ControllerTrait for Controller {
"Sending OpenAI text-to-speech API request"
);
let result = self.client.audio().speech(request).await?;
let result = self.client.audio().speech().create(request).await?;
Ok(TextToSpeechResult {
bytes: result.bytes.into(),
@@ -445,15 +603,15 @@ impl ControllerTrait for Controller {
}
fn response_format_to_mime_type(
response_format: &async_openai::types::SpeechResponseFormat,
response_format: &async_openai::types::audio::SpeechResponseFormat,
) -> Option<mxlink::mime::Mime> {
let content_type = match response_format {
async_openai::types::SpeechResponseFormat::Mp3 => "audio/mp3".to_owned(),
async_openai::types::SpeechResponseFormat::Wav => "audio/wav".to_owned(),
async_openai::types::SpeechResponseFormat::Opus => "audio/ogg".to_owned(),
async_openai::types::SpeechResponseFormat::Aac => "audio/aac".to_owned(),
async_openai::types::SpeechResponseFormat::Flac => "audio/flac".to_owned(),
async_openai::types::SpeechResponseFormat::Pcm => "audio/L8".to_owned(),
async_openai::types::audio::SpeechResponseFormat::Mp3 => "audio/mp3".to_owned(),
async_openai::types::audio::SpeechResponseFormat::Wav => "audio/wav".to_owned(),
async_openai::types::audio::SpeechResponseFormat::Opus => "audio/ogg".to_owned(),
async_openai::types::audio::SpeechResponseFormat::Aac => "audio/aac".to_owned(),
async_openai::types::audio::SpeechResponseFormat::Flac => "audio/flac".to_owned(),
async_openai::types::audio::SpeechResponseFormat::Pcm => "audio/L8".to_owned(),
};
match content_type.parse() {
@@ -481,3 +639,18 @@ fn audio_mime_type_to_file_name(mime_type: &mxlink::mime::Mime) -> Option<String
Some(format!("audio.{}", file_extension))
}
/// Returns the smallest supported size for stickers based on what the image model supports.
fn get_sticker_size(model: &ImageModel) -> async_openai::types::images::ImageSize {
use async_openai::types::images::ImageSize;
match model {
ImageModel::DallE2 => ImageSize::S256x256,
ImageModel::DallE3 => ImageSize::S1024x1024,
ImageModel::GptImage1 => ImageSize::S1024x1024,
ImageModel::GptImage1Mini => ImageSize::S1024x1024,
ImageModel::GptImage1dot5 => ImageSize::S1024x1024,
ImageModel::GptImage2 => ImageSize::S1024x1024,
ImageModel::Other(_) => ImageSize::S1024x1024,
}
}

View File

@@ -16,14 +16,16 @@ use super::super::AgentInstantiationResult;
use super::ConfigTrait;
use super::controller::ControllerType;
pub const OPENAI_IMAGE_MODEL_GPT_IMAGE_2: &str = "gpt-image-2";
pub fn create_controller_from_yaml_value_config(
agent_id: &str,
config: serde_yaml::Value,
config: serde_yaml_ng::Value,
) -> AgentInstantiationResult<ControllerType> {
let config = match &config {
serde_yaml::Value::Mapping(_) => {
serde_yaml_ng::Value::Mapping(_) => {
let config: Config =
serde_yaml::from_value(config).map_err(AgentInstantiationError::Yaml)?;
serde_yaml_ng::from_value(config).map_err(AgentInstantiationError::Yaml)?;
config
.validate()

View File

@@ -1,55 +1,64 @@
use async_openai::types::{
ChatCompletionRequestAssistantMessageArgs, ChatCompletionRequestMessage,
ChatCompletionRequestSystemMessageArgs, ChatCompletionRequestUserMessageArgs,
use async_openai::types::responses::{
EasyInputContent, EasyInputMessage, ImageDetail, InputContent, InputFileArgs,
InputImageContent, InputItem, InputParam, MessageType, Role,
};
use crate::conversation::llm::{Author as LLMAuthor, Message as LLMMessage};
use crate::conversation::llm::{
Author as LLMAuthor, Message as LLMMessage, MessageContent as LLMMessageContent,
};
use crate::utils::base64::base64_encode;
pub fn convert_llm_messages_to_openai_messages(
pub fn convert_llm_messages_to_openai_response_input(
conversation_messages: Vec<LLMMessage>,
) -> Vec<ChatCompletionRequestMessage> {
let mut openai_conversation_messages: Vec<ChatCompletionRequestMessage> =
Vec::with_capacity(conversation_messages.len());
) -> InputParam {
let mut items = Vec::with_capacity(conversation_messages.len());
for message in conversation_messages {
openai_conversation_messages.push(convert_llm_message_to_openai_message(message));
let role = match message.author {
LLMAuthor::Prompt => Role::System,
LLMAuthor::Assistant => Role::Assistant,
LLMAuthor::User => Role::User,
};
let content = match message.content {
LLMMessageContent::Text(text) => EasyInputContent::Text(text),
LLMMessageContent::Image(image_details) => {
let image_url = format!(
"data:{};base64,{}",
image_details.mime,
base64_encode(&image_details.data)
);
EasyInputContent::ContentList(vec![InputContent::InputImage(InputImageContent {
image_url: Some(image_url),
detail: ImageDetail::Auto,
file_id: None,
})])
}
LLMMessageContent::File(file_details) => {
let file_data = format!(
"data:{};base64,{}",
file_details.mime,
base64_encode(&file_details.data)
);
let file_content = InputFileArgs::default()
.file_data(file_data)
.filename(file_details.filename())
.build()
.expect("Failed to build InputFileContent");
EasyInputContent::ContentList(vec![InputContent::InputFile(file_content)])
}
};
items.push(InputItem::EasyMessage(EasyInputMessage {
r#type: MessageType::Message,
role,
content,
phase: None,
}));
}
openai_conversation_messages
}
fn convert_llm_message_to_openai_message(llm_message: LLMMessage) -> ChatCompletionRequestMessage {
match llm_message.author {
LLMAuthor::Prompt => ChatCompletionRequestSystemMessageArgs::default()
.content(llm_message.message_text)
.build()
.expect("Failed building OpenAI system message")
.into(),
LLMAuthor::Assistant => ChatCompletionRequestAssistantMessageArgs::default()
.content(llm_message.message_text)
.build()
.expect("Failed building OpenAI assistant message")
.into(),
LLMAuthor::User => ChatCompletionRequestUserMessageArgs::default()
.content(llm_message.message_text)
.build()
.expect("Failed building OpenAI user message")
.into(),
}
}
pub(super) fn convert_string_to_enum<T>(value: &str) -> Result<T, String>
where
T: serde::de::DeserializeOwned,
{
// This is a hacky way to construct an enum from the string we have.
let enum_result: serde_json::Result<T> = serde_json::from_str(&format!("\"{}\"", value));
match enum_result {
Ok(enum_result) => Ok(enum_result),
Err(err) => {
tracing::debug!(?err, "Failed to parse into enum");
Err(format!("The value ({}) is not supported.", value))
}
}
InputParam::Items(items)
}

View File

@@ -95,6 +95,7 @@ impl TryInto<OpenAITextGenerationConfig> for TextGenerationConfig {
max_response_tokens: self.max_response_tokens,
max_completion_tokens: None,
max_context_tokens: self.max_context_tokens,
tools: Default::default(),
})
}
}
@@ -161,13 +162,14 @@ impl TryInto<OpenAITextToSpeechConfig> for TextToSpeechConfig {
type Error = String;
fn try_into(self) -> Result<OpenAITextToSpeechConfig, Self::Error> {
let model_id = convert_string_to_enum::<async_openai::types::SpeechModel>(&self.model_id)?;
let model_id =
convert_string_to_enum::<async_openai::types::audio::SpeechModel>(&self.model_id)?;
let voice = convert_string_to_enum::<async_openai::types::Voice>(&self.voice)?;
let voice = convert_string_to_enum::<async_openai::types::audio::Voice>(&self.voice)?;
let response_format = convert_string_to_enum::<async_openai::types::SpeechResponseFormat>(
&self.response_format,
)?;
let response_format = convert_string_to_enum::<
async_openai::types::audio::SpeechResponseFormat,
>(&self.response_format)?;
Ok(OpenAITextToSpeechConfig {
model_id,
@@ -224,21 +226,27 @@ impl TryInto<OpenAIImageGenerationConfig> for ImageGenerationConfig {
fn try_into(self) -> Result<OpenAIImageGenerationConfig, Self::Error> {
let size = if let Some(size) = &self.size {
convert_string_to_enum::<async_openai::types::ImageSize>(size)?
Some(convert_string_to_enum::<
async_openai::types::images::ImageSize,
>(size)?)
} else {
async_openai::types::ImageSize::S1024x1024
None
};
let style = if let Some(style) = &self.style {
convert_string_to_enum::<async_openai::types::ImageStyle>(style)?
Some(convert_string_to_enum::<
async_openai::types::images::ImageStyle,
>(style)?)
} else {
async_openai::types::ImageStyle::Vivid
None
};
let quality = if let Some(quality) = &self.quality {
convert_string_to_enum::<async_openai::types::ImageQuality>(quality)?
Some(convert_string_to_enum::<
async_openai::types::images::ImageQuality,
>(quality)?)
} else {
async_openai::types::ImageQuality::Standard
None
};
Ok(OpenAIImageGenerationConfig {

View File

@@ -3,23 +3,27 @@ use etke_openai_api_rust::chat::{ChatApi, ChatBody};
use etke_openai_api_rust::images::{ImagesApi, ImagesBody};
use etke_openai_api_rust::{Auth, Message, OpenAI};
const SMALLEST_IMAGE_SIZE: &str = "256x256";
use super::super::ControllerTrait;
use crate::agent::utils::base64_decode;
use crate::utils::base64::base64_decode;
use crate::{
agent::provider::{
ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
ImageEditParams, ImageGenerationParams, ImageSource, SpeechToTextParams,
SpeechToTextResult,
entity::{TextGenerationParams, TextGenerationResult},
},
conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
shorten_messages_list_to_context_size,
MessageContent as LLMMessageContent, shorten_messages_list_to_context_size,
},
};
use crate::{
agent::{
AgentPurpose,
provider::entity::{
ImageGenerationResult, PingResult, TextToSpeechParams, TextToSpeechResult,
ImageEditResult, ImageGenerationResult, PingResult, TextToSpeechParams,
TextToSpeechResult,
},
},
strings,
@@ -60,7 +64,8 @@ impl ControllerTrait for Controller {
let messages = vec![LLMMessage {
author: LLMAuthor::User,
message_text: "Hello!".to_string(),
sender_id: None,
content: LLMMessageContent::Text("Hello!".to_string()),
timestamp: chrono::Utc::now(),
}];
@@ -97,7 +102,8 @@ impl ControllerTrait for Controller {
} else {
Some(LLMMessage {
author: LLMAuthor::Prompt,
message_text: prompt_text,
sender_id: None,
content: LLMMessageContent::Text(prompt_text),
timestamp: chrono::Utc::now(),
})
};
@@ -301,9 +307,11 @@ impl ControllerTrait for Controller {
// when they span multiple lines.
let prompt = prompt.replace("\n", " ");
let size: Option<String> = params
.size_override
.or_else(|| image_generation_config.size.clone());
let size: Option<String> = if params.smallest_size_possible {
Some(SMALLEST_IMAGE_SIZE.to_owned())
} else {
image_generation_config.size.clone()
};
let request = ImagesBody {
model: Some(image_generation_config.model_id.to_owned()),
@@ -366,6 +374,17 @@ impl ControllerTrait for Controller {
))
}
async fn create_image_edit(
&self,
_prompt: &str,
_images: Vec<ImageSource>,
_params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
Err(anyhow::anyhow!(
"The OpenAI image edit API is not supported by the OpenAI-compat provider"
))
}
async fn text_to_speech(
&self,
input: &str,

View File

@@ -26,12 +26,12 @@ use super::controller::ControllerType;
pub fn create_controller_from_yaml_value_config(
agent_id: &str,
config: serde_yaml::Value,
config: serde_yaml_ng::Value,
) -> AgentInstantiationResult<ControllerType> {
let config = match &config {
serde_yaml::Value::Mapping(_) => {
serde_yaml_ng::Value::Mapping(_) => {
let config: Config =
serde_yaml::from_value(config).map_err(AgentInstantiationError::Yaml)?;
serde_yaml_ng::from_value(config).map_err(AgentInstantiationError::Yaml)?;
config
.validate()

View File

@@ -2,7 +2,9 @@ use etke_openai_api_rust::{Message, Role};
use crate::agent::provider::openai::Config as OpenAIConfig;
use crate::conversation::llm::{Author as LLMAuthor, Message as LLMMessage};
use crate::conversation::llm::{
Author as LLMAuthor, Message as LLMMessage, MessageContent as LLMMessageContent,
};
pub fn convert_llm_messages_to_openai_messages(
conversation_messages: Vec<LLMMessage>,
@@ -11,22 +13,39 @@ pub fn convert_llm_messages_to_openai_messages(
Vec::with_capacity(conversation_messages.len());
for message in conversation_messages {
openai_conversation_messages.push(convert_llm_message_to_openai_message(message));
let openai_message = convert_llm_message_to_openai_message(message);
if let Some(openai_message) = openai_message {
openai_conversation_messages.push(openai_message);
}
}
openai_conversation_messages
}
fn convert_llm_message_to_openai_message(llm_message: LLMMessage) -> Message {
fn convert_llm_message_to_openai_message(llm_message: LLMMessage) -> Option<Message> {
let role = match llm_message.author {
LLMAuthor::Prompt => Role::System,
LLMAuthor::Assistant => Role::Assistant,
LLMAuthor::User => Role::User,
};
Message {
role,
content: llm_message.message_text,
match &llm_message.content {
LLMMessageContent::Text(text) => Some(Message {
role,
content: text.clone(),
}),
LLMMessageContent::Image(_image_details) => {
tracing::warn!(
"The OpenAI-compat provider's library does not support image content. This image message will be skipped."
);
None
}
LLMMessageContent::File(_file_details) => {
tracing::warn!(
"The OpenAI-compat provider's library does not support file content. This file message will be skipped."
);
None
}
}
}

View File

@@ -0,0 +1,156 @@
use crate::agent::AgentPurpose;
use crate::agent::provider::entity::{TextToSpeechParams, TextToSpeechResult};
use crate::agent::provider::{SpeechToTextParams, SpeechToTextResult};
use crate::strings;
use super::config::Config;
use super::wire::{SpeechRequest, TranscriptionResponse};
pub async fn speech_to_text(
config: &Config,
http: &reqwest::Client,
mime_type: &mxlink::mime::Mime,
media: Vec<u8>,
params: SpeechToTextParams,
) -> anyhow::Result<SpeechToTextResult> {
let Some(speech_to_text_config) = &config.speech_to_text else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::SpeechToText
),
));
};
// Unlike the openai_compat path (which writes the audio to a temp file because its library
// can't take bytes), reqwest's multipart takes the bytes directly.
let part = reqwest::multipart::Part::bytes(media)
.file_name("audio")
.mime_str(mime_type.as_ref())?;
let mut form = reqwest::multipart::Form::new()
.part("file", part)
.text("model", speech_to_text_config.model_id.clone())
.text("response_format", "json");
if let Some(language) = &params.language_override {
form = form.text("language", language.clone());
}
let url = format!(
"{}/audio/transcriptions",
config.base_url.trim_end_matches('/')
);
tracing::trace!(
model_id = speech_to_text_config.model_id,
language = ?params.language_override,
"Sending Venice audio transcription API request"
);
let response = http
.post(&url)
.bearer_auth(&config.api_key)
.multipart(form)
.send()
.await?;
let status = response.status();
if !status.is_success() {
// Body to the server log only, not into the returned error (which reaches the Matrix room).
let body = response.text().await.unwrap_or_default();
tracing::warn!(%status, body, "Venice audio transcription request failed");
return Err(anyhow::anyhow!(
"Venice audio transcription request failed with status {status}"
));
}
let response: TranscriptionResponse = response.json().await?;
Ok(SpeechToTextResult {
text: response.text,
})
}
pub async fn text_to_speech(
config: &Config,
http: &reqwest::Client,
input: &str,
params: TextToSpeechParams,
) -> anyhow::Result<TextToSpeechResult> {
let Some(text_to_speech_config) = &config.text_to_speech else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::TextToSpeech
),
));
};
// Per-call overrides win over the configured defaults.
let voice = params
.voice_override
.or_else(|| text_to_speech_config.voice.clone());
let speed = params.speed_override.or(text_to_speech_config.speed);
let response_format = text_to_speech_config.response_format.clone();
let mime_type = response_format_to_mime_type(response_format.as_deref());
let request = SpeechRequest {
model: text_to_speech_config.model_id.clone(),
input: input.to_owned(),
voice,
speed,
response_format,
prompt: text_to_speech_config.prompt.clone(),
temperature: text_to_speech_config.temperature,
top_p: text_to_speech_config.top_p,
};
let url = format!("{}/audio/speech", config.base_url.trim_end_matches('/'));
tracing::trace!(
model_id = text_to_speech_config.model_id,
voice = ?request.voice,
"Sending Venice text-to-speech API request"
);
let response = http
.post(&url)
.bearer_auth(&config.api_key)
.json(&request)
.send()
.await?;
let status = response.status();
if !status.is_success() {
// Body to the server log only, not into the returned error (which reaches the Matrix room).
let body = response.text().await.unwrap_or_default();
tracing::warn!(%status, body, "Venice text-to-speech request failed");
return Err(anyhow::anyhow!(
"Venice text-to-speech request failed with status {status}"
));
}
// The speech endpoint answers with raw binary audio; read the body directly.
let bytes = response.bytes().await?.to_vec();
Ok(TextToSpeechResult { bytes, mime_type })
}
/// Map a Venice TTS `response_format` to its MIME type. Defaults to `audio/mpeg` (the
/// IANA-registered MP3 type, RFC 3003) when the format is unset, matching Venice's own `mp3`
/// default. This deliberately uses `audio/mpeg` rather than the `audio/mp3` alias the openai
/// provider emits; baibot's downstream audio-filename mapping treats both as `.mp3`.
fn response_format_to_mime_type(response_format: Option<&str>) -> mxlink::mime::Mime {
let raw = match response_format.unwrap_or("mp3") {
"mp3" => "audio/mpeg",
"opus" => "audio/ogg",
"aac" => "audio/aac",
"flac" => "audio/flac",
"wav" => "audio/wav",
"pcm" => "audio/L8",
_ => "audio/mpeg",
};
raw.parse()
.unwrap_or(mxlink::mime::APPLICATION_OCTET_STREAM)
}

View File

@@ -0,0 +1,351 @@
use std::collections::hash_map::DefaultHasher;
use std::hash::{Hash, Hasher};
use std::sync::OnceLock;
use regex::Regex;
use crate::agent::AgentPurpose;
use crate::agent::provider::entity::{TextGenerationParams, TextGenerationResult};
use crate::conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
MessageContent as LLMMessageContent, shorten_messages_list_to_context_size,
};
use crate::strings;
use super::config::{Config, WebSearchMode};
use super::utils::convert_llm_messages_to_venice;
use super::wire::{ChatCompletionRequest, ChatCompletionResponse, WebSearchCitation};
pub async fn generate_text(
config: &Config,
http: &reqwest::Client,
unsupported: &super::recovery::UnsupportedFieldsCache,
conversation: LLMConversation,
params: TextGenerationParams,
) -> anyhow::Result<TextGenerationResult> {
let Some(text_generation_config) = &config.text_generation else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::TextGeneration
),
));
};
let prompt_text = params.prompt_variables.format(
params
.prompt_override
.unwrap_or(text_generation_config.prompt.clone().unwrap_or_default())
.trim(),
);
// Prompt-cache routing key. Hash ONLY conversation-stable inputs: the rendered system prompt
// and the conversation start time. Folding in anything per-turn (message content, the current
// time, the message count) would mint a fresh key every turn, miss the cache every lookup, and
// pay full price plus the hashing cost. The start time is rendered explicitly here so the key
// stays stable even when the user's prompt template never mentions the time variable; an
// unknown start time renders "unknown" and simply keys on the prompt alone.
let conversation_start_time = params
.prompt_variables
.format("{{ baibot_conversation_start_time_utc }}");
let prompt_cache_key = derive_prompt_cache_key(&prompt_text, &conversation_start_time);
let prompt_message = if prompt_text.is_empty() {
None
} else {
Some(LLMMessage {
author: LLMAuthor::Prompt,
sender_id: None,
content: LLMMessageContent::Text(prompt_text),
timestamp: chrono::Utc::now(),
})
};
let mut conversation_messages = conversation.messages;
if params.context_management_enabled {
conversation_messages = shorten_messages_list_to_context_size(
&text_generation_config.model_id,
&prompt_message,
conversation_messages,
text_generation_config.max_response_tokens,
text_generation_config.max_context_tokens,
);
}
if let Some(prompt_message) = prompt_message {
conversation_messages.insert(0, prompt_message);
}
let messages = convert_llm_messages_to_venice(conversation_messages)?;
let temperature = params
.temperature_override
.unwrap_or(text_generation_config.temperature);
// When web search is active, ask Venice to return structured search results so we can render
// readable citations from them. Respect an explicit user choice and only fill the flag when
// the user left it unset.
let venice_parameters = text_generation_config
.venice_parameters
.clone()
.map(|mut vp| {
let web_search_active = matches!(
vp.enable_web_search,
Some(WebSearchMode::On | WebSearchMode::Auto)
);
if web_search_active && vp.return_search_results_as_documents.is_none() {
vp.return_search_results_as_documents = Some(true);
}
vp
});
let mut request = ChatCompletionRequest {
model: text_generation_config.model_id.clone(),
messages,
temperature: Some(temperature),
// Web search rides entirely inside `venice_parameters`; there is no `tools` array here.
// `max_tokens` is deprecated on Venice in favor of `max_completion_tokens`.
max_completion_tokens: text_generation_config.max_response_tokens,
top_p: text_generation_config.top_p,
frequency_penalty: text_generation_config.frequency_penalty,
presence_penalty: text_generation_config.presence_penalty,
repetition_penalty: text_generation_config.repetition_penalty,
reasoning_effort: text_generation_config.reasoning_effort.clone(),
prompt_cache_key: Some(prompt_cache_key),
prompt_cache_retention: text_generation_config.prompt_cache_retention.clone(),
venice_parameters,
};
let url = format!("{}/chat/completions", config.base_url.trim_end_matches('/'));
let model_id = text_generation_config.model_id.clone();
// Proactively drop fields this model has already rejected earlier in this process, so a known
// mismatch costs zero wasted round-trips after the first discovery. Venice's body is
// `additionalProperties: false`, so sending a known-unsupported field would 400 again.
for field in unsupported.known_for(&model_id) {
super::recovery::strip_droppable_field(&mut request, &field);
}
// Send with bounded auto-recovery. When Venice 400s because a model does not support an optional
// knob, it names the field (`field: '...'`); if that field is one we may safely drop, we strip
// it, remember the rejection for this model, and retry. The loop is bounded: each retry clears a
// distinct droppable field (strip returns false once it is gone), so after at most
// `DROPPABLE_FIELDS.len()` retries the request either succeeds or surfaces the error.
let response = loop {
tracing::trace!(
model = model_id,
messages_count = request.messages.len(),
"Sending Venice chat completion API request"
);
let response = http
.post(&url)
.bearer_auth(&config.api_key)
.json(&request)
.send()
.await?;
let status = response.status();
if status.is_success() {
break response;
}
// Always log the body server-side: Venice explains a rejected strict body there.
let body = response.text().await.unwrap_or_default();
tracing::warn!(%status, body, "Venice chat completion request failed");
// Recover from a strict-body 400 over an unsupported optional knob: strip the named field
// and retry. Only fields in `DROPPABLE_FIELDS` are eligible, so a meaning-bearing knob (a
// sampling parameter) is never silently dropped; that case falls through to surface below.
if status == reqwest::StatusCode::BAD_REQUEST
&& let Some(field) = super::recovery::parse_rejected_field(&body)
&& super::recovery::strip_droppable_field(&mut request, &field)
{
unsupported.record(&model_id, &field);
tracing::info!(
model = model_id,
field,
"Venice rejected an unsupported field; dropping it and retrying"
);
continue;
}
// A 413 almost always means an attached file pushed the request past Venice's size limit.
if status == reqwest::StatusCode::PAYLOAD_TOO_LARGE {
return Err(anyhow::anyhow!(
"The request was too large for Venice, most likely an attached file over the 25MB limit."
));
}
// A 400 is a complaint about the request baibot built, so the body is safe and useful to
// surface: it tells the operator (e.g. at agent-create time) exactly which field or value
// Venice rejected, instead of an opaque status. Other statuses keep the body OUT of the
// returned error, since it can carry account / rate-limit details that shouldn't reach the
// room.
if status == reqwest::StatusCode::BAD_REQUEST {
return Err(anyhow::anyhow!(
"Venice rejected the request (400 Bad Request): {}",
super::recovery::extract_error_message(&body)
));
}
return Err(anyhow::anyhow!(
"Venice chat completion request failed with status {status}"
));
};
let response: ChatCompletionResponse = response.json().await?;
let citations = response
.venice_parameters
.map(|vp| vp.web_search_citations)
.unwrap_or_default();
let Some(choice) = response.choices.into_iter().next() else {
return Err(anyhow::anyhow!(
"No choices were returned from the Venice chat completion API"
));
};
let Some(content) = choice.message.content else {
return Err(anyhow::anyhow!(
"No message content was returned from the Venice chat completion API"
));
};
let text = render_with_citations(content, &citations);
let text = append_reasoning(
text,
choice.message.reasoning_content,
text_generation_config.show_reasoning,
);
Ok(TextGenerationResult { text })
}
/// Builds the prompt-cache routing key from conversation-stable inputs. `DefaultHasher::new()` is a
/// fixed-seed SipHasher (keys 0,0), so it is deterministic across processes and restarts: identical
/// inputs always produce the same key, which is what lets a restarted bot keep hitting the warm
/// cache. The algorithm is not guaranteed stable across Rust std versions, so a rebuild on a new
/// toolchain can shift every key once, a one-time cache warm-up with no correctness effect.
pub(super) fn derive_prompt_cache_key(prompt_text: &str, conversation_start_time: &str) -> String {
let mut hasher = DefaultHasher::new();
prompt_text.hash(&mut hasher);
conversation_start_time.hash(&mut hasher);
format!("{:016x}", hasher.finish())
}
/// Appends the model's thinking to the reply only when the deployment opts in via `show_reasoning`.
/// `reasoning_content` is a field separate from the answer `content` (it is unaffected by
/// `strip_thinking_response`, which only strips inline `<think>` blocks from `content`), so reading
/// it here is independent of that knob. Default-off matches today's behavior: thinking never reaches
/// a room that did not ask for it.
///
/// The thinking renders as a Matrix-native collapsible `<details>` block: folded by default, one
/// click to expand, so it stays out of the way of the answer instead of dumping a wall of reasoning
/// inline. This survives the send path: the reply goes through markdown (`send_text_markdown`),
/// whose pulldown-cmark pass writes raw HTML verbatim rather than escaping it, and ruma's HTML
/// sanitizer allow-lists `<details>`/`<summary>`. Clients that do not render `<details>` degrade to
/// showing the summary and reasoning inline, so nothing is lost there either.
pub(super) fn append_reasoning(
text: String,
reasoning_content: Option<String>,
show_reasoning: bool,
) -> String {
if !show_reasoning {
return text;
}
match reasoning_content {
Some(reasoning) if !reasoning.trim().is_empty() => {
// The blank lines around the trimmed reasoning keep it a separate markdown block from
// the surrounding `<details>`/`</details>` HTML blocks, so the reasoning itself still
// renders as markdown (lists, code, emphasis) inside the collapsible.
let reasoning = reasoning.trim();
format!(
"{text}\n\n<details><summary>💭 Reasoning</summary>\n\n{reasoning}\n\n</details>"
)
}
_ => text,
}
}
/// Rewrites Venice's inline `^n^` citation superscripts into readable `[n]` references and appends
/// a `Sources:` list of markdown links, one per citation in order. Returns the content unchanged
/// when web search returned no citations, so non-search replies are never touched.
///
/// Citation `title` and `url` come from scraped web pages, so they are attacker-influenced. The
/// title is escaped so it cannot break out of the markdown link label, and the URL is used as a
/// link target only when it is a clean `http(s)` URL with no markdown-breaking characters;
/// otherwise the citation renders as plain text. This stops a hostile page title or URL from
/// injecting a spoofed clickable link into the room.
pub(super) fn render_with_citations(content: String, citations: &[WebSearchCitation]) -> String {
if citations.is_empty() {
return content;
}
let mut text = rewrite_citation_superscripts(&content);
let mut sources = String::from("\n\nSources:");
for (index, citation) in citations.iter().enumerate() {
let n = index + 1;
let title = escape_markdown_link_text(&citation.title);
match sanitize_link_url(&citation.url) {
// A citation that arrived with no title still renders as a usable link by showing the
// URL as the link text, rather than an empty `[]( )` label.
Some(url) if title.is_empty() => sources.push_str(&format!("\n[{n}] [{url}]({url})")),
Some(url) => sources.push_str(&format!("\n[{n}] [{title}]({url})")),
None if !title.is_empty() => sources.push_str(&format!("\n[{n}] {title}")),
None => sources.push_str(&format!("\n[{n}] (source unavailable)")),
}
}
text.push_str(&sources);
text
}
/// Venice marks web-search citations with superscript runs in the reply text: a single `^1^`, a
/// comma list `^1,2^`, or a caret-chained run `^2^3^10^` where consecutive citations share a
/// caret. The whole run has to be matched at once: a per-citation pattern (string or regex)
/// consumes the shared caret on the first match and orphans the rest (`^2^3^` would leave `3^`).
/// So this matches each full run and expands it to one `[n]` per citation (`^2^3^` -> `[2][3]`).
fn rewrite_citation_superscripts(content: &str) -> String {
static RUN: OnceLock<Regex> = OnceLock::new();
let run = RUN.get_or_init(|| {
Regex::new(r"\^\d+(?:[,^]\d+)*\^").expect("citation superscript regex is valid")
});
run.replace_all(content, |caps: &regex::Captures| {
caps[0]
.split(['^', ','])
.filter(|piece| !piece.is_empty())
.map(|n| format!("[{n}]"))
.collect::<String>()
})
.into_owned()
}
/// Escapes the characters that would let citation title text break out of a markdown link label,
/// and folds newlines to spaces so a multi-line title cannot inject extra markdown structure.
fn escape_markdown_link_text(text: &str) -> String {
text.replace('\\', "\\\\")
.replace('[', "\\[")
.replace(']', "\\]")
.replace(['\r', '\n'], " ")
}
/// Returns the URL as a markdown link target only when it is a clean `http(s)` URL with no
/// characters that would break the `(...)` destination or smuggle a different scheme. Anything else
/// returns `None`, so the caller renders the citation as plain text instead of a link.
fn sanitize_link_url(url: &str) -> Option<String> {
let url = url.trim();
let is_http = url.starts_with("https://") || url.starts_with("http://");
let is_clean = !url.contains(['(', ')', '<', '>', ' ', '\t', '\r', '\n']);
if is_http && is_clean {
Some(url.to_owned())
} else {
None
}
}

View File

@@ -0,0 +1,410 @@
use serde::{Deserialize, Serialize};
use crate::agent::{default_prompt, provider::ConfigTrait};
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct Config {
pub base_url: String,
pub api_key: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub text_generation: Option<TextGenerationConfig>,
#[serde(skip_serializing_if = "Option::is_none")]
pub speech_to_text: Option<SpeechToTextConfig>,
#[serde(skip_serializing_if = "Option::is_none")]
pub text_to_speech: Option<TextToSpeechConfig>,
#[serde(skip_serializing_if = "Option::is_none")]
pub image_generation: Option<ImageGenerationConfig>,
}
impl Default for Config {
fn default() -> Self {
Self {
base_url: "https://api.venice.ai/api/v1".to_owned(),
api_key: "YOUR_API_KEY_HERE".to_owned(),
text_generation: Some(TextGenerationConfig::default()),
speech_to_text: Some(SpeechToTextConfig::default()),
text_to_speech: Some(TextToSpeechConfig::default()),
image_generation: Some(ImageGenerationConfig::default()),
}
}
}
impl ConfigTrait for Config {
fn validate(&self) -> Result<(), String> {
if self.base_url.is_empty() {
return Err("The base URL must not be empty.".to_owned());
}
Ok(())
}
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct TextGenerationConfig {
#[serde(default = "default_text_model_id")]
pub model_id: String,
#[serde(default)]
pub prompt: Option<String>,
#[serde(default = "super::super::default_temperature")]
pub temperature: f32,
#[serde(default)]
pub max_response_tokens: Option<u32>,
#[serde(default)]
pub max_context_tokens: u32,
/// Sampling and reasoning knobs that live at the top level of Venice's `/chat/completions`
/// body, not inside the `venice_parameters` bag. Venice silently ignores a top-level knob
/// placed in the bag, so these sit here as siblings and map straight to top-level wire fields
/// in `chat.rs`. Each is omitted from the request when unset.
#[serde(default)]
pub top_p: Option<f32>,
#[serde(default)]
pub frequency_penalty: Option<f32>,
#[serde(default)]
pub presence_penalty: Option<f32>,
#[serde(default)]
pub repetition_penalty: Option<f32>,
#[serde(default)]
pub reasoning_effort: Option<String>,
/// Prompt-cache retention window (`default`, `extended`, or `24h`). This carries a named
/// default rather than a bare `#[serde(default)]` (which would yield `None`), so a config that
/// omits the key still ships `24h` and keeps caching on. Caching is the per-deployment cost
/// lever, so the omitted-key case must not silently disable it. The value here must agree with
/// the `Default` impl below.
#[serde(default = "default_prompt_cache_retention")]
pub prompt_cache_retention: Option<String>,
/// When set, the model's `reasoning_content` (its thinking) is appended to the reply. Off by
/// default to match today's `strip_thinking_response: true` behavior, so existing deployments
/// see no change.
#[serde(default)]
pub show_reasoning: bool,
/// Venice-specific request knobs, serialized 1:1 into the `venice_parameters` bag on the
/// wire. Any unset field is omitted, so Venice applies its own server-side default.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub venice_parameters: Option<VeniceParameters>,
}
impl Default for TextGenerationConfig {
fn default() -> Self {
Self {
model_id: default_text_model_id(),
prompt: Some(default_prompt().to_owned()),
temperature: super::super::default_temperature(),
// Reserved output budget: sent as the response cap AND subtracted from the context
// window when trimming history. Mirrors the openai_compat sibling's default.
max_response_tokens: Some(4096),
// Matches Venice's own `availableContextTokens` (131072) and the non-OpenAI sibling
// providers (ollama/localai/mistral all default to 128_000).
max_context_tokens: 128_000,
// Sampling knobs stay None so Venice applies its own server-side default. Caching is
// the one exception: retention defaults to 24h here so a programmatic default caches
// out of the box, agreeing with the `#[serde(default = ...)]` on the field.
top_p: None,
frequency_penalty: None,
presence_penalty: None,
repetition_penalty: None,
reasoning_effort: None,
prompt_cache_retention: default_prompt_cache_retention(),
show_reasoning: false,
// A usable starting point, not an everything-set dump: only these three are sent;
// every other knob stays None so Venice applies its own default (omitting != false).
venice_parameters: Some(VeniceParameters {
enable_web_search: Some(WebSearchMode::Auto),
strip_thinking_response: Some(true),
enable_e2ee: Some(false),
..Default::default()
}),
}
}
}
fn default_text_model_id() -> String {
"kimi-k2-5".to_owned()
}
/// Defaults prompt-cache retention to 24h so caching is on unless a config explicitly opts out.
/// A bare `#[serde(default)]` would deserialize an omitted key to `None`, which disables caching;
/// this keeps the cost lever engaged for configs that never mention it.
fn default_prompt_cache_retention() -> Option<String> {
Some("24h".to_owned())
}
/// The full `venice_parameters` knob set, mirroring Venice's `ChatCompletionRequest`
/// schema field-for-field. Every field is optional with `skip_serializing_if`, so the
/// request never carries a knob the user didn't set (the body is `additionalProperties: false`,
/// and an unset knob simply omits rather than sending `null`).
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
pub struct VeniceParameters {
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_web_search: Option<WebSearchMode>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_web_citations: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_web_scraping: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub include_venice_system_prompt: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub include_search_results_in_stream: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub return_search_results_as_documents: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_x_search: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_e2ee: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub character_slug: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub strip_thinking_response: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub disable_thinking: Option<bool>,
/// Response verbosity (`low`, `medium`, `high`). Venice accepts this both top-level and inside
/// the bag; it lives here so the top-level config stays lean, and Venice reads it from the bag.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub verbosity: Option<String>,
}
#[derive(Debug, Clone, Copy, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum WebSearchMode {
Auto,
On,
Off,
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct SpeechToTextConfig {
#[serde(default = "default_speech_to_text_model_id")]
pub model_id: String,
}
impl Default for SpeechToTextConfig {
fn default() -> Self {
Self {
model_id: default_speech_to_text_model_id(),
}
}
}
fn default_speech_to_text_model_id() -> String {
"nvidia/parakeet-tdt-0.6b-v3".to_owned()
}
/// `/audio/speech` (`CreateSpeechRequestSchema`) request knobs. Only `model_id` is required on
/// the wire; everything else is optional with `skip_serializing_if` so an unset knob is omitted
/// rather than sent as `null` (the body is `additionalProperties: false`). `voice` is a free
/// `Option<String>`, not a closed enum: Venice's voice set spans dozens of model-specific names
/// plus arbitrary cloned-voice handles (`vv_<id>`), so an enum would reject valid handles.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct TextToSpeechConfig {
#[serde(default = "default_text_to_speech_model_id")]
pub model_id: String,
#[serde(
default = "default_text_to_speech_voice",
skip_serializing_if = "Option::is_none"
)]
pub voice: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub speed: Option<f32>,
#[serde(
default = "default_text_to_speech_response_format",
skip_serializing_if = "Option::is_none"
)]
pub response_format: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub prompt: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub temperature: Option<f32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub top_p: Option<f32>,
}
impl Default for TextToSpeechConfig {
fn default() -> Self {
Self {
model_id: default_text_to_speech_model_id(),
voice: default_text_to_speech_voice(),
speed: None,
response_format: default_text_to_speech_response_format(),
prompt: None,
temperature: None,
top_p: None,
}
}
}
fn default_text_to_speech_model_id() -> String {
"tts-kokoro".to_owned()
}
fn default_text_to_speech_voice() -> Option<String> {
Some("af_sky".to_owned())
}
fn default_text_to_speech_response_format() -> Option<String> {
Some("mp3".to_owned())
}
/// `/image/generate` (`GenerateImageRequest`) request knobs, mirroring Venice's schema
/// field-for-field. Only `model_id` is required; every other knob is optional with
/// `skip_serializing_if` so unset knobs are omitted (the body is `additionalProperties: false`).
/// The full knob set is deliberate: the native `/image/generate` endpoint is the flagship's
/// reason to exist over the knob-dropping OpenAI-compat path, so the knobs ARE the feature.
/// The deprecated `inpaint` knob is intentionally absent.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct ImageGenerationConfig {
#[serde(default = "default_image_generation_model_id")]
pub model_id: String,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub negative_prompt: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub cfg_scale: Option<f32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub steps: Option<u32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub style_preset: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub seed: Option<i64>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub safe_mode: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub hide_watermark: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub format: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub width: Option<u32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub height: Option<u32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub aspect_ratio: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub resolution: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub quality: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub lora_strength: Option<u32>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub embed_exif_metadata: Option<bool>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub enable_web_search: Option<bool>,
/// Image-edit settings, nested here because baibot has a single `ImageGeneration` purpose
/// and edit shares its config gate. The gen and edit model sets are disjoint, so edit
/// carries its own model field.
#[serde(default)]
pub edit: ImageEditSettings,
}
impl Default for ImageGenerationConfig {
fn default() -> Self {
Self {
model_id: default_image_generation_model_id(),
negative_prompt: None,
cfg_scale: None,
steps: None,
style_preset: None,
seed: None,
safe_mode: None,
hide_watermark: None,
format: None,
width: None,
height: None,
aspect_ratio: None,
resolution: None,
quality: None,
lora_strength: None,
embed_exif_metadata: None,
enable_web_search: None,
edit: ImageEditSettings::default(),
}
}
}
fn default_image_generation_model_id() -> String {
"chroma".to_owned()
}
/// `/image/edit` (`EditImageRequest`) request knobs, mirroring Venice's schema. The source image
/// and prompt are supplied per-call (not config), so only the model and the output-shaping knobs
/// live here. Each knob is optional with `skip_serializing_if`.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct ImageEditSettings {
#[serde(default = "default_image_edit_model_id")]
pub model_id: String,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub output_format: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub aspect_ratio: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub resolution: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub safe_mode: Option<bool>,
}
impl Default for ImageEditSettings {
fn default() -> Self {
Self {
model_id: default_image_edit_model_id(),
output_format: None,
aspect_ratio: None,
resolution: None,
safe_mode: None,
}
}
}
fn default_image_edit_model_id() -> String {
"firered-image-edit".to_owned()
}

View File

@@ -0,0 +1,167 @@
use crate::agent::AgentPurpose;
use crate::agent::provider::entity::{
ImageEditResult, ImageGenerationResult, ImageSource, PingResult, TextGenerationParams,
TextGenerationResult, TextToSpeechParams, TextToSpeechResult,
};
use crate::agent::provider::{
ImageEditParams, ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
};
use crate::conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
MessageContent as LLMMessageContent,
};
use super::super::ControllerTrait;
use super::config::Config;
use super::recovery::UnsupportedFieldsCache;
#[derive(Debug, Clone)]
pub struct Controller {
config: Config,
http: reqwest::Client,
// Per-model record of chat fields this Venice deployment has rejected as unsupported, learned at
// runtime. `Arc`-backed inside, so the `Clone` derive shares one cache across all clones of an
// agent's controller.
unsupported_fields: UnsupportedFieldsCache,
}
impl Controller {
pub fn new(config: Config) -> Self {
// Image generation and text-to-speech can run long, so give the client a generous timeout
// instead of reqwest's default (none). `build` only fails on TLS/system init; fall back to
// the infallible `Client::new()` so this constructor stays infallible.
let http = reqwest::Client::builder()
.timeout(std::time::Duration::from_secs(120))
.build()
.unwrap_or_else(|_| reqwest::Client::new());
// Process-global per-deployment cache rather than a fresh one: dynamic agents rebuild their
// controller on every message, so an instance-owned cache would never retain a learned
// rejection and the same unsupported field would 400 (and warn) on every turn.
let unsupported_fields = UnsupportedFieldsCache::shared_for(&config.base_url);
Self {
config,
http,
unsupported_fields,
}
}
}
impl ControllerTrait for Controller {
async fn ping(&self) -> anyhow::Result<PingResult> {
if !self.supports_purpose(AgentPurpose::TextGeneration) {
return Ok(PingResult::Inconclusive);
}
// Mirror the openai/openai_compat ping: a real "Hello!" round-trip exercises the strict
// /chat/completions body and auth, so a successful ping proves text generation works.
let messages = vec![LLMMessage {
author: LLMAuthor::User,
sender_id: None,
content: LLMMessageContent::Text("Hello!".to_string()),
timestamp: chrono::Utc::now(),
}];
let conversation = LLMConversation { messages };
self.generate_text(conversation, TextGenerationParams::default())
.await?;
Ok(PingResult::Successful)
}
async fn generate_text(
&self,
conversation: LLMConversation,
params: TextGenerationParams,
) -> anyhow::Result<TextGenerationResult> {
super::chat::generate_text(
&self.config,
&self.http,
&self.unsupported_fields,
conversation,
params,
)
.await
}
async fn speech_to_text(
&self,
mime_type: &mxlink::mime::Mime,
media: Vec<u8>,
params: SpeechToTextParams,
) -> anyhow::Result<SpeechToTextResult> {
super::audio::speech_to_text(&self.config, &self.http, mime_type, media, params).await
}
async fn generate_image(
&self,
prompt: &str,
params: ImageGenerationParams,
) -> anyhow::Result<ImageGenerationResult> {
super::images::generate_image(&self.config, &self.http, prompt, params).await
}
async fn create_image_edit(
&self,
prompt: &str,
images: Vec<ImageSource>,
params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
super::images::create_image_edit(&self.config, &self.http, prompt, images, params).await
}
async fn text_to_speech(
&self,
input: &str,
params: TextToSpeechParams,
) -> anyhow::Result<TextToSpeechResult> {
super::audio::text_to_speech(&self.config, &self.http, input, params).await
}
fn supports_purpose(&self, purpose: AgentPurpose) -> bool {
match purpose {
AgentPurpose::TextGeneration => self.config.text_generation.is_some(),
AgentPurpose::SpeechToText => self.config.speech_to_text.is_some(),
AgentPurpose::TextToSpeech => self.config.text_to_speech.is_some(),
AgentPurpose::ImageGeneration => self.config.image_generation.is_some(),
AgentPurpose::CatchAll => true,
}
}
fn text_generation_model_id(&self) -> Option<String> {
self.config
.text_generation
.as_ref()
.map(|config| config.model_id.to_owned())
}
fn text_generation_prompt(&self) -> Option<String> {
self.config
.text_generation
.as_ref()
.and_then(|config| config.prompt.clone())
}
fn text_generation_temperature(&self) -> Option<f32> {
self.config
.text_generation
.as_ref()
.map(|config| config.temperature)
}
fn text_to_speech_voice(&self) -> Option<String> {
self.config
.text_to_speech
.as_ref()
.and_then(|config| config.voice.clone())
}
fn text_to_speech_speed(&self) -> Option<f32> {
self.config
.text_to_speech
.as_ref()
.and_then(|config| config.speed)
}
}

View File

@@ -0,0 +1,190 @@
use crate::agent::AgentPurpose;
use crate::agent::provider::entity::{ImageEditResult, ImageGenerationResult, ImageSource};
use crate::agent::provider::{ImageEditParams, ImageGenerationParams};
use crate::strings;
use crate::utils::base64::{base64_decode, base64_encode};
use super::config::Config;
use super::wire::{EditImageRequest, GenerateImageRequest, GenerateImageResponse};
/// Generate an image via Venice's native `/image/generate` endpoint.
///
/// This is the base64-in-JSON path: we pin `return_binary: false` so Venice answers with a JSON
/// envelope (`GenerateImageResponse`) carrying the image as a base64 string, which we decode. The
/// sibling `create_image_edit` is the *other* response shape (raw binary); the two must not be
/// crossed. `params` is advisory only; the Venice config drives the request.
pub async fn generate_image(
config: &Config,
http: &reqwest::Client,
prompt: &str,
_params: ImageGenerationParams,
) -> anyhow::Result<ImageGenerationResult> {
let Some(image_generation_config) = &config.image_generation else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::ImageGeneration
),
));
};
let request = GenerateImageRequest {
model: image_generation_config.model_id.clone(),
prompt: prompt.to_owned(),
// Pinned: baibot wants exactly one image, returned as base64-in-JSON so `GenerateImageResponse`
// can decode it. Flipping `return_binary` would make Venice answer with raw binary and break
// the JSON decode below, so neither knob is configurable.
return_binary: false,
variants: 1,
negative_prompt: image_generation_config.negative_prompt.clone(),
cfg_scale: image_generation_config.cfg_scale,
steps: image_generation_config.steps,
style_preset: image_generation_config.style_preset.clone(),
seed: image_generation_config.seed,
safe_mode: image_generation_config.safe_mode,
hide_watermark: image_generation_config.hide_watermark,
format: image_generation_config.format.clone(),
width: image_generation_config.width,
height: image_generation_config.height,
aspect_ratio: image_generation_config.aspect_ratio.clone(),
resolution: image_generation_config.resolution.clone(),
quality: image_generation_config.quality.clone(),
lora_strength: image_generation_config.lora_strength,
embed_exif_metadata: image_generation_config.embed_exif_metadata,
enable_web_search: image_generation_config.enable_web_search,
};
let url = format!("{}/image/generate", config.base_url.trim_end_matches('/'));
// The prompt is user content; keep it out of logs (mirrors the STT/TTS paths).
tracing::trace!(
model_id = image_generation_config.model_id,
"Sending Venice image generation API request"
);
let response = http
.post(&url)
.bearer_auth(&config.api_key)
.json(&request)
.send()
.await?;
let status = response.status();
if !status.is_success() {
// Body to the server log only, not into the returned error (which reaches the Matrix room).
let body = response.text().await.unwrap_or_default();
tracing::warn!(%status, body, "Venice image generation request failed");
return Err(anyhow::anyhow!(
"Venice image generation request failed with status {status}"
));
}
let response: GenerateImageResponse = response.json().await?;
tracing::trace!(request_id = ?response.id, "Venice image generation succeeded");
let Some(image_base64) = response.images.into_iter().next() else {
return Err(anyhow::anyhow!(
"The Venice image generation API returned no images"
));
};
// Swallow the decode error's detail (it can echo input bytes/offsets); the returned error
// reaches the Matrix room, so it stays generic while the real cause goes to the server log.
let bytes = base64_decode(&image_base64).map_err(|decode_err| {
tracing::warn!(%decode_err, "Venice image generation returned undecodable base64");
anyhow::anyhow!("Venice image generation returned invalid base64 image data")
})?;
Ok(ImageGenerationResult {
bytes,
mime_type: image_format_to_mime_type(image_generation_config.format.as_deref()),
revised_prompt: None,
})
}
/// Edit an image via Venice's native `/image/edit` endpoint.
///
/// This is the raw-binary path: the request is JSON carrying the source image as a base64 string
/// (Venice's `image` field is `anyOf` upload/base64/URL; we send base64, no multipart), and the
/// response body IS the edited image bytes (no JSON envelope). `params` is advisory only.
pub async fn create_image_edit(
config: &Config,
http: &reqwest::Client,
prompt: &str,
images: Vec<ImageSource>,
_params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
let Some(image_generation_config) = &config.image_generation else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::ImageGeneration
),
));
};
let edit_config = &image_generation_config.edit;
let Some(source) = images.into_iter().next() else {
return Err(anyhow::anyhow!("No image sources provided"));
};
let request = EditImageRequest {
model: edit_config.model_id.clone(),
prompt: prompt.to_owned(),
image: base64_encode(&source.bytes),
output_format: edit_config.output_format.clone(),
aspect_ratio: edit_config.aspect_ratio.clone(),
resolution: edit_config.resolution.clone(),
safe_mode: edit_config.safe_mode,
};
let url = format!("{}/image/edit", config.base_url.trim_end_matches('/'));
// The prompt is user content; keep it out of logs (mirrors the STT/TTS paths).
tracing::trace!(
model_id = edit_config.model_id,
"Sending Venice image edit API request"
);
let response = http
.post(&url)
.bearer_auth(&config.api_key)
.json(&request)
.send()
.await?;
let status = response.status();
if !status.is_success() {
// Body to the server log only, not into the returned error (which reaches the Matrix room).
let body = response.text().await.unwrap_or_default();
tracing::warn!(%status, body, "Venice image edit request failed");
return Err(anyhow::anyhow!(
"Venice image edit request failed with status {status}"
));
}
// The edit endpoint answers with raw binary image bytes, so read the body directly instead of
// parsing JSON. The actual format comes from the response Content-Type header; fall back to the
// configured `output_format` when the header is missing or unparseable.
let mime_type = response
.headers()
.get(reqwest::header::CONTENT_TYPE)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<mxlink::mime::Mime>().ok())
.unwrap_or_else(|| image_format_to_mime_type(edit_config.output_format.as_deref()));
let bytes = response.bytes().await?.to_vec();
Ok(ImageEditResult { bytes, mime_type })
}
/// Map a Venice image `format`/`output_format` value (`jpeg`/`png`/`webp`) to its MIME type.
/// Venice defaults to `webp` when the format is unset, so an absent value maps to `image/webp`.
fn image_format_to_mime_type(format: Option<&str>) -> mxlink::mime::Mime {
match format.unwrap_or("webp") {
"jpeg" | "jpg" => mxlink::mime::IMAGE_JPEG,
"png" => mxlink::mime::IMAGE_PNG,
// No mxlink::mime constant for webp; parse it, falling back to PNG on any surprise value.
_ => "image/webp".parse().unwrap_or(mxlink::mime::IMAGE_PNG),
}
}

View File

@@ -0,0 +1,48 @@
mod audio;
mod chat;
mod config;
mod controller;
mod images;
mod recovery;
mod utils;
mod wire;
#[cfg(test)]
mod tests;
pub use config::Config;
pub use controller::Controller;
use super::super::AgentInstantiationError;
use super::super::AgentInstantiationResult;
use super::ConfigTrait;
use super::controller::ControllerType;
pub fn create_controller_from_yaml_value_config(
agent_id: &str,
config: serde_yaml_ng::Value,
) -> AgentInstantiationResult<ControllerType> {
let config = match &config {
serde_yaml_ng::Value::Mapping(_) => {
let config: Config =
serde_yaml_ng::from_value(config).map_err(AgentInstantiationError::Yaml)?;
config
.validate()
.map_err(AgentInstantiationError::ConfigFailsValidation)?;
config
}
_ => {
return Err(AgentInstantiationError::ConfigForAgentIsNotAMapping(
agent_id.to_owned(),
));
}
};
Ok(ControllerType::Venice(Box::new(Controller::new(config))))
}
pub fn default_config() -> Config {
Config::default()
}

View File

@@ -0,0 +1,312 @@
//! Auto-recovery for Venice's strict request bodies.
//!
//! Venice's `/chat/completions` body is `additionalProperties: false`, so a model that does not
//! support an optional knob rejects the whole request with a 400 instead of ignoring the field.
//! Some knobs are documented as model-specific ("for supported models") and are pure
//! optimization/tuning hints: dropping them changes nothing about the answer, only loses the
//! optimization. When such a field is the reason for a 400, we strip it and retry, then remember
//! the rejection per model so later requests skip the field (and the wasted round-trip) entirely.
//!
//! Universal sampling knobs (`temperature`, `top_p`, the penalties, `max_completion_tokens`) are
//! deliberately NOT recoverable here: dropping one silently changes the model's output, so a model
//! that rejects one is a real configuration problem the operator must see, not paper over.
use std::collections::{HashMap, HashSet};
use std::sync::{Arc, OnceLock, RwLock};
use regex::Regex;
use super::wire::ChatCompletionRequest;
/// Top-level chat-completion fields baibot may drop to recover from a 400. Each is documented by
/// Venice as model-specific or as a routing hint, and is meaning-preserving to omit (Venice falls
/// back to its server-side default):
/// - `prompt_cache_retention` — "extends retention ... for supported models" (cache TTL only)
/// - `reasoning_effort` — "control reasoning effort level for supported models"
/// - `prompt_cache_key` — cache-routing hint; dropping it only forfeits a cache-hit optimization
///
/// Recovery operates on TOP-LEVEL request fields only. Sub-fields inside the `venice_parameters`
/// bag (`disable_thinking`, `enable_e2ee`, `character_slug`, …) are intentionally absent: a model
/// that rejects one surfaces a clear 400 to the operator rather than being auto-stripped. Adding a
/// new bag field does not extend recovery to it; only a name listed here is droppable.
pub(super) const DROPPABLE_FIELDS: &[&str] = &[
"prompt_cache_retention",
"prompt_cache_key",
"reasoning_effort",
];
/// Per-model record of fields a Venice model has rejected as unsupported, learned at runtime from
/// 400 responses. The `Arc` is shared across `Controller` clones, so a rejection learned once is
/// seen by every clone of the same agent. The cache is process-lived only: a restart re-learns on
/// the first request to each model, which costs one extra round-trip and nothing else, so there is
/// no persistence to keep in sync with config changes.
#[derive(Debug, Clone, Default)]
pub(super) struct UnsupportedFieldsCache {
inner: Arc<RwLock<HashMap<String, HashSet<String>>>>,
}
/// Process-global registry of per-deployment caches, keyed by Venice base URL. A `Controller` is
/// rebuilt from its config on every message for dynamic (global/room-local) agents (see
/// `agent::manager::available_room_agents_by_room_config_context`), so a cache stored on the
/// controller would be reset each message and the "process-lived" intent above would never hold.
/// The registry lets a rebuilt controller re-attach to the same cache. Keyed by base URL because
/// "which fields a model rejects" is a property of the deployment+model, not of the agent instance:
/// agents on the same Venice deployment share learnings, while a distinct deployment keeps its own,
/// so a model name that supports a field on one deployment is never proactively stripped against
/// another.
fn registry() -> &'static RwLock<HashMap<String, UnsupportedFieldsCache>> {
static REGISTRY: OnceLock<RwLock<HashMap<String, UnsupportedFieldsCache>>> = OnceLock::new();
REGISTRY.get_or_init(|| RwLock::new(HashMap::new()))
}
impl UnsupportedFieldsCache {
/// The process-global cache for `base_url`, created on first use and shared by every controller
/// for that deployment thereafter. This is what makes a learned rejection survive the
/// per-message controller rebuild. A poisoned registry lock degrades to an isolated cache:
/// correctness holds, only cross-controller sharing is lost for that call.
pub(super) fn shared_for(base_url: &str) -> Self {
if let Ok(map) = registry().read()
&& let Some(cache) = map.get(base_url)
{
return cache.clone();
}
match registry().write() {
Ok(mut map) => map.entry(base_url.to_owned()).or_default().clone(),
Err(_) => Self::default(),
}
}
/// Fields already known unsupported for `model_id`. Returns an empty set on an unknown model or
/// a poisoned lock, so a cache failure degrades to "strip nothing proactively" rather than
/// breaking the request path.
pub(super) fn known_for(&self, model_id: &str) -> HashSet<String> {
self.inner
.read()
.ok()
.and_then(|map| map.get(model_id).cloned())
.unwrap_or_default()
}
/// Records that `model_id` rejected `field`. A poisoned lock is ignored: failing to memoize only
/// means the next request re-discovers the rejection, never a wrong result.
pub(super) fn record(&self, model_id: &str, field: &str) {
if let Ok(mut map) = self.inner.write() {
map.entry(model_id.to_owned())
.or_default()
.insert(field.to_owned());
}
}
}
/// Parses the offending field name out of a Venice 400 body. A field the schema does not allow is
/// reported as `... field: 'prompt_cache_retention', value: '...'`, so this matches the `field: '..'`
/// marker wherever it sits in the message. Returns `None` when the body carries no such marker (a
/// different 400 class, e.g. a missing required field), so the caller surfaces that error instead.
///
/// Returns only the FIRST `field: '..'` match by design. If Venice ever names several rejected
/// fields in one body, the retry loop strips this one, retries, and rediscovers the next on the
/// following 400 — bounded and correct. Do not switch to `captures_iter` to "batch" them without
/// re-checking the loop's per-field termination bound in `chat.rs`.
pub(super) fn parse_rejected_field(body: &str) -> Option<String> {
static RE: OnceLock<Regex> = OnceLock::new();
let re =
RE.get_or_init(|| Regex::new(r"field: '([^']+)'").expect("rejected-field regex is valid"));
re.captures(body)
.map(|caps| caps[1].to_owned())
.filter(|field| !field.is_empty())
}
/// Clears `field` from the request when it is one baibot may safely drop and it is currently set.
/// Returns `true` only when a value was actually removed, which is what bounds the retry loop: once
/// a field is `None`, a repeat rejection for the same name returns `false` and the caller stops
/// instead of retrying forever. A field outside [`DROPPABLE_FIELDS`] always returns `false`, so a
/// meaning-bearing knob is never silently dropped.
pub(super) fn strip_droppable_field(request: &mut ChatCompletionRequest, field: &str) -> bool {
if !DROPPABLE_FIELDS.contains(&field) {
return false;
}
match field {
"prompt_cache_retention" => request.prompt_cache_retention.take().is_some(),
"prompt_cache_key" => request.prompt_cache_key.take().is_some(),
"reasoning_effort" => request.reasoning_effort.take().is_some(),
_ => false,
}
}
/// Pulls a human-readable message out of a Venice error body for surfacing in the room. Venice's
/// usual envelope is `{"error": "..."}`; some OpenAI-compatible paths nest `{"error": {"message":
/// "..."}}`. Falls back to the trimmed raw body (length-capped so a large body cannot flood the
/// room) and finally to a fixed string for an empty body, so the caller always has something to
/// show.
pub(super) fn extract_error_message(body: &str) -> String {
let trimmed = body.trim();
if trimmed.is_empty() {
return "no response body".to_owned();
}
if let Ok(value) = serde_json::from_str::<serde_json::Value>(trimmed) {
if let Some(msg) = value.get("error").and_then(|e| e.as_str()) {
return msg.to_owned();
}
if let Some(msg) = value
.get("error")
.and_then(|e| e.get("message"))
.and_then(|m| m.as_str())
{
return msg.to_owned();
}
}
const MAX: usize = 500;
if trimmed.chars().count() > MAX {
trimmed.chars().take(MAX).collect::<String>() + "…"
} else {
trimmed.to_owned()
}
}
#[cfg(test)]
mod tests {
use super::*;
fn full_request() -> ChatCompletionRequest {
ChatCompletionRequest {
model: "venice-uncensored".to_owned(),
messages: vec![],
temperature: Some(0.7),
max_completion_tokens: Some(1024),
top_p: None,
frequency_penalty: None,
presence_penalty: None,
repetition_penalty: None,
reasoning_effort: Some("high".to_owned()),
prompt_cache_key: Some("cafef00d".to_owned()),
prompt_cache_retention: Some("24h".to_owned()),
venice_parameters: None,
}
}
#[test]
fn parses_the_rejected_field_from_a_real_venice_body() {
let body = r#"{"error":"Extra inputs are not permitted, field: 'prompt_cache_retention', value: 'default'","request_id":"qM_DmKSXKF07wRxmQJ-hc"}"#;
assert_eq!(
parse_rejected_field(body).as_deref(),
Some("prompt_cache_retention")
);
}
#[test]
fn returns_no_field_when_the_body_has_no_field_marker() {
// A different 400 class (e.g. a genuinely malformed request) carries no `field: '..'`
// marker, so there is nothing to strip and the caller must surface the error instead.
assert_eq!(parse_rejected_field(r#"{"error":"Invalid request"}"#), None);
assert_eq!(parse_rejected_field(""), None);
}
#[test]
fn strips_a_droppable_field_once_then_reports_no_progress() {
let mut request = full_request();
// First strip clears the field and reports progress, so the caller retries.
assert!(strip_droppable_field(
&mut request,
"prompt_cache_retention"
));
assert!(request.prompt_cache_retention.is_none());
// A repeat rejection for the same (now absent) field reports no progress: this is what
// stops the retry loop instead of spinning forever.
assert!(!strip_droppable_field(
&mut request,
"prompt_cache_retention"
));
}
#[test]
fn refuses_to_strip_a_meaning_bearing_field() {
let mut request = full_request();
// `temperature` is universal and changes the output; a rejection for it must surface, never
// be silently dropped. The whole droppable set is the only thing strip will touch.
assert!(!strip_droppable_field(&mut request, "temperature"));
assert_eq!(request.temperature, Some(0.7));
assert!(!strip_droppable_field(
&mut request,
"max_completion_tokens"
));
assert_eq!(request.max_completion_tokens, Some(1024));
for field in DROPPABLE_FIELDS {
assert!(
strip_droppable_field(&mut full_request(), field),
"every advertised droppable field must actually be strippable: {field}"
);
}
}
#[test]
fn cache_records_per_model_and_isolates_models() {
let cache = UnsupportedFieldsCache::default();
assert!(cache.known_for("venice-uncensored").is_empty());
cache.record("venice-uncensored", "prompt_cache_retention");
cache.record("venice-uncensored", "reasoning_effort");
let known = cache.known_for("venice-uncensored");
assert!(known.contains("prompt_cache_retention"));
assert!(known.contains("reasoning_effort"));
// A rejection learned for one model must not leak to another.
assert!(cache.known_for("kimi-k2-5").is_empty());
}
#[test]
fn shared_for_survives_controller_rebuild_and_isolates_deployments() {
// Distinct, test-only base URLs so the process-global registry can't collide with another
// test running in parallel.
let url_a = "https://shared-for-test-a.invalid/api/v1";
let url_b = "https://shared-for-test-b.invalid/api/v1";
// A controller rebuilt for the same deployment (a fresh `shared_for` call, as happens per
// message for dynamic agents) re-attaches to the SAME cache, so an earlier learning holds.
let first = UnsupportedFieldsCache::shared_for(url_a);
first.record("model-x", "reasoning_effort");
let rebuilt = UnsupportedFieldsCache::shared_for(url_a);
assert!(
rebuilt.known_for("model-x").contains("reasoning_effort"),
"a rebuilt controller for the same deployment must see the earlier rejection"
);
// A different deployment keeps its own learnings: a model name that rejects a field on one
// Venice must not silence/strip it on another.
let other_deployment = UnsupportedFieldsCache::shared_for(url_b);
assert!(
other_deployment.known_for("model-x").is_empty(),
"rejections must not leak across deployments"
);
}
#[test]
fn extracts_a_human_message_from_error_envelopes() {
assert_eq!(
extract_error_message(r#"{"error":"Extra inputs are not permitted","request_id":"x"}"#),
"Extra inputs are not permitted"
);
// OpenAI-style nested envelope.
assert_eq!(
extract_error_message(r#"{"error":{"message":"context length exceeded"}}"#),
"context length exceeded"
);
// Unknown shape falls back to the raw body; empty falls back to a fixed string.
assert_eq!(
extract_error_message("plain text failure"),
"plain text failure"
);
assert_eq!(extract_error_message(" "), "no response body");
}
}

View File

@@ -0,0 +1,549 @@
use mxlink::matrix_sdk::ruma::OwnedMxcUri;
use mxlink::matrix_sdk::ruma::events::room::message::{
FileMessageEventContent, ImageMessageEventContent,
};
use mxlink::mime;
use super::super::ControllerTrait;
use crate::agent::AgentPurpose;
use crate::conversation::llm::{
Author as LLMAuthor, FileDetails, ImageDetails, Message as LLMMessage,
MessageContent as LLMMessageContent,
};
use super::chat::{append_reasoning, derive_prompt_cache_key, render_with_citations};
use super::config::{Config, TextGenerationConfig, VeniceParameters, WebSearchMode};
use super::controller::Controller;
use super::utils::convert_llm_messages_to_venice;
use super::wire::{
ChatCompletionRequest, ContentPart, EditImageRequest, GenerateImageRequest, MessageContent,
SpeechRequest, WebSearchCitation,
};
#[test]
fn config_round_trips_with_venice_parameters() {
let yaml = r#"
base_url: https://api.venice.ai/api/v1
api_key: test-key
text_generation:
model_id: kimi-k2-5
temperature: 0.7
max_response_tokens: 1024
max_context_tokens: 65536
venice_parameters:
enable_web_search: "auto"
enable_web_citations: true
speech_to_text:
model_id: nvidia/parakeet-tdt-0.6b-v3
"#;
let config: Config = serde_yaml_ng::from_str(yaml).expect("config should deserialize");
let tg = config.text_generation.expect("text_generation present");
let vp = tg.venice_parameters.expect("venice_parameters present");
assert!(matches!(vp.enable_web_search, Some(WebSearchMode::Auto)));
assert_eq!(vp.enable_web_citations, Some(true));
assert_eq!(vp.character_slug, None);
// The bag must serialize the enum to the exact wire string, and an unset knob must be ABSENT
// (not `null`) so the strict `additionalProperties: false` body is honored.
let json = serde_json::to_string(&vp).expect("serialize venice_parameters");
assert!(
json.contains("\"enable_web_search\":\"auto\""),
"web search should be the literal \"auto\": {json}"
);
assert!(
!json.contains("character_slug"),
"an unset knob must be omitted entirely: {json}"
);
assert!(
!json.contains("null"),
"no nulls belong in the body: {json}"
);
}
#[test]
fn converts_text_image_and_file_to_content_parts() {
let messages = vec![
LLMMessage {
author: LLMAuthor::User,
sender_id: None,
timestamp: chrono::Utc::now(),
content: LLMMessageContent::Text("describe this".to_owned()),
},
LLMMessage {
author: LLMAuthor::User,
sender_id: None,
timestamp: chrono::Utc::now(),
content: LLMMessageContent::Image(ImageDetails::new(
ImageMessageEventContent::plain(
"pic.png".to_owned(),
OwnedMxcUri::from("mxc://example.com/abc"),
),
mime::IMAGE_PNG,
vec![1, 2, 3],
)),
},
LLMMessage {
author: LLMAuthor::User,
sender_id: None,
timestamp: chrono::Utc::now(),
content: LLMMessageContent::File(FileDetails::new(
FileMessageEventContent::plain(
"doc.pdf".to_owned(),
OwnedMxcUri::from("mxc://example.com/def"),
),
mime::APPLICATION_PDF,
vec![4, 5, 6],
)),
},
];
let converted = convert_llm_messages_to_venice(messages).expect("conversion should succeed");
// Text, image, AND file all survive now: the file is no longer warn-skipped.
assert_eq!(converted.len(), 3);
match &converted[0].content {
MessageContent::Text(text) => assert_eq!(text, "describe this"),
other => panic!("expected bare text, got {other:?}"),
}
match &converted[1].content {
MessageContent::Parts(parts) => match &parts[0] {
ContentPart::ImageUrl { image_url } => assert!(
image_url.url.starts_with("data:image/png;base64,"),
"image should be inlined as a data URI: {}",
image_url.url
),
other => panic!("expected an image part, got {other:?}"),
},
other => panic!("expected image parts, got {other:?}"),
}
match &converted[2].content {
MessageContent::Parts(parts) => match &parts[0] {
ContentPart::File { file } => {
assert!(
file.file_data.starts_with("data:application/pdf;base64,"),
"file should be inlined as a data URI: {}",
file.file_data
);
assert_eq!(file.filename.as_deref(), Some("doc.pdf"));
}
other => panic!("expected a file part, got {other:?}"),
},
other => panic!("expected file parts, got {other:?}"),
}
}
#[test]
fn supports_purpose_truth_table() {
let config: Config = serde_yaml_ng::from_str(
r#"
base_url: https://api.venice.ai/api/v1
api_key: test-key
text_generation:
model_id: kimi-k2-5
speech_to_text:
model_id: nvidia/parakeet-tdt-0.6b-v3
"#,
)
.expect("config should deserialize");
let controller = Controller::new(config);
assert!(controller.supports_purpose(AgentPurpose::TextGeneration));
assert!(controller.supports_purpose(AgentPurpose::SpeechToText));
assert!(controller.supports_purpose(AgentPurpose::CatchAll));
assert!(!controller.supports_purpose(AgentPurpose::TextToSpeech));
assert!(!controller.supports_purpose(AgentPurpose::ImageGeneration));
}
#[test]
fn supports_purpose_true_when_image_and_tts_blocks_present() {
let config: Config = serde_yaml_ng::from_str(
r#"
base_url: https://api.venice.ai/api/v1
api_key: test-key
text_to_speech:
model_id: tts-kokoro
image_generation:
model_id: chroma
"#,
)
.expect("config should deserialize");
let controller = Controller::new(config);
assert!(controller.supports_purpose(AgentPurpose::TextToSpeech));
assert!(controller.supports_purpose(AgentPurpose::ImageGeneration));
}
#[test]
fn speech_request_serializes_voice_and_omits_unset() {
let request = SpeechRequest {
model: "tts-kokoro".to_owned(),
input: "hello".to_owned(),
voice: Some("af_sky".to_owned()),
speed: None,
response_format: Some("mp3".to_owned()),
prompt: None,
temperature: None,
top_p: None,
};
let json = serde_json::to_string(&request).expect("serialize SpeechRequest");
assert!(
json.contains("\"voice\":\"af_sky\""),
"voice should be present: {json}"
);
assert!(
!json.contains("temperature"),
"an unset knob must be omitted (not null): {json}"
);
assert!(
!json.contains("null"),
"no nulls belong in the body: {json}"
);
}
#[test]
fn generate_image_request_pins_flags_and_omits_unset() {
let request = GenerateImageRequest {
model: "chroma".to_owned(),
prompt: "a cat".to_owned(),
return_binary: false,
variants: 1,
negative_prompt: None,
cfg_scale: None,
steps: None,
style_preset: None,
seed: None,
safe_mode: None,
hide_watermark: None,
format: None,
width: None,
height: None,
aspect_ratio: None,
resolution: None,
quality: None,
lora_strength: None,
embed_exif_metadata: None,
enable_web_search: None,
};
let json = serde_json::to_string(&request).expect("serialize GenerateImageRequest");
assert!(json.contains("\"model\":\"chroma\""), "{json}");
assert!(
json.contains("\"return_binary\":false"),
"return_binary must be pinned false: {json}"
);
assert!(
json.contains("\"variants\":1"),
"variants must be pinned 1: {json}"
);
assert!(
!json.contains("cfg_scale"),
"an unset knob must be omitted: {json}"
);
assert!(
!json.contains("null"),
"no nulls belong in the body: {json}"
);
}
#[test]
fn edit_image_request_carries_model_and_base64_image() {
let request = EditImageRequest {
model: "firered-image-edit".to_owned(),
prompt: "make it a sunrise".to_owned(),
image: "aGVsbG8=".to_owned(),
output_format: None,
aspect_ratio: None,
resolution: None,
safe_mode: None,
};
let json = serde_json::to_string(&request).expect("serialize EditImageRequest");
assert!(json.contains("\"model\":\"firered-image-edit\""), "{json}");
assert!(
json.contains("\"image\":\"aGVsbG8=\""),
"the base64 image string must be present: {json}"
);
assert!(
!json.contains("output_format"),
"an unset knob must be omitted: {json}"
);
}
#[test]
fn web_search_mode_off_deserializes_from_bare_yaml_off() {
// `off` is a YAML-1.1 boolean but a plain string under serde_yaml_ng's YAML-1.2 core schema,
// so it deserializes straight into the lowercase `WebSearchMode::Off`. This pins that the
// sample config and docs can use the bare, unquoted `off` without it parsing as a boolean.
let params: VeniceParameters =
serde_yaml_ng::from_str("enable_web_search: off").expect("bare `off` should deserialize");
assert!(matches!(params.enable_web_search, Some(WebSearchMode::Off)));
}
#[test]
fn request_places_sampling_top_level_and_verbosity_in_the_bag() {
// The whole config-shape decision in one assertion: top-level knobs serialize at the top
// level, the dual-position `verbosity` serializes inside the bag. Venice silently ignores a
// top-level knob misplaced into the bag, so this is the guard against a silent no-op.
let request = ChatCompletionRequest {
model: "kimi-k2-5".to_owned(),
messages: vec![],
temperature: Some(0.5),
max_completion_tokens: Some(1024),
top_p: Some(0.5),
frequency_penalty: None,
presence_penalty: None,
repetition_penalty: None,
reasoning_effort: Some("high".to_owned()),
prompt_cache_key: Some("00000000cafef00d".to_owned()),
prompt_cache_retention: Some("24h".to_owned()),
venice_parameters: Some(VeniceParameters {
verbosity: Some("high".to_owned()),
..Default::default()
}),
};
let json = serde_json::to_value(&request).expect("serialize request");
assert_eq!(json["top_p"], 0.5);
assert_eq!(json["reasoning_effort"], "high");
assert_eq!(json["prompt_cache_retention"], "24h");
assert_eq!(json["prompt_cache_key"], "00000000cafef00d");
assert!(
json.get("verbosity").is_none(),
"verbosity must not be a top-level field: {json}"
);
assert_eq!(json["venice_parameters"]["verbosity"], "high");
assert!(
json["venice_parameters"].get("top_p").is_none(),
"top_p must not be inside the bag: {json}"
);
}
#[test]
fn config_defaults_prompt_cache_retention_to_24h() {
// The programmatic default.
assert_eq!(
TextGenerationConfig::default()
.prompt_cache_retention
.as_deref(),
Some("24h")
);
// A config that omits the key must ALSO default to 24h, via the named serde default. A bare
// `#[serde(default)]` would yield None here and silently disable caching for such configs.
let tg: TextGenerationConfig = serde_yaml_ng::from_str("model_id: kimi-k2-5\n")
.expect("minimal config should deserialize");
assert_eq!(
tg.prompt_cache_retention.as_deref(),
Some("24h"),
"an omitted retention key must still default to 24h"
);
}
#[test]
fn cache_key_is_stable_for_same_inputs_and_varies_otherwise() {
let key = derive_prompt_cache_key("system prompt", "2024-09-20 (Friday), 18:34:15 UTC");
// Identical inputs produce an identical key: this is what keeps turn 5 routing to the warm
// server holding turns 1-4 (and what survives a process restart).
assert_eq!(
key,
derive_prompt_cache_key("system prompt", "2024-09-20 (Friday), 18:34:15 UTC"),
"identical inputs must produce an identical key"
);
assert_eq!(key.len(), 16, "the key is a 16-char hex string");
assert!(key.chars().all(|c| c.is_ascii_hexdigit()));
// A different conversation start time or a different prompt must change the key.
assert_ne!(
key,
derive_prompt_cache_key("system prompt", "2024-09-21 (Saturday), 09:00:00 UTC"),
"a different start time must change the key"
);
assert_ne!(
key,
derive_prompt_cache_key("other prompt", "2024-09-20 (Friday), 18:34:15 UTC"),
"a different prompt must change the key"
);
}
#[test]
fn citations_render_inline_refs_and_a_sources_block() {
let citations = vec![WebSearchCitation {
title: "Example Source".to_owned(),
url: "https://example.com/a".to_owned(),
}];
let rendered = render_with_citations("the sky is blue^1^".to_owned(), &citations);
assert!(
rendered.contains("the sky is blue[1]"),
"inline ^1^ becomes [1]: {rendered}"
);
assert!(
rendered.contains("Sources:"),
"a Sources block is appended: {rendered}"
);
assert!(
rendered.contains("[1] [Example Source](https://example.com/a)"),
"the source renders as a markdown link: {rendered}"
);
}
#[test]
fn citations_absent_leaves_content_untouched() {
let content = "plain answer, no web search".to_owned();
assert_eq!(render_with_citations(content.clone(), &[]), content);
}
#[test]
fn citation_title_and_url_cannot_inject_markdown() {
// A hostile page sets its title to break out of the link label and its URL to a non-http
// scheme. Neither may produce a spoofed clickable link in the room.
let citations = vec![WebSearchCitation {
title: "evil](http://phish.example) take".to_owned(),
url: "javascript:alert(1)".to_owned(),
}];
let rendered = render_with_citations("result^1^".to_owned(), &citations);
assert!(
rendered.contains("evil\\](http://phish.example) take"),
"the title's brackets must be escaped so it cannot close the link label: {rendered}"
);
assert!(
!rendered.contains("(javascript:alert(1))"),
"a non-http(s) URL must never become a markdown link target: {rendered}"
);
}
#[test]
fn chained_and_comma_citation_runs_each_expand_to_separate_refs() {
let citations = vec![
WebSearchCitation {
title: "One".to_owned(),
url: "https://example.com/1".to_owned(),
},
WebSearchCitation {
title: "Two".to_owned(),
url: "https://example.com/2".to_owned(),
},
WebSearchCitation {
title: "Three".to_owned(),
url: "https://example.com/3".to_owned(),
},
];
// Caret-chained run: Venice shares the caret between consecutive citations (`^2^3^`). The whole
// run must expand, not just the first, with no orphaned `3^` left behind.
let chained = render_with_citations("alpha^2^3^ and beta^1^".to_owned(), &citations);
assert!(
chained.contains("alpha[2][3] and beta[1]"),
"a chained ^2^3^ run must expand to [2][3] with no orphaned caret: {chained}"
);
// Comma run.
let comma = render_with_citations("gamma^1,3^".to_owned(), &citations);
assert!(
comma.contains("gamma[1][3]"),
"a comma ^1,3^ run must expand to [1][3]: {comma}"
);
// Multi-digit citation indices survive intact.
let multidigit = render_with_citations("delta^2^10^".to_owned(), &citations);
assert!(
multidigit.contains("delta[2][10]"),
"a multi-digit chained run must expand to [2][10]: {multidigit}"
);
}
#[test]
fn malformed_citation_degrades_instead_of_failing() {
// A citation arriving without a `url` must still deserialize (to an empty default) rather than
// failing the whole response parse and losing an otherwise-good answer.
let parsed: WebSearchCitation = serde_json::from_str(r#"{"title":"Only a title"}"#)
.expect("a citation missing `url` should still deserialize");
assert_eq!(parsed.url, "");
// Rendering citations with missing fields stays graceful: no empty `[]()` link, no panic.
let citations = vec![
WebSearchCitation {
title: String::new(),
url: "https://example.com/u".to_owned(),
},
WebSearchCitation {
title: String::new(),
url: String::new(),
},
];
let rendered = render_with_citations("answer^1^2^".to_owned(), &citations);
assert!(
rendered.contains("[1] [https://example.com/u](https://example.com/u)"),
"a citation with no title falls back to the URL as link text: {rendered}"
);
assert!(
rendered.contains("[2] (source unavailable)"),
"a citation with neither title nor URL renders a placeholder: {rendered}"
);
}
#[test]
fn reasoning_is_appended_only_when_show_reasoning_is_set() {
let base = "the answer".to_owned();
// Off (the default): thinking is dropped, never reaching the room.
let off = append_reasoning(base.clone(), Some("secret thinking".to_owned()), false);
assert_eq!(off, "the answer");
// On: thinking is appended below the answer in a collapsible <details> block (folded by
// default, expandable in clients that support it).
let on = append_reasoning(base.clone(), Some(" visible thinking ".to_owned()), true);
assert!(on.starts_with("the answer"));
assert!(on.contains("<details><summary>💭 Reasoning</summary>"));
assert!(on.contains("</details>"));
// The reasoning sits as its own markdown block (blank lines around it) and is trimmed.
assert!(on.contains("\n\nvisible thinking\n\n"));
// On but empty or missing reasoning: nothing is appended.
assert_eq!(
append_reasoning(base.clone(), Some(" ".to_owned()), true),
"the answer"
);
assert_eq!(append_reasoning(base.clone(), None, true), "the answer");
}
#[test]
fn oversized_file_is_rejected() {
let messages = vec![LLMMessage {
author: LLMAuthor::User,
sender_id: None,
timestamp: chrono::Utc::now(),
content: LLMMessageContent::File(FileDetails::new(
FileMessageEventContent::plain(
"big.pdf".to_owned(),
OwnedMxcUri::from("mxc://example.com/big"),
),
mime::APPLICATION_PDF,
vec![0u8; 25 * 1024 * 1024 + 1],
)),
}];
assert!(
convert_llm_messages_to_venice(messages).is_err(),
"a file over the 25MB limit must be rejected"
);
}

View File

@@ -0,0 +1,81 @@
use crate::conversation::llm::{
Author as LLMAuthor, Message as LLMMessage, MessageContent as LLMMessageContent,
};
use crate::utils::base64::base64_encode;
use super::wire::{ChatMessage, ContentPart, FilePart, ImageUrl, MessageContent};
/// Venice's documented file-input ceiling is 25MB on the decoded bytes (swagger `file_data`).
/// We check it here so an oversized file gets a clear message instead of an opaque 413 from the
/// API; the 413 status branch in `chat.rs` is the backstop if a file slips past this guard.
const MAX_FILE_BYTES: usize = 25 * 1024 * 1024;
pub fn convert_llm_messages_to_venice(
messages: Vec<LLMMessage>,
) -> anyhow::Result<Vec<ChatMessage>> {
let mut venice_messages: Vec<ChatMessage> = Vec::with_capacity(messages.len());
for message in messages {
venice_messages.push(convert_llm_message_to_venice(message)?);
}
Ok(venice_messages)
}
fn convert_llm_message_to_venice(message: LLMMessage) -> anyhow::Result<ChatMessage> {
let role = match message.author {
LLMAuthor::Prompt => "system",
LLMAuthor::Assistant => "assistant",
LLMAuthor::User => "user",
};
match message.content {
LLMMessageContent::Text(text) => Ok(ChatMessage {
role: role.to_owned(),
content: MessageContent::Text(text),
}),
LLMMessageContent::Image(image_details) => {
// Inline the image as a base64 data URI, the same shape the OpenAI vision content
// part uses. This is the gap the openai_compat provider can't fill (it drops images).
let data_uri = format!(
"data:{};base64,{}",
image_details.mime,
base64_encode(&image_details.data)
);
Ok(ChatMessage {
role: role.to_owned(),
content: MessageContent::Parts(vec![ContentPart::ImageUrl {
image_url: ImageUrl { url: data_uri },
}]),
})
}
LLMMessageContent::File(file_details) => {
// Inline the file as a base64 data URI in a `file` content part. This is the input
// type the openai_compat provider drops; baibot already extracts the bytes upstream.
// The message reaches the room, so it carries no user-controlled filename: a crafted
// name could otherwise inject markdown (a spoofed link) into the bot's reply.
if file_details.data.len() > MAX_FILE_BYTES {
return Err(anyhow::anyhow!(
"The attached file is too large for Venice (the limit is 25MB)."
));
}
let data_uri = format!(
"data:{};base64,{}",
file_details.mime,
base64_encode(&file_details.data)
);
Ok(ChatMessage {
role: role.to_owned(),
content: MessageContent::Parts(vec![ContentPart::File {
file: FilePart {
file_data: data_uri,
filename: Some(file_details.filename()),
},
}]),
})
}
}
}

View File

@@ -0,0 +1,273 @@
//! Serde structs modeling Venice's `/chat/completions`, `/audio/transcriptions`,
//! `/audio/speech`, `/image/generate`, and `/image/edit` wire shapes. Request types are
//! `Serialize`-only (we build them, Venice never sends them back); response types are
//! `Deserialize`-only. Keeping the split means the untagged request content enum is never on a
//! deserialize path, so a surprise response shape can't fail to match it.
//!
//! Field names match Venice's schema 1:1 (so the config's `model_id` becomes `model` here). Every
//! request body is `additionalProperties: false`, so optional knobs carry `skip_serializing_if`
//! to omit rather than send `null`. `/audio/speech` and `/image/edit` return raw binary (no
//! response struct); only `/image/generate` returns JSON (`GenerateImageResponse`).
use serde::{Deserialize, Serialize};
use super::config::VeniceParameters;
#[derive(Debug, Serialize)]
pub struct ChatCompletionRequest {
pub model: String,
pub messages: Vec<ChatMessage>,
#[serde(skip_serializing_if = "Option::is_none")]
pub temperature: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub max_completion_tokens: Option<u32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub top_p: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub frequency_penalty: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub presence_penalty: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub repetition_penalty: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub reasoning_effort: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub prompt_cache_key: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub prompt_cache_retention: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub venice_parameters: Option<VeniceParameters>,
}
#[derive(Debug, Serialize)]
pub struct ChatMessage {
pub role: String,
pub content: MessageContent,
}
/// A message body is either a bare string or a list of content parts. Venice accepts both; we
/// send the parts form when a message carries an image or a file (baibot keeps text, images, and
/// files in separate messages, so a parts list holds a single image part or file part).
#[derive(Debug, Serialize)]
#[serde(untagged)]
pub enum MessageContent {
Text(String),
Parts(Vec<ContentPart>),
}
#[derive(Debug, Serialize)]
#[serde(tag = "type", rename_all = "snake_case")]
pub enum ContentPart {
ImageUrl { image_url: ImageUrl },
File { file: FilePart },
}
#[derive(Debug, Serialize)]
pub struct ImageUrl {
/// A `data:<mime>;base64,<data>` URI for inline images.
pub url: String,
}
#[derive(Debug, Serialize)]
pub struct FilePart {
/// A `data:<mime>;base64,<data>` URI carrying the file bytes inline.
pub file_data: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub filename: Option<String>,
}
/// Standard OpenAI-shaped chat completion response. We read `choices[0].message.content` and,
/// when web search is on, the structured `venice_parameters.web_search_citations` (requested via
/// `return_search_results_as_documents`) to rewrite the inline `^n^` superscripts into readable
/// `[n]` references plus a `Sources:` block. `reasoning_content` carries the model's thinking when
/// the model exposes it; it is appended only when `show_reasoning` is set.
#[derive(Debug, Deserialize)]
pub struct ChatCompletionResponse {
pub choices: Vec<ChatChoice>,
#[serde(default)]
pub venice_parameters: Option<ResponseVeniceParameters>,
}
#[derive(Debug, Deserialize)]
pub struct ChatChoice {
pub message: ResponseMessage,
}
#[derive(Debug, Deserialize)]
pub struct ResponseMessage {
#[serde(default)]
pub content: Option<String>,
#[serde(default)]
pub reasoning_content: Option<String>,
}
/// The `venice_parameters` envelope on a chat-completion *response*, distinct from the request-side
/// `VeniceParameters` bag. Only the citation list is read; other response-side fields are ignored.
#[derive(Debug, Deserialize, Default)]
pub struct ResponseVeniceParameters {
#[serde(default)]
pub web_search_citations: Vec<WebSearchCitation>,
}
/// Only the `title` and `url` are read (for rendering the `Sources:` block). Venice also returns
/// `content` and `date` per citation; serde drops them, the same way the response structs above
/// ignore the response fields baibot does not use. Both fields default to empty so a single
/// citation that arrives without one (schema drift on scraped results) degrades gracefully in the
/// rendered list instead of failing the whole response deserialization.
#[derive(Debug, Deserialize)]
pub struct WebSearchCitation {
#[serde(default)]
pub title: String,
#[serde(default)]
pub url: String,
}
/// `/audio/transcriptions` response. We read `text`; the optional `duration`/`timestamps` the
/// API can return are not used in v1.
#[derive(Debug, Deserialize)]
pub struct TranscriptionResponse {
pub text: String,
}
/// `/audio/speech` (`CreateSpeechRequestSchema`) request. `input` and `model` are always sent;
/// the rest are omitted when unset. The response is raw binary audio, so there is no response
/// struct.
#[derive(Debug, Serialize)]
pub struct SpeechRequest {
pub model: String,
pub input: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub voice: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub speed: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub response_format: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub prompt: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub temperature: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub top_p: Option<f32>,
}
/// `/image/generate` (`GenerateImageRequest`) request. `return_binary` is pinned `false` and
/// `variants` to `1` by the builder: baibot wants exactly one image returned as base64-in-JSON,
/// which `GenerateImageResponse` then decodes. Flipping `return_binary` would make Venice answer
/// with raw binary and break that JSON decode, so it is not configurable.
#[derive(Debug, Serialize)]
pub struct GenerateImageRequest {
pub model: String,
pub prompt: String,
pub return_binary: bool,
pub variants: u32,
#[serde(skip_serializing_if = "Option::is_none")]
pub negative_prompt: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub cfg_scale: Option<f32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub steps: Option<u32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub style_preset: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub seed: Option<i64>,
#[serde(skip_serializing_if = "Option::is_none")]
pub safe_mode: Option<bool>,
#[serde(skip_serializing_if = "Option::is_none")]
pub hide_watermark: Option<bool>,
#[serde(skip_serializing_if = "Option::is_none")]
pub format: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub width: Option<u32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub height: Option<u32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub aspect_ratio: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub resolution: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub quality: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub lora_strength: Option<u32>,
#[serde(skip_serializing_if = "Option::is_none")]
pub embed_exif_metadata: Option<bool>,
#[serde(skip_serializing_if = "Option::is_none")]
pub enable_web_search: Option<bool>,
}
/// `/image/generate` response when `return_binary` is false: a JSON envelope carrying the images
/// as base64 strings. We read `images[0]`; `request`/`timing` and other fields are ignored. `id`
/// is telemetry only (logged, never used for correctness), so it is optional: a response that
/// carries usable `images` must not fail to deserialize just because the telemetry field drifted.
#[derive(Debug, Deserialize)]
pub struct GenerateImageResponse {
#[serde(default)]
pub id: Option<String>,
pub images: Vec<String>,
}
/// `/image/edit` (`EditImageRequest`) request. The source `image` is a base64-encoded string
/// (Venice's `image` field is `anyOf` upload/base64/URL; we send base64-in-JSON, no multipart).
/// The response is raw binary, so there is no response struct.
#[derive(Debug, Serialize)]
pub struct EditImageRequest {
pub model: String,
pub prompt: String,
/// Base64-encoded source image bytes.
pub image: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub output_format: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub aspect_ratio: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub resolution: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub safe_mode: Option<bool>,
}

View File

@@ -1,5 +1,3 @@
use base64::{Engine as _, engine::general_purpose::STANDARD};
use crate::{
agent::{
AgentInstance, AgentPurpose, ControllerTrait, Manager as AgentManager, PublicIdentifier,
@@ -140,7 +138,3 @@ async fn get_global_agent_id_for_purpose(
.handler
.get_by_purpose_with_catch_all_fallback(purpose)
}
pub(crate) fn base64_decode(base64_string: &str) -> Result<Vec<u8>, base64::DecodeError> {
STANDARD.decode(base64_string)
}

View File

@@ -1,8 +1,10 @@
use std::fs;
use std::sync::Arc;
use std::{future::Future, pin::Pin};
use mxlink::matrix_sdk::Room;
use mxlink::matrix_sdk::media::{MediaFormat, MediaRequestParameters};
use mxlink::matrix_sdk::ruma::api::client::profile::{AvatarUrl, DisplayName};
use mxlink::matrix_sdk::ruma::{
MilliSecondsSinceUnixEpoch, OwnedUserId, events::room::MediaSource,
};
@@ -17,12 +19,13 @@ use mxlink::helpers::account_data_config::{
RoomConfigManager as AccountDataRoomConfigManager,
};
use mxlink::helpers::encryption::Manager as EncryptionManager;
use mxlink::mime::Mime;
use crate::agent::Manager as AgentManager;
use crate::entity::catch_up_marker::{
CatchUpMarker, CatchUpMarkerManager, DelayedCatchUpMarkerManager,
};
use crate::entity::cfg::Config;
use crate::entity::cfg::{Avatar, Config, ConfigUserAuth};
use crate::entity::globalconfig::{GlobalConfig, GlobalConfigurationManager};
use crate::entity::roomconfig::{RoomConfig, RoomConfigurationManager};
@@ -285,24 +288,27 @@ impl Bot {
async fn do_prepare_profile(&self) -> anyhow::Result<()> {
tracing::debug!("Preparing profile..");
let desired_display_name = self.inner.config.user.name.clone();
let account = self.inner.matrix_link.client().account();
let media = self.inner.matrix_link.client().media();
let desired_display_name = self.inner.config.user.name.clone();
let profile = account
.fetch_user_profile()
.await
.map_err(|e| anyhow::anyhow!("Failed fetching profile: {:?}", e))?;
let should_update_display_name = match &profile.displayname {
let current_display_name = profile.get_static::<DisplayName>()?;
let current_avatar_url = profile.get_static::<AvatarUrl>()?;
let should_update_display_name = match &current_display_name {
Some(displayname) => displayname != &desired_display_name,
None => true,
};
if should_update_display_name {
tracing::info!(
?profile.displayname,
?current_display_name,
?desired_display_name,
"Updating display name.."
);
@@ -312,34 +318,72 @@ impl Bot {
}
}
let should_update_avatar = match &profile.avatar_url {
Some(avatar_url) => {
let request = MediaRequestParameters {
source: MediaSource::Plain(avatar_url.to_owned()),
format: MediaFormat::File,
};
let content = media
.get_media_content(&request, true)
.await
.map_err(|e| anyhow::anyhow!("Failed fetching existing avatar: {:?}", e))?;
content.as_slice() != LOGO_BYTES
let desired_avatar: Option<(Vec<u8>, Mime)> = match &self.inner.config.user.avatar {
Avatar::Keep => {
tracing::info!("Avatar configured to keep current, skipping avatar management");
None
}
Avatar::Default => {
tracing::info!("Avatar configured to use default");
Some((
LOGO_BYTES.to_vec(),
LOGO_MIME_TYPE
.parse()
.expect("Failed parsing mime type for logo"),
))
}
Avatar::Custom(avatar_path) => {
tracing::info!(?avatar_path, "Avatar configured to use custom path");
let bytes = fs::read(avatar_path).map_err(|e| {
anyhow::anyhow!("Failed reading avatar from {:?}: {:?}", avatar_path, e)
})?;
let mime = mime_guess::from_path(avatar_path).first_or_octet_stream();
tracing::debug!(?mime, bytes_len = bytes.len(), "Loaded custom avatar");
Some((bytes, mime))
}
None => true,
};
if should_update_avatar {
tracing::info!("Updating avatar..");
if let Some((desired_bytes, mime_type)) = desired_avatar {
let should_update_avatar = match &current_avatar_url {
Some(avatar_url) => {
tracing::debug!(?avatar_url, "Fetching current avatar to compare");
let request = MediaRequestParameters {
source: MediaSource::Plain(avatar_url.to_owned()),
format: MediaFormat::File,
};
let mime_type = LOGO_MIME_TYPE
.parse()
.expect("Failed parsing mime type for logo");
let content = media
.get_media_content(&request, true)
.await
.map_err(|e| anyhow::anyhow!("Failed fetching existing avatar: {:?}", e))?;
account
.upload_avatar(&mime_type, LOGO_BYTES.to_vec())
.await
.map_err(|e| anyhow::anyhow!("Failed uploading avatar: {:?}", e))?;
let needs_update = content.as_slice() != desired_bytes;
tracing::debug!(
current_bytes_len = content.len(),
desired_bytes_len = desired_bytes.len(),
?needs_update,
"Compared current and desired avatar"
);
needs_update
}
None => {
tracing::debug!("No current avatar set, will upload");
true
}
};
if should_update_avatar {
tracing::info!("Updating avatar..");
account
.upload_avatar(&mime_type, desired_bytes)
.await
.map_err(|e| anyhow::anyhow!("Failed uploading avatar: {:?}", e))?;
tracing::info!("Avatar updated successfully");
} else {
tracing::debug!("Avatar already up to date, skipping upload");
}
}
Ok(())
@@ -351,10 +395,22 @@ async fn create_matrix_link(config: &Config) -> anyhow::Result<MatrixLink> {
let session_encryption_key = config.persistence.session_encryption_key()?;
let db_dir_path: std::path::PathBuf = config.persistence.db_dir_path()?;
let login_creds = LoginCredentials::UserPassword(
config.user.mxid_localpart.to_owned(),
config.user.password.to_owned(),
);
let user_auth = config.user.auth_config(&config.homeserver.server_name)?;
let login_creds = match user_auth {
ConfigUserAuth::UserPassword { username, password } => {
LoginCredentials::UserPassword(username, password)
}
ConfigUserAuth::AccessToken {
user_id,
device_id,
access_token,
} => LoginCredentials::AccessToken {
user_id,
device_id,
access_token,
},
};
let login_encryption = LoginEncryption::new(
config.user.encryption.recovery_passphrase.clone(),

View File

@@ -5,7 +5,7 @@ use anyhow::anyhow;
use crate::agent::AgentPurpose;
pub use crate::entity::cfg::{Config, defaults as cfg_defaults, env as cfg_env};
pub use crate::entity::cfg::{Avatar, Config, defaults as cfg_defaults, env as cfg_env};
pub fn load() -> anyhow::Result<Config> {
let config_file_path = env::var(cfg_env::BAIBOT_CONFIG_FILE_PATH)
@@ -21,7 +21,7 @@ pub fn load() -> anyhow::Result<Config> {
}
let config_str = std::fs::read_to_string(config_file_path)?;
let mut config: Config = serde_yaml::from_str(&config_str)?;
let mut config: Config = serde_yaml_ng::from_str(&config_str)?;
// Allow environment variables to override some configuration keys
for (key, value) in env::vars() {
@@ -29,11 +29,25 @@ pub fn load() -> anyhow::Result<Config> {
cfg_env::BAIBOT_HOMESERVER_SERVER_NAME => config.homeserver.server_name = value,
cfg_env::BAIBOT_HOMESERVER_URL => config.homeserver.url = value,
cfg_env::BAIBOT_USER_MXID_LOCALPART => config.user.mxid_localpart = value,
cfg_env::BAIBOT_USER_PASSWORD => config.user.password = value,
cfg_env::BAIBOT_USER_PASSWORD => {
config.user.password = optional_non_empty(value);
}
cfg_env::BAIBOT_USER_ACCESS_TOKEN => {
config.user.access_token = optional_non_empty(value);
}
cfg_env::BAIBOT_USER_DEVICE_ID => {
config.user.device_id = optional_non_empty(value);
}
cfg_env::BAIBOT_USER_ENCRYPTION_RECOVERY_PASSPHRASE => {
config.user.encryption.recovery_passphrase = Some(value);
}
cfg_env::BAIBOT_USER_ENCRYPTION_RECOVERY_RESET_ALLOWED => {
config.user.encryption.recovery_reset_allowed = value.parse::<bool>()?;
}
cfg_env::BAIBOT_USER_NAME => config.user.name = value,
cfg_env::BAIBOT_USER_AVATAR => {
config.user.avatar = Avatar::from_string(value);
}
cfg_env::BAIBOT_COMMAND_PREFIX => config.command_prefix = value,
cfg_env::BAIBOT_ROOM_POST_JOIN_SELF_INTRODUCTION_ENABLED => {
config.room.post_join_self_introduction_enabled = value.parse::<bool>()?;
@@ -51,6 +65,9 @@ pub fn load() -> anyhow::Result<Config> {
cfg_env::BAIBOT_PERSISTENCE_DATA_DIR_PATH => {
config.persistence.data_dir_path = Some(value);
}
cfg_env::BAIBOT_PERSISTENCE_SESSION_ENCRYPTION_KEY => {
config.persistence.session_encryption_key = Some(value);
}
cfg_env::BAIBOT_PERSISTENCE_CONFIG_ENCRYPTION_KEY => {
config.persistence.config_encryption_key = Some(value);
}
@@ -111,3 +128,7 @@ pub fn load() -> anyhow::Result<Config> {
Ok(config)
}
fn optional_non_empty(value: String) -> Option<String> {
if value.is_empty() { None } else { Some(value) }
}

View File

@@ -1,8 +1,11 @@
use mxlink::matrix_sdk::{
Room,
room::edit::EditedContent,
ruma::{
OwnedEventId, api::client::receipt::create_receipt::v3::ReceiptType,
events::room::message::OriginalSyncRoomMessageEvent,
EventId, OwnedEventId, api::client::receipt::create_receipt::v3::ReceiptType,
events::room::message::{
OriginalSyncRoomMessageEvent, RoomMessageEventContentWithoutRelation,
},
},
};
@@ -52,6 +55,83 @@ impl Messaging {
}
}
/// Like `send_text_markdown_no_fail`, but logs send failures at `warn` instead of `error`.
/// For cosmetic, best-effort sends (the thinking-notice placeholder) where a failure is
/// acceptable-impact: the real response still ships, so this must NOT page the team.
pub async fn send_text_markdown_no_fail_quietly(
&self,
room: &Room,
message: String,
response_type: MessageResponseType,
) -> Option<mxlink::matrix_sdk::ruma::api::client::message::send_message_event::v3::Response>
{
let result = self
.bot
.matrix_link()
.messaging()
.send_text_markdown(room, message, response_type)
.await;
match result {
Ok(result) => Some(result),
Err(err) => {
tracing::warn!(
room_id = format!("{:?}", room.room_id()),
?err,
"Failed to send thinking-notice placeholder to room",
);
None
}
}
}
/// Edits an existing message's text in place (`m.replace`), best-effort.
///
/// Built and sent directly via matrix-sdk's `make_edit_event` + `room.send`,
/// deliberately bypassing the mxlink messaging layer: that layer overwrites
/// `relates_to` from a `MessageResponseType`, which would clobber the
/// `m.replace` relation and turn the edit into a brand-new message. The edit
/// stays in the original's thread by inheritance, so no `response_type` is needed.
/// Failures are warned and swallowed so a flickered notice never breaks the real response.
pub async fn edit_text_markdown_no_fail(
&self,
room: &Room,
event_id: &EventId,
markdown: String,
) -> Option<mxlink::matrix_sdk::ruma::api::client::message::send_message_event::v3::Response>
{
let new_content = RoomMessageEventContentWithoutRelation::text_markdown(markdown);
let edit_content = match room
.make_edit_event(event_id, EditedContent::RoomMessage(new_content))
.await
{
Ok(edit_content) => edit_content,
Err(err) => {
tracing::warn!(
room_id = format!("{:?}", room.room_id()),
?event_id,
?err,
"Failed to build edit event",
);
return None;
}
};
match room.send(edit_content).await {
Ok(result) => Some(result.response),
Err(err) => {
tracing::warn!(
room_id = format!("{:?}", room.room_id()),
?event_id,
?err,
"Failed to send edit to room",
);
None
}
}
}
pub async fn send_notice_markdown_no_fail(
&self,
room: &Room,

View File

@@ -28,18 +28,18 @@ pub async fn handle_set(
message_context: &MessageContext,
patterns: &Option<Vec<String>>,
) -> anyhow::Result<()> {
if let Some(patterns) = patterns {
if let Err(err) = mxidwc::parse_patterns_vector(patterns) {
bot.messaging()
.send_error_markdown_no_fail(
message_context.room(),
&strings::access::failed_to_parse_patterns(&err.to_string()),
MessageResponseType::Reply(message_context.thread_info().root_event_id.clone()),
)
.await;
if let Some(patterns) = patterns
&& let Err(err) = mxidwc::parse_patterns_vector(patterns)
{
bot.messaging()
.send_error_markdown_no_fail(
message_context.room(),
&strings::access::failed_to_parse_patterns(&err.to_string()),
MessageResponseType::Reply(message_context.thread_info().root_event_id.clone()),
)
.await;
return Ok(());
}
return Ok(());
}
let mut global_config_manager_guard = bot.global_config_manager().lock().await;

View File

@@ -24,18 +24,18 @@ pub async fn handle_set(
message_context: &MessageContext,
patterns: &Option<Vec<String>>,
) -> anyhow::Result<()> {
if let Some(patterns) = patterns {
if let Err(err) = mxidwc::parse_patterns_vector(patterns) {
bot.messaging()
.send_error_markdown_no_fail(
message_context.room(),
&strings::access::failed_to_parse_patterns(&err.to_string()),
MessageResponseType::Reply(message_context.thread_info().root_event_id.clone()),
)
.await;
if let Some(patterns) = patterns
&& let Err(err) = mxidwc::parse_patterns_vector(patterns)
{
bot.messaging()
.send_error_markdown_no_fail(
message_context.room(),
&strings::access::failed_to_parse_patterns(&err.to_string()),
MessageResponseType::Reply(message_context.thread_info().root_event_id.clone()),
)
.await;
return Ok(());
}
return Ok(());
}
let mut global_config_manager_guard = bot.global_config_manager().lock().await;

View File

@@ -15,7 +15,7 @@ use crate::{Bot, entity::MessageContext};
struct ParsedAgentConfig {
agent: AgentInstance,
config: serde_yaml::Value,
config: serde_yaml_ng::Value,
}
pub async fn handle_room_local(
@@ -250,7 +250,7 @@ async fn send_guide(
provider: &AgentProvider,
) -> anyhow::Result<()> {
let sample_config = crate::agent::default_config_for_provider(provider);
let sample_config_pretty_yaml = serde_yaml::to_string(&sample_config)?;
let sample_config_pretty_yaml = serde_yaml_ng::to_string(&sample_config)?;
bot.messaging()
.send_text_markdown_no_fail(
@@ -263,7 +263,7 @@ async fn send_guide(
Ok(())
}
fn parse_from_message_to_yaml_value(text: &str) -> Result<serde_yaml::Value, String> {
fn parse_from_message_to_yaml_value(text: &str) -> Result<serde_yaml_ng::Value, String> {
let mut text = text.trim();
if text.starts_with("```") {
@@ -274,10 +274,10 @@ fn parse_from_message_to_yaml_value(text: &str) -> Result<serde_yaml::Value, Str
text = text.trim_end_matches("```");
}
let config: serde_yaml::Value = serde_yaml::from_str(text).map_err(|e| e.to_string())?;
let config: serde_yaml_ng::Value = serde_yaml_ng::from_str(text).map_err(|e| e.to_string())?;
match config {
serde_yaml::Value::Mapping(_) => {}
serde_yaml_ng::Value::Mapping(_) => {}
_ => {
return Err("Not a valid YAML hashmap".to_owned());
}

View File

@@ -2,14 +2,14 @@
fn agent_config_parsing_works() {
struct TestCase {
input: String,
expected: Option<serde_yaml::Value>,
expected: Option<serde_yaml_ng::Value>,
}
let provider = crate::agent::AgentProvider::OpenAI;
let sample_config = crate::agent::default_config_for_provider(&provider);
let sample_config_pretty_yaml = serde_yaml::to_string(&sample_config).unwrap();
let sample_config_pretty_yaml = serde_yaml_ng::to_string(&sample_config).unwrap();
let test_cases = vec![
let test_cases = [
// Invalid input
TestCase {
input: r#"Hello"#.to_owned(),

View File

@@ -64,7 +64,7 @@ pub async fn handle(
PublicIdentifier::Static(_) => {}
};
let config_yaml_pretty = serde_yaml::to_string(&agent.definition().config)?;
let config_yaml_pretty = serde_yaml_ng::to_string(&agent.definition().config)?;
bot.messaging()
.send_text_markdown_no_fail(

View File

@@ -3,7 +3,8 @@ use crate::{
entity::roomconfig::{
SpeechToTextFlowType, SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages,
TextGenerationAutoUsage, TextGenerationPrefixRequirementType,
TextToSpeechBotMessagesFlowType, TextToSpeechUserMessagesFlowType,
TextGenerationSenderContextMode, TextToSpeechBotMessagesFlowType,
TextToSpeechUserMessagesFlowType,
},
};
@@ -37,6 +38,9 @@ pub enum ConfigTextGenerationSettingRelatedControllerType {
GetContextManagementEnabled,
SetContextManagementEnabled(Option<bool>),
GetThinkingNoticeEnabled,
SetThinkingNoticeEnabled(Option<bool>),
GetPrefixRequirementType,
SetPrefixRequirementType(Option<TextGenerationPrefixRequirementType>),
@@ -48,6 +52,9 @@ pub enum ConfigTextGenerationSettingRelatedControllerType {
GetTemperatureOverride,
SetTemperatureOverride(Option<f32>),
GetSenderContextMode,
SetSenderContextMode(Option<TextGenerationSenderContextMode>),
}
#[derive(Debug, PartialEq)]

View File

@@ -163,6 +163,26 @@ fn determine_controller() {
),
)),
},
TestCase {
name: "per-room text-generation/sender-context-mode getter",
input: "room text-generation sender-context-mode",
expected: super::ControllerType::Config(controller_type::ConfigControllerType::SettingsRelated(
controller_type::SettingsStorageSource::Room,
controller_type::ConfigSettingRelatedControllerType::TextGeneration(
controller_type::ConfigTextGenerationSettingRelatedControllerType::GetSenderContextMode,
),
)),
},
TestCase {
name: "global text-generation/sender-context-mode getter",
input: "global text-generation sender-context-mode",
expected: super::ControllerType::Config(controller_type::ConfigControllerType::SettingsRelated(
controller_type::SettingsStorageSource::Global,
controller_type::ConfigSettingRelatedControllerType::TextGeneration(
controller_type::ConfigTextGenerationSettingRelatedControllerType::GetSenderContextMode,
),
)),
},
TestCase {
name: "per-room text-to-speech/speed-override getter",
input: "room text-to-speech speed-override",

View File

@@ -3,7 +3,10 @@ mod tests;
use crate::{
controller::ControllerType,
entity::roomconfig::{TextGenerationAutoUsage, TextGenerationPrefixRequirementType},
entity::roomconfig::{
TextGenerationAutoUsage, TextGenerationPrefixRequirementType,
TextGenerationSenderContextMode,
},
strings,
};
@@ -52,6 +55,44 @@ pub(super) fn determine(
);
}
if let Some(remaining_text) = text.strip_prefix("thinking-notice-enabled") {
let remaining_text = remaining_text.trim();
if !remaining_text.is_empty() {
return Err(ControllerType::Error(
strings::cfg::configuration_getter_used_with_extra_text(
"thinking-notice-enabled",
remaining_text,
)
.to_owned(),
));
}
return Ok(ConfigTextGenerationSettingRelatedControllerType::GetThinkingNoticeEnabled);
}
if let Some(value_string) = text.strip_prefix("set-thinking-notice-enabled") {
let value_string = value_string.trim().to_owned();
let value_opt = if value_string.is_empty() {
None
} else {
let value_string_lowercase = value_string.to_lowercase();
Some(match value_string_lowercase.as_str() {
"true" => true,
"false" => false,
_ => {
return Err(ControllerType::Error(
strings::cfg::configuration_value_unrecognized(&value_string).to_owned(),
));
}
})
};
return Ok(
ConfigTextGenerationSettingRelatedControllerType::SetThinkingNoticeEnabled(value_opt),
);
}
if let Some(remaining_text) = text.strip_prefix("prefix-requirement-type") {
let remaining_text = remaining_text.trim();
@@ -197,5 +238,43 @@ pub(super) fn determine(
);
}
if let Some(remaining_text) = text.strip_prefix("sender-context-mode") {
let remaining_text = remaining_text.trim();
if !remaining_text.is_empty() {
return Err(ControllerType::Error(
strings::cfg::configuration_getter_used_with_extra_text(
"sender-context-mode",
remaining_text,
)
.to_owned(),
));
}
return Ok(ConfigTextGenerationSettingRelatedControllerType::GetSenderContextMode);
}
if let Some(value_string) = text.strip_prefix("set-sender-context-mode") {
let value_string = value_string.trim().to_owned();
let value_choice = if value_string.is_empty() {
None
} else {
let value_choice =
TextGenerationSenderContextMode::from_str(&value_string.to_lowercase());
if value_choice.is_none() {
return Err(ControllerType::Error(
strings::cfg::configuration_value_unrecognized(&value_string).to_owned(),
));
}
value_choice
};
return Ok(
ConfigTextGenerationSettingRelatedControllerType::SetSenderContextMode(value_choice),
);
}
Err(ControllerType::Unknown)
}

View File

@@ -90,6 +90,74 @@ fn determine_controller_context_management() {
}
}
#[test]
fn determine_controller_sender_context() {
use super::ConfigTextGenerationSettingRelatedControllerType;
use super::ControllerType;
use crate::entity::roomconfig::TextGenerationSenderContextMode;
struct TestCase {
name: &'static str,
input: &'static str,
expected: Result<ConfigTextGenerationSettingRelatedControllerType, ControllerType>,
}
let test_cases = vec![
TestCase {
name: "sender-context-mode getter ok",
input: "sender-context-mode",
expected: Ok(ConfigTextGenerationSettingRelatedControllerType::GetSenderContextMode),
},
TestCase {
name: "sender-context-mode getter extra args",
input: "sender-context-mode some values here",
expected: Err(ControllerType::Error(
crate::strings::cfg::configuration_getter_used_with_extra_text(
"sender-context-mode",
"some values here",
),
)),
},
TestCase {
name: "sender-context-mode setter matrix_user_id",
input: "set-sender-context-mode matrix_user_id",
expected: Ok(
ConfigTextGenerationSettingRelatedControllerType::SetSenderContextMode(Some(
TextGenerationSenderContextMode::MatrixUserId,
)),
),
},
TestCase {
name: "sender-context-mode setter uppercase",
input: "set-sender-context-mode MATRIX_USER_ID_AND_TIMESTAMP",
expected: Ok(
ConfigTextGenerationSettingRelatedControllerType::SetSenderContextMode(Some(
TextGenerationSenderContextMode::MatrixUserIdAndTimestamp,
)),
),
},
TestCase {
name: "sender-context-mode setter invalid",
input: "set-sender-context-mode non-Enum-Value",
expected: Err(ControllerType::Error(
crate::strings::cfg::configuration_value_unrecognized("non-Enum-Value"),
)),
},
TestCase {
name: "sender-context-mode unsetter",
input: "set-sender-context-mode",
expected: Ok(
ConfigTextGenerationSettingRelatedControllerType::SetSenderContextMode(None),
),
},
];
for test_case in test_cases {
let result = super::determine(test_case.input);
assert_eq!(result, test_case.expected, "Test case: {}", test_case.name);
}
}
#[test]
fn determine_controller_prefix_requirement_type() {
use super::ConfigTextGenerationSettingRelatedControllerType;

View File

@@ -39,18 +39,18 @@ async fn dispatch_config_related_handler(
message_context: &MessageContext,
bot: &Bot,
) -> anyhow::Result<()> {
if let SettingsStorageSource::Global = config_type {
if !message_context.sender_can_manage_global_config() {
bot.messaging()
.send_error_markdown_no_fail(
message_context.room(),
strings::global_config::no_permissions_to_administrate(),
MessageResponseType::Reply(message_context.thread_info().root_event_id.clone()),
)
.await;
return Ok(());
}
};
if let SettingsStorageSource::Global = config_type
&& !message_context.sender_can_manage_global_config()
{
bot.messaging()
.send_error_markdown_no_fail(
message_context.room(),
strings::global_config::no_permissions_to_administrate(),
MessageResponseType::Reply(message_context.thread_info().root_event_id.clone()),
)
.await;
return Ok(());
}
let room_settings = match config_type {
SettingsStorageSource::Room => &message_context.room_config().settings,

View File

@@ -1,5 +1,6 @@
use crate::entity::roomconfig::{
RoomSettings, TextGenerationAutoUsage, TextGenerationPrefixRequirementType,
TextGenerationSenderContextMode,
};
use crate::{Bot, entity::MessageContext};
@@ -42,6 +43,27 @@ pub(super) async fn dispatch(
}
}
ConfigTextGenerationSettingRelatedControllerType::GetThinkingNoticeEnabled => {
let value = &room_settings.text_generation.thinking_notice_enabled;
setting_get::<bool>(bot, message_context, value).await
}
ConfigTextGenerationSettingRelatedControllerType::SetThinkingNoticeEnabled(value) => {
let value = value.to_owned();
let setter_callback = Box::new(move |room_settings: &mut RoomSettings| {
room_settings.text_generation.thinking_notice_enabled = value;
});
match config_type {
SettingsStorageSource::Room => {
room_setting_set::<bool>(bot, message_context, &value, setter_callback).await
}
SettingsStorageSource::Global => {
global_setting_set::<bool>(bot, message_context, &value, setter_callback).await
}
}
}
ConfigTextGenerationSettingRelatedControllerType::GetPrefixRequirementType => {
let value = &room_settings.text_generation.prefix_requirement_type;
setting_get::<TextGenerationPrefixRequirementType>(bot, message_context, value).await
@@ -151,5 +173,38 @@ pub(super) async fn dispatch(
}
}
}
ConfigTextGenerationSettingRelatedControllerType::GetSenderContextMode => {
let value = &room_settings.text_generation.sender_context_mode;
setting_get::<TextGenerationSenderContextMode>(bot, message_context, value).await
}
ConfigTextGenerationSettingRelatedControllerType::SetSenderContextMode(value) => {
let value = value.to_owned();
let setter_callback = Box::new(move |room_settings: &mut RoomSettings| {
room_settings.text_generation.sender_context_mode = value;
});
match config_type {
SettingsStorageSource::Room => {
room_setting_set::<TextGenerationSenderContextMode>(
bot,
message_context,
&value,
setter_callback,
)
.await
}
SettingsStorageSource::Global => {
global_setting_set::<TextGenerationSenderContextMode>(
bot,
message_context,
&value,
setter_callback,
)
.await
}
}
}
}
}

View File

@@ -7,7 +7,8 @@ use crate::{
roomconfig::{
SpeechToTextFlowType, SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages,
TextGenerationAutoUsage, TextGenerationPrefixRequirementType,
TextToSpeechBotMessagesFlowType, TextToSpeechUserMessagesFlowType,
TextGenerationSenderContextMode, TextToSpeechBotMessagesFlowType,
TextToSpeechUserMessagesFlowType,
},
},
strings,
@@ -233,6 +234,84 @@ fn build_section_text_generation(command_prefix: &str, bot_username: &str) -> St
));
message.push_str("\n\n");
// Thinking Notice
message.push_str(&format!(
"#### {}",
strings::help::cfg::text_generation_thinking_notice_heading()
));
message.push_str("\n\n");
message.push_str(&strings::help::cfg::text_generation_thinking_notice_intro());
message.push('\n');
message.push_str(
&strings::help::cfg::the_following_configuration_values_are_recognized(vec![true, false]),
);
message.push_str("\n\n");
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_show(
command_prefix,
"text-generation thinking-notice-enabled"
)
));
message.push('\n');
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_set(
command_prefix,
"text-generation set-thinking-notice-enabled VALUE"
)
));
message.push('\n');
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_unset(
command_prefix,
"text-generation set-thinking-notice-enabled"
)
));
message.push_str("\n\n");
// Sender Context
message.push_str(&format!(
"#### {}",
strings::help::cfg::text_generation_sender_context_heading()
));
message.push_str("\n\n");
message.push_str(&strings::help::cfg::text_generation_sender_context_intro());
message.push('\n');
message.push_str(
&strings::help::cfg::the_following_configuration_values_are_recognized(
TextGenerationSenderContextMode::choices(),
),
);
message.push_str("\n\n");
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_show(
command_prefix,
"text-generation sender-context-mode"
)
));
message.push('\n');
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_set(
command_prefix,
"text-generation set-sender-context-mode VALUE"
)
));
message.push('\n');
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_unset(
command_prefix,
"text-generation set-sender-context-mode"
)
));
message.push_str("\n\n");
// Prompt override
message.push_str(&format!(

View File

@@ -69,7 +69,7 @@ pub async fn handle(bot: &Bot, message_context: &MessageContext) -> anyhow::Resu
);
message.push_str("\n\n");
// Image Generation
// Image Creation
message.push_str(
&generate_image_generation_section(agent_manager, message_context.room_config_context())
.await,
@@ -359,6 +359,60 @@ async fn generate_text_generation_section(
),
);
// Thinking Notice
let effective_thinking_notice = room_config_context.text_generation_thinking_notice_enabled();
let room_config_thinking_notice = room_config_context
.room_config
.settings
.text_generation
.thinking_notice_enabled;
let global_config_thinking_notice = room_config_context
.global_config
.fallback_room_settings
.text_generation
.thinking_notice_enabled;
let thinking_notice_set_where = if room_config_thinking_notice.is_some() {
strings::cfg::status_badge_set_in_room_config()
} else if global_config_thinking_notice.is_some() {
strings::cfg::status_badge_set_in_global_config()
} else {
strings::cfg::status_badge_using_hardcoded_default()
};
message.push_str(&strings::cfg::status_text_generation_entry_thinking_notice(
effective_thinking_notice,
thinking_notice_set_where,
));
// Sender Context
let effective_sender_context = room_config_context.text_generation_sender_context_mode();
let room_config_sender_context = room_config_context
.room_config
.settings
.text_generation
.sender_context_mode;
let global_config_sender_context = room_config_context
.global_config
.fallback_room_settings
.text_generation
.sender_context_mode;
let sender_context_set_where = if room_config_sender_context.is_some() {
strings::cfg::status_badge_set_in_room_config()
} else if global_config_sender_context.is_some() {
strings::cfg::status_badge_set_in_global_config()
} else {
strings::cfg::status_badge_using_hardcoded_default()
};
message.push_str(&strings::cfg::status_text_generation_entry_sender_context(
effective_sender_context,
sender_context_set_where,
));
// Prompt override
let text_agent_prompt = if let Some(text_generation_agent) = &text_generation_agent {

View File

@@ -15,7 +15,8 @@ use crate::conversation::matrix::MatrixMessageProcessingParams;
use crate::entity::MessagePayload;
use crate::entity::roomconfig::{
SpeechToTextFlowType, SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages,
TextToSpeechBotMessagesFlowType, TextToSpeechUserMessagesFlowType,
TextGenerationSenderContextMode, TextToSpeechBotMessagesFlowType,
TextToSpeechUserMessagesFlowType,
};
use crate::strings;
use crate::utils::text_to_speech::create_transcribed_message_text;
@@ -23,11 +24,19 @@ use crate::{
Bot,
conversation::{
create_llm_conversation_for_matrix_reply_chain, create_llm_conversation_for_matrix_thread,
llm::{Author, Conversation, MessageContent},
matrix::create_list_of_bot_user_prefixes_to_strip,
},
entity::MessageContext,
};
/// How long text generation must run before the first "thinking…" placeholder appears.
/// Fast responses (under this threshold) never get a placeholder.
const THINKING_NOTICE_FIRST_DELAY: std::time::Duration = std::time::Duration::from_secs(3);
/// How often the "thinking…" placeholder is refreshed once it has appeared.
const THINKING_NOTICE_INTERVAL: std::time::Duration = std::time::Duration::from_secs(10);
#[derive(Debug, PartialEq)]
pub enum ChatCompletionControllerType {
// Invoked via a command prefix (e.g. `!bai Hello!`)
@@ -39,6 +48,10 @@ pub enum ChatCompletionControllerType {
Audio,
Image,
File,
ThreadMention,
ReplyMention,
}
@@ -416,7 +429,9 @@ async fn handle_stage_text_generation(
ChatCompletionControllerType::TextCommand
| ChatCompletionControllerType::TextMention
| ChatCompletionControllerType::TextDirect
| ChatCompletionControllerType::Audio => {
| ChatCompletionControllerType::Audio
| ChatCompletionControllerType::Image
| ChatCompletionControllerType::File => {
Some(message_context.combined_admin_and_user_regexes())
}
@@ -438,6 +453,7 @@ async fn handle_stage_text_generation(
// When we're triggered via a reply mention, the context is the whole reply chain upward of the message that triggered us.
ChatCompletionControllerType::ReplyMention => {
create_llm_conversation_for_matrix_reply_chain(
&matrix_link,
&bot.room_event_fetcher().clone(),
message_context.room(),
message_context.thread_info().last_event_id.clone(),
@@ -449,7 +465,7 @@ async fn handle_stage_text_generation(
// Everything else is happening in a thread, so the context is the whole thread.
_ => {
create_llm_conversation_for_matrix_thread(
matrix_link.clone(),
&matrix_link,
message_context.room(),
message_context.thread_info().root_event_id.clone(),
&params,
@@ -479,6 +495,13 @@ async fn handle_stage_text_generation(
}
};
let conversation = inject_sender_context(
conversation,
message_context
.room_config_context()
.text_generation_sender_context_mode(),
);
tracing::debug!(
agent_id = agent.identifier().as_string(),
provider = format!("{}", agent.definition().provider.clone()),
@@ -504,6 +527,16 @@ async fn handle_stage_text_generation(
conversation.start_time(),
);
// Cloned only when the thinking-notice is enabled; the original is moved into `params` below.
let notice_prompt_variables = if message_context
.room_config_context()
.text_generation_thinking_notice_enabled()
{
Some(prompt_variables.clone())
} else {
None
};
let params = TextGenerationParams {
context_management_enabled: message_context
.room_config_context()
@@ -520,10 +553,73 @@ async fn handle_stage_text_generation(
prompt_variables,
};
let result = controller
.generate_text(conversation, params)
.instrument(span)
.await;
// When the thinking-notice is enabled, race generation against a timer that posts and then
// periodically edits a "thinking…" placeholder. `biased;` makes generation win a tie, and the
// loop exits the instant generation resolves, so there is no detached task and no late edit can
// ever clobber the real answer. `placeholder` is the event we must finalize in every exit path.
let (result, placeholder) = if let Some(notice_prompt_variables) = notice_prompt_variables {
let generation = controller.generate_text(conversation, params).instrument(span);
tokio::pin!(generation);
let mut placeholder: Option<OwnedEventId> = None;
// Seed the flavor sequence per-generation so different turns don't all open on the same
// line; the monotonic increment then guarantees consecutive notices differ.
let mut notice_sequence: usize = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|since| since.subsec_nanos() as usize)
.unwrap_or(0);
let mut tick = tokio::time::interval_at(
tokio::time::Instant::now() + THINKING_NOTICE_FIRST_DELAY,
THINKING_NOTICE_INTERVAL,
);
// If an edit runs long, hold ~INTERVAL spacing rather than bursting the missed ticks.
tick.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Delay);
let result = loop {
tokio::select! {
biased;
generation_result = &mut generation => break generation_result,
_ = tick.tick() => {
let message = notice_prompt_variables.format(
strings::thinking::pick_message(start_time.elapsed(), notice_sequence),
);
notice_sequence = notice_sequence.wrapping_add(1);
match &placeholder {
None => {
placeholder = bot
.messaging()
.send_text_markdown_no_fail_quietly(
message_context.room(),
message,
response_type.clone(),
)
.await
.map(|response| response.event_id);
}
Some(event_id) => {
bot.messaging()
.edit_text_markdown_no_fail(
message_context.room(),
event_id,
message,
)
.await;
}
}
}
}
};
(result, placeholder)
} else {
let result = controller
.generate_text(conversation, params)
.instrument(span)
.await;
(result, None)
};
let duration = std::time::Instant::now().duration_since(start_time);
@@ -544,17 +640,20 @@ async fn handle_stage_text_generation(
err,
);
bot.messaging()
.send_error_markdown_no_fail(
message_context.room(),
&strings::agent::error_while_serving_purpose(
agent.identifier(),
&AgentPurpose::TextGeneration,
&err,
),
response_type,
)
.await;
let error_text = strings::agent::error_while_serving_purpose(
agent.identifier(),
&AgentPurpose::TextGeneration,
&err,
);
finalize_thinking_notice_with_error(
bot,
message_context,
placeholder.as_ref(),
&error_text,
response_type,
)
.await;
return None;
}
@@ -567,26 +666,70 @@ async fn handle_stage_text_generation(
"Agent returned empty text",
);
bot.messaging()
.send_error_markdown_no_fail(
message_context.room(),
&strings::agent::empty_response_returned(agent.identifier()),
response_type,
)
.await;
let empty_text = strings::agent::empty_response_returned(agent.identifier());
finalize_thinking_notice_with_error(
bot,
message_context,
placeholder.as_ref(),
&empty_text,
response_type,
)
.await;
return None;
}
let send_message_response = bot
.messaging()
.send_text_markdown_no_fail(message_context.room(), text.clone(), response_type)
.await?;
// Finalize the answer into a single message. With a placeholder, edit it in place (so the
// "thinking…" message becomes the answer); the TTS payload then points at that same event.
// If the edit fails, fall back to a fresh send so the real answer is never lost.
let event_id = match &placeholder {
Some(event_id)
if bot
.messaging()
.edit_text_markdown_no_fail(message_context.room(), event_id, text.clone())
.await
.is_some() =>
{
event_id.clone()
}
_ => {
bot.messaging()
.send_text_markdown_no_fail(message_context.room(), text.clone(), response_type)
.await?
.event_id
}
};
Some(TextToSpeechEligiblePayload {
text,
event_id: send_message_response.event_id,
})
Some(TextToSpeechEligiblePayload { text, event_id })
}
/// Finalizes a thinking-notice placeholder (if one was posted) with error/notice text, so a
/// failed or empty generation never leaves an orphaned "thinking…" message behind. With no
/// placeholder, this is the original behavior: a fresh error notice.
async fn finalize_thinking_notice_with_error(
bot: &Bot,
message_context: &MessageContext,
placeholder: Option<&OwnedEventId>,
text: &str,
response_type: MessageResponseType,
) {
match placeholder {
Some(event_id) => {
bot.messaging()
.edit_text_markdown_no_fail(
message_context.room(),
event_id,
crate::utils::status::create_error_message_text(text),
)
.await;
}
None => {
bot.messaging()
.send_error_markdown_no_fail(message_context.room(), text, response_type)
.await;
}
}
}
async fn handle_stage_speech_to_text_actual_transcribing(
@@ -754,3 +897,238 @@ async fn generate_and_send_tts_for_message(
)
.await
}
fn inject_sender_context(
conversation: Conversation,
sender_context_mode: TextGenerationSenderContextMode,
) -> Conversation {
if sender_context_mode == TextGenerationSenderContextMode::Disabled {
return conversation;
}
let include_timestamp =
sender_context_mode == TextGenerationSenderContextMode::MatrixUserIdAndTimestamp;
let messages = conversation
.messages
.into_iter()
.map(|mut message| {
if message.author == Author::Prompt {
return message;
}
let Some(sender_id) = &message.sender_id else {
return message;
};
if let MessageContent::Text(ref mut text) = message.content {
*text = if include_timestamp {
let timestamp = message.timestamp.format("%Y-%m-%dT%H:%M:%SZ");
format!("[sender={} sent_at={}] {}", sender_id, timestamp, text)
} else {
format!("[sender={}] {}", sender_id, text)
};
}
message
})
.collect();
Conversation { messages }
}
#[cfg(test)]
mod sender_context_tests {
use super::inject_sender_context;
use crate::conversation::llm::{Author, Conversation, ImageDetails, Message, MessageContent};
use crate::entity::roomconfig::TextGenerationSenderContextMode;
use chrono::{TimeZone, Utc};
use mxlink::matrix_sdk::ruma::events::room::message::ImageMessageEventContent;
use mxlink::matrix_sdk::ruma::{OwnedMxcUri, OwnedUserId};
use mxlink::mime;
#[test]
fn test_inject_sender_context_prefixes_text_messages() {
let timestamp = Utc.with_ymd_and_hms(2026, 3, 23, 14, 30, 0).unwrap();
let user_id = OwnedUserId::try_from("@alice:example.com").unwrap();
let conversation = Conversation {
messages: vec![Message {
author: Author::User,
sender_id: Some(user_id),
timestamp,
content: MessageContent::Text("Hello bot".to_string()),
}],
};
let result = inject_sender_context(
conversation,
TextGenerationSenderContextMode::MatrixUserIdAndTimestamp,
);
assert_eq!(result.messages.len(), 1);
assert_eq!(
result.messages[0].content,
MessageContent::Text(
"[sender=@alice:example.com sent_at=2026-03-23T14:30:00Z] Hello bot".to_string()
)
);
}
#[test]
fn test_inject_sender_context_can_prefix_without_timestamp() {
let timestamp = Utc.with_ymd_and_hms(2026, 3, 23, 14, 30, 0).unwrap();
let user_id = OwnedUserId::try_from("@alice:example.com").unwrap();
let conversation = Conversation {
messages: vec![Message {
author: Author::User,
sender_id: Some(user_id),
timestamp,
content: MessageContent::Text("Hello bot".to_string()),
}],
};
let result =
inject_sender_context(conversation, TextGenerationSenderContextMode::MatrixUserId);
assert_eq!(result.messages.len(), 1);
assert_eq!(
result.messages[0].content,
MessageContent::Text("[sender=@alice:example.com] Hello bot".to_string())
);
}
#[test]
fn test_inject_sender_context_prefixes_assistant_messages() {
let timestamp = Utc.with_ymd_and_hms(2026, 3, 23, 14, 30, 0).unwrap();
let user_id = OwnedUserId::try_from("@baibot:example.com").unwrap();
let conversation = Conversation {
messages: vec![Message {
author: Author::Assistant,
sender_id: Some(user_id),
timestamp,
content: MessageContent::Text("Hello human".to_string()),
}],
};
let result = inject_sender_context(
conversation,
TextGenerationSenderContextMode::MatrixUserIdAndTimestamp,
);
assert_eq!(result.messages.len(), 1);
assert_eq!(
result.messages[0].content,
MessageContent::Text(
"[sender=@baibot:example.com sent_at=2026-03-23T14:30:00Z] Hello human".to_string()
)
);
}
#[test]
fn test_inject_sender_context_skips_prompt_messages() {
let timestamp = Utc.with_ymd_and_hms(2026, 3, 23, 14, 30, 0).unwrap();
let conversation = Conversation {
messages: vec![Message {
author: Author::Prompt,
sender_id: None,
timestamp,
content: MessageContent::Text("You are a bot".to_string()),
}],
};
let result =
inject_sender_context(conversation, TextGenerationSenderContextMode::MatrixUserId);
assert_eq!(
result.messages[0].content,
MessageContent::Text("You are a bot".to_string())
);
}
#[test]
fn test_inject_sender_context_skips_messages_without_sender_id() {
let timestamp = Utc.with_ymd_and_hms(2026, 3, 23, 14, 30, 0).unwrap();
let conversation = Conversation {
messages: vec![Message {
author: Author::User,
sender_id: None,
timestamp,
content: MessageContent::Text("Transcribed text".to_string()),
}],
};
let result = inject_sender_context(
conversation,
TextGenerationSenderContextMode::MatrixUserIdAndTimestamp,
);
assert_eq!(
result.messages[0].content,
MessageContent::Text("Transcribed text".to_string())
);
}
#[test]
fn test_inject_sender_context_leaves_non_text_content_unchanged() {
let timestamp = Utc.with_ymd_and_hms(2026, 3, 23, 14, 30, 0).unwrap();
let user_id = OwnedUserId::try_from("@alice:example.com").unwrap();
let image_event_content = ImageMessageEventContent::plain(
"image.png".to_string(),
OwnedMxcUri::from("mxc://example.com/1234567890"),
);
let conversation = Conversation {
messages: vec![Message {
author: Author::User,
sender_id: Some(user_id),
timestamp,
content: MessageContent::Image(ImageDetails::new(
image_event_content.clone(),
mime::IMAGE_PNG,
vec![],
)),
}],
};
let result = inject_sender_context(
conversation,
TextGenerationSenderContextMode::MatrixUserIdAndTimestamp,
);
assert_eq!(
result.messages[0].content,
MessageContent::Image(ImageDetails::new(
image_event_content,
mime::IMAGE_PNG,
vec![]
))
);
}
#[test]
fn test_inject_sender_context_none_leaves_text_unchanged() {
let timestamp = Utc.with_ymd_and_hms(2026, 3, 23, 14, 30, 0).unwrap();
let user_id = OwnedUserId::try_from("@alice:example.com").unwrap();
let conversation = Conversation {
messages: vec![Message {
author: Author::User,
sender_id: Some(user_id),
timestamp,
content: MessageContent::Text("Hello bot".to_string()),
}],
};
let result = inject_sender_context(conversation, TextGenerationSenderContextMode::Disabled);
assert_eq!(
result.messages[0].content,
MessageContent::Text("Hello bot".to_string())
);
}
}

View File

@@ -23,5 +23,6 @@ pub enum ControllerType {
ChatCompletion(super::chat_completion::ChatCompletionControllerType),
ImageGeneration(String),
ImageEdit(String),
StickerGeneration(String),
}

View File

@@ -36,6 +36,18 @@ pub fn determine_controller(
first_thread_message.is_mentioning_bot,
)
}
MessagePayload::Image(_image_message_content) => {
let prefix_requirement_type = message_context
.room_config_context()
.text_generation_prefix_requirement_type();
match prefix_requirement_type {
TextGenerationPrefixRequirementType::CommandPrefix => ControllerType::Ignore,
TextGenerationPrefixRequirementType::No => {
ControllerType::ChatCompletion(ChatCompletionControllerType::Image)
}
}
}
MessagePayload::Encrypted(thread_info) => {
if thread_info.is_thread_root_only() {
ControllerType::Error(strings::error::message_is_encrypted().to_owned())
@@ -46,6 +58,18 @@ pub fn determine_controller(
)
}
}
MessagePayload::File(_file_message_content) => {
let prefix_requirement_type = message_context
.room_config_context()
.text_generation_prefix_requirement_type();
match prefix_requirement_type {
TextGenerationPrefixRequirementType::CommandPrefix => ControllerType::Ignore,
TextGenerationPrefixRequirementType::No => {
ControllerType::ChatCompletion(ChatCompletionControllerType::File)
}
}
}
MessagePayload::Audio(_) => {
ControllerType::ChatCompletion(ChatCompletionControllerType::Audio)
}
@@ -84,7 +108,7 @@ fn determine_text_controller(
}
if let Some(prompt) = text.strip_prefix(&format!("{command_prefix} image")) {
return ControllerType::ImageGeneration(prompt.trim().to_owned());
return super::image::determine_controller(prompt.trim());
}
if let Some(prompt) = text.strip_prefix(&format!("{command_prefix} sticker")) {

View File

@@ -84,9 +84,17 @@ fn determine_text_controller() {
expected: ControllerType::Config(controller::cfg::ConfigControllerType::Help),
},
TestCase {
name: "Image generation",
name: "Generic image command causes usage help",
input: "!bai image Draw a cat!",
is_mentioning_bot: false,
room_text_generation_prefix_requirement_type:
super::TextGenerationPrefixRequirementType::No,
expected: ControllerType::UsageHelp,
},
TestCase {
name: "Image generation",
input: "!bai image create Draw a cat!",
is_mentioning_bot: false,
room_text_generation_prefix_requirement_type:
super::TextGenerationPrefixRequirementType::No,
expected: ControllerType::ImageGeneration("Draw a cat!".to_owned()),

View File

@@ -77,6 +77,10 @@ pub async fn dispatch_controller(
)
.await
}
ControllerType::ImageEdit(prompt) => {
super::image::edit::handle(bot, bot.matrix_link().clone(), message_context, prompt)
.await
}
ControllerType::StickerGeneration(prompt) => {
super::image::generation::handle_sticker(
bot,

View File

@@ -0,0 +1,16 @@
use crate::controller::ControllerType;
mod tests;
pub fn determine_controller(text: &str) -> ControllerType {
let text = text.trim();
if let Some(prompt) = text.strip_prefix("create") {
return ControllerType::ImageGeneration(prompt.trim().to_owned());
}
if let Some(prompt) = text.strip_prefix("edit") {
return ControllerType::ImageEdit(prompt.trim().to_owned());
}
ControllerType::UsageHelp
}

Some files were not shown because too many files have changed in this diff Show More