Compare commits

..

81 Commits

Author SHA1 Message Date
Slavi Pantaleev
ce81fe69bd Release 1.7.0 2025-05-10 12:22:58 +03:00
Slavi Pantaleev
1162636b88 Upgrade Rust (1.85.1 -> 1.86.0) 2025-05-10 12:22:46 +03:00
Slavi Pantaleev
8c90e13a79 Upgrade services 2025-05-10 12:11:08 +03:00
Slavi Pantaleev
274b614d25 Update dependencies 2025-05-10 11:58:06 +03:00
Slavi Pantaleev
7bd46821dc Update README to mention the images editing feature 2025-05-10 11:49:43 +03:00
Slavi Pantaleev
a84135ff32 fmt 2025-05-10 11:47:50 +03:00
Slavi Pantaleev
231528a0d8 Document which providers support vision 2025-05-10 11:47:00 +03:00
Slavi Pantaleev
d8e47b0578 Document vision support for text-generation 2025-05-10 11:39:33 +03:00
Slavi Pantaleev
96c1542f4a Add an image-editing screenshot that demos multiple image support 2025-05-10 11:36:14 +03:00
Slavi Pantaleev
2f9c3dfce0 Use patched anthropic-rs library to add Vision support to text conversations
Related to: https://github.com/AbdelStark/anthropic-rs/pull/11
2025-05-10 10:25:15 +03:00
Slavi Pantaleev
de958208b2 Use patched async-openai library to work around a few upstream issues
Related to:

- CreateImageEditRequest forces an application/octet-stream content type
  for images (https://github.com/64bit/async-openai/issues/364)

- CreateImageEditRequest only deals with a single image
  (https://github.com/64bit/async-openai/issues/363)

This is a continuation of 8f86289373
and fixes the Image Editing feature for OpenAI.
2025-05-10 10:00:36 +03:00
Slavi Pantaleev
ac4f2080ce Make some improvements as suggested by clippy 2025-05-10 09:33:44 +03:00
Slavi Pantaleev
3ffa50b7b9 Use ImageInput|AudioInput::from_vec_u8 helper 2025-05-10 09:29:06 +03:00
Slavi Pantaleev
8f86289373 Initial work on Vision support in text conversations and Image Editing
This is a huge patch which does some major refactoring like:

- renaming "Image Generation" to "Image Creation" in most places,
  to better match its new command (`!bai image create`)

- relocating image creation command (`!bai image` -> `!bai image create`),
  so it wouldn't conflict with the new image editing command (`!bai image edit`)

- introducing a new image editing command (`!bai image edit`), which
  is meant to work only with the OpenAI provider, but doesn't fully work yet
  due to https://github.com/64bit/async-openai/issues/364, though a next patch will fix it

- adding support for reading images off of Matrix conversations and forwarding them to
  text conversations. Works for OpenAI, but not for Anthropic yet
  (requires custom patches) and not for OpenAI-Compat (no support for
  images there)

- relocating some utils around (base64, mime)
2025-05-10 09:18:01 +03:00
Slavi Pantaleev
e0dcc39a72 Default OpenAI image-generation model to gpt-image-1 (previously dall-e-3) 2025-05-03 09:42:46 +03:00
Slavi Pantaleev
c94376109c Avoid passing response_format to OpenAI's image generation API for the gpt-image-1 model
The API reference for `response_format` says:

> This parameter isn't supported for gpt-image-1 which will always return base64-encoded images.

Related to https://github.com/etkecc/baibot/issues/40
2025-05-03 09:42:44 +03:00
Slavi Pantaleev
256ed05662 Make style, quality and internal response_format image generation parameters optional
Some OpenAI models (like `gpt-image-1`) either don't support these or
only support specific other values.

Related to https://github.com/etkecc/baibot/issues/40
2025-05-03 09:42:39 +03:00
Slavi Pantaleev
8222681e27 Default OpenAI text-generation model to gpt-4.1 (previously gpt-4o) 2025-05-03 09:42:37 +03:00
Slavi Pantaleev
f304b93c68 Improve in-room error reporting details when image generation fails
What previously was a generic error message like:

> ⚠️ Error: An error occurred while processing your message. Please try again.

.. now becomes a much more helpful error message like:

> ⚠️ Error: There was a problem performing image-generation via the room-local/my-openai-agent agent:
>
> invalid_request_error: Invalid value: 'standard'. Supported values are: 'low', 'medium', 'high', and 'auto'. (param: quality) (code: invalid_value)

Related to https://github.com/etkecc/baibot/issues/40
2025-05-03 09:42:28 +03:00
Slavi Pantaleev
889d8a1d04 Upgrade Synapse (v1.127.1 -> v1.128.0) and Element Web (v1.11.96 -> v1.11.99) 2025-05-03 08:50:00 +03:00
Slavi Pantaleev
6082bfaf56 Release 1.6.0 2025-04-12 08:04:23 +03:00
Slavi Pantaleev
1d629e0859 Update dependencies 2025-04-12 07:57:12 +03:00
Slavi Pantaleev
49471c1df0 Upgrade mxlink (1.6.0 -> 1.7.0) and matrix-sdk (0.10.0 -> 0.11.0) 2025-04-12 07:49:26 +03:00
Slavi Pantaleev
8f956d2329 Release 1.5.1 2025-03-31 14:37:23 +03:00
Slavi Pantaleev
ba4aa35987 Upgrade Rust compiler in container image (1.85.0 -> 1.85.1) 2025-03-31 14:33:49 +03:00
Slavi Pantaleev
06d699a17d Update dependencies 2025-03-31 14:33:49 +03:00
Slavi Pantaleev
77d41fb7eb Update services 2025-03-31 14:33:49 +03:00
Slavi Pantaleev
aaf283dde3 Upgrade Element Web (v1.11.93 -> v1.11.96) and adapt to its own way of running as a non-privileged user 2025-03-31 14:33:49 +03:00
Slavi Pantaleev
17eafa86af Release 1.5.0 2025-02-27 11:29:26 +02:00
Slavi Pantaleev
4704934b06 Run amd64 builds on a self-hosted runner 2025-02-27 11:06:18 +02:00
Slavi Pantaleev
6719538530 Switch to using native ARM64 builders
The workflow code is inspired by https://github.com/matrix-org/matrix-hookshot/pull/1007

Ref: https://github.blog/news-insights/product-news/arm64-on-github-actions-powering-faster-more-efficient-build-systems/

This decreases build times from ~40 minutes (on powerful self-hosted runners) to ~8 minutes (on public runners).
2025-02-27 10:31:23 +02:00
Slavi Pantaleev
59e2746578 Remove explicit lifetime to fix clippy-reported warning 2025-02-27 09:58:58 +02:00
Slavi Pantaleev
47d8edea70 Add support for configuring max_completion_tokens for OpenAI
Related to db9422740c
2025-02-27 09:58:58 +02:00
Slavi Pantaleev
692d61b239 Replace Anthropic library (anthropic-rs -> anthropic) and switch default recommended model (claude-3-5-sonnet-20240620 -> claude-3-7-sonnet-20250219)
Fixes https://github.com/etkecc/baibot/issues/22

Ultimate related to `anthropic-rs` hardcoding models as enumeration variants in the code
and not updating them. See:
- https://github.com/roushou/mesh/issues/1
- https://github.com/roushou/mesh/pull/2

https://github.com/cortesi/misanthropy was also considered as an
alternative, but it did not allow configuring the base API URL like our
old Anthropic library (`anthropic-rs`) and like our new choice (`anthropic`).
We'd rather not lose support for this, so we're going with the `anthropic` library.
2025-02-27 09:44:49 +02:00
Slavi Pantaleev
05902f4c17 Add fmt justfile recipe 2025-02-27 09:31:26 +02:00
Slavi Pantaleev
7e66068b16 Update dependencies 2025-02-27 07:48:10 +02:00
Slavi Pantaleev
406141cd7d fmt 2025-02-27 07:46:16 +02:00
Slavi Pantaleev
c051da2f4a Add config setting controlling if a self-introduction message is posted after joining a room
Fixes https://github.com/etkecc/baibot/issues/32
2025-02-26 20:51:24 +02:00
Slavi Pantaleev
1ff7e8cf79 Switch Rust edition (2021 -> 2024) 2025-02-26 20:51:24 +02:00
Slavi Pantaleev
b3bca98e84 Upgrade Rust compiler in container image (1.82.0 -> 1.85.0) 2025-02-26 20:51:24 +02:00
Slavi Pantaleev
c07b712318 Upgrade mxlink (1.5.0 -> 1.6.0) and matrix-sdk (0.9.0 -> 0.10.0) 2025-02-26 20:51:24 +02:00
Slavi Pantaleev
6741483056 Update dependencies 2025-02-26 20:51:24 +02:00
Slavi Pantaleev
06b2b6d776 Use progress indicator emoji (⏳), not 🦻 to indicate that speech-to-text is happening
🦻 is used for another purpose - to denote that a message is one coming
from speech-to-text, by:

- having the bot react to its own speech-to-text transcription message
  with the 🦻 emoji when it's posted in a non-thread

- having the bot prefix its speech-to-text transcription messages with
  `> 🦻` when it's posted in a thread

⏳ is already used as a progress indicator for other features, so it
makes sense to use it for indicating that speech-to-text is happening for a given audio message as well.
While 🦻 was an even more descriptive illustration of what's actually happening to the audio message
("it's being heard by the bot"), us using the 🦻 emoji for different things didn't seem good.
2025-02-26 20:51:24 +02:00
Slavi Pantaleev
a1bd292752 Add support for making Text-To-Speech send regular text messages instead of notices
When speech-to-text/flow-type = `only_transcribe`, the bot will now send
text messages by default, not notices.

While notice messages may be less desirable with other bots in the room,
it's probably a better default for most people who enable "transcribe-only" mode.

This is an improvement related to https://github.com/etkecc/baibot/issues/14
2025-02-26 20:51:08 +02:00
Slavi Pantaleev
e4e1fe0e7b Upgrade some components 2025-02-26 09:39:48 +02:00
Slavi Pantaleev
45a2d96029 Enable necessary feature for matrix-sdk which affects us if used as a library
Fixes https://github.com/etkecc/rust-mxlink/issues/1
2025-02-02 07:51:25 +02:00
Slavi Pantaleev
ec1879d212 Populate image/audio attachment body with a filename, not with text
Various clients (including newer versions of Element Web), do not like
it when the `body` field of the attachment is not a file name.

For images, a preview may not be shown and downloading the attachment
may suggest that the whole long text is used as a filename (which is odd).

There is value (improved accessibility, etc.)
in adding better descriptions (especially to generated images),
but given that it's currently problematic, I'm getting rid of it.
It's better and safer if we stick to using filenames.
2025-01-24 11:35:08 +02:00
Slavi Pantaleev
5e6a600895 Update dependencies 2025-01-24 10:55:12 +02:00
Slavi Pantaleev
3db924b124 Upgrade services 2025-01-24 10:01:50 +02:00
Slavi Pantaleev
ff7a5ef7af Install SQLite 3 for test-and-clippy CI job 2024-12-12 12:25:56 +02:00
Slavi Pantaleev
cd7d9137e8 Fix broken commit link in changelog entry 2024-12-12 12:03:25 +02:00
Slavi Pantaleev
0d509b2d0e Release 1.4.1 2024-12-12 12:01:22 +02:00
Slavi Pantaleev
3c47d40781 Upgrade mxlink to fix incorrect on-being-last-member determination
Related to ec69da79ff
2024-12-12 12:00:47 +02:00
Slavi Pantaleev
78893247e7 Mention Matrix authenticated media support for the 1.4.0 release
Related to https://github.com/etkecc/baibot/issues/12
2024-11-20 07:41:05 +02:00
Slavi Pantaleev
4847bd8ba8 Enable authenticated media for the local Synapse service
This is now supported thanks to us upgrading mxlink/matrix-sdk
in 39a184e5d0.
2024-11-20 07:40:08 +02:00
Slavi Pantaleev
c8abf0e316 Release 1.4.0 2024-11-19 21:05:56 +02:00
Slavi Pantaleev
39a184e5d0 Adapt to mxlink 1.4.0 (matrix-sdk 0.8.0) 2024-11-19 20:57:35 +02:00
Slavi Pantaleev
9d166e35ba Add missing typing notices sending functionality while generating images 2024-11-19 20:46:23 +02:00
Slavi Pantaleev
4a5966401c Release 1.3.2 2024-11-12 10:37:49 +02:00
Slavi Pantaleev
d92dfba2bf Upgrade Rust compiler in container image (1.81.0 -> 1.82.0) 2024-11-12 10:37:16 +02:00
Slavi Pantaleev
8538d6b2b8 Upgrade component services 2024-11-12 10:36:54 +02:00
Slavi Pantaleev
23f763ba72 Update dependencies 2024-11-12 10:22:26 +02:00
Slavi Pantaleev
a9e4ab1bdb Release 1.3.1 2024-10-03 16:30:58 +03:00
Slavi Pantaleev
d9a045a5e4 Make fallback user mentions support also match against the bot's room-specific username
It seems like Element iOS benefits from this.
2024-10-03 16:28:49 +03:00
Slavi Pantaleev
393be9be5a Remove strip_rich_reply_fallback_text in favor of remove_plain_reply_fallback from ruma events
No need to reinvent the wheel.
2024-10-03 16:02:06 +03:00
Slavi Pantaleev
a7b016a3d3 Release 1.3.0 2024-10-03 12:08:00 +03:00
Slavi Pantaleev
85e66406dc Allow for prompt caching to work by using baibot_conversation_start_time_utc instead of baibot_now_utc
This patch introduces a new `baibot_conversation_start_time_utc`
variable which indicates the time the conversation got started.

Using `baibot_now_utc` is still possible, but given that the current
time is a moving target, its use is in conflict with prompt caching.

Because the new `baibot_conversation_start_time_utc` prompt variable
is a more reasonable default, we're now using it in all sample configs.
2024-10-03 11:48:14 +03:00
Slavi Pantaleev
db9422740c Add support for OpenAI's o1 models by making max_response_tokens optional
The other prerequisite seems to be not using a `prompt` (`prompt: null`),
but we already supported this.

It'd be nice to add an optional `max_completion_tokens` parameter as
well, for the benefit of the o1 models, but this is not yet supported by
async-openai.
Possibly tracked here: https://github.com/64bit/async-openai/issues/272
2024-10-03 10:36:28 +03:00
Slavi Pantaleev
90fbad5b64 Update sample & default OpenAI provider configs to use gpt-4o (instead of gpt-4o-2024-08-06)
Since 2024-10-02, `gpt-4o` is actually the same as `gpt-4o-2024-08-06`.

We previously used `gpt-4o-2024-08-06`, because it was pointing to a
much better (longer context) model. Since they're both the same now,
we'd better stick to the unpinned model and make it easier for future
users to get upgrades.
2024-10-03 09:26:41 +03:00
Slavi Pantaleev
b40226826f Restore fallback support for user mentions
Fallback support was intentionally removed in 9908512968,
because it was deemed OK to do so.

It turns out that Element iOS still doesn't properly do user mentions
(and likely never will, until Element X replaces it), so we can't just
drop the fallback user mentions logic without affecting all these
clients. It's possible that the Element Android is no better (unverified claim).
2024-10-03 09:18:02 +03:00
Slavi Pantaleev
b89f0db71a Relocate "On-demand involvement" feature description section
[skip ci]
2024-10-02 09:12:12 +03:00
Slavi Pantaleev
36fdb46633 Release 1.2.0 2024-10-01 22:01:17 +03:00
Slavi Pantaleev
04ce8db1fc Update dependencies 2024-10-01 21:35:49 +03:00
Slavi Pantaleev
9908512968 Add support for on-demand involvement
Fixes https://github.com/etkecc/baibot/issues/15
2024-10-01 21:06:54 +03:00
Slavi Pantaleev
eae6472c7a Upgrade services 2024-10-01 18:33:08 +03:00
Slavi Pantaleev
e6aa956423 Do not send blockquote-formatted transcription when replying without a thread
This actually fixes 2 issues.

Fixes https://github.com/etkecc/baibot/issues/14
Fixes https://github.com/etkecc/baibot/issues/17

When people enable transcribe-only mode and the bot replies outside of a
thread, messages will no longer look like this: `> 🦻 Transcribed text`.

Instead, they will:

- look like this: `Transcribed text`

- get an emoji reaction (🦻) sent by the bot itself,
  to indicate that the message is a transcription

---------------------------------------

As https://github.com/etkecc/baibot/issues/14 discusses,
the `> 🦻` prefixing of messages also served the purpose of indicating
to the bot that this is not its own message, but rather something it
"heard" from a user.

Given that out-of-thread replies no longer include this, they could be
mistaken for bot messages.

Because transcribed messages are posted as notice messages, we can
easily tell them apart from regular text-generated messages by the bot
itself, so we can (and do) treat them differently.

Thankfully, the bot does not yet support building a text-generation
conversation from arbitrary messages (something discussed in
https://github.com/etkecc/baibot/issues/15), so these out-of-thread
replies having the wrong owner are not an issue for now.

If we do land support for this, we'll probably need to make the bot inspect such notice messages
posted by it, inspect their reactons and attribute them properly (🦻 -> user message).
2024-09-30 17:35:14 +03:00
Slavi Pantaleev
7a38216192 Upgrade services 2024-09-26 22:32:34 +03:00
Slavi Pantaleev
d522d268e2 Use the same variable-infused prompt for the ollama agent in etc/app/config.yml.dist 2024-09-26 22:32:28 +03:00
Slavi Pantaleev
72120c5dc2 Explicitly keep authenticated media disabled in the Synapse configuration
[skip ci]

Related to https://github.com/etkecc/baibot/issues/12
2024-09-23 06:09:32 +00:00
Slavi Pantaleev
a2c35238c2 Upgrade services
[skip ci]
2024-09-23 06:08:17 +00:00
Slavi Pantaleev
97f5cbb00b Fix example for baibot_now_utc prompt variable
This feature went through a few iterations. At some point,
a `(local timezone/time: unknown)` suffix was part of the
`baibot_now_utc` variable (hoping it improves the model's awareness that
it doesn't know the current local time), but the suffix was ultimately removed
as unnecessary.
2024-09-22 08:52:19 +00:00
163 changed files with 4506 additions and 1847 deletions

View File

@@ -3,13 +3,14 @@ on:
push:
branches: [ "main" ]
tags: [ "v*" ]
schedule:
- cron: '0 0 * * 1'
permissions:
checks: write
contents: write
packages: write
pull-requests: read
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: false
jobs:
test-and-clippy:
name: Unit testing and linting
@@ -17,20 +18,46 @@ jobs:
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- name: Install SQLite3
run: sudo apt-get update && sudo apt-get install -y libsqlite3-dev
- run: cargo test --all-features
- run: cargo clippy
build-publish:
name: Build and Publish
runs-on: self-hosted
docker-clean-metadata:
runs-on: ubuntu-latest
outputs:
json: ${{ steps.meta.outputs.json }}
steps:
- name: Set up QEMU
uses: docker/setup-qemu-action@v3
- name: Extract metadata (tags, labels) for Docker
id: meta
uses: docker/metadata-action@v5
with:
platforms: arm64
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v1
- name: Login to ghcr.io
images: |
ghcr.io/${{ github.repository }}
tags: |
type=raw,value=latest,enable=${{ github.ref_name == 'main' }}
type=semver,pattern={{raw}}
docker-build:
permissions:
contents: read
packages: write
attestations: write
id-token: write
strategy:
matrix:
include:
- os: self-hosted
arch: amd64
- os: ubuntu-24.04-arm
arch: arm64
runs-on: ${{ matrix.os }}
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Log in to the GitHub Container registry
uses: docker/login-action@v3
with:
registry: ghcr.io
@@ -40,17 +67,41 @@ jobs:
id: meta
uses: docker/metadata-action@v5
with:
images: |
ghcr.io/${{ github.repository }}
registry.etke.cc/${{ github.repository }}
tags: |
type=raw,value=latest,enable=${{ github.ref_name == 'main' }}
type=semver,pattern={{raw}}
- name: Build and push
flavor: |
latest=auto
suffix=-${{ matrix.arch }},onlatest=true
images: |
ghcr.io/${{ github.repository }}
- name: Build and push Docker images
uses: docker/build-push-action@v6
with:
platforms: linux/amd64,linux/arm64
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
file: Dockerfile.ci
docker-manifest:
needs:
- docker-build
- docker-clean-metadata
runs-on: ubuntu-latest
strategy:
matrix:
image: ${{ fromJson(needs.docker-clean-metadata.outputs.json).tags }}
steps:
- name: Log in to the GitHub Container registry
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Create and push manifest
run: |
docker manifest create ${{ matrix.image }} ${{ matrix.image }}-amd64 ${{ matrix.image }}-arm64
docker manifest push ${{ matrix.image }}

View File

@@ -1,3 +1,93 @@
# (2025-05-10) Version 1.7.0
- (**Feature**) Add vision support to the OpenAI and Anthropic providers. You can now mix text and images in your conversations - fixes [issue #5](https://github.com/etkecc/baibot/issues/5)
- (**Feature**) Add [image-editing](./docs/features.md#-image-editing) support to the OpenAI provider
- (**Improvement**) Add compatibility with OpenAI's `gpt-image-1` model - fixes [issue #40](https://github.com/etkecc/baibot/issues/40)
- (**Change**) Rework [image-creation](./docs/features.md#-image-creation) to avoid command conflicts with [image-editing](./docs/features.md#-image-editing). The image-creation command syntax is now `!bai image create <prompt>` (previously: `!bai image <prompt>`).
- (**Internal Improvement**) Dependency and compiler updates
> [!WARNING]
> Unlike other releases, this release is not published to [crates.io](https://crates.io), because it relies on multiple library forks (`async-openai` and `anthropic-rs`) sourced from Github.
# (2025-04-12) Version 1.6.0
- (**Internal Improvement**) Dependency updates. This version is based on [mxlink](https://crates.io/crates/mxlink)@1.7.0 (which is based on the newly released [matrix-sdk](https://crates.io/crates/matrix-sdk)@[0.11.0](https://github.com/matrix-org/matrix-rust-sdk/releases/tag/matrix-sdk-0.11.0))
# (2025-03-31) Version 1.5.1
- (**Internal Improvement**) Dependency updates
# (2025-02-27) Version 1.5.0
- (**Feature**) Add support for sending Speech-to-Text replies for [Transcribe-only mode](./docs/features.md#transcribe-only-mode) as regular text messages instead of notices and doing it so by default ([a1bd292752](https://github.com/etkecc/baibot/commit/a1bd292752bdd37a196788c73d00b5619e843a78)) - improvement for [issue #14](https://github.com/etkecc/baibot/issues/14). See [🦻 Speech-to-Text / 🪄 Message Type for non-threaded only-transcribed messages](./docs/configuration/speech-to-text.md#-message-type-for-non-threaded-only-transcribed-messages) for details.
- (**Feature**) Add config setting controlling if a self-introduction message is posted after joining a room ([c051da2f4a](https://github.com/etkecc/baibot/commit/c051da2f4a161de0974ebb917f7a52d01f5a001f)) - fixes [issue #32](https://github.com/etkecc/baibot/issues/32). You may wish to add a `room.post_join_self_introduction_enabled` property to your configuration. See the [sample config](./etc/app/config.yml.dist) for details. If unspecified, it defaults to `true` anyway which preserves the old behavior.
- (**Feature**) Add support for configuring `max_completion_tokens` for OpenAI ([47d8edea70](https://github.com/etkecc/baibot/commit/47d8edea705a44aa25a9bfaec4888c0f9ea8700e))
- (**Improvement**) Dependency updates. This version is based on [mxlink](https://crates.io/crates/mxlink)@1.6.1 (which is based on the newly released [matrix-sdk](https://crates.io/crates/matrix-sdk)@[0.10.0](https://github.com/matrix-org/matrix-rust-sdk/releases/tag/matrix-sdk-0.10.0))
- (**Improvement**) Populate image/audio attachment `body` with a filename, not with text to avoid incorrect rendering in Element Web, etc. ([ec1879d212](https://github.com/etkecc/baibot/commit/ec1879d212fa8d6e5f8590486e94c72abfcb75a5))
- (**Improvement**) Replace Anthropic library ([anthropic-rs](https://crates.io/crates/anthropic-rs) -> [anthropic](https://crates.io/crates/anthropic)) and switch default recommended model (`claude-3-5-sonnet-20240620` -> `claude-3-7-sonnet-20250219`) ([692d61b239](https://github.com/etkecc/baibot/commit/692d61b2398f073b81d32d4cbe8145ab3929e48c)) - fixes [issue #22](https://github.com/etkecc/baibot/issues/22)
- (**Internal Improvement**) Switch to native building of `arm64` container images to decrease total build times from ~40 minutes to ~8 minutes ([6719538530b](https://github.com/etkecc/baibot/commit/6719538530bf76b3ff2d24077b2a7fa868276b79))
- (**Internal Improvement**) Various other internal changes, including upgrading [Rust from 1.82 to 1.85 and switching to Rust edition 2024](https://blog.rust-lang.org/2025/02/20/Rust-1.85.0.html)
# (2024-12-12) Version 1.4.1
- (**Bugfix**) Fix detection for whether the bot is the last member in a room, to avoid incorrectly leaving multi-user rooms that have had at least one person `leave` ([3c47d40781](https://github.com/etkecc/baibot/commit/3c47d407819aa9c0121117a411858238724f06da))
# (2024-11-19) Version 1.4.0
- (**Improvement**) Dependency updates. This version is based on [mxlink](https://crates.io/crates/mxlink)@1.4.0 (which is based on the newly released [matrix-sdk](https://crates.io/crates/matrix-sdk)@[0.8.0](https://github.com/matrix-org/matrix-rust-sdk/releases/tag/matrix-sdk-0.8.0)). Once you run this version at least once and your matrix-sdk datastore gets upgraded to the new schema, **you will not be able to downgrade to older baibot versions** (based on the older matrix-sdk), unless you start with an empty datastore.
- (**Bugfix**) Add missing typing notices sending functionality while generating images ([9d166e35ba](https://github.com/etkecc/baibot/commit/9d166e35ba6fc0daaf69318870e92436f3302056))
- (**Feature**) Support for [Matrix authenticated media](https://matrix.org/docs/spec-guides/authed-media-servers/), thanks to upgrading [mxlink](https://crates.io/crates/mxlink) / [matrix-sdk](https://crates.io/crates/matrix-sdk) - fixes [issue #12](https://github.com/etkecc/baibot/issues/12)
# (2024-11-12) Version 1.3.2
Dependency updates.
# (2024-10-03) Version 1.3.1
- (**Improvement**) Improves fallback user mentions support for old clients (like Element iOS) which use the bot's display name (not its full Matrix User ID). ([d9a045a5e4](https://github.com/etkecc/baibot/commit/d9a045a5e41d2b99694f92ec9e90f47529546d89))
# (2024-10-03) Version 1.3.0
**TLDR**: you can now use OpenAI's [o1](https://platform.openai.com/docs/models/o1) models, benefit from [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) and mention the bot again from old clients lacking proper [user mentions support](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) (like Element iOS).
- (**Feature**) Introduces a new `baibot_conversation_start_time_utc` [prompt variable](./docs/configuration/text-generation.md#️-prompt-override) which is not a moving target (like the `baibot_now_utc` variable) and allows [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) to work. All default/sample configs have been adjusted to make use of this new variable, but users need to adjust your existing dynamically-created agents to start using it. ([85e66406dc](https://github.com/etkecc/baibot/commit/85e66406dc6f430741c7819f420e2df4ae6e8d3b))
- (**Improvement**) Allows for the `max_response_tokens` configuration value for the [OpenAI provider](./docs/providers.md#openai) to be set to `null` to allow [o1](https://platform.openai.com/docs/models/o1) models (which do not support `max_response_tokens`) to be used. See the new o1 sample config [here](./docs/sample-provider-configs/openai-o1.yml). ([db9422740c](https://github.com/etkecc/baibot/commit/db9422740ceca32956d9628b6326b8be206344e2))
- (**Improvement**) Switches the sample configs for the [OpenAI provider](./docs/providers.md#openai) to point to the `gpt-4o` model, which since 2024-10-02 is the same as the `gpt-4o-2024-08-06` model. We previously explicitly pointed the bot to the `gpt-4o-2024-08-06` model, because it was much better (longer context window). Now that `gpt-4o` points to the same powerful model, we don't need to pin its version anymore. Existing users may wish to adjust their configuration to match. ([90fbad5b64](https://github.com/etkecc/baibot/commit/90fbad5b643cd06c23179f055a309ec6a7cba161))
- (**Bugfix**) Restores fallback user mentions support (via regular text, not via the [user mentions spec](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions)) to allow certain old clients (like Element iOS) to be able to mention the bot again. Support for this was intentionally removed recently (in [v1.2.0](#2024-10-01-version-120)), but it turned out to be too early to do this. ([b40226826f](https://github.com/etkecc/baibot/commit/b40226826fe914d0d5d265230ebc5bac8058b6f7))
# (2024-10-01) Version 1.2.0
- (**Feature**) Adds support for [on-demand involvement](./docs/features.md#on-demand-involvement) of the bot (via mention) in arbitrary threads and reply chains ([9908512968](https://github.com/etkecc/baibot/commit/990851296828168c2106eb3f4668833e9e5a7463)) - fixes [issue #15](https://github.com/etkecc/baibot/issues/15)
- (**Improvement**) Simplifies [Transcribe-only mode](./docs/features.md#transcribe-only-mode) reply format (removing `> 🦻` prefixing) to allow easier forwarding, etc. ([e6aa956423](https://github.com/etkecc/baibot/commit/e6aa95642376ee7d87932d0e66dcfedf261b188b)) - fixes [issue #14](https://github.com/etkecc/baibot/issues/14)
- (**Bugfix**) Fixes speech-to-text replies rendering incorrectly in certain clients, due to them confusing our old reply format with [fallback for rich replies](https://spec.matrix.org/v1.11/client-server-api/#fallbacks-for-rich-replies) ([e6aa956423](https://github.com/etkecc/baibot/commit/e6aa95642376ee7d87932d0e66dcfedf261b188b)) - fixes [issue #17](https://github.com/etkecc/baibot/issues/17)
# (2024-09-22) Version 1.1.1
- (**Bugfix**) Fix thread messages being lost due to lack of pagination support ([d4ddd29660](https://github.com/etkecc/baibot/commit/d4ddd29660d9f51d248119dd6032e68ab29e7d35)) - fixes [issue #13](https://github.com/etkecc/baibot/issues/13)

2142
Cargo.lock generated

File diff suppressed because it is too large Load Diff

View File

@@ -7,32 +7,33 @@ license = "AGPL-3.0-or-later"
readme = "README.md"
keywords = ["matrix", "chat", "bot", "AI", "LLM"]
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
version = "1.1.1"
edition = "2021"
version = "1.7.0"
edition = "2024"
[lib]
name = "baibot"
path = "src/lib.rs"
[dependencies]
anthropic-rs = "0.1.*"
anthropic = { git = "https://github.com/etkecc/anthropic-rs.git", branch = "fix-content-block-image" }
anyhow = "1.0.*"
async-openai = "0.24.*"
async-openai = { git = "https://github.com/etkecc/async-openai", branch = "async-openai-v0.28.1-patched" }
base64 = "0.22.*"
chrono = { version = "0.4.*", default-features = false, features = ["std", "now"] }
# We'd rather not depend on this, but we cannot use the ruma-events EventContent macro without it.
matrix-sdk = { version = "0.7.1", default-features = false }
# We add the `native-tls` feature, because of https://github.com/etkecc/rust-mxlink/issues/1
matrix-sdk = { version = "0.11.0", default-features = false, features = ["native-tls"] }
mxidwc = "1.0.*"
mxlink = ">=1.3.0"
mxlink = ">=1.7.0"
etke_openai_api_rust = "0.1.*"
quick_cache = "0.6.*"
regex = "1.10.*"
regex = "1.11.*"
serde = { version = "1.0.*", features = ["derive"], default-features = false }
serde_json = "1.0.*"
serde_yaml = "0.9.*"
tempfile = "3.12.*"
tiktoken-rs = { version = "0.5.*", features = ["async-openai"] }
tokio = { version = "1.40.*", features = ["rt", "rt-multi-thread", "macros"] }
tempfile = "3.19.*"
tiktoken-rs = { version = "0.6.*", features = ["async-openai"] }
tokio = { version = "1.45.*", features = ["rt", "rt-multi-thread", "macros"] }
tracing = "0.1.*"
tracing-subscriber = { version = "0.3.*", features = ["env-filter"] }
url = "2.5.*"

View File

@@ -4,7 +4,7 @@
# #
#######################################
FROM docker.io/rust:1.81.0-slim-bookworm AS build
FROM docker.io/rust:1.86.0-slim-bookworm AS build
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev

View File

@@ -4,7 +4,7 @@
# #
#######################################
FROM docker.io/rust:1.81.0-slim-bookworm AS build
FROM docker.io/rust:1.86.0-slim-bookworm AS build
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev

View File

@@ -17,10 +17,10 @@ It's influenced by [chaz](https://github.com/arcuru/chaz), but does **not** use
- Supports **different use purposes** (depending on the [☁️ provider](./docs/providers.md) & model):
- [💬 text-generation](./docs/features.md#-text-generation): communicating with you via text
- [💬 text-generation](./docs/features.md#-text-generation): communicating with you via text (though certain models may "see" images as well)
- [🦻 speech-to-text](./docs/features.md#-speech-to-text): turning your voice messages into text
- [🗣️ text-to-speech](./docs/features.md#%EF%B8%8F-text-to-speech): turning bot or users text messages into voice messages
- [🖌️ image-generation](./docs/features.md#%EF%B8%8F-image-generation): generating images based on instructions
- [🖌️ image-generation](./docs/features.md#image-generation): creating and editing images based on instructions
- 🪄 Supports [seamless voice interaction](./docs/features.md#seamless-voice-interaction) (turning user voice messages into text, answering in text, then turning that text back into voice)
@@ -41,7 +41,7 @@ It's influenced by [chaz](https://github.com/arcuru/chaz), but does **not** use
![Introduction and general usage](./docs/screenshots/introduction-and-general-usage.webp)
You can find more screenshots on the the [🌟 Features](./docs/features.md) and other [📚 Documentation](./docs/README.md) pages, as well as in the [docs/screenshots](./docs/screenshots) directory.
You can find more screenshots on the [🌟 Features](./docs/features.md) and other [📚 Documentation](./docs/README.md) pages, as well as in the [docs/screenshots](./docs/screenshots) directory.
## 🚀 Getting Started

View File

@@ -16,6 +16,7 @@ Users:
- ✅ can **invite the bot to rooms**
- ✅ can **use all the bot's [features](./features.md)** ([💬 Text Generation](./features.md#-text-generation), [🦻 Speech-to-Text](./features.md#-speech-to-text), etc.) by sending room messages
- ✅ can **mention the bot** in threads and reply chains to provoke it to respond to non-user messages (see [🌟 Features / 💬 Text Generation / On-demand involvement](./features.md#on-demand-involvement))
- ✅ can **change the bot's configuration in a room** (e.g. `!bai config room ...` commands)
- ❌ cannot **change the bot's global configuration** (e.g. `!bai config global ...` commands)
- ❌ cannot **create new [🤖 Agents](./agents.md)** (neither in rooms, nor globally). See [💼 Room-local agent managers](#-room-local-agent-managers) for controlling which users can create agents.

View File

@@ -35,7 +35,7 @@ Depending on where the agent is defined (within a room, globally, or [statically
When creating an agent, you will be given some sample [YAML](https://en.wikipedia.org/wiki/YAML) configuration which you can use to customize the agent's behavior.
This configuration varies depending on the [☁️ provider](./providers.md) used and the capabilities of the agent. Based on the configuration keys you pass, certain features will be enabled or disabled. For example, if you skip the `image_generation` key for an [OpenAI](./providers.md#openai) agent, it won't be able to generate images (see [🖌️ Image Generation](./features.md#-image-generation)).
This configuration varies depending on the [☁️ provider](./providers.md) used and the capabilities of the agent. Based on the configuration keys you pass, certain features will be enabled or disabled. For example, if you skip the `image_generation` key for an [OpenAI](./providers.md#openai) agent, it won't be able to generate images (see [🖌️ Image Creation](./features.md#-image-creation), [🎨 Image Editing](./features.md#-image-editing), [🫵 Sticker Creation](./features.md#-sticker-creation)).
After making your modifications to the sample YAML, you submit it back to the bot and the new agent will be created.

View File

@@ -40,7 +40,7 @@ You can adjust the following settings per room and/or globally:
- [💬 Text Generation](text-generation.md)
- [🦻 Speech-to-Text](speech-to-text.md)
- [🗣️ Text-to-Speech](text-to-speech.md)
- [🖌️ Image Generation](image-generation.md)
- [🖌️ Image Creation](image-generation.md)
- [🤝 Handlers](handlers.md)
Refer to the bot's help messages (as a response to a `!bai config` help command) for the most up-to-date information on what Room Settings can be configured.

View File

@@ -8,10 +8,10 @@ You can also use **different models within the same room** (e.g. [💬 text-gene
The bot supports the following use-purposes:
- [💬 text-generation](../features.md#-text-generation): communicating with you via text
- [💬 text-generation](../features.md#-text-generation): communicating with you via text (though certain models may "see" images as well)
- [🦻 speech-to-text](../features.md#-speech-to-text): turning your voice messages into text
- [🗣️ text-to-speech](../features.md#️-text-to-speech): turning bot or users text messages into voice messages
- [🖌️ image-generation](../features.md#-image-generation): generating images based on instructions
- [🖌️ image-generation](../features.md#image-generation): generating images based on instructions
In a given room, each different purpose can be served by a different [provider](../providers.md) and model. This combination of provider and model configuration is called an [🤖 agent](../agents.md). Each purpose can be served by a different **handler** agent.

View File

@@ -1,9 +1,11 @@
## 🖌️ Image Generation
## Image Generation
The Image Generation feature is not configurable at this moment.
The Image Creation and Image Editing features are not configurable at this moment.
You may also wish to see:
- [🌟 Features / 🖌️ Image Generation](../features.md#-image-generation) for a higher-level introduction to the Image Generation features
- [📖 Usage / 🖌️ Image Generation](../usage.md#-image-generation) section for more details on how to use the bot for Image Generation in a room
- [🌟 Features / Image Generation / 🖌️ Image Creation](../features.md#-image-creation) for a higher-level introduction to the Image Creation features
- [🌟 Features / Image Generation / 🎨 Image Editing](../features.md#-image-editing) for a higher-level introduction to the Image Editing features
- [📖 Usage / Image Generation / 🖌️ Creating Images](../usage.md#-creating-images) section for more details on how to use the bot for Image Creation in a room
- [📖 Usage / Image Generation / 🎨 Editing images](../usage.md#-editing-images) section for more details on how to use the bot for Image Editing in a room

View File

@@ -23,6 +23,19 @@ The following configuration values are recognized:
Example: `!bai config room speech-to-text set-flow-type ignore` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
### 🪄 Message Type for non-threaded only-transcribed messages
Controls how the transcribed text of voice messages is sent to the chat when Flow Type = `only_transcribe`.
The following configuration values are recognized:
- (default) `text`: the transcribed text is sent as a regular message. This is more convenient if you'd like to forward the transcribed message to other rooms.
- `notice`: the transcribed text is sent as a notice message. This provides better compatibility with other bots in the room, as they are less likely to interact with messages of type notice.
Example: `!bai config room speech-to-text set-msg-type-for-non-threaded-only-transcribed-messages notice` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
### 🔤 Language
Lets you specify the language of the input voice messages, to avoid using auto-detection.

View File

@@ -13,7 +13,7 @@ You may also wish to see:
In Direct Message rooms with the bot (1:1 rooms), it most usually makes sense for the bot to respond to **all** of your messages, as shown on this [🖼️ screenshot](../screenshots/text-generation.webp).
In group rooms (with multiple users), it may be more appropriate for the bot to only respond to messages that are **prefixed** with the command prefix (e.g. `!bai`), so that other chat exchange in the room will not trigger it. Such a setup is shown on this [🖼️ screenshot](../screenshots/text-generation-prefix-requirement.webp).
In group rooms (with multiple users), it may be more appropriate for the bot to only respond to messages that are **prefixed** with the command prefix (e.g. `!bai`) or which are [mentioning](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) the bot (e.g. `@baibot`), so that other chat exchange in the room will not trigger it. Such a setup is shown on the [🖼️ On-demand involvement in the room](../screenshots/text-generation-prefix-requirement.webp) screenshot.
There are exceptions to these rules, and you can configure the bot to respond only to prefixed messages in a 1:1 room, or to respond to all messages even in a multi-user group room.
@@ -27,7 +27,10 @@ By default, the bot is **auto-configured (upon joining a new room)** to use the
Example: `!bai config room text-generation set-prefix-requirement-type command_prefix` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
Regardless of this configuration, **the bot will also respond to messages which directly [mention](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) the bot** (e.g. `@baibot`), even if they are not prefixed. An example of this can be seen on this [🖼️ screenshot](../screenshots/text-generation-prefix-requirement.webp).
Regardless of this configuration, **the bot will also respond to messages by allowed [👥 Users](../access.md#-users) which directly [mention](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) the bot** (e.g. `@baibot`), even if they are not prefixed. An example of this can be seen on these screenshots:
- [🖼️ On-demand involvement in a thread](../screenshots/text-generation-on-demand-thread-involvement.webp)
- [🖼️ On-demand involvement in a reply chain](../screenshots/text-generation-on-demand-reply-involvement.webp)
### 🪄 Auto Usage
@@ -74,11 +77,14 @@ Prompts may contain the following **placeholder variables** which will be replac
|---------------------------|-------------|---------|
| `{{ baibot_name }}` | Name of the bot as configured in the `user.name` field in the [Static configuration](./README.md#static-configuration) | `Baibot` |
| `{{ baibot_model_id }}` | Text-Generation model ID as configured in the [🤖 agent](../agents.md)'s configuration | `gpt-4o` |
| `{{ baibot_now_utc }}` | Current date and time in UTC | `2024-09-20 (Friday), 14:26:42 UTC (local timezone/time: unknown)` |
| `{{ baibot_now_utc }}` | Current date and time in UTC (⚠️ usage may break prompt caching - see below) | `2024-09-20 (Friday), 14:26:42 UTC` |
| `{{ baibot_conversation_start_time_utc }}` | The date and time in UTC that the conversation started | `2024-09-20 (Friday), 14:26:42 UTC` |
💡 `{{ baibot_now_utc }}` changes as time goes on, which prevents [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) from working. It's better to use `{{ baibot_conversation_start_time_utc }}` in prompts, as its value doesn't change yet still orients the bot to the current date/time.
Here's a prompt that combines some of the above variables:
> You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
> You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
### 🌡️ Temperature Override

View File

@@ -93,7 +93,7 @@ For getting started most quickly (and locally), we recommend using [LocalAI](#lo
**Ollama is most lightweight** (~2GB for the container image + ~1.6GB for the model), but supports only [💬 text-generation](./features.md#-text-generation).
**LocalAI requires 4x more disk space** (~6GB for the container image + ~12GB for the models), but supports [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text) and [🖼️ image-generation](./features.md#️-image-generation).
**LocalAI requires 4x more disk space** (~6GB for the container image + ~12GB for the models), but supports [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text) and [🖼️ image-generation](./features.md#️-image-creation).
**OpenAI supports all of these capabilities** as well and does not require powerful hardware or lots of disk space. However, it requires signup and an API key.

View File

@@ -8,7 +8,7 @@ You can also use **different models within the same room** (e.g. [💬 text-gene
The bot supports the following use-purposes:
- [💬 text-generation](#-text-generation): communicating with you via text
- [💬 text-generation](#-text-generation): communicating with you via text (though certain models may "see" images as well)
- [🦻 speech-to-text](#-speech-to-text): turning your voice messages into text
- [🗣️ text-to-speech](#%EF%B8%8F-text-to-speech): turning bot or users text messages into voice messages
- [🖌️ image-generation](#%EF%B8%8F-image-generation): generating images based on instructions
@@ -22,12 +22,16 @@ For more information about configuring handlers, see the [🤝 Handlers / Config
### 💬 Text Generation
Text Generation is the bot's ability to **respond to users' text messages with text**.
Text Generation is the bot's ability to **respond to users' messages with text**.
![Screenshot of Text Generation - a user sends a message and the bot replies in a new conversation thread](./screenshots/text-generation.webp)
Some models also support vision, so you may be able to mix text and images in the same conversation.
In multi-user (group) rooms, to avoid disturbing the normal conversation between people, the bot is auto-configured to only respond to messages starting with the command prefix (`!bai`) or direct mentions via the [💬 Text Generation / 🗟 Prefix Requirement Type](./configuration/text-generation.md#-prefix-requirement-type) setting.
Normally, the bot only responds to allowed [👥 Users](./access.md#-users). In certain cases, it's useful for an allowed user to provoke the bot to respond even in foreign threads or reply chains. You can learn more about this feature in the [On-demand involvement](./features.md#on-demand-involvement) section below.
A few other features (like [🗣️ Text-to-Speech](#️-text-to-speech) and [🦻 Speech-to-Text](#-speech-to-text)) combine well with Text Generation, so you **don't necessarily need to communicate with the bot via text** (with [Seamless voice interaction](#seamless-voice-interaction), you can communicate only with voice).
You may also wish to see:
@@ -36,6 +40,22 @@ You may also wish to see:
- [📖 Usage / 💬 Text Generation](./usage.md#-text-generation) section for more details on how to use the bot for Text Generation in a room
#### On-demand involvement
In the following 2 cases, it's useful to involve the bot in conversations on-demand:
1. In multi-user rooms (with the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting set to "required")
2. In rooms with foreign users (users that are not authorized bot [👥 users](./access.md#-users))
In these instances, an allowed [👥 user](./access.md#-users) can also provoke the bot to respond to **any** thread or reply chain by [mentioning](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) the bot (e.g. `@baibot Hello!`). The following screenshots demonstrate this behavior:
- [🖼️ On-demand involvement in the room](./screenshots/text-generation-prefix-requirement.webp)
- [🖼️ On-demand involvement in a thread](./screenshots/text-generation-on-demand-thread-involvement.webp) (the Alice user in this example is not an allowed user, yet her messages are still considered as part of the conversation context)
- [🖼️ On-demand involvement in a reply chain](./screenshots/text-generation-on-demand-reply-involvement.webp) (the Alice user in this example is not an allowed user, yet her messages are still considered as part of the conversation context)
💡 **NOTE**: Normally, the bot **only considers messages from allowed [👥 Users](./access.md#-users)** and ignores all other messages when responding. However, **when the bot is explicitly invoked (via mention)** in a thread or reply chain, **it will consider all messages** in the thread and reply chain (even those from foreign users) as part of the conversation context.
### 🗣️ Text-to-Speech
Text-to-Speech is the bot's ability to **turn text messages into voice messages**.
@@ -118,27 +138,45 @@ To operate in this mode, you can:
- adjust the [🦻 Speech-to-Text / 🪄 Flow Type](./configuration/speech-to-text.md#-flow-type) setting to make the bot only transcribe (without doing [💬 Text Generation](#-text-generation)): `!bai config room speech-to-text set-flow-type only_transcribe`
- optionally adjust [🦻 Speech-to-Text / 🪄 Message Type for non-threaded only-transcribed messages](./configuration/speech-to-text.md#-message-type-for-non-threaded-only-transcribed-messages), if you'd like to bot to send messages of type `notice` (for better compatibility with other bots in the room) instead of sending regular `text` messages (default)
### 🖌️ Image Generation
Image generation is the bot's ability to **generate images** based on text prompts.
### Image Generation
See a [🖼️ Screenshot of the Image Generation feature](./screenshots/image-generation.webp).
#### 🖌️ Image Creation
Image creation is the bot's ability to **create images** based on text prompts.
See a [🖼️ Screenshot of the Image Creation feature](./screenshots/image-creation.webp).
You may also wish to see:
- [🛠️ Configuration / 🖌️ Image Generation](./configuration/image-generation.md) for configuration options related to Image Generation
- [📖 Usage / 🖌️ Image Generation](./usage.md#-image-generation) section for more details on how to use the bot for Image Generation in a room
- [🫵 Sticker Generation](#-sticker-generation) - a special case of Image Generation
- [📖 Usage / Image Generation / 🖌️ Creating Images](./usage.md#-creating-images) section for more details on how to use the bot for Image Creation in a room
- [🖌️ Image Editing](#️-image-editing) - another image generation feature
- [🫵 Sticker Creation](#-sticker-creation) - a special case of Image Creation
### 🫵 Sticker Generation
#### 🎨 Image Editing
Sticker generation is the bot's ability to **generate sticker** images based on text prompts. It's a special case of [🖌️ Image Generation](#️-image-generation).
Image editing is the bot's ability to **edit images** based on a prompt and one or more existing images.
See a [🖼️ Screenshot of the Sticker Generation feature](./screenshots/sticker-generation.webp).
See a [🖼️ Screenshot of the Image Editing feature (manipulating a single image)](./screenshots/image-editing-single-image.webp) and a [🖼️ Screenshot of the Image Editing feature (manipulating multiple images)](./screenshots/image-editing-multiple-images.webp).
See [📖 Usage / 🖌️ Image Generation / Generating Stickers](./usage.md#generating-stickers) for details.
You may also wish to see:
- [🛠️ Configuration / 🖌️ Image Generation](./configuration/image-generation.md) for configuration options related to Image Generation
- [📖 Usage / Image Generation / 🎨 Editing images](./usage.md#-editing-images) section for more details on how to use the bot for Image Editing in a room
- [🖌️ Image Creation](#️-image-creation) - another image generation feature
#### 🫵 Sticker Creation
Sticker generation is the bot's ability to **generate sticker** images based on text prompts. It's a special case of [🖌️ Image Creation](#️-image-creation).
See a [🖼️ Screenshot of the Sticker Creation feature](./screenshots/sticker-generation.webp).
See [📖 Usage / Image Generation / 🫵 Creating Stickers](./usage.md#-creating-stickers) for details.
### 🔒 Encryption

View File

@@ -23,7 +23,7 @@ The list of supported providers is below.
### How to choose a provider
If you're not sure which provider to start with, **we recommend [OpenAI](#openai)** as it's the most popular and has the **widest range of capabilities**: [💬 text-generation](./features.md#-text-generation), [🖌️ image-generation](./features.md#️-image-generation), [🦻 speech-to-text](./features.md#-speech-to-text), [🗣️ text-to-speech](./features.md#️-text-to-speech).
If you're not sure which provider to start with, **we recommend [OpenAI](#openai)** as it's the most popular and has the **widest range of capabilities**: [💬 text-generation](./features.md#-text-generation) (no vision), [🖌️ image-generation](./features.md#️image-generation), [🦻 speech-to-text](./features.md#-speech-to-text), [🗣️ text-to-speech](./features.md#️-text-to-speech).
You don't need to choose just one though. The bot supports [mixing & matching models](./features.md#-mixing--matching-models), so you can use multiple providers at the same time.
@@ -47,7 +47,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `anthropic`
- 🔗 Links: [🏠 Home page](https://www.anthropic.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Anthropic), [👤 Sign up](https://console.anthropic.com/), [📋 Models list](https://docs.anthropic.com/en/docs/about-claude/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (incl. vision)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local anthropic my-anthropic-agent`
- create a global agent: `!bai agent create-global anthropic my-anthropic-agent`
@@ -61,7 +61,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `groq`
- 🔗 Links: [🏠 Home page](https://groq.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Groq), [👤 Sign up](https://console.groq.com/login), [📋 Models list](https://console.groq.com/docs/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local groq my-groq-agent`
- create a global agent: `!bai agent create-global groq my-groq-agent`
@@ -75,7 +75,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `localai`
- 🔗 Links: [🏠 Home page](https://localai.io/), [📋 Models list](https://localai.io/gallery.html)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local localai my-localai-agent`
- create a global agent: `!bai agent create-global localai my-localai-agent`
@@ -89,7 +89,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `mistral`
- 🔗 Links: [🏠 Home page](https://mistral.ai/), [🌐 Wiki](https://en.wikipedia.org/wiki/Mistral_AI), [👤 Sign up](https://auth.mistral.ai/ui/registration), [📋 Models list](https://docs.mistral.ai/getting-started/models/)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local mistral my-mistral-agent`
- create a global agent: `!bai agent create-global mistral my-mistral-agent`
@@ -103,7 +103,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `ollama`
- 🔗 Links: [🏠 Home page](https://ollama.com/), [📋 Models list](https://ollama.com/library)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local ollama my-ollama-agent`
- create a global agent: `!bai agent create-global ollama my-ollama-agent`
@@ -120,12 +120,15 @@ For services which are not fully compatible with the OpenAI API, consider using
- 🆔 Identifier: `openai`
- 🔗 Links: [🏠 Home page](https://openai.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/OpenAI), [👤 Sign up](https://platform.openai.com/signup), [📋 Models list](https://platform.openai.com/docs/models)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-generation), [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (incl. vision), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local openai my-openai-agent`
- create a global agent: `!bai agent create-global openai my-openai-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/openai.yml).
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which:
- in the general case looks [like this](./sample-provider-configs/openai.yml)
- for the [o1](https://platform.openai.com/docs/models/o1) models needs to look [like this](./sample-provider-configs/openai-o1.yml)
### OpenAI Compatible
@@ -137,7 +140,7 @@ Some of these popular services already have **shortcut** providers (leading to t
This provider is just as featureful as the [OpenAI](#openai) provider, but is more compatible with services which do not fully adhere to the [OpenAI API spec](https://github.com/openai/openai-openapi/).
- 🆔 Identifier: `openai-compatible`
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-generation), [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (no vision), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local openai-compatible my-openai-compatible-agent`
- create a global agent: `!bai agent create-global openai-compatible my-openai-compatible-agent`
@@ -151,7 +154,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- 🆔 Identifier: `openrouter`
- 🔗 Links: [🏠 Home page](https://openrouter.ai/), [👤 Sign up](https://openrouter.ai/), [📋 Models list](https://openrouter.ai/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local openrouter my-openrouter-agent`
- create a global agent: `!bai agent create-global openrouter my-openrouter-agent`
@@ -165,7 +168,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- 🆔 Identifier: `together-ai`
- 🔗 Links: [🏠 Home page](https://www.together.ai/), [👤 Sign up](https://api.together.ai/signup), [📋 Models list](https://api.together.xyz/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local together-ai my-together-ai-agent`
- create a global agent: `!bai agent create-global together-ai my-together-ai-agent`

View File

@@ -1,8 +1,8 @@
base_url: https://api.anthropic.com/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: claude-3-5-sonnet-20240620
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
model_id: claude-3-7-sonnet-20250219
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 8192
max_context_tokens: 204800

View File

@@ -2,7 +2,7 @@ base_url: https://api.groq.com/openai/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: llama3-70b-8192
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 4096
max_context_tokens: 131072

View File

@@ -2,7 +2,7 @@ base_url: http://my-localai-self-hosted-service:8080/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: gpt-4
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 4096
max_context_tokens: 128000

View File

@@ -2,7 +2,7 @@ base_url: https://api.mistral.ai/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: mistral-large-latest
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 4096
max_context_tokens: 128000

View File

@@ -2,7 +2,7 @@ base_url: http://my-ollama-self-hosted-service:11434/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: gemma2:2b
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 4096
max_context_tokens: 128000

View File

@@ -2,7 +2,7 @@ base_url: ''
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: some-model
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 4096
max_context_tokens: 128000

View File

@@ -0,0 +1,24 @@
base_url: https://api.openai.com/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: o1-mini
# o1 models do not support a system prompt
prompt: null
temperature: 1.0
# o1 models do not support max_response_tokens.
# They use `max_completion_tokens` as an alternative
max_response_tokens: null
max_completion_tokens: 16384
max_context_tokens: 128000
speech_to_text:
model_id: whisper-1
text_to_speech:
model_id: tts-1-hd
voice: onyx
speed: 1.0
response_format: opus
image_generation:
model_id: gpt-image-1
style: null
size: 1024x1024
quality: null

View File

@@ -1,8 +1,8 @@
base_url: https://api.openai.com/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: gpt-4o-2024-08-06
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
model_id: gpt-4.1
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 16384
max_context_tokens: 128000
@@ -14,7 +14,7 @@ text_to_speech:
speed: 1.0
response_format: opus
image_generation:
model_id: dall-e-3
style: vivid
model_id: gpt-image-1
style: null
size: 1024x1024
quality: standard
quality: null

View File

@@ -2,7 +2,7 @@ base_url: https://openrouter.ai/api/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: mattshumer/reflection-70b:free
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 2048
max_context_tokens: 8192

View File

@@ -2,7 +2,7 @@ base_url: https://api.together.xyz/v1
api_key: YOUR_API_KEY_HERE
text_generation:
model_id: meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
temperature: 1.0
max_response_tokens: 2048
max_context_tokens: 8192

Binary file not shown.

After

Width:  |  Height:  |  Size: 298 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 339 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 285 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 684 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 13 KiB

After

Width:  |  Height:  |  Size: 22 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 92 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 60 KiB

View File

@@ -11,10 +11,13 @@ This is related to the [💬 Text Generation](./features.md#-text-generation) fe
If there's a text-generation handler agent configured, the bot **may** respond to messages sent in the room.
🖼️ See screenshots of:
Some models also support vision, so you may be able to mix text and images in the same conversation.
- the [default Text Generation flow](./screenshots/text-generation.webp) for 1:1 rooms
- the [Text Generation flow in multi-user rooms](./screenshots/text-generation-prefix-requirement.webp) (where the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting is auto-configured to "required")
See screenshots of:
- 🖼️ [the default Text Generation flow](./screenshots/text-generation.webp) in 1:1 rooms
- 🖼️ [the Text Generation flow in multi-user rooms](./screenshots/text-generation-prefix-requirement.webp) (where the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting is auto-configured to "required")
- the [on-demand involvement](./features.md#on-demand-involvement) feature
Whether the bot responds depends on:
@@ -24,9 +27,9 @@ Whether the bot responds depends on:
- (🎨 agent capabilities) whether the configured `text-generation` (or `catch-all`) handler agent actually supports text-generation. The provider may lack support for this feature or it may be disabled in the [🤖 agents](./agents.md) configuration
- (the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting) whether a prefix (e.g. `!bai`) is required in front of messages sent to the room. For multi-user rooms, this setting defaults to "required"
- (the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting) whether a prefix (e.g. `!bai`) or user mention (e.g. `@baibot`) is required for messages sent to the room. For multi-user rooms, this setting defaults to "required". See [🌟 Features / 💬 Text Generation / On-demand involvement](./features.md#on-demand-involvement) for details.
Room messages start a threaded conversation where you can continue back-and-forth communication with the bot.
Room messages start a threaded conversation where you can continue back-and-forth communication with the bot. Using [on-demand involvement](./features.md#on-demand-involvement), you can can also mention the bot to provoke it to get involved in any conversation thread or reply chain.
Unless you've enabled the [♻️ Context Management](./features.md#️-context-management) feature, all messages will be sent to the agent's API each time. If the context management feature is enabled, older messages may be dropped.
@@ -63,18 +66,18 @@ The speech-to-text feature triggers automatically by default, but can be adjuste
If all your messages are in the same language, you can improve accuracy & latency by configuring the language (see [🦻 Speech-to-Text / 🔤 Language](./configuration/speech-to-text.md#-language)).
### 🖌️ Image Generation
This is related to the [🖌️ Image Generation](./features.md#️-image-generation) feature.
### Image Generation
This feature is not configurable at the moment. The configuration (size, quality, style) specified at the [🤖 agent](./agents.md) level will be used.
Capabilities depend on the [☁️ provider](./providers.md) and model used.
#### Generating images
Simply send a command like `!bai image A beautiful sunset over the ocean` and the bot will start a threaded conversation and post an image based on your prompt.
#### 🖌️ Creating images
See a [🖼️ Screenshot of the Image Generation feature](./screenshots/image-generation.webp).
Simply send a command like `!bai image create A beautiful sunset over the ocean` and the bot will start a threaded conversation and post an image based on your prompt.
See a [🖼️ Screenshot of the Image Creation feature](./screenshots/image-creation.webp).
You can then, respond in the same message thread with:
@@ -82,15 +85,29 @@ You can then, respond in the same message thread with:
- a message saying `again`, to generate one more image with the current prompt.
#### Generating stickers
#### 🎨 Editing images
A variation of [generating images](#generating-images) is to generate "sticker images".
Simply send a command like `!bai image edit Turn the following image into an anime-style drawing` and the bot will start a threaded conversation asking for more details.
See a [🖼️ Screenshot of the Sticker Generation feature](./screenshots/sticker-generation.webp).
See a [🖼️ Screenshot of the Image Editing feature (manipulating a single image)](./screenshots/image-editing-single-image.webp) and a [🖼️ Screenshot of the Image Editing feature (manipulating multiple images)](./screenshots/image-editing-multiple-images.webp).
To generate a sticker, send a command like `!bai sticker A huge ramen bowl with lots of chashu and a mountain of beansprouts on top`.
You can then, respond in the same message thread with:
The difference from [generating images](#generating-images) is that the bot will:
- more messages, to add more criteria to your prompt.
- one or more images, to provide the images that the bot will operate on.
- a message saying `go`, to start the image generation process.
- a message saying `again`, to prompt the bot to generate one more image edit with the current prompt.
#### 🫵 Creating stickers
A variation of [creating images](#creating-images) is to create "sticker images".
See a [🖼️ Screenshot of the Sticker Creation feature](./screenshots/sticker-generation.webp).
To create a sticker, send a command like `!bai sticker A huge ramen bowl with lots of chashu and a mountain of beansprouts on top`.
The difference from [creating images](#creating-images) is that the bot will:
- generate a smaller-resolution image (currently hardcoded to `256x256`) - smaller/quicker, but still good enough for a sticker
- potentially switch to a different (cheaper or otherwise more suitable) model, if available

View File

@@ -32,6 +32,10 @@ user:
# Command prefix. Leave empty to use the default (!bai).
command_prefix: "!bai"
room:
# Whether the bot should send an introduction message after joining a room.
post_join_self_introduction_enabled: true
access:
# Space-separated list of MXID patterns which specify who is an admin.
admin_patterns:
@@ -72,10 +76,12 @@ agents:
# base_url: https://api.openai.com/v1
# api_key: ""
# text_generation:
# model_id: gpt-4o-2024-08-06
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
# model_id: gpt-4o
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
# temperature: 1.0
# max_response_tokens: 16384
# # Reasoning models need to use `max_completion_tokens` instead of `max_response_tokens`.
# max_completion_tokens: ~
# max_context_tokens: 128000
# speech_to_text:
# model_id: whisper-1
@@ -97,7 +103,7 @@ agents:
# api_key: null
# text_generation:
# model_id: gpt-4
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
# temperature: 1.0
# max_response_tokens: 16384
# max_context_tokens: 128000
@@ -122,7 +128,7 @@ agents:
# api_key: null
# text_generation:
# model_id: "gemma2:2b"
# prompt: "You are an assistant based on the gemma2:2b model. Be brief in your responses."
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
# temperature: 1.0
# max_response_tokens: 4096
# max_context_tokens: 128000

View File

@@ -1,6 +1,6 @@
services:
postgres:
image: docker.io/postgres:16.3-alpine
image: docker.io/postgres:16.8-alpine
user: ${UID}:${GID}
restart: unless-stopped
environment:
@@ -13,7 +13,7 @@ services:
- /etc/passwd:/etc/passwd:ro
synapse:
image: ghcr.io/element-hq/synapse:v1.114.0
image: ghcr.io/element-hq/synapse:v1.129.0
user: "${UID}:${GID}"
restart: unless-stopped
entrypoint: python
@@ -26,14 +26,20 @@ services:
- ./synapse/media-store:/media-store
element-web:
image: docker.io/vectorim/element-web:v1.11.77
image: ghcr.io/element-hq/element-web:v1.11.100
user: "${UID}:${GID}"
restart: unless-stopped
environment:
ELEMENT_WEB_PORT: 8080
ports:
- "${SERVICE_ELEMENT_WEB_BIND_PORT_HTTP}:8080"
volumes:
- ../../etc/services/core/element-web/nginx.conf:/etc/nginx/nginx.conf:ro
- ../../etc/services/core/element-web/config.json:/app/config.json:ro
tmpfs:
- /var/cache/nginx:rw,mode=777
- /var/run:rw,mode=777
- /tmp/element-web-config:rw,mode=777
- /etc/nginx/conf.d:rw,mode=777
networks:
default:

View File

@@ -3,7 +3,7 @@
"default_is_url": "https://vector.im",
"integrations_ui_url": "https://scalar.vector.im/",
"integrations_rest_url": "https://scalar.vector.im/api",
"bug_report_endpoint_url": "https://riot.im/bugreports/submit",
"bug_report_endpoint_url": "https://element.io/bugreports/submit",
"enableLabs": true,
"roomDirectory": {
"servers": [

View File

@@ -1,60 +0,0 @@
# This is a custom nginx configuration file that we use in the container (instead of the default one),
# because it allows us to run nginx with a non-root user.
#
# For this to work, the default vhost file (`/etc/nginx/conf.d/default.conf`) also needs to be removed.
# (mounting `/dev/null` over `/etc/nginx/conf.d/default.conf` works well)
#
# The following changes have been done compared to a default nginx configuration file:
# - default server port is changed (80 -> 8080), so that a non-root user can bind it
# - various temp paths are changed to `/tmp`, so that a non-root user can write to them
# - the `user` directive was removed, as we don't want nginx to switch users
worker_processes 1;
error_log /var/log/nginx/error.log warn;
pid /tmp/nginx.pid;
events {
worker_connections 1024;
}
http {
client_body_temp_path /tmp/client_body_temp;
proxy_temp_path /tmp/proxy_temp;
fastcgi_temp_path /tmp/fastcgi_temp;
uwsgi_temp_path /tmp/uwsgi_temp;
scgi_temp_path /tmp/scgi_temp;
include /etc/nginx/mime.types;
default_type application/octet-stream;
log_format main '$remote_addr - $remote_user [$time_local] "$request" '
'$status $body_bytes_sent "$http_referer" '
'"$http_user_agent" "$http_x_forwarded_for"';
access_log /var/log/nginx/access.log main;
sendfile on;
#tcp_nopush on;
keepalive_timeout 65;
#gzip on;
server {
listen 8080;
server_name localhost;
location / {
root /usr/share/nginx/html;
index index.html index.htm;
}
error_page 500 502 503 504 /50x.html;
location = /50x.html {
root /usr/share/nginx/html;
}
}
}

View File

@@ -579,7 +579,7 @@ rc_login:
#
#federation_rr_transactions_per_room_per_second: 50
enable_authenticated_media: true
# Directory where uploaded images and attachments are stored.
#

View File

@@ -1,6 +1,6 @@
services:
ollama:
image: docker.io/ollama/ollama:0.3.9
image: docker.io/ollama/ollama:0.6.8
restart: unless-stopped
ports:
- "${SERVICE_OLLAMA_BIND_PORT_HTTP}:11434"

View File

@@ -32,6 +32,10 @@ run-in-container *extra_args: app-container-prepare build-container-image-debug
test *extra_args:
RUST_BACKTRACE=1 cargo test {{ extra_args }}
# Formats the code
fmt:
RUST_BACKTRACE=1 cargo fmt --all
# Builds a debug binary (target/debug/*)
build-debug *extra_args:
RUST_BACKTRACE=1 cargo build {{ extra_args }}

View File

@@ -1,6 +1,6 @@
use super::{
provider::{self, ControllerType},
AgentDefinition, AgentProvider, PublicIdentifier,
provider::{self, ControllerType},
};
// Dead-code is allowed. We do not use these enum struct payloads directly,

View File

@@ -1,7 +1,7 @@
use super::instantiation;
use super::instantiation::AgentInstance;
use super::AgentDefinition;
use super::PublicIdentifier;
use super::instantiation;
use super::instantiation::AgentInstance;
use crate::entity::RoomConfigContext;
#[derive(Debug)]

View File

@@ -11,15 +11,15 @@ pub use manager::Manager;
pub use definition::AgentDefinition;
pub use instantiation::create_from_provider_and_yaml_value_config;
pub use instantiation::default_config_for_provider;
pub use instantiation::AgentInstance;
pub use instantiation::Error as AgentInstantiationError;
pub use instantiation::Result as AgentInstantiationResult;
pub use instantiation::create_from_provider_and_yaml_value_config;
pub use instantiation::default_config_for_provider;
pub use provider::{AgentProvider, AgentProviderInfo, ControllerTrait};
pub use purpose::AgentPurpose;
pub(super) fn default_prompt() -> &'static str {
"You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
"You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
}

View File

@@ -1,7 +1,5 @@
use serde::{Deserialize, Serialize};
use anthropic_rs::models::claude::ClaudeModel;
use crate::agent::{default_prompt, provider::ConfigTrait};
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -28,6 +26,9 @@ impl ConfigTrait for Config {
if self.base_url.is_empty() {
return Err("The base URL must not be empty.".to_owned());
}
if !self.base_url.ends_with("/v1") {
return Err("The base URL must end with '/v1'.".to_owned());
}
if self.api_key.is_empty() {
return Err("The API key must not be empty.".to_owned());
}
@@ -67,5 +68,5 @@ impl Default for TextGenerationConfig {
}
fn default_text_model_id() -> String {
ClaudeModel::Claude35Sonnet.as_str().to_owned()
"claude-3-7-sonnet-20250219".to_owned()
}

View File

@@ -1,30 +1,28 @@
use std::fmt::Debug;
use std::str::FromStr;
use std::sync::Arc;
use anthropic_rs::completion::message::ContentType;
use anthropic_rs::{
client::Client as AnthropicClient, config::Config as AnthropicConfig,
models::claude::ClaudeModel,
};
use anthropic::client::{Client, ClientBuilder};
use anthropic::types::ContentBlock;
use super::super::ControllerTrait;
use crate::agent::provider::entity::{
ImageGenerationResult, PingResult, TextGenerationParams, TextGenerationResult,
TextToSpeechParams, TextToSpeechResult,
};
use crate::agent::provider::{ImageGenerationParams, SpeechToTextParams, SpeechToTextResult};
use crate::agent::AgentPurpose;
use crate::agent::provider::entity::{
ImageEditResult, ImageGenerationResult, ImageSource, PingResult, TextGenerationParams,
TextGenerationResult, TextToSpeechParams, TextToSpeechResult,
};
use crate::agent::provider::{
ImageEditParams, ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
};
use crate::conversation::llm::{
shorten_messages_list_to_context_size, Author as LLMAuthor, Conversation as LLMConversation,
Message as LLMMessage,
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
MessageContent as LLMMessageContent, shorten_messages_list_to_context_size,
};
use crate::strings;
use super::config::Config;
struct ControllerInner {
client: AnthropicClient,
client: Client,
}
#[derive(Clone)]
@@ -43,18 +41,20 @@ impl Debug for Controller {
impl Controller {
pub fn new(config: Config) -> anyhow::Result<Self> {
let anthropic_config =
AnthropicConfig::new(config.api_key.clone()).with_base_url(config.base_url.clone());
// The previous library that we used expected a base URL that ends with "/v1"
// (e.g. "https://api.anthropic.com/v1"), while the new one doesn't.
//
// To keep backward compatibility, we don't ask people to change their configuration
// and rather adapt by removing the "/v1" from the base URL.
if !config.base_url.ends_with("/v1") {
return Err(anyhow::anyhow!("base_url must end with '/v1'"));
}
let client = match AnthropicClient::new(anthropic_config) {
Ok(client) => client,
Err(err) => {
return Err(anyhow::anyhow!(
"Failed to create Anthropic client: {}",
err.to_string()
));
}
};
let base_url = &config.base_url[..config.base_url.len() - 3];
let client = ClientBuilder::default()
.api_base(base_url.to_string())
.api_key(config.api_key.clone())
.build()?;
Ok(Self {
config,
@@ -71,7 +71,8 @@ impl ControllerTrait for Controller {
let messages = vec![LLMMessage {
author: LLMAuthor::User,
message_text: "Hello!".to_string(),
content: LLMMessageContent::Text("Hello!".to_string()),
timestamp: chrono::Utc::now(),
}];
let conversation = LLMConversation { messages };
@@ -107,7 +108,8 @@ impl ControllerTrait for Controller {
} else {
Some(LLMMessage {
author: LLMAuthor::Prompt,
message_text: prompt_text,
content: LLMMessageContent::Text(prompt_text),
timestamp: chrono::Utc::now(),
})
};
@@ -129,7 +131,7 @@ impl ControllerTrait for Controller {
&text_generation_config.model_id,
&prompt_message,
conversation_messages,
text_generation_config.max_response_tokens,
Some(text_generation_config.max_response_tokens),
text_generation_config.max_context_tokens,
);
@@ -140,29 +142,19 @@ impl ControllerTrait for Controller {
let mut request = super::utils::create_anthropic_message_request(conversation_messages);
let model = match ClaudeModel::from_str(&text_generation_config.model_id) {
Ok(model) => model,
Err(err) => {
tracing::debug!(?err, "Failed to parse model ID");
return Err(anyhow::anyhow!(
"Failed to parse model ID: {}",
&text_generation_config.model_id
));
}
};
let temperature = params
.temperature_override
.unwrap_or(text_generation_config.temperature);
if let Some(prompt_message) = prompt_message {
request.system = Some(prompt_message.message_text);
if let LLMMessageContent::Text(text) = &prompt_message.content {
request.system = text.clone();
}
}
request.model = model;
request.temperature = Some(temperature);
request.max_tokens = text_generation_config.max_response_tokens;
request.model = text_generation_config.model_id.clone();
request.temperature = Some(temperature as f64);
request.max_tokens = text_generation_config.max_response_tokens as usize;
if let Ok(request_as_json) = serde_json::to_string(&request) {
tracing::trace!(
@@ -173,19 +165,20 @@ impl ControllerTrait for Controller {
);
}
let response = self.inner.client.create_message(request).await?;
let response = self.inner.client.messages(request).await?;
tracing::trace!(?response, "Got response from Anthropic create message API");
// response.content usually contains a single element, but we support handling multiple to account for all possibilities
let mut text_parts = vec![];
for content in response.content {
let content_type = content.content_type;
match content_type {
ContentType::Text => {
text_parts.push(content.text);
} // There are no other content types to handle yet, but there may be in the future
match content {
ContentBlock::Text { text } => {
text_parts.push(text);
}
ContentBlock::Image { .. } => {
text_parts.push("The model responded with an image".to_string());
}
}
}
@@ -217,6 +210,15 @@ impl ControllerTrait for Controller {
Err(anyhow::anyhow!("Image generation not supported"))
}
async fn create_image_edit(
&self,
_prompt: &str,
_images: Vec<ImageSource>,
_params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
Err(anyhow::anyhow!("Image editing is not supported"))
}
async fn text_to_speech(
&self,
_input: &str,

View File

@@ -7,8 +7,8 @@ pub use controller::Controller;
use super::super::AgentInstantiationError;
use super::super::AgentInstantiationResult;
use super::controller::ControllerType;
use super::ConfigTrait;
use super::controller::ControllerType;
pub fn create_controller_from_yaml_value_config(
agent_id: &str,

View File

@@ -1,8 +1,12 @@
use anthropic_rs::completion::message::{Content, ContentType, Message, MessageRequest, Role};
use anthropic::types::{
ContentBlock, ImageSource, Message, MessagesRequest, MessagesRequestBuilder, Role,
};
use crate::conversation::llm::{Author as LLMAuthor, Message as LLMMessage};
use crate::conversation::llm::{
Author as LLMAuthor, Message as LLMMessage, MessageContent as LLMMessageContent,
};
pub(super) fn create_anthropic_message_request(llm_messages: Vec<LLMMessage>) -> MessageRequest {
pub(super) fn create_anthropic_message_request(llm_messages: Vec<LLMMessage>) -> MessagesRequest {
let mut messages = vec![];
for message in llm_messages {
@@ -14,19 +18,26 @@ pub(super) fn create_anthropic_message_request(llm_messages: Vec<LLMMessage>) ->
}
};
let content = vec![Content {
content_type: ContentType::Text,
text: message.message_text,
}];
let content = match &message.content {
LLMMessageContent::Text(text) => vec![ContentBlock::Text { text: text.clone() }],
LLMMessageContent::Image(image_details) => {
vec![ContentBlock::Image {
source: ImageSource::Base64 {
media_type: image_details.mime.to_string(),
data: crate::utils::base64::base64_encode(&image_details.data),
},
}]
}
};
let message = Message { role, content };
messages.push(message);
}
MessageRequest {
stream: false,
messages,
..Default::default()
}
MessagesRequestBuilder::default()
.messages(messages)
.stream(false)
.build()
.expect("Failed to build messages request")
}

View File

@@ -1,11 +1,11 @@
use crate::{agent::AgentPurpose, conversation::llm::Conversation};
use super::{
ImageEditParams, ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
entity::{
ImageGenerationResult, PingResult, TextGenerationParams, TextGenerationResult,
TextToSpeechParams, TextToSpeechResult,
ImageEditResult, ImageGenerationResult, ImageSource, PingResult, TextGenerationParams,
TextGenerationResult, TextToSpeechParams, TextToSpeechResult,
},
ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
};
pub trait ControllerTrait {
@@ -42,6 +42,13 @@ pub trait ControllerTrait {
params: ImageGenerationParams,
) -> impl std::future::Future<Output = anyhow::Result<ImageGenerationResult>> + Send;
fn create_image_edit(
&self,
prompt: &str,
images: Vec<ImageSource>,
params: ImageEditParams,
) -> impl std::future::Future<Output = anyhow::Result<ImageEditResult>> + Send;
fn text_to_speech(
&self,
text: &str,
@@ -166,6 +173,25 @@ impl ControllerTrait for ControllerType {
}
}
async fn create_image_edit(
&self,
prompt: &str,
images: Vec<ImageSource>,
params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
match &self {
ControllerType::OpenAI(controller) => {
controller.create_image_edit(prompt, images, params).await
}
ControllerType::OpenAICompat(controller) => {
controller.create_image_edit(prompt, images, params).await
}
ControllerType::Anthropic(controller) => {
controller.create_image_edit(prompt, images, params).await
}
}
}
async fn text_to_speech(
&self,
text: &str,

View File

@@ -67,9 +67,8 @@ impl AgentProvider {
wiki_url: Some("https://en.wikipedia.org/wiki/Anthropic"),
sign_up_url: Some("https://console.anthropic.com/"),
models_list_url: Some("https://docs.anthropic.com/en/docs/about-claude/models"),
supported_purposes: vec![
AgentPurpose::TextGeneration,
],
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: true,
},
Self::Groq => AgentProviderInfo {
id: Self::Groq.to_static_str(),
@@ -79,10 +78,8 @@ impl AgentProvider {
wiki_url: Some("https://en.wikipedia.org/wiki/Groq"),
sign_up_url: Some("https://console.groq.com/login"),
models_list_url: Some("https://console.groq.com/docs/models"),
supported_purposes: vec![
AgentPurpose::TextGeneration,
AgentPurpose::SpeechToText,
],
supported_purposes: vec![AgentPurpose::TextGeneration, AgentPurpose::SpeechToText],
text_generation_supports_vision: false,
},
Self::LocalAI => AgentProviderInfo {
id: Self::LocalAI.to_static_str(),
@@ -97,6 +94,7 @@ impl AgentProvider {
AgentPurpose::TextToSpeech,
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: false,
},
Self::Mistral => AgentProviderInfo {
id: Self::Mistral.to_static_str(),
@@ -106,9 +104,8 @@ impl AgentProvider {
wiki_url: Some("https://en.wikipedia.org/wiki/Mistral_AI"),
sign_up_url: Some("https://auth.mistral.ai/ui/registration"),
models_list_url: Some("https://docs.mistral.ai/getting-started/models/"),
supported_purposes: vec![
AgentPurpose::TextGeneration,
],
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
},
Self::Ollama => AgentProviderInfo {
id: Self::Ollama.to_static_str(),
@@ -118,9 +115,8 @@ impl AgentProvider {
wiki_url: None,
sign_up_url: None,
models_list_url: Some("https://ollama.com/library"),
supported_purposes: vec![
AgentPurpose::TextGeneration,
],
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
},
Self::OpenAI => AgentProviderInfo {
id: Self::OpenAI.to_static_str(),
@@ -136,6 +132,7 @@ impl AgentProvider {
AgentPurpose::TextToSpeech,
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: true,
},
Self::OpenAICompat => AgentProviderInfo {
id: Self::OpenAICompat.to_static_str(),
@@ -151,6 +148,7 @@ impl AgentProvider {
AgentPurpose::TextToSpeech,
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: false,
},
Self::OpenRouter => AgentProviderInfo {
id: Self::OpenRouter.to_static_str(),
@@ -160,9 +158,8 @@ impl AgentProvider {
wiki_url: None,
sign_up_url: Some("https://openrouter.ai/"),
models_list_url: Some("https://openrouter.ai/models"),
supported_purposes: vec![
AgentPurpose::TextGeneration,
],
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
},
Self::TogetherAI => AgentProviderInfo {
id: Self::TogetherAI.to_static_str(),
@@ -172,9 +169,8 @@ impl AgentProvider {
wiki_url: None,
sign_up_url: Some("https://api.together.ai/signup"),
models_list_url: Some("https://api.together.xyz/models"),
supported_purposes: vec![
AgentPurpose::TextGeneration,
],
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
},
}
}
@@ -195,4 +191,5 @@ pub struct AgentProviderInfo {
pub sign_up_url: Option<&'static str>,
pub models_list_url: Option<&'static str>,
pub supported_purposes: Vec<AgentPurpose>,
pub text_generation_supports_vision: bool,
}

View File

@@ -1,3 +1,5 @@
use mxlink::mime;
#[derive(Default)]
pub struct ImageGenerationParams {
pub size_override: Option<String>,
@@ -26,6 +28,40 @@ impl ImageGenerationParams {
pub struct ImageGenerationResult {
pub bytes: Vec<u8>,
pub mime_type: mxlink::mime::Mime,
pub mime_type: mime::Mime,
pub revised_prompt: Option<String>,
}
#[derive(Default)]
pub struct ImageEditParams {}
pub struct ImageEditResult {
pub bytes: Vec<u8>,
pub mime_type: mime::Mime,
}
pub struct ImageSource {
pub filename: String,
pub bytes: Vec<u8>,
pub mime_type: mime::Mime,
}
impl ImageSource {
pub fn new(filename: String, bytes: Vec<u8>, mime_type: mime::Mime) -> Self {
Self {
filename,
bytes,
mime_type,
}
}
}
impl From<ImageSource> for async_openai::types::ImageInput {
fn from(value: ImageSource) -> Self {
async_openai::types::ImageInput::from_vec_u8(
value.filename,
value.bytes,
value.mime_type.to_string(),
)
}
}

View File

@@ -1,12 +1,14 @@
mod agent_provider;
mod image_generation;
mod image;
mod ping;
mod speech_to_text;
mod text_generation;
mod text_to_speech;
pub use agent_provider::{AgentProvider, AgentProviderInfo};
pub use image_generation::{ImageGenerationParams, ImageGenerationResult};
pub use image::{
ImageEditParams, ImageEditResult, ImageGenerationParams, ImageGenerationResult, ImageSource,
};
pub use ping::PingResult;
pub use speech_to_text::{SpeechToTextParams, SpeechToTextResult};
pub use text_generation::{

View File

@@ -7,17 +7,33 @@ pub struct TextGenerationPromptVariables {
impl Default for TextGenerationPromptVariables {
fn default() -> Self {
Self::new("unnamed", "unknown-model", Utc::now())
let now = Utc::now();
Self::new("unnamed", "unknown-model", now, Some(now))
}
}
impl TextGenerationPromptVariables {
pub fn new(bot_name: &str, model_id: &str, utc_time: DateTime<Utc>) -> Self {
pub fn new(
bot_name: &str,
model_id: &str,
now_time: DateTime<Utc>,
conversation_start_time: Option<DateTime<Utc>>,
) -> Self {
let mut map = HashMap::new();
map.insert("baibot_name".to_string(), bot_name.to_string());
map.insert("baibot_model_id".to_string(), model_id.to_string());
map.insert("baibot_now_utc".to_string(), format_utc_time(utc_time));
map.insert("baibot_now_utc".to_string(), format_utc_time(now_time));
let baibot_conversation_start_time_utc = match conversation_start_time {
Some(conversation_start_time) => format_utc_time(conversation_start_time),
None => "unknown".to_string(),
};
map.insert(
"baibot_conversation_start_time_utc".to_string(),
baibot_conversation_start_time_utc,
);
Self { map }
}
@@ -52,7 +68,18 @@ mod tests {
.with_nanosecond(250000000)
.unwrap();
let variables = TextGenerationPromptVariables::new("baibot", "gpt-4o", now_utc);
let conversation_start_time_utc = Utc
.with_ymd_and_hms(2024, 9, 19, 18, 34, 15)
.unwrap()
.with_nanosecond(250000000)
.unwrap();
let variables = TextGenerationPromptVariables::new(
"baibot",
"gpt-4o",
now_utc,
Some(conversation_start_time_utc),
);
assert_eq!(
variables.map.get("baibot_name"),
@@ -66,9 +93,13 @@ mod tests {
variables.map.get("baibot_now_utc"),
Some(&format_utc_time(now_utc))
);
assert_eq!(
variables.map.get("baibot_conversation_start_time_utc"),
Some(&format_utc_time(conversation_start_time_utc))
);
let prompt = "Hello, I'm {{ baibot_name }} using {{ baibot_model_id }}. The date/time now is {{ baibot_now_utc }}.";
let expected = "Hello, I'm baibot using gpt-4o. The date/time now is 2024-09-20 (Friday), 18:34:15 UTC.";
let prompt = "Hello, I'm {{ baibot_name }} using {{ baibot_model_id }}. The date/time now is {{ baibot_now_utc }} and this conversation started at {{ baibot_conversation_start_time_utc }}.";
let expected = "Hello, I'm baibot using gpt-4o. The date/time now is 2024-09-20 (Friday), 18:34:15 UTC and this conversation started at 2024-09-19 (Thursday), 18:34:15 UTC.";
assert_eq!(variables.format(prompt), expected);
}

View File

@@ -15,7 +15,7 @@ pub fn default_config() -> Config {
if let Some(ref mut config) = config.text_generation.as_mut() {
config.model_id = "llama3-70b-8192".to_owned();
config.max_context_tokens = 131_072;
config.max_response_tokens = 4096;
config.max_response_tokens = Some(4096);
}
if let Some(ref mut config) = config.speech_to_text.as_mut() {

View File

@@ -13,7 +13,7 @@ pub fn default_config() -> Config {
if let Some(ref mut config) = config.text_generation.as_mut() {
config.model_id = "gpt-4".to_owned();
config.max_context_tokens = 128_000;
config.max_response_tokens = 4096;
config.max_response_tokens = Some(4096);
}
if let Some(ref mut config) = config.text_to_speech.as_mut() {

View File

@@ -20,6 +20,7 @@ pub use controller::{ControllerTrait, ControllerType};
pub use config::ConfigTrait;
pub use entity::{
AgentProvider, AgentProviderInfo, ImageGenerationParams, PingResult, SpeechToTextParams,
SpeechToTextResult, TextGenerationParams, TextGenerationPromptVariables, TextToSpeechParams,
AgentProvider, AgentProviderInfo, ImageEditParams, ImageGenerationParams, ImageSource,
PingResult, SpeechToTextParams, SpeechToTextResult, TextGenerationParams,
TextGenerationPromptVariables, TextToSpeechParams,
};

View File

@@ -17,7 +17,7 @@ pub fn default_config() -> Config {
if let Some(ref mut config) = config.text_generation.as_mut() {
config.model_id = "gemma2:2b".to_owned();
config.max_context_tokens = 128_000;
config.max_response_tokens = 4096;
config.max_response_tokens = Some(4096);
}
config

View File

@@ -1,5 +1,6 @@
use serde::{Deserialize, Serialize};
use super::OPENAI_IMAGE_MODEL_GPT_IMAGE_1;
use crate::agent::{default_prompt, provider::ConfigTrait};
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -56,7 +57,10 @@ pub struct TextGenerationConfig {
pub temperature: f32,
#[serde(default)]
pub max_response_tokens: u32,
pub max_response_tokens: Option<u32>,
#[serde(default)]
pub max_completion_tokens: Option<u32>,
#[serde(default)]
pub max_context_tokens: u32,
@@ -68,14 +72,15 @@ impl Default for TextGenerationConfig {
model_id: default_text_model_id(),
prompt: Some(default_prompt().to_owned()),
temperature: super::super::default_temperature(),
max_response_tokens: 16_384,
max_response_tokens: Some(16_384),
max_completion_tokens: None,
max_context_tokens: 128_000,
}
}
}
fn default_text_model_id() -> String {
"gpt-4o-2024-08-06".to_owned()
"gpt-4.1".to_owned()
}
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -145,19 +150,19 @@ pub struct ImageGenerationConfig {
pub model_id: String,
#[serde(default = "default_image_style")]
pub style: async_openai::types::ImageStyle,
pub style: Option<async_openai::types::ImageStyle>,
#[serde(default = "default_image_size")]
pub size: async_openai::types::ImageSize,
#[serde(default = "default_image_quality")]
pub quality: async_openai::types::ImageQuality,
pub quality: Option<async_openai::types::ImageQuality>,
}
impl Default for ImageGenerationConfig {
fn default() -> Self {
Self {
model_id: "dall-e-3".to_owned(),
model_id: OPENAI_IMAGE_MODEL_GPT_IMAGE_1.to_owned(),
style: default_image_style(),
size: default_image_size(),
quality: default_image_quality(),
@@ -177,14 +182,14 @@ impl ImageGenerationConfig {
}
}
fn default_image_style() -> async_openai::types::ImageStyle {
async_openai::types::ImageStyle::Vivid
fn default_image_style() -> Option<async_openai::types::ImageStyle> {
None
}
fn default_image_size() -> async_openai::types::ImageSize {
async_openai::types::ImageSize::S1024x1024
}
fn default_image_quality() -> async_openai::types::ImageQuality {
async_openai::types::ImageQuality::Standard
fn default_image_quality() -> Option<async_openai::types::ImageQuality> {
None
}

View File

@@ -1,41 +1,45 @@
use std::ops::Deref;
use async_openai::{
Client as OpenAIClient,
config::OpenAIConfig,
types::{
ChatCompletionRequestMessage, CreateChatCompletionRequestArgs, CreateImageRequestArgs,
CreateSpeechRequestArgs, CreateTranscriptionRequestArgs,
ChatCompletionRequestMessage, CreateChatCompletionRequestArgs, CreateImageEditRequestArgs,
CreateImageRequestArgs, CreateSpeechRequestArgs, CreateTranscriptionRequestArgs,
DallE2ImageSize, Image, ImageModel, ImageResponseFormat,
},
Client as OpenAIClient,
};
use super::super::ControllerTrait;
use crate::{
agent::{
provider::{
entity::{ImageGenerationResult, PingResult, TextToSpeechParams, TextToSpeechResult},
openai::utils::convert_string_to_enum,
},
AgentPurpose,
agent::provider::{
ImageEditParams, ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
entity::{TextGenerationParams, TextGenerationResult},
},
strings,
conversation::llm::{
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
MessageContent as LLMMessageContent, shorten_messages_list_to_context_size,
},
utils::base64::base64_decode,
};
use crate::{
agent::{
AgentPurpose,
provider::{
entity::{TextGenerationParams, TextGenerationResult},
ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
entity::{
ImageEditResult, ImageGenerationResult, ImageSource, PingResult,
TextToSpeechParams, TextToSpeechResult,
},
openai::utils::convert_string_to_enum,
},
utils::base64_decode,
},
conversation::llm::{
shorten_messages_list_to_context_size, Author as LLMAuthor,
Conversation as LLMConversation, Message as LLMMessage,
},
strings,
};
use super::config::Config;
use super::OPENAI_IMAGE_MODEL_GPT_IMAGE_1;
#[derive(Debug, Clone)]
pub struct Controller {
config: Config,
@@ -62,7 +66,8 @@ impl ControllerTrait for Controller {
let messages = vec![LLMMessage {
author: LLMAuthor::User,
message_text: "Hello!".to_string(),
content: LLMMessageContent::Text("Hello!".to_string()),
timestamp: chrono::Utc::now(),
}];
let conversation = LLMConversation { messages };
@@ -98,7 +103,8 @@ impl ControllerTrait for Controller {
} else {
Some(LLMMessage {
author: LLMAuthor::Prompt,
message_text: prompt_text,
content: LLMMessageContent::Text(prompt_text),
timestamp: chrono::Utc::now(),
})
};
@@ -131,12 +137,22 @@ impl ControllerTrait for Controller {
.temperature_override
.unwrap_or(text_generation_config.temperature);
let request = CreateChatCompletionRequestArgs::default()
.max_tokens(text_generation_config.max_response_tokens)
let mut request_builder = CreateChatCompletionRequestArgs::default();
request_builder
.model(&text_generation_config.model_id)
.temperature(temperature)
.messages(openai_conversation_messages)
.build()?;
.messages(openai_conversation_messages);
if let Some(max_response_tokens) = text_generation_config.max_response_tokens {
request_builder.max_tokens(max_response_tokens);
}
if let Some(max_completion_tokens) = text_generation_config.max_completion_tokens {
request_builder.max_completion_tokens(max_completion_tokens);
}
let request = request_builder.build()?;
if let Ok(request_as_json) = serde_json::to_string(&request) {
tracing::trace!(
@@ -193,12 +209,11 @@ impl ControllerTrait for Controller {
let request = CreateTranscriptionRequestArgs::default()
.model(&speech_to_text_config.model_id)
.file(async_openai::types::AudioInput {
source: async_openai::types::InputSource::VecU8 {
filename,
vec: media,
},
})
.file(async_openai::types::AudioInput::from_vec_u8(
filename,
media,
mime_type.to_string(),
))
.language(language.clone())
.build()?;
@@ -253,12 +268,15 @@ impl ControllerTrait for Controller {
let quality = if params.cheaper_quality_switching_allowed {
// Switch to a cheaper quality
match &image_generation_config.quality {
async_openai::types::ImageQuality::Standard => {
async_openai::types::ImageQuality::Standard
}
async_openai::types::ImageQuality::HD => {
async_openai::types::ImageQuality::Standard
}
Some(quality) => match quality {
async_openai::types::ImageQuality::Standard => {
Some(async_openai::types::ImageQuality::Standard)
}
async_openai::types::ImageQuality::HD => {
Some(async_openai::types::ImageQuality::Standard)
}
},
None => None,
}
} else {
image_generation_config.quality.clone()
@@ -272,14 +290,37 @@ impl ControllerTrait for Controller {
})
.unwrap_or(image_generation_config.size);
let request = CreateImageRequestArgs::default()
let response_format = match model.clone() {
ImageModel::DallE2 => Some(ImageResponseFormat::B64Json),
ImageModel::DallE3 => Some(ImageResponseFormat::B64Json),
ImageModel::Other(model_str) => match model_str.as_str() {
// gpt-image-1 only outputs base64 and we don't need to specify the response format.
// In fact, specifying the response format results in an error.
OPENAI_IMAGE_MODEL_GPT_IMAGE_1 => None,
_ => Some(ImageResponseFormat::B64Json),
},
};
let mut request_builder = CreateImageRequestArgs::default();
request_builder
.model(model)
.prompt(prompt.to_owned())
.response_format(async_openai::types::ImageResponseFormat::B64Json)
.size(size)
.style(image_generation_config.style.clone())
.quality(quality)
.build()?;
.size(size);
if let Some(response_format) = response_format {
request_builder.response_format(response_format);
}
if let Some(style) = &image_generation_config.style {
request_builder.style(style.clone());
}
if let Some(quality) = quality {
request_builder.quality(quality.clone());
}
let request = request_builder.build()?;
tracing::trace!(
?prompt,
@@ -317,6 +358,104 @@ impl ControllerTrait for Controller {
))
}
async fn create_image_edit(
&self,
prompt: &str,
images: Vec<ImageSource>,
_params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
let Some(image_generation_config) = &self.config.image_generation else {
return Err(anyhow::anyhow!(
strings::agent::no_configuration_for_purpose_so_cannot_be_used(
&AgentPurpose::ImageGeneration
),
));
};
if images.is_empty() {
return Err(anyhow::anyhow!("No image sources provided"));
}
let mut image_inputs = Vec::new();
for image in images {
image_inputs.push(image.into());
}
let dalle2_size = match image_generation_config.size {
async_openai::types::ImageSize::S256x256 => Some(DallE2ImageSize::S256x256),
async_openai::types::ImageSize::S512x512 => Some(DallE2ImageSize::S512x512),
async_openai::types::ImageSize::S1024x1024 => Some(DallE2ImageSize::S1024x1024),
_ => None,
};
let model = image_generation_config
.model_id_as_openai_image_model()
.map_err(|err| anyhow::anyhow!(err))?;
let response_format = match model.clone() {
async_openai::types::ImageModel::DallE2 => {
Some(async_openai::types::ImageResponseFormat::B64Json)
}
async_openai::types::ImageModel::DallE3 => {
Some(async_openai::types::ImageResponseFormat::B64Json)
}
async_openai::types::ImageModel::Other(model_str) => match model_str.as_str() {
OPENAI_IMAGE_MODEL_GPT_IMAGE_1 => None,
_ => Some(async_openai::types::ImageResponseFormat::B64Json),
},
};
let mut request_builder = CreateImageEditRequestArgs::default();
request_builder
.image(image_inputs)
.prompt(prompt.to_owned())
.model(model);
if let Some(size) = dalle2_size {
request_builder.size(size);
}
if let Some(response_format) = response_format {
request_builder.response_format(response_format);
}
let request = request_builder
.build()
.map_err(|e| anyhow::anyhow!("Failed to build CreateImageEditRequest: {}", e))?;
tracing::trace!(
model = format!("{:?}", request.model),
size = format!("{:?}", request.size),
response_format = format!("{:?}", request.response_format),
"Sending OpenAI image edit API request"
);
let response = self.client.images().create_edit(request).await?;
if let Some(image_data) = response.data.into_iter().next() {
match image_data.deref() {
Image::B64Json { b64_json, .. } => {
let bytes = base64_decode(b64_json)?;
return Ok(ImageEditResult {
bytes,
mime_type: mxlink::mime::IMAGE_PNG,
});
}
Image::Url { url, .. } => {
tracing::warn!(?url, "Received URL instead of B64Json for image edit");
return Err(anyhow::anyhow!(
"Unexpected image type (URL) when B64Json was requested"
));
}
}
}
Err(anyhow::anyhow!(
"The OpenAI image edit API returned no images"
))
}
async fn text_to_speech(
&self,
input: &str,

View File

@@ -13,8 +13,10 @@ pub(super) use config::TextToSpeechConfig;
use super::super::AgentInstantiationError;
use super::super::AgentInstantiationResult;
use super::controller::ControllerType;
use super::ConfigTrait;
use super::controller::ControllerType;
pub const OPENAI_IMAGE_MODEL_GPT_IMAGE_1: &str = "gpt-image-1";
pub fn create_controller_from_yaml_value_config(
agent_id: &str,

View File

@@ -1,9 +1,14 @@
use async_openai::types::{
ChatCompletionRequestAssistantMessageArgs, ChatCompletionRequestMessage,
ChatCompletionRequestSystemMessageArgs, ChatCompletionRequestUserMessageArgs,
ChatCompletionRequestMessageContentPartImage, ChatCompletionRequestSystemMessageArgs,
ChatCompletionRequestUserMessageArgs, ChatCompletionRequestUserMessageContent,
ChatCompletionRequestUserMessageContentPart, ImageUrlArgs,
};
use crate::conversation::llm::{Author as LLMAuthor, Message as LLMMessage};
use crate::conversation::llm::{
Author as LLMAuthor, Message as LLMMessage, MessageContent as LLMMessageContent,
};
use crate::utils::base64::base64_encode;
pub fn convert_llm_messages_to_openai_messages(
conversation_messages: Vec<LLMMessage>,
@@ -12,29 +17,71 @@ pub fn convert_llm_messages_to_openai_messages(
Vec::with_capacity(conversation_messages.len());
for message in conversation_messages {
openai_conversation_messages.push(convert_llm_message_to_openai_message(message));
let openai_message = convert_llm_message_to_openai_message(message);
if let Some(openai_message) = openai_message {
openai_conversation_messages.push(openai_message);
}
}
openai_conversation_messages
}
fn convert_llm_message_to_openai_message(llm_message: LLMMessage) -> ChatCompletionRequestMessage {
match llm_message.author {
LLMAuthor::Prompt => ChatCompletionRequestSystemMessageArgs::default()
.content(llm_message.message_text)
.build()
.expect("Failed building OpenAI system message")
.into(),
LLMAuthor::Assistant => ChatCompletionRequestAssistantMessageArgs::default()
.content(llm_message.message_text)
.build()
.expect("Failed building OpenAI assistant message")
.into(),
LLMAuthor::User => ChatCompletionRequestUserMessageArgs::default()
.content(llm_message.message_text)
.build()
.expect("Failed building OpenAI user message")
.into(),
fn convert_llm_message_to_openai_message(
llm_message: LLMMessage,
) -> Option<ChatCompletionRequestMessage> {
match &llm_message.content {
LLMMessageContent::Text(text) => Some(match llm_message.author {
LLMAuthor::Prompt => ChatCompletionRequestSystemMessageArgs::default()
.content(text.clone())
.build()
.expect("Failed building OpenAI system message")
.into(),
LLMAuthor::Assistant => ChatCompletionRequestAssistantMessageArgs::default()
.content(text.clone())
.build()
.expect("Failed building OpenAI assistant message")
.into(),
LLMAuthor::User => ChatCompletionRequestUserMessageArgs::default()
.content(text.clone())
.build()
.expect("Failed building OpenAI user message")
.into(),
}),
LLMMessageContent::Image(image_details) => {
let image_url = format!(
"data:{};base64,{}",
image_details.mime,
base64_encode(&image_details.data)
);
let part = ChatCompletionRequestUserMessageContentPart::ImageUrl(
ChatCompletionRequestMessageContentPartImage {
image_url: ImageUrlArgs::default()
.url(image_url)
.build()
.expect("Failed building OpenAI image url"),
},
);
let message_content = ChatCompletionRequestUserMessageContent::Array(vec![part]);
match llm_message.author {
LLMAuthor::User => Some(
ChatCompletionRequestUserMessageArgs::default()
.content(message_content)
.build()
.expect("Failed building OpenAI user message")
.into(),
),
_ => {
tracing::warn!(
"OpenAI API does not support image content for messages authored by {:?}. This message part will be skipped.",
llm_message.author
);
None
}
}
}
}
}

View File

@@ -66,7 +66,7 @@ pub struct TextGenerationConfig {
pub temperature: f32,
#[serde(default)]
pub max_response_tokens: u32,
pub max_response_tokens: Option<u32>,
#[serde(default)]
pub max_context_tokens: u32,
@@ -78,7 +78,7 @@ impl Default for TextGenerationConfig {
model_id: default_text_model_id(),
prompt: Some(default_prompt().to_owned()),
temperature: super::super::default_temperature(),
max_response_tokens: 4096,
max_response_tokens: Some(4096),
max_context_tokens: 128_000,
}
}
@@ -93,6 +93,7 @@ impl TryInto<OpenAITextGenerationConfig> for TextGenerationConfig {
prompt: self.prompt,
temperature: self.temperature,
max_response_tokens: self.max_response_tokens,
max_completion_tokens: None,
max_context_tokens: self.max_context_tokens,
})
}
@@ -229,15 +230,19 @@ impl TryInto<OpenAIImageGenerationConfig> for ImageGenerationConfig {
};
let style = if let Some(style) = &self.style {
convert_string_to_enum::<async_openai::types::ImageStyle>(style)?
Some(convert_string_to_enum::<async_openai::types::ImageStyle>(
style,
)?)
} else {
async_openai::types::ImageStyle::Vivid
None
};
let quality = if let Some(quality) = &self.quality {
convert_string_to_enum::<async_openai::types::ImageQuality>(quality)?
Some(convert_string_to_enum::<async_openai::types::ImageQuality>(
quality,
)?)
} else {
async_openai::types::ImageQuality::Standard
None
};
Ok(OpenAIImageGenerationConfig {

View File

@@ -4,23 +4,25 @@ use etke_openai_api_rust::images::{ImagesApi, ImagesBody};
use etke_openai_api_rust::{Auth, Message, OpenAI};
use super::super::ControllerTrait;
use crate::agent::utils::base64_decode;
use crate::utils::base64::base64_decode;
use crate::{
agent::provider::{
ImageEditParams, ImageGenerationParams, ImageSource, SpeechToTextParams,
SpeechToTextResult,
entity::{TextGenerationParams, TextGenerationResult},
ImageGenerationParams, SpeechToTextParams, SpeechToTextResult,
},
conversation::llm::{
shorten_messages_list_to_context_size, Author as LLMAuthor,
Conversation as LLMConversation, Message as LLMMessage,
Author as LLMAuthor, Conversation as LLMConversation, Message as LLMMessage,
MessageContent as LLMMessageContent, shorten_messages_list_to_context_size,
},
};
use crate::{
agent::{
provider::entity::{
ImageGenerationResult, PingResult, TextToSpeechParams, TextToSpeechResult,
},
AgentPurpose,
provider::entity::{
ImageEditResult, ImageGenerationResult, PingResult, TextToSpeechParams,
TextToSpeechResult,
},
},
strings,
};
@@ -60,7 +62,8 @@ impl ControllerTrait for Controller {
let messages = vec![LLMMessage {
author: LLMAuthor::User,
message_text: "Hello!".to_string(),
content: LLMMessageContent::Text("Hello!".to_string()),
timestamp: chrono::Utc::now(),
}];
let conversation = LLMConversation { messages };
@@ -96,7 +99,8 @@ impl ControllerTrait for Controller {
} else {
Some(LLMMessage {
author: LLMAuthor::Prompt,
message_text: prompt_text,
content: LLMMessageContent::Text(prompt_text),
timestamp: chrono::Utc::now(),
})
};
@@ -131,12 +135,15 @@ impl ControllerTrait for Controller {
let max_tokens = text_generation_config
.max_response_tokens
.try_into()
.expect("Failed converting max_response_tokens from u32 to i32");
.map(|max_response_tokens| {
max_response_tokens
.try_into()
.expect("Failed converting max_response_tokens from u32 to i32")
});
let request = ChatBody {
model: text_generation_config.model_id.clone(),
max_tokens: Some(max_tokens),
max_tokens,
temperature: Some(temperature),
top_p: None,
n: Some(1),
@@ -361,6 +368,17 @@ impl ControllerTrait for Controller {
))
}
async fn create_image_edit(
&self,
_prompt: &str,
_images: Vec<ImageSource>,
_params: ImageEditParams,
) -> anyhow::Result<ImageEditResult> {
Err(anyhow::anyhow!(
"The OpenAI image edit API is not supported by the OpenAI-compat provider"
))
}
async fn text_to_speech(
&self,
input: &str,

View File

@@ -21,8 +21,8 @@ pub use controller::Controller;
use super::super::AgentInstantiationError;
use super::super::AgentInstantiationResult;
use super::controller::ControllerType;
use super::ConfigTrait;
use super::controller::ControllerType;
pub fn create_controller_from_yaml_value_config(
agent_id: &str,
@@ -56,7 +56,7 @@ pub fn default_config() -> Config {
if let Some(text_generation) = &mut config.text_generation {
text_generation.model_id = "some-model".to_string();
text_generation.max_response_tokens = 4096;
text_generation.max_response_tokens = Some(4096);
text_generation.max_context_tokens = 128_000;
}

View File

@@ -2,7 +2,9 @@ use etke_openai_api_rust::{Message, Role};
use crate::agent::provider::openai::Config as OpenAIConfig;
use crate::conversation::llm::{Author as LLMAuthor, Message as LLMMessage};
use crate::conversation::llm::{
Author as LLMAuthor, Message as LLMMessage, MessageContent as LLMMessageContent,
};
pub fn convert_llm_messages_to_openai_messages(
conversation_messages: Vec<LLMMessage>,
@@ -11,22 +13,33 @@ pub fn convert_llm_messages_to_openai_messages(
Vec::with_capacity(conversation_messages.len());
for message in conversation_messages {
openai_conversation_messages.push(convert_llm_message_to_openai_message(message));
let openai_message = convert_llm_message_to_openai_message(message);
if let Some(openai_message) = openai_message {
openai_conversation_messages.push(openai_message);
}
}
openai_conversation_messages
}
fn convert_llm_message_to_openai_message(llm_message: LLMMessage) -> Message {
fn convert_llm_message_to_openai_message(llm_message: LLMMessage) -> Option<Message> {
let role = match llm_message.author {
LLMAuthor::Prompt => Role::System,
LLMAuthor::Assistant => Role::Assistant,
LLMAuthor::User => Role::User,
};
Message {
role,
content: llm_message.message_text,
match &llm_message.content {
LLMMessageContent::Text(text) => Some(Message {
role,
content: text.clone(),
}),
LLMMessageContent::Image(_image_details) => {
tracing::warn!(
"The OpenAI-compat provider's library does not support image content. This image message will be skipped."
);
None
}
}
}

View File

@@ -14,7 +14,7 @@ pub fn default_config() -> Config {
if let Some(ref mut config) = config.text_generation.as_mut() {
config.model_id = "mattshumer/reflection-70b:free".to_owned();
config.max_context_tokens = 8192;
config.max_response_tokens = 2048;
config.max_response_tokens = Some(2048);
}
config

View File

@@ -14,7 +14,7 @@ pub fn default_config() -> Config {
if let Some(ref mut config) = config.text_generation.as_mut() {
config.model_id = "meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo".to_owned();
config.max_context_tokens = 8192;
config.max_response_tokens = 2048;
config.max_response_tokens = Some(2048);
}
config

View File

@@ -1,5 +1,3 @@
use base64::{engine::general_purpose::STANDARD, Engine as _};
use crate::{
agent::{
AgentInstance, AgentPurpose, ControllerTrait, Manager as AgentManager, PublicIdentifier,
@@ -140,7 +138,3 @@ async fn get_global_agent_id_for_purpose(
.handler
.get_by_purpose_with_catch_all_fallback(purpose)
}
pub(crate) fn base64_decode(base64_string: &str) -> Result<Vec<u8>, base64::DecodeError> {
STANDARD.decode(base64_string)
}

View File

@@ -1,11 +1,11 @@
use std::sync::Arc;
use std::{future::Future, pin::Pin};
use mxlink::matrix_sdk::media::{MediaFormat, MediaRequest};
use mxlink::matrix_sdk::ruma::{
events::room::MediaSource, MilliSecondsSinceUnixEpoch, OwnedUserId,
};
use mxlink::matrix_sdk::Room;
use mxlink::matrix_sdk::media::{MediaFormat, MediaRequestParameters};
use mxlink::matrix_sdk::ruma::{
MilliSecondsSinceUnixEpoch, OwnedUserId, events::room::MediaSource,
};
use mxlink::{
InitConfig, LoginConfig, LoginCredentials, LoginEncryption, MatrixLink, PersistenceConfig,
@@ -140,6 +140,10 @@ impl Bot {
&self.inner.config.command_prefix
}
pub(crate) fn post_join_self_introduction_enabled(&self) -> bool {
self.inner.config.room.post_join_self_introduction_enabled
}
pub(crate) fn homeserver_name(&self) -> &str {
&self.inner.config.homeserver.server_name
}
@@ -172,6 +176,24 @@ impl Bot {
self.matrix_link().user_id()
}
pub(crate) async fn user_display_name_in_room(&self, room: &Room) -> Option<String> {
let bot_display_name = self
.room_display_name_fetcher()
.own_display_name_in_room(room)
.await;
match bot_display_name {
Ok(value) => value,
Err(err) => {
tracing::warn!(
?err,
"Failed to fetch bot display name. Proceeding without it"
);
None
}
}
}
pub(crate) fn reacting(&self) -> super::reacting::Reacting {
super::reacting::Reacting::new(self.clone())
}
@@ -269,7 +291,7 @@ impl Bot {
let desired_display_name = self.inner.config.user.name.clone();
let profile = account
.get_profile()
.fetch_user_profile()
.await
.map_err(|e| anyhow::anyhow!("Failed fetching profile: {:?}", e))?;
@@ -292,7 +314,7 @@ impl Bot {
let should_update_avatar = match &profile.avatar_url {
Some(avatar_url) => {
let request = MediaRequest {
let request = MediaRequestParameters {
source: MediaSource::Plain(avatar_url.to_owned()),
format: MediaFormat::File,
};

View File

@@ -5,7 +5,7 @@ use anyhow::anyhow;
use crate::agent::AgentPurpose;
pub use crate::entity::cfg::{defaults as cfg_defaults, env as cfg_env, Config};
pub use crate::entity::cfg::{Config, defaults as cfg_defaults, env as cfg_env};
pub fn load() -> anyhow::Result<Config> {
let config_file_path = env::var(cfg_env::BAIBOT_CONFIG_FILE_PATH)
@@ -35,6 +35,9 @@ pub fn load() -> anyhow::Result<Config> {
}
cfg_env::BAIBOT_USER_NAME => config.user.name = value,
cfg_env::BAIBOT_COMMAND_PREFIX => config.command_prefix = value,
cfg_env::BAIBOT_ROOM_POST_JOIN_SELF_INTRODUCTION_ENABLED => {
config.room.post_join_self_introduction_enabled = value.parse::<bool>()?;
}
cfg_env::BAIBOT_LOGGING => {
config.logging = value;
}

View File

@@ -1,9 +1,9 @@
use mxlink::matrix_sdk::{
ruma::{
api::client::receipt::create_receipt::v3::ReceiptType,
events::room::message::OriginalSyncRoomMessageEvent, OwnedEventId,
},
Room,
ruma::{
OwnedEventId, api::client::receipt::create_receipt::v3::ReceiptType,
events::room::message::OriginalSyncRoomMessageEvent,
},
};
use mxlink::{CallbackError, MessageResponseType};
@@ -11,7 +11,7 @@ use mxlink::{CallbackError, MessageResponseType};
use tracing::Instrument;
use crate::{
conversation::matrix::determine_thread_context_for_room_event,
conversation::matrix::determine_interaction_context_for_room_event,
entity::{MessageContext, MessagePayload, RoomConfigContext, TriggerEventInfo},
};
@@ -239,8 +239,11 @@ impl Messaging {
}
};
let thread_context = determine_thread_context_for_room_event(
let bot_display_name = self.bot.user_display_name_in_room(&room).await;
let interaction_context = determine_interaction_context_for_room_event(
self.bot.user_id(),
&bot_display_name,
&room,
&event,
&payload,
@@ -248,16 +251,18 @@ impl Messaging {
)
.await;
let thread_context = match thread_context {
let interaction_context = match interaction_context {
Ok(value) => value,
Err(err) => {
tracing::error!(?err, "Failed to determine thread context for event");
tracing::error!(?err, "Failed to determine interaction context for event");
return Ok(());
}
};
let Some(thread_context) = thread_context else {
tracing::debug!("Ignoring message with unknown thread context (likely not a threaded message or a top-level message)");
let Some(interaction_context) = interaction_context else {
tracing::debug!(
"Ignoring message with unknown interaction context (likely not a message for us)"
);
return Ok(());
};
@@ -276,33 +281,14 @@ impl Messaging {
room_config_context,
self.bot.admin_pattern_regexes().clone(),
trigger_event_info,
thread_context.info.clone(),
);
interaction_context.thread_info.clone(),
)
.with_bot_display_name(bot_display_name);
let bot_display_name = self
.bot
.room_display_name_fetcher()
.own_display_name_in_room(message_context.room())
.await;
let bot_display_name = match bot_display_name {
Ok(value) => value,
Err(err) => {
tracing::warn!(
?err,
"Failed to fetch bot display name. Proceeding without it"
);
None
}
};
// The first event in the thread determines which handler processes the current event.
let controller_type = crate::controller::determine_controller(
self.bot.command_prefix(),
&thread_context.first_message,
&interaction_context.trigger,
&message_context,
self.bot.user_id(),
&bot_display_name,
);
tracing::info!(?controller_type, "Determined controller");
@@ -310,7 +296,7 @@ impl Messaging {
let _ = room
.send_single_receipt(
ReceiptType::Read,
thread_context.info.clone().into(),
interaction_context.thread_info.clone().into(),
event.event_id.clone(),
)
.await;

View File

@@ -1,12 +1,12 @@
use mxlink::matrix_sdk::{
ruma::{
events::{
room::message::Relation, AnyMessageLikeEvent, AnySyncTimelineEvent, AnyTimelineEvent,
MessageLikeEvent,
},
OwnedEventId, OwnedUserId,
},
Room,
ruma::{
OwnedEventId, OwnedUserId,
events::{
AnySyncMessageLikeEvent, AnySyncTimelineEvent, SyncMessageLikeEvent,
room::message::Relation,
},
},
};
use mxlink::CallbackError;
@@ -139,7 +139,7 @@ impl Reacting {
}
};
let reacted_to_event_any_timeline_event = match reacted_to_event.event.deserialize() {
let reacted_to_event_any_timeline_event = match reacted_to_event.raw().deserialize() {
Ok(value) => value,
Err(err) => {
tracing::error!(
@@ -154,7 +154,7 @@ impl Reacting {
let reacted_to_event_sender_id: OwnedUserId =
reacted_to_event_any_timeline_event.sender().to_owned();
let AnyTimelineEvent::MessageLike(reacted_to_event_message_like) =
let AnySyncTimelineEvent::MessageLike(reacted_to_event_message_like) =
reacted_to_event_any_timeline_event
else {
tracing::debug!(
@@ -164,7 +164,7 @@ impl Reacting {
return Ok(());
};
let AnyMessageLikeEvent::RoomMessage(reacted_to_event_room_message) =
let AnySyncMessageLikeEvent::RoomMessage(reacted_to_event_room_message) =
reacted_to_event_message_like
else {
tracing::debug!(
@@ -174,7 +174,7 @@ impl Reacting {
return Ok(());
};
let MessageLikeEvent::Original(reacted_to_event_room_message_original) =
let SyncMessageLikeEvent::Original(reacted_to_event_room_message_original) =
reacted_to_event_room_message
else {
tracing::debug!(?reacted_to_event_id, "Ignoring redacted reacted-to event",);

View File

@@ -1,9 +1,9 @@
use mxlink::{
matrix_sdk::{
ruma::events::{room::member::StrippedRoomMemberEvent, AnySyncTimelineEvent},
Room,
},
InvitationDecision,
matrix_sdk::{
Room,
ruma::events::{AnySyncTimelineEvent, room::member::StrippedRoomMemberEvent},
},
};
use mxlink::CallbackError;

View File

@@ -1,6 +1,6 @@
use mxlink::MessageResponseType;
use crate::{entity::MessageContext, strings, Bot};
use crate::{Bot, entity::MessageContext, strings};
use super::AccessControllerType;

View File

@@ -1,6 +1,6 @@
use mxlink::MessageResponseType;
use crate::{entity::MessageContext, strings, Bot};
use crate::{Bot, entity::MessageContext, strings};
pub async fn handle(bot: &Bot, message_context: &MessageContext) -> anyhow::Result<()> {
let mut message = String::new();

View File

@@ -4,5 +4,5 @@ pub mod help;
mod room_local_agent_managers;
mod users;
pub use determination::{determine_controller, AccessControllerType};
pub use determination::{AccessControllerType, determine_controller};
pub use dispatching::dispatch_controller;

View File

@@ -1,6 +1,6 @@
use mxlink::MessageResponseType;
use crate::{entity::MessageContext, strings, Bot};
use crate::{Bot, entity::MessageContext, strings};
pub async fn handle_get(bot: &Bot, message_context: &MessageContext) -> anyhow::Result<()> {
let message = match &message_context

View File

@@ -1,6 +1,6 @@
use mxlink::MessageResponseType;
use crate::{entity::MessageContext, strings, Bot};
use crate::{Bot, entity::MessageContext, strings};
pub async fn handle_get(bot: &Bot, message_context: &MessageContext) -> anyhow::Result<()> {
let message = match &message_context.global_config().access.user_patterns {

View File

@@ -3,15 +3,15 @@ mod tests;
use mxlink::MessageResponseType;
use crate::agent::provider::{ControllerTrait, PingResult};
use crate::agent::PublicIdentifier;
use crate::agent::{create_from_provider_and_yaml_value_config, AgentDefinition};
use crate::agent::provider::{ControllerTrait, PingResult};
use crate::agent::{AgentDefinition, create_from_provider_and_yaml_value_config};
use crate::agent::{AgentInstance, AgentProvider};
use crate::controller::utils::get_text_body_or_complain;
use crate::entity::globalconfig::GlobalConfigurationManager;
use crate::entity::roomconfig::RoomConfigurationManager;
use crate::strings;
use crate::{entity::MessageContext, Bot};
use crate::{Bot, entity::MessageContext};
struct ParsedAgentConfig {
agent: AgentInstance,

View File

@@ -1,9 +1,9 @@
use mxlink::MessageResponseType;
use crate::entity::{
globalconfig::GlobalConfigurationManager, roomconfig::RoomConfigurationManager, MessageContext,
MessageContext, globalconfig::GlobalConfigurationManager, roomconfig::RoomConfigurationManager,
};
use crate::{agent::PublicIdentifier, strings, Bot};
use crate::{Bot, agent::PublicIdentifier, strings};
pub async fn handle(
bot: &Bot,

View File

@@ -1,6 +1,6 @@
use mxlink::MessageResponseType;
use crate::{agent::PublicIdentifier, entity::MessageContext, strings, Bot};
use crate::{Bot, agent::PublicIdentifier, entity::MessageContext, strings};
pub async fn handle(
bot: &Bot,

View File

@@ -1,6 +1,6 @@
use mxlink::MessageResponseType;
use crate::{entity::MessageContext, strings, Bot};
use crate::{Bot, entity::MessageContext, strings};
pub async fn handle(bot: &Bot, message_context: &MessageContext) -> anyhow::Result<()> {
// Anyone can access this help command, because certain subcommands ("list")

View File

@@ -2,7 +2,7 @@ use mxlink::MessageResponseType;
use crate::agent::AgentPurpose;
use crate::strings;
use crate::{entity::MessageContext, Bot};
use crate::{Bot, entity::MessageContext};
pub async fn handle(bot: &Bot, message_context: &MessageContext) -> anyhow::Result<()> {
let agents = bot

View File

@@ -1,4 +1,4 @@
use crate::{entity::MessageContext, Bot};
use crate::{Bot, entity::MessageContext};
pub mod create;
pub mod delete;
@@ -7,7 +7,7 @@ pub mod determination;
pub mod help;
pub mod list;
pub use determination::{determine_controller, AgentControllerType};
pub use determination::{AgentControllerType, determine_controller};
pub async fn dispatch_controller(
handler: &AgentControllerType,

View File

@@ -1,6 +1,6 @@
use mxlink::MessageResponseType;
use crate::{entity::MessageContext, strings, Bot};
use crate::{Bot, entity::MessageContext, strings};
pub async fn handle_get<T>(
bot: &Bot,

View File

@@ -1,7 +1,8 @@
use crate::{
agent::{AgentPurpose, PublicIdentifier},
entity::roomconfig::{
SpeechToTextFlowType, TextGenerationAutoUsage, TextGenerationPrefixRequirementType,
SpeechToTextFlowType, SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages,
TextGenerationAutoUsage, TextGenerationPrefixRequirementType,
TextToSpeechBotMessagesFlowType, TextToSpeechUserMessagesFlowType,
},
};
@@ -54,6 +55,11 @@ pub enum ConfigSpeechToTextSettingRelatedControllerType {
GetFlowType,
SetFlowType(Option<SpeechToTextFlowType>),
GetMsgTypeForNonThreadedOnlyTranscribedMessages,
SetMsgTypeForNonThreadedOnlyTranscribedMessages(
Option<SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages>,
),
GetLanguage,
SetLanguage(Option<String>),
}

View File

@@ -1,7 +1,13 @@
#[cfg(test)]
mod tests;
use crate::{controller::ControllerType, entity::roomconfig::SpeechToTextFlowType, strings};
use crate::{
controller::ControllerType,
entity::roomconfig::{
SpeechToTextFlowType, SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages,
},
strings,
};
use super::super::controller_type::ConfigSpeechToTextSettingRelatedControllerType;
@@ -48,6 +54,53 @@ pub(super) fn determine(
));
}
// msg_type_for_non_threaded_only_transcribed_messages
if let Some(remaining_text) =
text.strip_prefix("msg-type-for-non-threaded-only-transcribed-messages")
{
let remaining_text = remaining_text.trim();
if !remaining_text.is_empty() {
return Err(ControllerType::Error(
strings::cfg::configuration_getter_used_with_extra_text(
"msg-type-for-non-threaded-only-transcribed-messages",
remaining_text,
)
.to_owned(),
));
}
return Ok(ConfigSpeechToTextSettingRelatedControllerType::GetMsgTypeForNonThreadedOnlyTranscribedMessages);
}
if let Some(value_string) =
text.strip_prefix("set-msg-type-for-non-threaded-only-transcribed-messages")
{
let value_string = value_string.trim().to_owned();
let value_choice = if value_string.is_empty() {
None
} else {
let value_choice =
SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages::from_str(
&value_string.to_lowercase(),
);
if value_choice.is_none() {
return Err(ControllerType::Error(
strings::cfg::configuration_value_unrecognized(&value_string).to_owned(),
));
}
value_choice
};
return Ok(ConfigSpeechToTextSettingRelatedControllerType::SetMsgTypeForNonThreadedOnlyTranscribedMessages(
value_choice,
));
}
// Language
if let Some(remaining_text) = text.strip_prefix("language") {

View File

@@ -1,5 +1,5 @@
use crate::strings;
use crate::{entity::MessageContext, Bot};
use crate::{Bot, entity::MessageContext};
use mxlink::MessageResponseType;
use super::controller_type::{

View File

@@ -1,5 +1,8 @@
use crate::entity::roomconfig::{RoomSettings, SpeechToTextFlowType};
use crate::{entity::MessageContext, Bot};
use crate::entity::roomconfig::{
RoomSettings, SpeechToTextFlowType,
SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages,
};
use crate::{Bot, entity::MessageContext};
use super::super::controller_type::{
ConfigSpeechToTextSettingRelatedControllerType, SettingsStorageSource,
@@ -52,6 +55,39 @@ pub(super) async fn dispatch(
}
}
ConfigSpeechToTextSettingRelatedControllerType::GetMsgTypeForNonThreadedOnlyTranscribedMessages => {
let value = &room_settings.speech_to_text.msg_type_for_non_threaded_only_transcribed_messages;
setting_get::<SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages>(bot, message_context, value).await
}
ConfigSpeechToTextSettingRelatedControllerType::SetMsgTypeForNonThreadedOnlyTranscribedMessages(value) => {
let value = value.to_owned();
let setter_callback = Box::new(move |room_settings: &mut RoomSettings| {
room_settings.speech_to_text.msg_type_for_non_threaded_only_transcribed_messages = value;
});
match config_type {
SettingsStorageSource::Room => {
room_setting_set::<SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages>(
bot,
message_context,
&value,
setter_callback,
)
.await
}
SettingsStorageSource::Global => {
global_setting_set::<SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages>(
bot,
message_context,
&value,
setter_callback,
)
.await
}
}
}
ConfigSpeechToTextSettingRelatedControllerType::GetLanguage => {
let value = &room_settings.speech_to_text.language;
setting_get::<String>(bot, message_context, value).await

View File

@@ -1,7 +1,7 @@
use crate::entity::roomconfig::{
RoomSettings, TextGenerationAutoUsage, TextGenerationPrefixRequirementType,
};
use crate::{entity::MessageContext, Bot};
use crate::{Bot, entity::MessageContext};
use super::super::controller_type::{
ConfigTextGenerationSettingRelatedControllerType, SettingsStorageSource,

View File

@@ -1,7 +1,7 @@
use crate::entity::roomconfig::{
RoomSettings, TextToSpeechBotMessagesFlowType, TextToSpeechUserMessagesFlowType,
};
use crate::{entity::MessageContext, Bot};
use crate::{Bot, entity::MessageContext};
use super::super::controller_type::{
ConfigTextToSpeechSettingRelatedControllerType, SettingsStorageSource,

View File

@@ -1,7 +1,7 @@
use mxlink::MessageResponseType;
use crate::entity::{roomconfig::RoomSettings, MessageContext};
use crate::{strings, Bot};
use crate::entity::{MessageContext, roomconfig::RoomSettings};
use crate::{Bot, strings};
pub async fn handle_set<T>(
bot: &Bot,

View File

@@ -1,9 +1,10 @@
use mxlink::MessageResponseType;
use crate::{
Bot,
agent::{AgentPurpose, PublicIdentifier},
entity::{globalconfig::GlobalConfigurationManager, MessageContext},
strings, Bot,
entity::{MessageContext, globalconfig::GlobalConfigurationManager},
strings,
};
pub async fn handle_get(

View File

@@ -1,14 +1,16 @@
use mxlink::MessageResponseType;
use crate::{
Bot,
entity::{
MessageContext,
roomconfig::{
SpeechToTextFlowType, TextGenerationAutoUsage, TextGenerationPrefixRequirementType,
SpeechToTextFlowType, SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages,
TextGenerationAutoUsage, TextGenerationPrefixRequirementType,
TextToSpeechBotMessagesFlowType, TextToSpeechUserMessagesFlowType,
},
MessageContext,
},
strings, Bot,
strings,
};
pub async fn handle(bot: &Bot, message_context: &MessageContext) -> anyhow::Result<()> {
@@ -346,6 +348,46 @@ fn build_section_speech_to_text(command_prefix: &str) -> String {
));
message.push_str("\n\n");
// Msg Type For Non Threaded Only Transcribed Messages
message.push_str(&format!(
"#### {}",
strings::help::cfg::speech_to_text_msg_type_for_non_threaded_only_transcribed_messages_heading()
));
message.push_str("\n\n");
message.push_str(strings::help::cfg::speech_to_text_msg_type_for_non_threaded_only_transcribed_messages_intro());
message.push('\n');
message.push_str(
&strings::help::cfg::the_following_configuration_values_are_recognized(
SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages::choices(),
),
);
message.push_str("\n\n");
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_show(
command_prefix,
"speech-to-text msg-type-for-non-threaded-only-transcribed-messages"
)
));
message.push('\n');
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_set(
command_prefix,
"speech-to-text set-msg-type-for-non-threaded-only-transcribed-messages VALUE"
)
));
message.push('\n');
message.push_str(&format!(
"- {}",
&strings::help::cfg::current_setting_unset(
command_prefix,
"speech-to-text set-msg-type-for-non-threaded-only-transcribed-messages"
)
));
message.push_str("\n\n");
// Language
message.push_str(&format!(

View File

@@ -1,7 +1,7 @@
use mxlink::MessageResponseType;
use crate::entity::{roomconfig::RoomSettings, MessageContext};
use crate::{strings, Bot};
use crate::entity::{MessageContext, roomconfig::RoomSettings};
use crate::{Bot, strings};
pub async fn handle_set<T>(
bot: &Bot,

View File

@@ -1,9 +1,10 @@
use mxlink::MessageResponseType;
use crate::{
Bot,
agent::{AgentPurpose, PublicIdentifier},
entity::MessageContext,
strings, Bot,
strings,
};
use crate::entity::roomconfig::RoomConfigurationManager;

View File

@@ -1,15 +1,16 @@
use mxlink::MessageResponseType;
use crate::{
Bot,
agent::{
utils::get_effective_agent_for_purpose, AgentInstance, AgentPurpose, ControllerTrait,
Manager as AgentManager, PublicIdentifier,
AgentInstance, AgentPurpose, ControllerTrait, Manager as AgentManager, PublicIdentifier,
utils::get_effective_agent_for_purpose,
},
entity::{
roomconfig::{RoomConfig, RoomSettingsHandler},
MessageContext, RoomConfigContext,
roomconfig::{RoomConfig, RoomSettingsHandler},
},
strings, Bot,
strings,
};
pub async fn handle(bot: &Bot, message_context: &MessageContext) -> anyhow::Result<()> {
@@ -68,7 +69,7 @@ pub async fn handle(bot: &Bot, message_context: &MessageContext) -> anyhow::Resu
);
message.push_str("\n\n");
// Image Generation
// Image Creation
message.push_str(
&generate_image_generation_section(agent_manager, message_context.room_config_context())
.await,
@@ -507,6 +508,35 @@ async fn generate_speech_to_text_section(
flow_type_set_where,
));
// Msg Type For Non Threaded Only Transcribed Messages
let effective_msg_type_for_non_threaded_only_transcribed_messages =
room_config_context.speech_to_text_msg_type_for_non_threaded_only_transcribed_messages();
let room_config_msg_type_for_non_threaded_only_transcribed_messages = room_config_context
.room_config
.settings
.speech_to_text
.msg_type_for_non_threaded_only_transcribed_messages;
let global_config_msg_type_for_non_threaded_only_transcribed_messages = room_config_context
.global_config
.fallback_room_settings
.speech_to_text
.msg_type_for_non_threaded_only_transcribed_messages;
let msg_type_for_non_threaded_only_transcribed_messages_set_where =
if room_config_msg_type_for_non_threaded_only_transcribed_messages.is_some() {
strings::cfg::status_badge_set_in_room_config()
} else if global_config_msg_type_for_non_threaded_only_transcribed_messages.is_some() {
strings::cfg::status_badge_set_in_global_config()
} else {
strings::cfg::status_badge_using_hardcoded_default()
};
message.push_str(&strings::cfg::status_speech_to_text_entry_msg_type_for_non_threaded_only_transcribed_messages(
effective_msg_type_for_non_threaded_only_transcribed_messages,
msg_type_for_non_threaded_only_transcribed_messages_set_where,
));
// Language
let effective_language = room_config_context.speech_to_text_language();

View File

@@ -1,30 +1,48 @@
use mxlink::matrix_sdk::ruma::events::room::message::AudioMessageEventContent;
use mxlink::matrix_sdk::ruma::OwnedEventId;
use mxlink::matrix_sdk::ruma::events::room::message::AudioMessageEventContent;
use mxlink::{MatrixLink, MessageResponseType};
use tracing::Instrument;
use crate::agent::provider::{
SpeechToTextParams, TextGenerationParams, TextGenerationPromptVariables,
};
use crate::agent::AgentInstance;
use crate::agent::AgentPurpose;
use crate::agent::ControllerTrait;
use crate::agent::provider::{
SpeechToTextParams, TextGenerationParams, TextGenerationPromptVariables,
};
use crate::controller::utils::agent::get_effective_agent_for_purpose_or_complain;
use crate::conversation::matrix::MatrixMessageProcessingParams;
use crate::entity::roomconfig::{
SpeechToTextFlowType, TextToSpeechBotMessagesFlowType, TextToSpeechUserMessagesFlowType,
};
use crate::entity::MessagePayload;
use crate::entity::roomconfig::{
SpeechToTextFlowType, SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages,
TextToSpeechBotMessagesFlowType, TextToSpeechUserMessagesFlowType,
};
use crate::strings;
use crate::utils::text_to_speech::create_transcribed_message_text;
use crate::{conversation::create_llm_conversation_for_matrix_thread, entity::MessageContext, Bot};
use crate::{
Bot,
conversation::{
create_llm_conversation_for_matrix_reply_chain, create_llm_conversation_for_matrix_thread,
matrix::create_list_of_bot_user_prefixes_to_strip,
},
entity::MessageContext,
};
#[derive(Debug, PartialEq)]
pub enum ChatCompletionControllerType {
ViaText { prefixes_to_strip: Vec<String> },
// Invoked via a command prefix (e.g. `!bai Hello!`)
TextCommand,
// Invoked via a mention (e.g. `@baibot Hello!`)
TextMention,
// Invoked via a direct message (e.g. `Hello!`)
TextDirect,
ViaAudio,
Audio,
Image,
ThreadMention,
ReplyMention,
}
struct TextToSpeechEligiblePayload {
@@ -56,21 +74,35 @@ pub async fn handle(
if let MessagePayload::Audio(audio_content) = &message_context.payload() {
original_message_is_audio = true;
let response_type = match speech_to_text_flow_type {
let (response_type, msg_type) = match speech_to_text_flow_type {
SpeechToTextFlowType::Ignore => {
tracing::debug!("Intentionally ignoring audio message");
return Ok(());
}
SpeechToTextFlowType::TranscribeAndGenerateText => {
tracing::debug!("Will be transcribing and possibly generating text..");
MessageResponseType::InThread(message_context.thread_info().clone())
(
MessageResponseType::InThread(message_context.thread_info().clone()),
SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages::Notice,
)
}
SpeechToTextFlowType::OnlyTranscribe => {
tracing::debug!("Will only be transcribing audio to text..");
if message_context.thread_info().is_thread_root_only() {
MessageResponseType::Reply(message_context.thread_info().root_event_id.clone())
let msg_type = message_context
.room_config_context()
.speech_to_text_msg_type_for_non_threaded_only_transcribed_messages();
(
MessageResponseType::Reply(
message_context.thread_info().root_event_id.clone(),
),
msg_type,
)
} else {
MessageResponseType::InThread(message_context.thread_info().clone())
(
MessageResponseType::InThread(message_context.thread_info().clone()),
SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages::Notice,
)
}
}
};
@@ -79,8 +111,14 @@ pub async fn handle(
_typing_notice_guard = Some(bot.start_typing_notice(message_context.room()).await);
}
let Some(speech_to_text_created_event_id_result) =
handle_stage_speech_to_text(bot, message_context, audio_content, response_type).await
let Some(speech_to_text_created_event_id_result) = handle_stage_speech_to_text(
bot,
message_context,
audio_content,
response_type,
msg_type,
)
.await
else {
return Ok(());
};
@@ -125,7 +163,15 @@ pub async fn handle(
None
};
let response_type = MessageResponseType::InThread(message_context.thread_info().clone());
let response_type = match controller_type {
// When we're triggered via a reply mention, we reply to the message that triggered us.
ChatCompletionControllerType::ReplyMention => {
MessageResponseType::Reply(message_context.thread_info().last_event_id.clone())
}
// In all other cases, we're dealing with a threaded conversation, so we reply in the thread.
_ => MessageResponseType::InThread(message_context.thread_info().clone()),
};
let text_to_speech_eligible_payload = handle_stage_text_generation(
bot,
@@ -259,6 +305,7 @@ async fn handle_stage_speech_to_text(
message_context: &MessageContext,
audio_content: &AudioMessageEventContent,
response_type: MessageResponseType,
msg_type: SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages,
) -> Option<OwnedEventId> {
let agent = get_effective_agent_for_purpose_or_complain(
bot,
@@ -279,7 +326,7 @@ async fn handle_stage_speech_to_text(
.react_no_fail(
message_context.room(),
message_context.event_id().clone(),
AgentPurpose::SpeechToText.emoji().to_owned(),
strings::PROGRESS_INDICATOR_EMOJI.to_owned(),
)
.await;
@@ -289,6 +336,7 @@ async fn handle_stage_speech_to_text(
&agent,
audio_content,
response_type.clone(),
msg_type,
)
.await;
@@ -353,24 +401,66 @@ async fn handle_stage_text_generation(
)
.await?;
let prefixes_to_strip = match controller_type {
ChatCompletionControllerType::ViaText { prefixes_to_strip } => prefixes_to_strip.clone(),
ChatCompletionControllerType::ViaAudio => vec![],
// We only strip text from the first message if we're invoked via a command prefix.
// Otherwise, we do bot-user mentions stripping on all messages below.
let first_message_prefixes_to_strip = match controller_type {
ChatCompletionControllerType::TextCommand => vec![bot.command_prefix().to_owned()],
_ => vec![],
};
let params = MatrixMessageProcessingParams::new(
bot.user_id().as_str().to_owned(),
message_context.combined_admin_and_user_regexes(),
)
.with_first_message_stripped_prefixes(prefixes_to_strip);
let bot_user_prefixes_to_strip = create_list_of_bot_user_prefixes_to_strip(
bot.user_id(),
message_context.bot_display_name(),
);
let conversation = create_llm_conversation_for_matrix_thread(
matrix_link.clone(),
message_context.room(),
message_context.thread_info().root_event_id.clone(),
&params,
)
.await;
let allowed_users = match controller_type {
// Regular chat completion only operates on messages from allowed users.
ChatCompletionControllerType::TextCommand
| ChatCompletionControllerType::TextMention
| ChatCompletionControllerType::TextDirect
| ChatCompletionControllerType::Audio
| ChatCompletionControllerType::Image => {
Some(message_context.combined_admin_and_user_regexes())
}
// When we're triggered via an explicit mention (thread or reply), we wish to operate against the mention's whole context
// (the whole thread or the whole reply chain upward of the message that triggered us).
//
// This is to allow admins and users to trigger text-generation for other users' messages.
// When we're dragged into a conversation by a known (to us) user, we'd like to process all messages in the conversation,
// not just those from allowed users.
ChatCompletionControllerType::ThreadMention
| ChatCompletionControllerType::ReplyMention => None,
};
let params = MatrixMessageProcessingParams::new(bot.user_id().to_owned(), allowed_users)
.with_first_message_prefixes_to_strip(first_message_prefixes_to_strip)
.with_bot_user_prefixes_to_strip(bot_user_prefixes_to_strip);
let conversation = match controller_type {
// When we're triggered via a reply mention, the context is the whole reply chain upward of the message that triggered us.
ChatCompletionControllerType::ReplyMention => {
create_llm_conversation_for_matrix_reply_chain(
&matrix_link,
&bot.room_event_fetcher().clone(),
message_context.room(),
message_context.thread_info().last_event_id.clone(),
&params,
)
.await
}
// Everything else is happening in a thread, so the context is the whole thread.
_ => {
create_llm_conversation_for_matrix_thread(
&matrix_link,
message_context.room(),
message_context.thread_info().root_event_id.clone(),
&params,
)
.await
}
};
let conversation = match conversation {
Ok(conversation) => conversation,
@@ -415,6 +505,7 @@ async fn handle_stage_text_generation(
.text_generation_model_id()
.unwrap_or("unknown-model".to_owned()),
chrono::Utc::now(),
conversation.start_time(),
);
let params = TextGenerationParams {
@@ -508,10 +599,11 @@ async fn handle_stage_speech_to_text_actual_transcribing(
agent: &AgentInstance,
audio_content: &AudioMessageEventContent,
response_type: MessageResponseType,
msg_type: SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages,
) -> anyhow::Result<OwnedEventId> {
let src = &audio_content.source;
let media_request = mxlink::matrix_sdk::media::MediaRequest {
let media_request = mxlink::matrix_sdk::media::MediaRequestParameters {
source: src.to_owned(),
format: mxlink::matrix_sdk::media::MediaFormat::File,
};
@@ -548,16 +640,62 @@ async fn handle_stage_speech_to_text_actual_transcribing(
.instrument(span)
.await?;
let transcribed_text = create_transcribed_message_text(&speech_to_text_result.text);
// Only use the `> 🦻 Transcribed text` format if we're posting in a thread.
//
// If we're dealing with a regular reply (which would be the case in "Transcribe-only mode" = speech-to-text/flow-type=only_transcribe),
// we don't want to use the `> 🦻 Transcribed text` format for 2 reasons:
//
// 1. This kind of blockquote-formatting can be confused by clients for a fallback-for-rich-replies
// (see https://spec.matrix.org/v1.11/client-server-api/#fallbacks-for-rich-replies).
// It makes certain clients render our messages incorrectly.
//
// 2. Transcribe-only mode is typically used for memos. Sticking to a plain-text format
// allows people to copy-paste the text or forward it to another room more easily (without having to strip formatting, etc.)
//
// When sending a bare reply, we'd better annotate the message with a 🦻 reaction instead,
// to make it clear to users that it's a transcription.
let (transcribed_text, annotate_message_with_reaction) =
if let MessageResponseType::InThread(_) = response_type {
(
create_transcribed_message_text(&speech_to_text_result.text),
false,
)
} else {
(speech_to_text_result.text, true)
};
let result = bot
.messaging()
.send_notice_markdown_no_fail(message_context.room(), transcribed_text, response_type)
.await;
let result = match msg_type {
SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages::Text => {
bot.messaging()
.send_text_markdown_no_fail(message_context.room(), transcribed_text, response_type)
.await
}
SpeechToTextMessageTypeForNonThreadedOnlyTranscribedMessages::Notice => {
bot.messaging()
.send_notice_markdown_no_fail(
message_context.room(),
transcribed_text,
response_type,
)
.await
}
};
result
let event_id = result
.map(|result| result.event_id)
.ok_or_else(|| anyhow::anyhow!("Failed to send transcribed text"))
.ok_or_else(|| anyhow::anyhow!("Failed to send transcribed text"))?;
if annotate_message_with_reaction {
bot.reacting()
.react_no_fail(
message_context.room(),
event_id.clone(),
AgentPurpose::SpeechToText.emoji().to_owned(),
)
.await;
}
Ok(event_id)
}
async fn send_tts_offer_for_message(

View File

@@ -23,5 +23,6 @@ pub enum ControllerType {
ChatCompletion(super::chat_completion::ChatCompletionControllerType),
ImageGeneration(String),
ImageEdit(String),
StickerGeneration(String),
}

Some files were not shown because too many files have changed in this diff Show More