Compare commits
19 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a9e4ab1bdb | ||
|
|
d9a045a5e4 | ||
|
|
393be9be5a | ||
|
|
a7b016a3d3 | ||
|
|
85e66406dc | ||
|
|
db9422740c | ||
|
|
90fbad5b64 | ||
|
|
b40226826f | ||
|
|
b89f0db71a | ||
|
|
36fdb46633 | ||
|
|
04ce8db1fc | ||
|
|
9908512968 | ||
|
|
eae6472c7a | ||
|
|
e6aa956423 | ||
|
|
7a38216192 | ||
|
|
d522d268e2 | ||
|
|
72120c5dc2 | ||
|
|
a2c35238c2 | ||
|
|
97f5cbb00b |
27
CHANGELOG.md
27
CHANGELOG.md
@@ -1,3 +1,30 @@
|
|||||||
|
# (2024-10-03) Version 1.3.1
|
||||||
|
|
||||||
|
- (**Improvement**) Improves fallback user mentions support for old clients (like Element iOS) which use the bot's display name (not its full Matrix User ID). ([d9a045a5e4](https://github.com/etkecc/baibot/commit/d9a045a5e41d2b99694f92ec9e90f47529546d89))
|
||||||
|
|
||||||
|
|
||||||
|
# (2024-10-03) Version 1.3.0
|
||||||
|
|
||||||
|
**TLDR**: you can now use OpenAI's [o1](https://platform.openai.com/docs/models/o1) models, benefit from [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) and mention the bot again from old clients lacking proper [user mentions support](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) (like Element iOS).
|
||||||
|
|
||||||
|
- (**Feature**) Introduces a new `baibot_conversation_start_time_utc` [prompt variable](./docs/configuration/text-generation.md#️-prompt-override) which is not a moving target (like the `baibot_now_utc` variable) and allows [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) to work. All default/sample configs have been adjusted to make use of this new variable, but users need to adjust your existing dynamically-created agents to start using it. ([85e66406dc](https://github.com/etkecc/baibot/commit/85e66406dc6f430741c7819f420e2df4ae6e8d3b))
|
||||||
|
|
||||||
|
- (**Improvement**) Allows for the `max_response_tokens` configuration value for the [OpenAI provider](./docs/providers.md#openai) to be set to `null` to allow [o1](https://platform.openai.com/docs/models/o1) models (which do not support `max_response_tokens`) to be used. See the new o1 sample config [here](./docs/sample-provider-configs/openai-o1.yml). ([db9422740c](https://github.com/etkecc/baibot/commit/db9422740ceca32956d9628b6326b8be206344e2))
|
||||||
|
|
||||||
|
- (**Improvement**) Switches the sample configs for the [OpenAI provider](./docs/providers.md#openai) to point to the `gpt-4o` model, which since 2024-10-02 is the same as the `gpt-4o-2024-08-06` model. We previously explicitly pointed the bot to the `gpt-4o-2024-08-06` model, because it was much better (longer context window). Now that `gpt-4o` points to the same powerful model, we don't need to pin its version anymore. Existing users may wish to adjust their configuration to match. ([90fbad5b64](https://github.com/etkecc/baibot/commit/90fbad5b643cd06c23179f055a309ec6a7cba161))
|
||||||
|
|
||||||
|
- (**Bugfix**) Restores fallback user mentions support (via regular text, not via the [user mentions spec](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions)) to allow certain old clients (like Element iOS) to be able to mention the bot again. Support for this was intentionally removed recently (in [v1.2.0](#2024-10-01-version-120)), but it turned out to be too early to do this. ([b40226826f](https://github.com/etkecc/baibot/commit/b40226826fe914d0d5d265230ebc5bac8058b6f7))
|
||||||
|
|
||||||
|
|
||||||
|
# (2024-10-01) Version 1.2.0
|
||||||
|
|
||||||
|
- (**Feature**) Adds support for [on-demand involvement](./docs/features.md#on-demand-involvement) of the bot (via mention) in arbitrary threads and reply chains ([9908512968](https://github.com/etkecc/baibot/commit/990851296828168c2106eb3f4668833e9e5a7463)) - fixes [issue #15](https://github.com/etkecc/baibot/issues/15)
|
||||||
|
|
||||||
|
- (**Improvement**) Simplifies [Transcribe-only mode](./docs/features.md#transcribe-only-mode) reply format (removing `> 🦻` prefixing) to allow easier forwarding, etc. ([e6aa956423](https://github.com/etkecc/baibot/commit/e6aa95642376ee7d87932d0e66dcfedf261b188b)) - fixes [issue #14](https://github.com/etkecc/baibot/issues/14)
|
||||||
|
|
||||||
|
- (**Bugfix**) Fixes speech-to-text replies rendering incorrectly in certain clients, due to them confusing our old reply format with [fallback for rich replies](https://spec.matrix.org/v1.11/client-server-api/#fallbacks-for-rich-replies) ([e6aa956423](https://github.com/etkecc/baibot/commit/e6aa95642376ee7d87932d0e66dcfedf261b188b)) - fixes [issue #17](https://github.com/etkecc/baibot/issues/17)
|
||||||
|
|
||||||
|
|
||||||
# (2024-09-22) Version 1.1.1
|
# (2024-09-22) Version 1.1.1
|
||||||
|
|
||||||
- (**Bugfix**) Fix thread messages being lost due to lack of pagination support ([d4ddd29660](https://github.com/etkecc/baibot/commit/d4ddd29660d9f51d248119dd6032e68ab29e7d35)) - fixes [issue #13](https://github.com/etkecc/baibot/issues/13)
|
- (**Bugfix**) Fix thread messages being lost due to lack of pagination support ([d4ddd29660](https://github.com/etkecc/baibot/commit/d4ddd29660d9f51d248119dd6032e68ab29e7d35)) - fixes [issue #13](https://github.com/etkecc/baibot/issues/13)
|
||||||
|
|||||||
413
Cargo.lock
generated
413
Cargo.lock
generated
File diff suppressed because it is too large
Load Diff
@@ -7,7 +7,7 @@ license = "AGPL-3.0-or-later"
|
|||||||
readme = "README.md"
|
readme = "README.md"
|
||||||
keywords = ["matrix", "chat", "bot", "AI", "LLM"]
|
keywords = ["matrix", "chat", "bot", "AI", "LLM"]
|
||||||
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
|
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
|
||||||
version = "1.1.1"
|
version = "1.3.1"
|
||||||
edition = "2021"
|
edition = "2021"
|
||||||
|
|
||||||
[lib]
|
[lib]
|
||||||
@@ -26,11 +26,11 @@ mxidwc = "1.0.*"
|
|||||||
mxlink = ">=1.3.0"
|
mxlink = ">=1.3.0"
|
||||||
etke_openai_api_rust = "0.1.*"
|
etke_openai_api_rust = "0.1.*"
|
||||||
quick_cache = "0.6.*"
|
quick_cache = "0.6.*"
|
||||||
regex = "1.10.*"
|
regex = "1.11.*"
|
||||||
serde = { version = "1.0.*", features = ["derive"], default-features = false }
|
serde = { version = "1.0.*", features = ["derive"], default-features = false }
|
||||||
serde_json = "1.0.*"
|
serde_json = "1.0.*"
|
||||||
serde_yaml = "0.9.*"
|
serde_yaml = "0.9.*"
|
||||||
tempfile = "3.12.*"
|
tempfile = "3.13.*"
|
||||||
tiktoken-rs = { version = "0.5.*", features = ["async-openai"] }
|
tiktoken-rs = { version = "0.5.*", features = ["async-openai"] }
|
||||||
tokio = { version = "1.40.*", features = ["rt", "rt-multi-thread", "macros"] }
|
tokio = { version = "1.40.*", features = ["rt", "rt-multi-thread", "macros"] }
|
||||||
tracing = "0.1.*"
|
tracing = "0.1.*"
|
||||||
|
|||||||
@@ -41,7 +41,7 @@ It's influenced by [chaz](https://github.com/arcuru/chaz), but does **not** use
|
|||||||
|
|
||||||

|

|
||||||
|
|
||||||
You can find more screenshots on the the [🌟 Features](./docs/features.md) and other [📚 Documentation](./docs/README.md) pages, as well as in the [docs/screenshots](./docs/screenshots) directory.
|
You can find more screenshots on the [🌟 Features](./docs/features.md) and other [📚 Documentation](./docs/README.md) pages, as well as in the [docs/screenshots](./docs/screenshots) directory.
|
||||||
|
|
||||||
|
|
||||||
## 🚀 Getting Started
|
## 🚀 Getting Started
|
||||||
|
|||||||
@@ -16,6 +16,7 @@ Users:
|
|||||||
|
|
||||||
- ✅ can **invite the bot to rooms**
|
- ✅ can **invite the bot to rooms**
|
||||||
- ✅ can **use all the bot's [features](./features.md)** ([💬 Text Generation](./features.md#-text-generation), [🦻 Speech-to-Text](./features.md#-speech-to-text), etc.) by sending room messages
|
- ✅ can **use all the bot's [features](./features.md)** ([💬 Text Generation](./features.md#-text-generation), [🦻 Speech-to-Text](./features.md#-speech-to-text), etc.) by sending room messages
|
||||||
|
- ✅ can **mention the bot** in threads and reply chains to provoke it to respond to non-user messages (see [🌟 Features / 💬 Text Generation / On-demand involvement](./features.md#on-demand-involvement))
|
||||||
- ✅ can **change the bot's configuration in a room** (e.g. `!bai config room ...` commands)
|
- ✅ can **change the bot's configuration in a room** (e.g. `!bai config room ...` commands)
|
||||||
- ❌ cannot **change the bot's global configuration** (e.g. `!bai config global ...` commands)
|
- ❌ cannot **change the bot's global configuration** (e.g. `!bai config global ...` commands)
|
||||||
- ❌ cannot **create new [🤖 Agents](./agents.md)** (neither in rooms, nor globally). See [💼 Room-local agent managers](#-room-local-agent-managers) for controlling which users can create agents.
|
- ❌ cannot **create new [🤖 Agents](./agents.md)** (neither in rooms, nor globally). See [💼 Room-local agent managers](#-room-local-agent-managers) for controlling which users can create agents.
|
||||||
|
|||||||
@@ -13,7 +13,7 @@ You may also wish to see:
|
|||||||
|
|
||||||
In Direct Message rooms with the bot (1:1 rooms), it most usually makes sense for the bot to respond to **all** of your messages, as shown on this [🖼️ screenshot](../screenshots/text-generation.webp).
|
In Direct Message rooms with the bot (1:1 rooms), it most usually makes sense for the bot to respond to **all** of your messages, as shown on this [🖼️ screenshot](../screenshots/text-generation.webp).
|
||||||
|
|
||||||
In group rooms (with multiple users), it may be more appropriate for the bot to only respond to messages that are **prefixed** with the command prefix (e.g. `!bai`), so that other chat exchange in the room will not trigger it. Such a setup is shown on this [🖼️ screenshot](../screenshots/text-generation-prefix-requirement.webp).
|
In group rooms (with multiple users), it may be more appropriate for the bot to only respond to messages that are **prefixed** with the command prefix (e.g. `!bai`) or which are [mentioning](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) the bot (e.g. `@baibot`), so that other chat exchange in the room will not trigger it. Such a setup is shown on the [🖼️ On-demand involvement in the room](../screenshots/text-generation-prefix-requirement.webp) screenshot.
|
||||||
|
|
||||||
There are exceptions to these rules, and you can configure the bot to respond only to prefixed messages in a 1:1 room, or to respond to all messages even in a multi-user group room.
|
There are exceptions to these rules, and you can configure the bot to respond only to prefixed messages in a 1:1 room, or to respond to all messages even in a multi-user group room.
|
||||||
|
|
||||||
@@ -27,7 +27,10 @@ By default, the bot is **auto-configured (upon joining a new room)** to use the
|
|||||||
|
|
||||||
Example: `!bai config room text-generation set-prefix-requirement-type command_prefix` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
Example: `!bai config room text-generation set-prefix-requirement-type command_prefix` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
||||||
|
|
||||||
Regardless of this configuration, **the bot will also respond to messages which directly [mention](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) the bot** (e.g. `@baibot`), even if they are not prefixed. An example of this can be seen on this [🖼️ screenshot](../screenshots/text-generation-prefix-requirement.webp).
|
Regardless of this configuration, **the bot will also respond to messages by allowed [👥 Users](../access.md#-users) which directly [mention](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) the bot** (e.g. `@baibot`), even if they are not prefixed. An example of this can be seen on these screenshots:
|
||||||
|
|
||||||
|
- [🖼️ On-demand involvement in a thread](../screenshots/text-generation-on-demand-thread-involvement.webp)
|
||||||
|
- [🖼️ On-demand involvement in a reply chain](../screenshots/text-generation-on-demand-reply-involvement.webp)
|
||||||
|
|
||||||
|
|
||||||
### 🪄 Auto Usage
|
### 🪄 Auto Usage
|
||||||
@@ -74,11 +77,14 @@ Prompts may contain the following **placeholder variables** which will be replac
|
|||||||
|---------------------------|-------------|---------|
|
|---------------------------|-------------|---------|
|
||||||
| `{{ baibot_name }}` | Name of the bot as configured in the `user.name` field in the [Static configuration](./README.md#static-configuration) | `Baibot` |
|
| `{{ baibot_name }}` | Name of the bot as configured in the `user.name` field in the [Static configuration](./README.md#static-configuration) | `Baibot` |
|
||||||
| `{{ baibot_model_id }}` | Text-Generation model ID as configured in the [🤖 agent](../agents.md)'s configuration | `gpt-4o` |
|
| `{{ baibot_model_id }}` | Text-Generation model ID as configured in the [🤖 agent](../agents.md)'s configuration | `gpt-4o` |
|
||||||
| `{{ baibot_now_utc }}` | Current date and time in UTC | `2024-09-20 (Friday), 14:26:42 UTC (local timezone/time: unknown)` |
|
| `{{ baibot_now_utc }}` | Current date and time in UTC (⚠️ usage may break prompt caching - see below) | `2024-09-20 (Friday), 14:26:42 UTC` |
|
||||||
|
| `{{ baibot_conversation_start_time_utc }}` | The date and time in UTC that the conversation started | `2024-09-20 (Friday), 14:26:42 UTC` |
|
||||||
|
|
||||||
|
💡 `{{ baibot_now_utc }}` changes as time goes on, which prevents [prompt caching](https://platform.openai.com/docs/guides/prompt-caching) from working. It's better to use `{{ baibot_conversation_start_time_utc }}` in prompts, as its value doesn't change yet still orients the bot to the current date/time.
|
||||||
|
|
||||||
Here's a prompt that combines some of the above variables:
|
Here's a prompt that combines some of the above variables:
|
||||||
|
|
||||||
> You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
> You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
|
|
||||||
|
|
||||||
### 🌡️ Temperature Override
|
### 🌡️ Temperature Override
|
||||||
|
|||||||
@@ -28,6 +28,8 @@ Text Generation is the bot's ability to **respond to users' text messages with t
|
|||||||
|
|
||||||
In multi-user (group) rooms, to avoid disturbing the normal conversation between people, the bot is auto-configured to only respond to messages starting with the command prefix (`!bai`) or direct mentions via the [💬 Text Generation / 🗟 Prefix Requirement Type](./configuration/text-generation.md#-prefix-requirement-type) setting.
|
In multi-user (group) rooms, to avoid disturbing the normal conversation between people, the bot is auto-configured to only respond to messages starting with the command prefix (`!bai`) or direct mentions via the [💬 Text Generation / 🗟 Prefix Requirement Type](./configuration/text-generation.md#-prefix-requirement-type) setting.
|
||||||
|
|
||||||
|
Normally, the bot only responds to allowed [👥 Users](./access.md#-users). In certain cases, it's useful for an allowed user to provoke the bot to respond even in foreign threads or reply chains. You can learn more about this feature in the [On-demand involvement](./features.md#on-demand-involvement) section below.
|
||||||
|
|
||||||
A few other features (like [🗣️ Text-to-Speech](#️-text-to-speech) and [🦻 Speech-to-Text](#-speech-to-text)) combine well with Text Generation, so you **don't necessarily need to communicate with the bot via text** (with [Seamless voice interaction](#seamless-voice-interaction), you can communicate only with voice).
|
A few other features (like [🗣️ Text-to-Speech](#️-text-to-speech) and [🦻 Speech-to-Text](#-speech-to-text)) combine well with Text Generation, so you **don't necessarily need to communicate with the bot via text** (with [Seamless voice interaction](#seamless-voice-interaction), you can communicate only with voice).
|
||||||
|
|
||||||
You may also wish to see:
|
You may also wish to see:
|
||||||
@@ -36,6 +38,22 @@ You may also wish to see:
|
|||||||
- [📖 Usage / 💬 Text Generation](./usage.md#-text-generation) section for more details on how to use the bot for Text Generation in a room
|
- [📖 Usage / 💬 Text Generation](./usage.md#-text-generation) section for more details on how to use the bot for Text Generation in a room
|
||||||
|
|
||||||
|
|
||||||
|
#### On-demand involvement
|
||||||
|
|
||||||
|
In the following 2 cases, it's useful to involve the bot in conversations on-demand:
|
||||||
|
|
||||||
|
1. In multi-user rooms (with the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting set to "required")
|
||||||
|
2. In rooms with foreign users (users that are not authorized bot [👥 users](./access.md#-users))
|
||||||
|
|
||||||
|
In these instances, an allowed [👥 user](./access.md#-users) can also provoke the bot to respond to **any** thread or reply chain by [mentioning](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) the bot (e.g. `@baibot Hello!`). The following screenshots demonstrate this behavior:
|
||||||
|
|
||||||
|
- [🖼️ On-demand involvement in the room](./screenshots/text-generation-prefix-requirement.webp)
|
||||||
|
- [🖼️ On-demand involvement in a thread](./screenshots/text-generation-on-demand-thread-involvement.webp) (the Alice user in this example is not an allowed user, yet her messages are still considered as part of the conversation context)
|
||||||
|
- [🖼️ On-demand involvement in a reply chain](./screenshots/text-generation-on-demand-reply-involvement.webp) (the Alice user in this example is not an allowed user, yet her messages are still considered as part of the conversation context)
|
||||||
|
|
||||||
|
💡 **NOTE**: Normally, the bot **only considers messages from allowed [👥 Users](./access.md#-users)** and ignores all other messages when responding. However, **when the bot is explicitly invoked (via mention)** in a thread or reply chain, **it will consider all messages** in the thread and reply chain (even those from foreign users) as part of the conversation context.
|
||||||
|
|
||||||
|
|
||||||
### 🗣️ Text-to-Speech
|
### 🗣️ Text-to-Speech
|
||||||
|
|
||||||
Text-to-Speech is the bot's ability to **turn text messages into voice messages**.
|
Text-to-Speech is the bot's ability to **turn text messages into voice messages**.
|
||||||
|
|||||||
@@ -125,7 +125,10 @@ For services which are not fully compatible with the OpenAI API, consider using
|
|||||||
- create a room-local agent: `!bai agent create-room-local openai my-openai-agent`
|
- create a room-local agent: `!bai agent create-room-local openai my-openai-agent`
|
||||||
- create a global agent: `!bai agent create-global openai my-openai-agent`
|
- create a global agent: `!bai agent create-global openai my-openai-agent`
|
||||||
|
|
||||||
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/openai.yml).
|
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which:
|
||||||
|
|
||||||
|
- in the general case looks [like this](./sample-provider-configs/openai.yml)
|
||||||
|
- for the [o1](https://platform.openai.com/docs/models/o1) models needs to look [like this](./sample-provider-configs/openai-o1.yml)
|
||||||
|
|
||||||
|
|
||||||
### OpenAI Compatible
|
### OpenAI Compatible
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ base_url: https://api.anthropic.com/v1
|
|||||||
api_key: YOUR_API_KEY_HERE
|
api_key: YOUR_API_KEY_HERE
|
||||||
text_generation:
|
text_generation:
|
||||||
model_id: claude-3-5-sonnet-20240620
|
model_id: claude-3-5-sonnet-20240620
|
||||||
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
max_response_tokens: 8192
|
max_response_tokens: 8192
|
||||||
max_context_tokens: 204800
|
max_context_tokens: 204800
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ base_url: https://api.groq.com/openai/v1
|
|||||||
api_key: YOUR_API_KEY_HERE
|
api_key: YOUR_API_KEY_HERE
|
||||||
text_generation:
|
text_generation:
|
||||||
model_id: llama3-70b-8192
|
model_id: llama3-70b-8192
|
||||||
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
max_response_tokens: 4096
|
max_response_tokens: 4096
|
||||||
max_context_tokens: 131072
|
max_context_tokens: 131072
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ base_url: http://my-localai-self-hosted-service:8080/v1
|
|||||||
api_key: YOUR_API_KEY_HERE
|
api_key: YOUR_API_KEY_HERE
|
||||||
text_generation:
|
text_generation:
|
||||||
model_id: gpt-4
|
model_id: gpt-4
|
||||||
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
max_response_tokens: 4096
|
max_response_tokens: 4096
|
||||||
max_context_tokens: 128000
|
max_context_tokens: 128000
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ base_url: https://api.mistral.ai/v1
|
|||||||
api_key: YOUR_API_KEY_HERE
|
api_key: YOUR_API_KEY_HERE
|
||||||
text_generation:
|
text_generation:
|
||||||
model_id: mistral-large-latest
|
model_id: mistral-large-latest
|
||||||
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
max_response_tokens: 4096
|
max_response_tokens: 4096
|
||||||
max_context_tokens: 128000
|
max_context_tokens: 128000
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ base_url: http://my-ollama-self-hosted-service:11434/v1
|
|||||||
api_key: YOUR_API_KEY_HERE
|
api_key: YOUR_API_KEY_HERE
|
||||||
text_generation:
|
text_generation:
|
||||||
model_id: gemma2:2b
|
model_id: gemma2:2b
|
||||||
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
max_response_tokens: 4096
|
max_response_tokens: 4096
|
||||||
max_context_tokens: 128000
|
max_context_tokens: 128000
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ base_url: ''
|
|||||||
api_key: YOUR_API_KEY_HERE
|
api_key: YOUR_API_KEY_HERE
|
||||||
text_generation:
|
text_generation:
|
||||||
model_id: some-model
|
model_id: some-model
|
||||||
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
max_response_tokens: 4096
|
max_response_tokens: 4096
|
||||||
max_context_tokens: 128000
|
max_context_tokens: 128000
|
||||||
|
|||||||
24
docs/sample-provider-configs/openai-o1.yml
Normal file
24
docs/sample-provider-configs/openai-o1.yml
Normal file
@@ -0,0 +1,24 @@
|
|||||||
|
base_url: https://api.openai.com/v1
|
||||||
|
api_key: YOUR_API_KEY_HERE
|
||||||
|
text_generation:
|
||||||
|
model_id: o1-mini
|
||||||
|
# o1 models do not support a system prompt
|
||||||
|
prompt: null
|
||||||
|
temperature: 1.0
|
||||||
|
# o1 models do not support max_response_tokens.
|
||||||
|
# They use `max_completion_tokens` as an alternative,
|
||||||
|
# but we don't support it yet (see https://github.com/64bit/async-openai/issues/272).
|
||||||
|
max_response_tokens: null
|
||||||
|
max_context_tokens: 128000
|
||||||
|
speech_to_text:
|
||||||
|
model_id: whisper-1
|
||||||
|
text_to_speech:
|
||||||
|
model_id: tts-1-hd
|
||||||
|
voice: onyx
|
||||||
|
speed: 1.0
|
||||||
|
response_format: opus
|
||||||
|
image_generation:
|
||||||
|
model_id: dall-e-3
|
||||||
|
style: vivid
|
||||||
|
size: 1024x1024
|
||||||
|
quality: standard
|
||||||
@@ -1,8 +1,8 @@
|
|||||||
base_url: https://api.openai.com/v1
|
base_url: https://api.openai.com/v1
|
||||||
api_key: YOUR_API_KEY_HERE
|
api_key: YOUR_API_KEY_HERE
|
||||||
text_generation:
|
text_generation:
|
||||||
model_id: gpt-4o-2024-08-06
|
model_id: gpt-4o
|
||||||
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
max_response_tokens: 16384
|
max_response_tokens: 16384
|
||||||
max_context_tokens: 128000
|
max_context_tokens: 128000
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ base_url: https://openrouter.ai/api/v1
|
|||||||
api_key: YOUR_API_KEY_HERE
|
api_key: YOUR_API_KEY_HERE
|
||||||
text_generation:
|
text_generation:
|
||||||
model_id: mattshumer/reflection-70b:free
|
model_id: mattshumer/reflection-70b:free
|
||||||
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
max_response_tokens: 2048
|
max_response_tokens: 2048
|
||||||
max_context_tokens: 8192
|
max_context_tokens: 8192
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ base_url: https://api.together.xyz/v1
|
|||||||
api_key: YOUR_API_KEY_HERE
|
api_key: YOUR_API_KEY_HERE
|
||||||
text_generation:
|
text_generation:
|
||||||
model_id: meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo
|
model_id: meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo
|
||||||
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
max_response_tokens: 2048
|
max_response_tokens: 2048
|
||||||
max_context_tokens: 8192
|
max_context_tokens: 8192
|
||||||
|
|||||||
Binary file not shown.
|
After Width: | Height: | Size: 92 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 60 KiB |
@@ -11,10 +11,11 @@ This is related to the [💬 Text Generation](./features.md#-text-generation) fe
|
|||||||
|
|
||||||
If there's a text-generation handler agent configured, the bot **may** respond to messages sent in the room.
|
If there's a text-generation handler agent configured, the bot **may** respond to messages sent in the room.
|
||||||
|
|
||||||
🖼️ See screenshots of:
|
See screenshots of:
|
||||||
|
|
||||||
- the [default Text Generation flow](./screenshots/text-generation.webp) for 1:1 rooms
|
- 🖼️ [the default Text Generation flow](./screenshots/text-generation.webp) in 1:1 rooms
|
||||||
- the [Text Generation flow in multi-user rooms](./screenshots/text-generation-prefix-requirement.webp) (where the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting is auto-configured to "required")
|
- 🖼️ [the Text Generation flow in multi-user rooms](./screenshots/text-generation-prefix-requirement.webp) (where the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting is auto-configured to "required")
|
||||||
|
- the [on-demand involvement](./features.md#on-demand-involvement) feature
|
||||||
|
|
||||||
Whether the bot responds depends on:
|
Whether the bot responds depends on:
|
||||||
|
|
||||||
@@ -24,9 +25,9 @@ Whether the bot responds depends on:
|
|||||||
|
|
||||||
- (🎨 agent capabilities) whether the configured `text-generation` (or `catch-all`) handler agent actually supports text-generation. The provider may lack support for this feature or it may be disabled in the [🤖 agents](./agents.md) configuration
|
- (🎨 agent capabilities) whether the configured `text-generation` (or `catch-all`) handler agent actually supports text-generation. The provider may lack support for this feature or it may be disabled in the [🤖 agents](./agents.md) configuration
|
||||||
|
|
||||||
- (the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting) whether a prefix (e.g. `!bai`) is required in front of messages sent to the room. For multi-user rooms, this setting defaults to "required"
|
- (the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting) whether a prefix (e.g. `!bai`) or user mention (e.g. `@baibot`) is required for messages sent to the room. For multi-user rooms, this setting defaults to "required". See [🌟 Features / 💬 Text Generation / On-demand involvement](./features.md#on-demand-involvement) for details.
|
||||||
|
|
||||||
Room messages start a threaded conversation where you can continue back-and-forth communication with the bot.
|
Room messages start a threaded conversation where you can continue back-and-forth communication with the bot. Using [on-demand involvement](./features.md#on-demand-involvement), you can can also mention the bot to provoke it to get involved in any conversation thread or reply chain.
|
||||||
|
|
||||||
Unless you've enabled the [♻️ Context Management](./features.md#️-context-management) feature, all messages will be sent to the agent's API each time. If the context management feature is enabled, older messages may be dropped.
|
Unless you've enabled the [♻️ Context Management](./features.md#️-context-management) feature, all messages will be sent to the agent's API each time. If the context management feature is enabled, older messages may be dropped.
|
||||||
|
|
||||||
|
|||||||
@@ -72,8 +72,8 @@ agents:
|
|||||||
# base_url: https://api.openai.com/v1
|
# base_url: https://api.openai.com/v1
|
||||||
# api_key: ""
|
# api_key: ""
|
||||||
# text_generation:
|
# text_generation:
|
||||||
# model_id: gpt-4o-2024-08-06
|
# model_id: gpt-4o
|
||||||
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
# temperature: 1.0
|
# temperature: 1.0
|
||||||
# max_response_tokens: 16384
|
# max_response_tokens: 16384
|
||||||
# max_context_tokens: 128000
|
# max_context_tokens: 128000
|
||||||
@@ -97,7 +97,7 @@ agents:
|
|||||||
# api_key: null
|
# api_key: null
|
||||||
# text_generation:
|
# text_generation:
|
||||||
# model_id: gpt-4
|
# model_id: gpt-4
|
||||||
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
# temperature: 1.0
|
# temperature: 1.0
|
||||||
# max_response_tokens: 16384
|
# max_response_tokens: 16384
|
||||||
# max_context_tokens: 128000
|
# max_context_tokens: 128000
|
||||||
@@ -122,7 +122,7 @@ agents:
|
|||||||
# api_key: null
|
# api_key: null
|
||||||
# text_generation:
|
# text_generation:
|
||||||
# model_id: "gemma2:2b"
|
# model_id: "gemma2:2b"
|
||||||
# prompt: "You are an assistant based on the gemma2:2b model. Be brief in your responses."
|
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
# temperature: 1.0
|
# temperature: 1.0
|
||||||
# max_response_tokens: 4096
|
# max_response_tokens: 4096
|
||||||
# max_context_tokens: 128000
|
# max_context_tokens: 128000
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
services:
|
services:
|
||||||
postgres:
|
postgres:
|
||||||
image: docker.io/postgres:16.3-alpine
|
image: docker.io/postgres:16.4-alpine
|
||||||
user: ${UID}:${GID}
|
user: ${UID}:${GID}
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
environment:
|
environment:
|
||||||
@@ -13,7 +13,7 @@ services:
|
|||||||
- /etc/passwd:/etc/passwd:ro
|
- /etc/passwd:/etc/passwd:ro
|
||||||
|
|
||||||
synapse:
|
synapse:
|
||||||
image: ghcr.io/element-hq/synapse:v1.114.0
|
image: ghcr.io/element-hq/synapse:v1.116.0
|
||||||
user: "${UID}:${GID}"
|
user: "${UID}:${GID}"
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
entrypoint: python
|
entrypoint: python
|
||||||
@@ -26,7 +26,7 @@ services:
|
|||||||
- ./synapse/media-store:/media-store
|
- ./synapse/media-store:/media-store
|
||||||
|
|
||||||
element-web:
|
element-web:
|
||||||
image: docker.io/vectorim/element-web:v1.11.77
|
image: docker.io/vectorim/element-web:v1.11.79
|
||||||
user: "${UID}:${GID}"
|
user: "${UID}:${GID}"
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
ports:
|
ports:
|
||||||
|
|||||||
@@ -579,7 +579,9 @@ rc_login:
|
|||||||
#
|
#
|
||||||
#federation_rr_transactions_per_room_per_second: 50
|
#federation_rr_transactions_per_room_per_second: 50
|
||||||
|
|
||||||
|
# Authenticated media is not supported yet.
|
||||||
|
# See: https://github.com/etkecc/baibot/issues/12
|
||||||
|
enable_authenticated_media: false
|
||||||
|
|
||||||
# Directory where uploaded images and attachments are stored.
|
# Directory where uploaded images and attachments are stored.
|
||||||
#
|
#
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
services:
|
services:
|
||||||
ollama:
|
ollama:
|
||||||
image: docker.io/ollama/ollama:0.3.9
|
image: docker.io/ollama/ollama:0.3.11
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
ports:
|
ports:
|
||||||
- "${SERVICE_OLLAMA_BIND_PORT_HTTP}:11434"
|
- "${SERVICE_OLLAMA_BIND_PORT_HTTP}:11434"
|
||||||
|
|||||||
@@ -21,5 +21,5 @@ pub use provider::{AgentProvider, AgentProviderInfo, ControllerTrait};
|
|||||||
pub use purpose::AgentPurpose;
|
pub use purpose::AgentPurpose;
|
||||||
|
|
||||||
pub(super) fn default_prompt() -> &'static str {
|
pub(super) fn default_prompt() -> &'static str {
|
||||||
"You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time now is: {{ baibot_now_utc }}."
|
"You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ use std::fmt::Debug;
|
|||||||
use std::str::FromStr;
|
use std::str::FromStr;
|
||||||
use std::sync::Arc;
|
use std::sync::Arc;
|
||||||
|
|
||||||
use anthropic_rs::completion::message::ContentType;
|
use anthropic_rs::completion::message::{ContentType, System};
|
||||||
use anthropic_rs::{
|
use anthropic_rs::{
|
||||||
client::Client as AnthropicClient, config::Config as AnthropicConfig,
|
client::Client as AnthropicClient, config::Config as AnthropicConfig,
|
||||||
models::claude::ClaudeModel,
|
models::claude::ClaudeModel,
|
||||||
@@ -72,6 +72,7 @@ impl ControllerTrait for Controller {
|
|||||||
let messages = vec![LLMMessage {
|
let messages = vec![LLMMessage {
|
||||||
author: LLMAuthor::User,
|
author: LLMAuthor::User,
|
||||||
message_text: "Hello!".to_string(),
|
message_text: "Hello!".to_string(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
}];
|
}];
|
||||||
|
|
||||||
let conversation = LLMConversation { messages };
|
let conversation = LLMConversation { messages };
|
||||||
@@ -108,6 +109,7 @@ impl ControllerTrait for Controller {
|
|||||||
Some(LLMMessage {
|
Some(LLMMessage {
|
||||||
author: LLMAuthor::Prompt,
|
author: LLMAuthor::Prompt,
|
||||||
message_text: prompt_text,
|
message_text: prompt_text,
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
})
|
})
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -129,7 +131,7 @@ impl ControllerTrait for Controller {
|
|||||||
&text_generation_config.model_id,
|
&text_generation_config.model_id,
|
||||||
&prompt_message,
|
&prompt_message,
|
||||||
conversation_messages,
|
conversation_messages,
|
||||||
text_generation_config.max_response_tokens,
|
Some(text_generation_config.max_response_tokens),
|
||||||
text_generation_config.max_context_tokens,
|
text_generation_config.max_context_tokens,
|
||||||
);
|
);
|
||||||
|
|
||||||
@@ -157,7 +159,7 @@ impl ControllerTrait for Controller {
|
|||||||
.unwrap_or(text_generation_config.temperature);
|
.unwrap_or(text_generation_config.temperature);
|
||||||
|
|
||||||
if let Some(prompt_message) = prompt_message {
|
if let Some(prompt_message) = prompt_message {
|
||||||
request.system = Some(prompt_message.message_text);
|
request.system = Some(System::Text(prompt_message.message_text));
|
||||||
}
|
}
|
||||||
|
|
||||||
request.model = model;
|
request.model = model;
|
||||||
|
|||||||
@@ -7,17 +7,33 @@ pub struct TextGenerationPromptVariables {
|
|||||||
|
|
||||||
impl Default for TextGenerationPromptVariables {
|
impl Default for TextGenerationPromptVariables {
|
||||||
fn default() -> Self {
|
fn default() -> Self {
|
||||||
Self::new("unnamed", "unknown-model", Utc::now())
|
let now = Utc::now();
|
||||||
|
Self::new("unnamed", "unknown-model", now, Some(now))
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
impl TextGenerationPromptVariables {
|
impl TextGenerationPromptVariables {
|
||||||
pub fn new(bot_name: &str, model_id: &str, utc_time: DateTime<Utc>) -> Self {
|
pub fn new(
|
||||||
|
bot_name: &str,
|
||||||
|
model_id: &str,
|
||||||
|
now_time: DateTime<Utc>,
|
||||||
|
conversation_start_time: Option<DateTime<Utc>>,
|
||||||
|
) -> Self {
|
||||||
let mut map = HashMap::new();
|
let mut map = HashMap::new();
|
||||||
|
|
||||||
map.insert("baibot_name".to_string(), bot_name.to_string());
|
map.insert("baibot_name".to_string(), bot_name.to_string());
|
||||||
map.insert("baibot_model_id".to_string(), model_id.to_string());
|
map.insert("baibot_model_id".to_string(), model_id.to_string());
|
||||||
map.insert("baibot_now_utc".to_string(), format_utc_time(utc_time));
|
map.insert("baibot_now_utc".to_string(), format_utc_time(now_time));
|
||||||
|
|
||||||
|
let baibot_conversation_start_time_utc = match conversation_start_time {
|
||||||
|
Some(conversation_start_time) => format_utc_time(conversation_start_time),
|
||||||
|
None => "unknown".to_string(),
|
||||||
|
};
|
||||||
|
|
||||||
|
map.insert(
|
||||||
|
"baibot_conversation_start_time_utc".to_string(),
|
||||||
|
baibot_conversation_start_time_utc,
|
||||||
|
);
|
||||||
|
|
||||||
Self { map }
|
Self { map }
|
||||||
}
|
}
|
||||||
@@ -52,7 +68,18 @@ mod tests {
|
|||||||
.with_nanosecond(250000000)
|
.with_nanosecond(250000000)
|
||||||
.unwrap();
|
.unwrap();
|
||||||
|
|
||||||
let variables = TextGenerationPromptVariables::new("baibot", "gpt-4o", now_utc);
|
let conversation_start_time_utc = Utc
|
||||||
|
.with_ymd_and_hms(2024, 9, 19, 18, 34, 15)
|
||||||
|
.unwrap()
|
||||||
|
.with_nanosecond(250000000)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
let variables = TextGenerationPromptVariables::new(
|
||||||
|
"baibot",
|
||||||
|
"gpt-4o",
|
||||||
|
now_utc,
|
||||||
|
Some(conversation_start_time_utc),
|
||||||
|
);
|
||||||
|
|
||||||
assert_eq!(
|
assert_eq!(
|
||||||
variables.map.get("baibot_name"),
|
variables.map.get("baibot_name"),
|
||||||
@@ -66,9 +93,13 @@ mod tests {
|
|||||||
variables.map.get("baibot_now_utc"),
|
variables.map.get("baibot_now_utc"),
|
||||||
Some(&format_utc_time(now_utc))
|
Some(&format_utc_time(now_utc))
|
||||||
);
|
);
|
||||||
|
assert_eq!(
|
||||||
|
variables.map.get("baibot_conversation_start_time_utc"),
|
||||||
|
Some(&format_utc_time(conversation_start_time_utc))
|
||||||
|
);
|
||||||
|
|
||||||
let prompt = "Hello, I'm {{ baibot_name }} using {{ baibot_model_id }}. The date/time now is {{ baibot_now_utc }}.";
|
let prompt = "Hello, I'm {{ baibot_name }} using {{ baibot_model_id }}. The date/time now is {{ baibot_now_utc }} and this conversation started at {{ baibot_conversation_start_time_utc }}.";
|
||||||
let expected = "Hello, I'm baibot using gpt-4o. The date/time now is 2024-09-20 (Friday), 18:34:15 UTC.";
|
let expected = "Hello, I'm baibot using gpt-4o. The date/time now is 2024-09-20 (Friday), 18:34:15 UTC and this conversation started at 2024-09-19 (Thursday), 18:34:15 UTC.";
|
||||||
|
|
||||||
assert_eq!(variables.format(prompt), expected);
|
assert_eq!(variables.format(prompt), expected);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -15,7 +15,7 @@ pub fn default_config() -> Config {
|
|||||||
if let Some(ref mut config) = config.text_generation.as_mut() {
|
if let Some(ref mut config) = config.text_generation.as_mut() {
|
||||||
config.model_id = "llama3-70b-8192".to_owned();
|
config.model_id = "llama3-70b-8192".to_owned();
|
||||||
config.max_context_tokens = 131_072;
|
config.max_context_tokens = 131_072;
|
||||||
config.max_response_tokens = 4096;
|
config.max_response_tokens = Some(4096);
|
||||||
}
|
}
|
||||||
|
|
||||||
if let Some(ref mut config) = config.speech_to_text.as_mut() {
|
if let Some(ref mut config) = config.speech_to_text.as_mut() {
|
||||||
|
|||||||
@@ -13,7 +13,7 @@ pub fn default_config() -> Config {
|
|||||||
if let Some(ref mut config) = config.text_generation.as_mut() {
|
if let Some(ref mut config) = config.text_generation.as_mut() {
|
||||||
config.model_id = "gpt-4".to_owned();
|
config.model_id = "gpt-4".to_owned();
|
||||||
config.max_context_tokens = 128_000;
|
config.max_context_tokens = 128_000;
|
||||||
config.max_response_tokens = 4096;
|
config.max_response_tokens = Some(4096);
|
||||||
}
|
}
|
||||||
|
|
||||||
if let Some(ref mut config) = config.text_to_speech.as_mut() {
|
if let Some(ref mut config) = config.text_to_speech.as_mut() {
|
||||||
|
|||||||
@@ -17,7 +17,7 @@ pub fn default_config() -> Config {
|
|||||||
if let Some(ref mut config) = config.text_generation.as_mut() {
|
if let Some(ref mut config) = config.text_generation.as_mut() {
|
||||||
config.model_id = "gemma2:2b".to_owned();
|
config.model_id = "gemma2:2b".to_owned();
|
||||||
config.max_context_tokens = 128_000;
|
config.max_context_tokens = 128_000;
|
||||||
config.max_response_tokens = 4096;
|
config.max_response_tokens = Some(4096);
|
||||||
}
|
}
|
||||||
|
|
||||||
config
|
config
|
||||||
|
|||||||
@@ -56,7 +56,7 @@ pub struct TextGenerationConfig {
|
|||||||
pub temperature: f32,
|
pub temperature: f32,
|
||||||
|
|
||||||
#[serde(default)]
|
#[serde(default)]
|
||||||
pub max_response_tokens: u32,
|
pub max_response_tokens: Option<u32>,
|
||||||
|
|
||||||
#[serde(default)]
|
#[serde(default)]
|
||||||
pub max_context_tokens: u32,
|
pub max_context_tokens: u32,
|
||||||
@@ -68,14 +68,14 @@ impl Default for TextGenerationConfig {
|
|||||||
model_id: default_text_model_id(),
|
model_id: default_text_model_id(),
|
||||||
prompt: Some(default_prompt().to_owned()),
|
prompt: Some(default_prompt().to_owned()),
|
||||||
temperature: super::super::default_temperature(),
|
temperature: super::super::default_temperature(),
|
||||||
max_response_tokens: 16_384,
|
max_response_tokens: Some(16_384),
|
||||||
max_context_tokens: 128_000,
|
max_context_tokens: 128_000,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn default_text_model_id() -> String {
|
fn default_text_model_id() -> String {
|
||||||
"gpt-4o-2024-08-06".to_owned()
|
"gpt-4o".to_owned()
|
||||||
}
|
}
|
||||||
|
|
||||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||||
|
|||||||
@@ -63,6 +63,7 @@ impl ControllerTrait for Controller {
|
|||||||
let messages = vec![LLMMessage {
|
let messages = vec![LLMMessage {
|
||||||
author: LLMAuthor::User,
|
author: LLMAuthor::User,
|
||||||
message_text: "Hello!".to_string(),
|
message_text: "Hello!".to_string(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
}];
|
}];
|
||||||
|
|
||||||
let conversation = LLMConversation { messages };
|
let conversation = LLMConversation { messages };
|
||||||
@@ -99,6 +100,7 @@ impl ControllerTrait for Controller {
|
|||||||
Some(LLMMessage {
|
Some(LLMMessage {
|
||||||
author: LLMAuthor::Prompt,
|
author: LLMAuthor::Prompt,
|
||||||
message_text: prompt_text,
|
message_text: prompt_text,
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
})
|
})
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -131,12 +133,18 @@ impl ControllerTrait for Controller {
|
|||||||
.temperature_override
|
.temperature_override
|
||||||
.unwrap_or(text_generation_config.temperature);
|
.unwrap_or(text_generation_config.temperature);
|
||||||
|
|
||||||
let request = CreateChatCompletionRequestArgs::default()
|
let mut request_builder = CreateChatCompletionRequestArgs::default();
|
||||||
.max_tokens(text_generation_config.max_response_tokens)
|
|
||||||
|
request_builder
|
||||||
.model(&text_generation_config.model_id)
|
.model(&text_generation_config.model_id)
|
||||||
.temperature(temperature)
|
.temperature(temperature)
|
||||||
.messages(openai_conversation_messages)
|
.messages(openai_conversation_messages);
|
||||||
.build()?;
|
|
||||||
|
if let Some(max_response_tokens) = text_generation_config.max_response_tokens {
|
||||||
|
request_builder.max_tokens(max_response_tokens);
|
||||||
|
}
|
||||||
|
|
||||||
|
let request = request_builder.build()?;
|
||||||
|
|
||||||
if let Ok(request_as_json) = serde_json::to_string(&request) {
|
if let Ok(request_as_json) = serde_json::to_string(&request) {
|
||||||
tracing::trace!(
|
tracing::trace!(
|
||||||
|
|||||||
@@ -66,7 +66,7 @@ pub struct TextGenerationConfig {
|
|||||||
pub temperature: f32,
|
pub temperature: f32,
|
||||||
|
|
||||||
#[serde(default)]
|
#[serde(default)]
|
||||||
pub max_response_tokens: u32,
|
pub max_response_tokens: Option<u32>,
|
||||||
|
|
||||||
#[serde(default)]
|
#[serde(default)]
|
||||||
pub max_context_tokens: u32,
|
pub max_context_tokens: u32,
|
||||||
@@ -78,7 +78,7 @@ impl Default for TextGenerationConfig {
|
|||||||
model_id: default_text_model_id(),
|
model_id: default_text_model_id(),
|
||||||
prompt: Some(default_prompt().to_owned()),
|
prompt: Some(default_prompt().to_owned()),
|
||||||
temperature: super::super::default_temperature(),
|
temperature: super::super::default_temperature(),
|
||||||
max_response_tokens: 4096,
|
max_response_tokens: Some(4096),
|
||||||
max_context_tokens: 128_000,
|
max_context_tokens: 128_000,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -61,6 +61,7 @@ impl ControllerTrait for Controller {
|
|||||||
let messages = vec![LLMMessage {
|
let messages = vec![LLMMessage {
|
||||||
author: LLMAuthor::User,
|
author: LLMAuthor::User,
|
||||||
message_text: "Hello!".to_string(),
|
message_text: "Hello!".to_string(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
}];
|
}];
|
||||||
|
|
||||||
let conversation = LLMConversation { messages };
|
let conversation = LLMConversation { messages };
|
||||||
@@ -97,6 +98,7 @@ impl ControllerTrait for Controller {
|
|||||||
Some(LLMMessage {
|
Some(LLMMessage {
|
||||||
author: LLMAuthor::Prompt,
|
author: LLMAuthor::Prompt,
|
||||||
message_text: prompt_text,
|
message_text: prompt_text,
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
})
|
})
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -131,12 +133,15 @@ impl ControllerTrait for Controller {
|
|||||||
|
|
||||||
let max_tokens = text_generation_config
|
let max_tokens = text_generation_config
|
||||||
.max_response_tokens
|
.max_response_tokens
|
||||||
|
.map(|max_response_tokens| {
|
||||||
|
max_response_tokens
|
||||||
.try_into()
|
.try_into()
|
||||||
.expect("Failed converting max_response_tokens from u32 to i32");
|
.expect("Failed converting max_response_tokens from u32 to i32")
|
||||||
|
});
|
||||||
|
|
||||||
let request = ChatBody {
|
let request = ChatBody {
|
||||||
model: text_generation_config.model_id.clone(),
|
model: text_generation_config.model_id.clone(),
|
||||||
max_tokens: Some(max_tokens),
|
max_tokens,
|
||||||
temperature: Some(temperature),
|
temperature: Some(temperature),
|
||||||
top_p: None,
|
top_p: None,
|
||||||
n: Some(1),
|
n: Some(1),
|
||||||
|
|||||||
@@ -56,7 +56,7 @@ pub fn default_config() -> Config {
|
|||||||
|
|
||||||
if let Some(text_generation) = &mut config.text_generation {
|
if let Some(text_generation) = &mut config.text_generation {
|
||||||
text_generation.model_id = "some-model".to_string();
|
text_generation.model_id = "some-model".to_string();
|
||||||
text_generation.max_response_tokens = 4096;
|
text_generation.max_response_tokens = Some(4096);
|
||||||
text_generation.max_context_tokens = 128_000;
|
text_generation.max_context_tokens = 128_000;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -14,7 +14,7 @@ pub fn default_config() -> Config {
|
|||||||
if let Some(ref mut config) = config.text_generation.as_mut() {
|
if let Some(ref mut config) = config.text_generation.as_mut() {
|
||||||
config.model_id = "mattshumer/reflection-70b:free".to_owned();
|
config.model_id = "mattshumer/reflection-70b:free".to_owned();
|
||||||
config.max_context_tokens = 8192;
|
config.max_context_tokens = 8192;
|
||||||
config.max_response_tokens = 2048;
|
config.max_response_tokens = Some(2048);
|
||||||
}
|
}
|
||||||
|
|
||||||
config
|
config
|
||||||
|
|||||||
@@ -14,7 +14,7 @@ pub fn default_config() -> Config {
|
|||||||
if let Some(ref mut config) = config.text_generation.as_mut() {
|
if let Some(ref mut config) = config.text_generation.as_mut() {
|
||||||
config.model_id = "meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo".to_owned();
|
config.model_id = "meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo".to_owned();
|
||||||
config.max_context_tokens = 8192;
|
config.max_context_tokens = 8192;
|
||||||
config.max_response_tokens = 2048;
|
config.max_response_tokens = Some(2048);
|
||||||
}
|
}
|
||||||
|
|
||||||
config
|
config
|
||||||
|
|||||||
@@ -172,6 +172,24 @@ impl Bot {
|
|||||||
self.matrix_link().user_id()
|
self.matrix_link().user_id()
|
||||||
}
|
}
|
||||||
|
|
||||||
|
pub(crate) async fn user_display_name_in_room(&self, room: &Room) -> Option<String> {
|
||||||
|
let bot_display_name = self
|
||||||
|
.room_display_name_fetcher()
|
||||||
|
.own_display_name_in_room(room)
|
||||||
|
.await;
|
||||||
|
|
||||||
|
match bot_display_name {
|
||||||
|
Ok(value) => value,
|
||||||
|
Err(err) => {
|
||||||
|
tracing::warn!(
|
||||||
|
?err,
|
||||||
|
"Failed to fetch bot display name. Proceeding without it"
|
||||||
|
);
|
||||||
|
None
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
pub(crate) fn reacting(&self) -> super::reacting::Reacting {
|
pub(crate) fn reacting(&self) -> super::reacting::Reacting {
|
||||||
super::reacting::Reacting::new(self.clone())
|
super::reacting::Reacting::new(self.clone())
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -11,7 +11,7 @@ use mxlink::{CallbackError, MessageResponseType};
|
|||||||
use tracing::Instrument;
|
use tracing::Instrument;
|
||||||
|
|
||||||
use crate::{
|
use crate::{
|
||||||
conversation::matrix::determine_thread_context_for_room_event,
|
conversation::matrix::determine_interaction_context_for_room_event,
|
||||||
entity::{MessageContext, MessagePayload, RoomConfigContext, TriggerEventInfo},
|
entity::{MessageContext, MessagePayload, RoomConfigContext, TriggerEventInfo},
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -239,8 +239,11 @@ impl Messaging {
|
|||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
let thread_context = determine_thread_context_for_room_event(
|
let bot_display_name = self.bot.user_display_name_in_room(&room).await;
|
||||||
|
|
||||||
|
let interaction_context = determine_interaction_context_for_room_event(
|
||||||
self.bot.user_id(),
|
self.bot.user_id(),
|
||||||
|
&bot_display_name,
|
||||||
&room,
|
&room,
|
||||||
&event,
|
&event,
|
||||||
&payload,
|
&payload,
|
||||||
@@ -248,16 +251,18 @@ impl Messaging {
|
|||||||
)
|
)
|
||||||
.await;
|
.await;
|
||||||
|
|
||||||
let thread_context = match thread_context {
|
let interaction_context = match interaction_context {
|
||||||
Ok(value) => value,
|
Ok(value) => value,
|
||||||
Err(err) => {
|
Err(err) => {
|
||||||
tracing::error!(?err, "Failed to determine thread context for event");
|
tracing::error!(?err, "Failed to determine interaction context for event");
|
||||||
return Ok(());
|
return Ok(());
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
let Some(thread_context) = thread_context else {
|
let Some(interaction_context) = interaction_context else {
|
||||||
tracing::debug!("Ignoring message with unknown thread context (likely not a threaded message or a top-level message)");
|
tracing::debug!(
|
||||||
|
"Ignoring message with unknown interaction context (likely not a message for us)"
|
||||||
|
);
|
||||||
return Ok(());
|
return Ok(());
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -276,33 +281,14 @@ impl Messaging {
|
|||||||
room_config_context,
|
room_config_context,
|
||||||
self.bot.admin_pattern_regexes().clone(),
|
self.bot.admin_pattern_regexes().clone(),
|
||||||
trigger_event_info,
|
trigger_event_info,
|
||||||
thread_context.info.clone(),
|
interaction_context.thread_info.clone(),
|
||||||
);
|
)
|
||||||
|
.with_bot_display_name(bot_display_name);
|
||||||
|
|
||||||
let bot_display_name = self
|
|
||||||
.bot
|
|
||||||
.room_display_name_fetcher()
|
|
||||||
.own_display_name_in_room(message_context.room())
|
|
||||||
.await;
|
|
||||||
|
|
||||||
let bot_display_name = match bot_display_name {
|
|
||||||
Ok(value) => value,
|
|
||||||
Err(err) => {
|
|
||||||
tracing::warn!(
|
|
||||||
?err,
|
|
||||||
"Failed to fetch bot display name. Proceeding without it"
|
|
||||||
);
|
|
||||||
None
|
|
||||||
}
|
|
||||||
};
|
|
||||||
|
|
||||||
// The first event in the thread determines which handler processes the current event.
|
|
||||||
let controller_type = crate::controller::determine_controller(
|
let controller_type = crate::controller::determine_controller(
|
||||||
self.bot.command_prefix(),
|
self.bot.command_prefix(),
|
||||||
&thread_context.first_message,
|
&interaction_context.trigger,
|
||||||
&message_context,
|
&message_context,
|
||||||
self.bot.user_id(),
|
|
||||||
&bot_display_name,
|
|
||||||
);
|
);
|
||||||
|
|
||||||
tracing::info!(?controller_type, "Determined controller");
|
tracing::info!(?controller_type, "Determined controller");
|
||||||
@@ -310,7 +296,7 @@ impl Messaging {
|
|||||||
let _ = room
|
let _ = room
|
||||||
.send_single_receipt(
|
.send_single_receipt(
|
||||||
ReceiptType::Read,
|
ReceiptType::Read,
|
||||||
thread_context.info.clone().into(),
|
interaction_context.thread_info.clone().into(),
|
||||||
event.event_id.clone(),
|
event.event_id.clone(),
|
||||||
)
|
)
|
||||||
.await;
|
.await;
|
||||||
|
|||||||
@@ -18,13 +18,28 @@ use crate::entity::roomconfig::{
|
|||||||
use crate::entity::MessagePayload;
|
use crate::entity::MessagePayload;
|
||||||
use crate::strings;
|
use crate::strings;
|
||||||
use crate::utils::text_to_speech::create_transcribed_message_text;
|
use crate::utils::text_to_speech::create_transcribed_message_text;
|
||||||
use crate::{conversation::create_llm_conversation_for_matrix_thread, entity::MessageContext, Bot};
|
use crate::{
|
||||||
|
conversation::{
|
||||||
|
create_llm_conversation_for_matrix_reply_chain, create_llm_conversation_for_matrix_thread,
|
||||||
|
matrix::create_list_of_bot_user_prefixes_to_strip,
|
||||||
|
},
|
||||||
|
entity::MessageContext,
|
||||||
|
Bot,
|
||||||
|
};
|
||||||
|
|
||||||
#[derive(Debug, PartialEq)]
|
#[derive(Debug, PartialEq)]
|
||||||
pub enum ChatCompletionControllerType {
|
pub enum ChatCompletionControllerType {
|
||||||
ViaText { prefixes_to_strip: Vec<String> },
|
// Invoked via a command prefix (e.g. `!bai Hello!`)
|
||||||
|
TextCommand,
|
||||||
|
// Invoked via a mention (e.g. `@baibot Hello!`)
|
||||||
|
TextMention,
|
||||||
|
// Invoked via a direct message (e.g. `Hello!`)
|
||||||
|
TextDirect,
|
||||||
|
|
||||||
ViaAudio,
|
Audio,
|
||||||
|
|
||||||
|
ThreadMention,
|
||||||
|
ReplyMention,
|
||||||
}
|
}
|
||||||
|
|
||||||
struct TextToSpeechEligiblePayload {
|
struct TextToSpeechEligiblePayload {
|
||||||
@@ -125,7 +140,15 @@ pub async fn handle(
|
|||||||
None
|
None
|
||||||
};
|
};
|
||||||
|
|
||||||
let response_type = MessageResponseType::InThread(message_context.thread_info().clone());
|
let response_type = match controller_type {
|
||||||
|
// When we're triggered via a reply mention, we reply to the message that triggered us.
|
||||||
|
ChatCompletionControllerType::ReplyMention => {
|
||||||
|
MessageResponseType::Reply(message_context.thread_info().last_event_id.clone())
|
||||||
|
}
|
||||||
|
|
||||||
|
// In all other cases, we're dealing with a threaded conversation, so we reply in the thread.
|
||||||
|
_ => MessageResponseType::InThread(message_context.thread_info().clone()),
|
||||||
|
};
|
||||||
|
|
||||||
let text_to_speech_eligible_payload = handle_stage_text_generation(
|
let text_to_speech_eligible_payload = handle_stage_text_generation(
|
||||||
bot,
|
bot,
|
||||||
@@ -353,24 +376,64 @@ async fn handle_stage_text_generation(
|
|||||||
)
|
)
|
||||||
.await?;
|
.await?;
|
||||||
|
|
||||||
let prefixes_to_strip = match controller_type {
|
// We only strip text from the first message if we're invoked via a command prefix.
|
||||||
ChatCompletionControllerType::ViaText { prefixes_to_strip } => prefixes_to_strip.clone(),
|
// Otherwise, we do bot-user mentions stripping on all messages below.
|
||||||
ChatCompletionControllerType::ViaAudio => vec![],
|
let first_message_prefixes_to_strip = match controller_type {
|
||||||
|
ChatCompletionControllerType::TextCommand => vec![bot.command_prefix().to_owned()],
|
||||||
|
_ => vec![],
|
||||||
};
|
};
|
||||||
|
|
||||||
let params = MatrixMessageProcessingParams::new(
|
let bot_user_prefixes_to_strip = create_list_of_bot_user_prefixes_to_strip(
|
||||||
bot.user_id().as_str().to_owned(),
|
bot.user_id(),
|
||||||
message_context.combined_admin_and_user_regexes(),
|
message_context.bot_display_name(),
|
||||||
)
|
);
|
||||||
.with_first_message_stripped_prefixes(prefixes_to_strip);
|
|
||||||
|
|
||||||
let conversation = create_llm_conversation_for_matrix_thread(
|
let allowed_users = match controller_type {
|
||||||
|
// Regular chat completion only operates on messages from allowed users.
|
||||||
|
ChatCompletionControllerType::TextCommand
|
||||||
|
| ChatCompletionControllerType::TextMention
|
||||||
|
| ChatCompletionControllerType::TextDirect
|
||||||
|
| ChatCompletionControllerType::Audio => {
|
||||||
|
Some(message_context.combined_admin_and_user_regexes())
|
||||||
|
}
|
||||||
|
|
||||||
|
// When we're triggered via an explicit mention (thread or reply), we wish to operate against the mention's whole context
|
||||||
|
// (the whole thread or the whole reply chain upward of the message that triggered us).
|
||||||
|
//
|
||||||
|
// This is to allow admins and users to trigger text-generation for other users' messages.
|
||||||
|
// When we're dragged into a conversation by a known (to us) user, we'd like to process all messages in the conversation,
|
||||||
|
// not just those from allowed users.
|
||||||
|
ChatCompletionControllerType::ThreadMention
|
||||||
|
| ChatCompletionControllerType::ReplyMention => None,
|
||||||
|
};
|
||||||
|
|
||||||
|
let params = MatrixMessageProcessingParams::new(bot.user_id().to_owned(), allowed_users)
|
||||||
|
.with_first_message_prefixes_to_strip(first_message_prefixes_to_strip)
|
||||||
|
.with_bot_user_prefixes_to_strip(bot_user_prefixes_to_strip);
|
||||||
|
|
||||||
|
let conversation = match controller_type {
|
||||||
|
// When we're triggered via a reply mention, the context is the whole reply chain upward of the message that triggered us.
|
||||||
|
ChatCompletionControllerType::ReplyMention => {
|
||||||
|
create_llm_conversation_for_matrix_reply_chain(
|
||||||
|
&bot.room_event_fetcher().clone(),
|
||||||
|
message_context.room(),
|
||||||
|
message_context.thread_info().last_event_id.clone(),
|
||||||
|
¶ms,
|
||||||
|
)
|
||||||
|
.await
|
||||||
|
}
|
||||||
|
|
||||||
|
// Everything else is happening in a thread, so the context is the whole thread.
|
||||||
|
_ => {
|
||||||
|
create_llm_conversation_for_matrix_thread(
|
||||||
matrix_link.clone(),
|
matrix_link.clone(),
|
||||||
message_context.room(),
|
message_context.room(),
|
||||||
message_context.thread_info().root_event_id.clone(),
|
message_context.thread_info().root_event_id.clone(),
|
||||||
¶ms,
|
¶ms,
|
||||||
)
|
)
|
||||||
.await;
|
.await
|
||||||
|
}
|
||||||
|
};
|
||||||
|
|
||||||
let conversation = match conversation {
|
let conversation = match conversation {
|
||||||
Ok(conversation) => conversation,
|
Ok(conversation) => conversation,
|
||||||
@@ -415,6 +478,7 @@ async fn handle_stage_text_generation(
|
|||||||
.text_generation_model_id()
|
.text_generation_model_id()
|
||||||
.unwrap_or("unknown-model".to_owned()),
|
.unwrap_or("unknown-model".to_owned()),
|
||||||
chrono::Utc::now(),
|
chrono::Utc::now(),
|
||||||
|
conversation.start_time(),
|
||||||
);
|
);
|
||||||
|
|
||||||
let params = TextGenerationParams {
|
let params = TextGenerationParams {
|
||||||
@@ -548,16 +612,53 @@ async fn handle_stage_speech_to_text_actual_transcribing(
|
|||||||
.instrument(span)
|
.instrument(span)
|
||||||
.await?;
|
.await?;
|
||||||
|
|
||||||
let transcribed_text = create_transcribed_message_text(&speech_to_text_result.text);
|
// Only use the `> 🦻 Transcribed text` format if we're posting in a thread.
|
||||||
|
//
|
||||||
|
// If we're dealing with a regular reply (which would be the case in "Transcribe-only mode" = speech-to-text/flow-type=only_transcribe),
|
||||||
|
// we don't want to use the `> 🦻 Transcribed text` format for 2 reasons:
|
||||||
|
//
|
||||||
|
// 1. This kind of blockquote-formatting can be confused by clients for a fallback-for-rich-replies
|
||||||
|
// (see https://spec.matrix.org/v1.11/client-server-api/#fallbacks-for-rich-replies).
|
||||||
|
// It makes certain clients render our messages incorrectly.
|
||||||
|
//
|
||||||
|
// 2. Transcribe-only mode is typically used for memos. Sticking to a plain-text format
|
||||||
|
// allows people to copy-paste the text or forward it to another room more easily (without having to strip formatting, etc.)
|
||||||
|
//
|
||||||
|
// When sending a bare reply, we'd better annotate the message with a 🦻 reaction instead,
|
||||||
|
// to make it clear to users that it's a transcription.
|
||||||
|
//
|
||||||
|
// Regardless of how we post this message, it will be posted as a notice,
|
||||||
|
// which can indicate to the bot (for potential future text-generation purposes) that this message is not a bot message.
|
||||||
|
let (transcribed_text, annotate_message_with_reaction) =
|
||||||
|
if let MessageResponseType::InThread(_) = response_type {
|
||||||
|
(
|
||||||
|
create_transcribed_message_text(&speech_to_text_result.text),
|
||||||
|
false,
|
||||||
|
)
|
||||||
|
} else {
|
||||||
|
(speech_to_text_result.text, true)
|
||||||
|
};
|
||||||
|
|
||||||
let result = bot
|
let result = bot
|
||||||
.messaging()
|
.messaging()
|
||||||
.send_notice_markdown_no_fail(message_context.room(), transcribed_text, response_type)
|
.send_notice_markdown_no_fail(message_context.room(), transcribed_text, response_type)
|
||||||
.await;
|
.await;
|
||||||
|
|
||||||
result
|
let event_id = result
|
||||||
.map(|result| result.event_id)
|
.map(|result| result.event_id)
|
||||||
.ok_or_else(|| anyhow::anyhow!("Failed to send transcribed text"))
|
.ok_or_else(|| anyhow::anyhow!("Failed to send transcribed text"))?;
|
||||||
|
|
||||||
|
if annotate_message_with_reaction {
|
||||||
|
bot.reacting()
|
||||||
|
.react_no_fail(
|
||||||
|
message_context.room(),
|
||||||
|
event_id.clone(),
|
||||||
|
AgentPurpose::SpeechToText.emoji().to_owned(),
|
||||||
|
)
|
||||||
|
.await;
|
||||||
|
}
|
||||||
|
|
||||||
|
Ok(event_id)
|
||||||
}
|
}
|
||||||
|
|
||||||
async fn send_tts_offer_for_message(
|
async fn send_tts_offer_for_message(
|
||||||
|
|||||||
@@ -1,13 +1,11 @@
|
|||||||
#[cfg(test)]
|
#[cfg(test)]
|
||||||
mod tests;
|
mod tests;
|
||||||
|
|
||||||
use mxlink::matrix_sdk::ruma::OwnedUserId;
|
|
||||||
|
|
||||||
use super::chat_completion::ChatCompletionControllerType;
|
use super::chat_completion::ChatCompletionControllerType;
|
||||||
use crate::{
|
use crate::{
|
||||||
entity::{
|
entity::{
|
||||||
roomconfig::TextGenerationPrefixRequirementType, MessageContext, MessagePayload,
|
roomconfig::TextGenerationPrefixRequirementType, InteractionTrigger, MessageContext,
|
||||||
ThreadContextFirstMessage,
|
MessagePayload,
|
||||||
},
|
},
|
||||||
strings,
|
strings,
|
||||||
};
|
};
|
||||||
@@ -16,12 +14,16 @@ use super::ControllerType;
|
|||||||
|
|
||||||
pub fn determine_controller(
|
pub fn determine_controller(
|
||||||
command_prefix: &str,
|
command_prefix: &str,
|
||||||
first_thread_message: &ThreadContextFirstMessage,
|
first_thread_message: &InteractionTrigger,
|
||||||
message_context: &MessageContext,
|
message_context: &MessageContext,
|
||||||
bot_user_id: &OwnedUserId,
|
|
||||||
bot_display_name: &Option<String>,
|
|
||||||
) -> ControllerType {
|
) -> ControllerType {
|
||||||
match &first_thread_message.payload {
|
match &first_thread_message.payload {
|
||||||
|
MessagePayload::SynthethicChatCompletionTriggerInThread => {
|
||||||
|
ControllerType::ChatCompletion(ChatCompletionControllerType::ThreadMention)
|
||||||
|
}
|
||||||
|
MessagePayload::SynthethicChatCompletionTriggerForReply => {
|
||||||
|
ControllerType::ChatCompletion(ChatCompletionControllerType::ReplyMention)
|
||||||
|
}
|
||||||
MessagePayload::Text(text_message_content) => {
|
MessagePayload::Text(text_message_content) => {
|
||||||
let prefix_requirement_type = message_context
|
let prefix_requirement_type = message_context
|
||||||
.room_config_context()
|
.room_config_context()
|
||||||
@@ -32,8 +34,6 @@ pub fn determine_controller(
|
|||||||
&text_message_content.body,
|
&text_message_content.body,
|
||||||
prefix_requirement_type,
|
prefix_requirement_type,
|
||||||
first_thread_message.is_mentioning_bot,
|
first_thread_message.is_mentioning_bot,
|
||||||
bot_user_id,
|
|
||||||
bot_display_name,
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
MessagePayload::Encrypted(thread_info) => {
|
MessagePayload::Encrypted(thread_info) => {
|
||||||
@@ -47,7 +47,7 @@ pub fn determine_controller(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
MessagePayload::Audio(_) => {
|
MessagePayload::Audio(_) => {
|
||||||
ControllerType::ChatCompletion(ChatCompletionControllerType::ViaAudio)
|
ControllerType::ChatCompletion(ChatCompletionControllerType::Audio)
|
||||||
}
|
}
|
||||||
MessagePayload::Reaction { .. } => {
|
MessagePayload::Reaction { .. } => {
|
||||||
panic!("Handling reaction as first message in thread does not make sense")
|
panic!("Handling reaction as first message in thread does not make sense")
|
||||||
@@ -60,8 +60,6 @@ fn determine_text_controller(
|
|||||||
text: &str,
|
text: &str,
|
||||||
room_text_generation_prefix_requirement_type: TextGenerationPrefixRequirementType,
|
room_text_generation_prefix_requirement_type: TextGenerationPrefixRequirementType,
|
||||||
is_mentioning_bot: bool,
|
is_mentioning_bot: bool,
|
||||||
bot_user_id: &OwnedUserId,
|
|
||||||
bot_display_name: &Option<String>,
|
|
||||||
) -> ControllerType {
|
) -> ControllerType {
|
||||||
let text = text.trim();
|
let text = text.trim();
|
||||||
|
|
||||||
@@ -102,53 +100,26 @@ fn determine_text_controller(
|
|||||||
// Otherwise, it depends on the prefix requirement for text generation - it may be routed for chat completion or ignored.
|
// Otherwise, it depends on the prefix requirement for text generation - it may be routed for chat completion or ignored.
|
||||||
|
|
||||||
if is_mentioning_bot {
|
if is_mentioning_bot {
|
||||||
// Different clients do mentions differently.
|
return ControllerType::ChatCompletion(ChatCompletionControllerType::TextMention);
|
||||||
// The body text containing the mention usually contains one of:
|
|
||||||
// - the full user ID (includes a @ prefix by default)
|
|
||||||
// - the localpart (with a @ prefix)
|
|
||||||
// - the localpart (without a @ prefix)
|
|
||||||
// - the display name (with a @ prefix)
|
|
||||||
// - the display name (without a @ prefix)
|
|
||||||
//
|
|
||||||
// Some add a `: ` suffix after the mention.
|
|
||||||
//
|
|
||||||
// There's no guarantee that the mention is at the start even.
|
|
||||||
// It being there is most common and we try to strip it from there
|
|
||||||
// as best as we can.
|
|
||||||
let bot_user_id_localpart = bot_user_id.localpart();
|
|
||||||
|
|
||||||
let mut prefixes_to_strip = vec![
|
|
||||||
bot_user_id.as_str().to_owned(),
|
|
||||||
format!("@{}", bot_user_id_localpart),
|
|
||||||
bot_user_id_localpart.to_owned(),
|
|
||||||
];
|
|
||||||
|
|
||||||
if let Some(bot_display_name) = bot_display_name {
|
|
||||||
prefixes_to_strip.push(format!("@{}", bot_display_name));
|
|
||||||
prefixes_to_strip.push(bot_display_name.to_owned());
|
|
||||||
}
|
}
|
||||||
|
|
||||||
prefixes_to_strip.push(":".to_owned());
|
// Regardless of what the prefix requirement is, if we encounter a command prefix, we'll consider it a chat completion via command prefix invokation.
|
||||||
|
// This is to correctly indicate to the chat completion controller that a command prefix was used,
|
||||||
return ControllerType::ChatCompletion(ChatCompletionControllerType::ViaText {
|
// so that it can be stripped from the beginning of the message.
|
||||||
prefixes_to_strip,
|
if text.starts_with(command_prefix) {
|
||||||
});
|
return ControllerType::ChatCompletion(ChatCompletionControllerType::TextCommand);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// We're dealing with a regular message that does not start with a command prefix.
|
||||||
|
|
||||||
match room_text_generation_prefix_requirement_type {
|
match room_text_generation_prefix_requirement_type {
|
||||||
TextGenerationPrefixRequirementType::CommandPrefix => {
|
TextGenerationPrefixRequirementType::CommandPrefix => {
|
||||||
if text.starts_with(command_prefix) {
|
// A prefix is required, but we've already checked (above) that the message does not start with a command prefix.
|
||||||
ControllerType::ChatCompletion(ChatCompletionControllerType::ViaText {
|
// It's to be ignored.
|
||||||
prefixes_to_strip: vec![command_prefix.to_owned()],
|
|
||||||
})
|
|
||||||
} else {
|
|
||||||
ControllerType::Ignore
|
ControllerType::Ignore
|
||||||
}
|
}
|
||||||
}
|
|
||||||
TextGenerationPrefixRequirementType::No => {
|
TextGenerationPrefixRequirementType::No => {
|
||||||
ControllerType::ChatCompletion(ChatCompletionControllerType::ViaText {
|
ControllerType::ChatCompletion(ChatCompletionControllerType::TextDirect)
|
||||||
prefixes_to_strip: vec![],
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -4,9 +4,6 @@ fn determine_text_controller() {
|
|||||||
use super::ControllerType;
|
use super::ControllerType;
|
||||||
use crate::controller;
|
use crate::controller;
|
||||||
|
|
||||||
let bot_user_id = mxlink::matrix_sdk::ruma::owned_user_id!("@bot:example.com");
|
|
||||||
let bot_display_name = "Bot";
|
|
||||||
|
|
||||||
let command_prefix = "!bai";
|
let command_prefix = "!bai";
|
||||||
|
|
||||||
struct TestCase {
|
struct TestCase {
|
||||||
@@ -44,9 +41,7 @@ fn determine_text_controller() {
|
|||||||
is_mentioning_bot: false,
|
is_mentioning_bot: false,
|
||||||
room_text_generation_prefix_requirement_type:
|
room_text_generation_prefix_requirement_type:
|
||||||
super::TextGenerationPrefixRequirementType::No,
|
super::TextGenerationPrefixRequirementType::No,
|
||||||
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::ViaText {
|
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::TextCommand),
|
||||||
prefixes_to_strip: vec![],
|
|
||||||
}),
|
|
||||||
},
|
},
|
||||||
TestCase {
|
TestCase {
|
||||||
name: "Access top-level",
|
name: "Access top-level",
|
||||||
@@ -110,9 +105,7 @@ fn determine_text_controller() {
|
|||||||
is_mentioning_bot: false,
|
is_mentioning_bot: false,
|
||||||
room_text_generation_prefix_requirement_type:
|
room_text_generation_prefix_requirement_type:
|
||||||
super::TextGenerationPrefixRequirementType::No,
|
super::TextGenerationPrefixRequirementType::No,
|
||||||
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::ViaText {
|
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::TextDirect),
|
||||||
prefixes_to_strip: vec![],
|
|
||||||
}),
|
|
||||||
},
|
},
|
||||||
TestCase {
|
TestCase {
|
||||||
name: "Regular text is ignored when prefix is required",
|
name: "Regular text is ignored when prefix is required",
|
||||||
@@ -128,9 +121,7 @@ fn determine_text_controller() {
|
|||||||
is_mentioning_bot: false,
|
is_mentioning_bot: false,
|
||||||
room_text_generation_prefix_requirement_type:
|
room_text_generation_prefix_requirement_type:
|
||||||
super::TextGenerationPrefixRequirementType::CommandPrefix,
|
super::TextGenerationPrefixRequirementType::CommandPrefix,
|
||||||
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::ViaText {
|
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::TextCommand),
|
||||||
prefixes_to_strip: vec!["!bai".to_owned()],
|
|
||||||
}),
|
|
||||||
},
|
},
|
||||||
TestCase {
|
TestCase {
|
||||||
name: "Command-prefixed text triggers completion even when prefix is not required",
|
name: "Command-prefixed text triggers completion even when prefix is not required",
|
||||||
@@ -138,58 +129,35 @@ fn determine_text_controller() {
|
|||||||
is_mentioning_bot: false,
|
is_mentioning_bot: false,
|
||||||
room_text_generation_prefix_requirement_type:
|
room_text_generation_prefix_requirement_type:
|
||||||
super::TextGenerationPrefixRequirementType::No,
|
super::TextGenerationPrefixRequirementType::No,
|
||||||
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::ViaText {
|
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::TextCommand),
|
||||||
prefixes_to_strip: vec![],
|
|
||||||
}),
|
|
||||||
},
|
},
|
||||||
TestCase {
|
TestCase {
|
||||||
name: "Regular message with bot mention triggers completion stripping bot id and display name (no prefix requirement)",
|
name: "Regular message with bot mention triggers completion (no prefix requirement)",
|
||||||
input: "Regular text goes here",
|
input: "Regular text goes here",
|
||||||
is_mentioning_bot: true,
|
is_mentioning_bot: true,
|
||||||
room_text_generation_prefix_requirement_type:
|
room_text_generation_prefix_requirement_type:
|
||||||
super::TextGenerationPrefixRequirementType::No,
|
super::TextGenerationPrefixRequirementType::No,
|
||||||
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::ViaText {
|
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::TextMention),
|
||||||
prefixes_to_strip: vec![
|
|
||||||
"@bot:example.com".to_owned(),
|
|
||||||
"@bot".to_owned(),
|
|
||||||
"bot".to_owned(),
|
|
||||||
"@Bot".to_owned(),
|
|
||||||
"Bot".to_owned(),
|
|
||||||
":".to_owned(),
|
|
||||||
],
|
|
||||||
}),
|
|
||||||
},
|
},
|
||||||
// This test case is the same as the one above, just with a different prefix requirement.
|
// This test case is the same as the one above, just with a different prefix requirement setting.
|
||||||
// We expect the same result.
|
// We expect the same result.
|
||||||
TestCase {
|
TestCase {
|
||||||
name: "Regular message with bot mention triggers completion stripping bot id and display name (command_prefix requirement)",
|
name:
|
||||||
|
"Regular message with bot mention triggers completion (command prefix requirement)",
|
||||||
input: "Regular text goes here",
|
input: "Regular text goes here",
|
||||||
is_mentioning_bot: true,
|
is_mentioning_bot: true,
|
||||||
room_text_generation_prefix_requirement_type:
|
room_text_generation_prefix_requirement_type:
|
||||||
super::TextGenerationPrefixRequirementType::CommandPrefix,
|
super::TextGenerationPrefixRequirementType::CommandPrefix,
|
||||||
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::ViaText {
|
expected: ControllerType::ChatCompletion(ChatCompletionControllerType::TextMention),
|
||||||
prefixes_to_strip: vec![
|
|
||||||
"@bot:example.com".to_owned(),
|
|
||||||
"@bot".to_owned(),
|
|
||||||
"bot".to_owned(),
|
|
||||||
"@Bot".to_owned(),
|
|
||||||
"Bot".to_owned(),
|
|
||||||
":".to_owned(),
|
|
||||||
],
|
|
||||||
}),
|
|
||||||
},
|
},
|
||||||
];
|
];
|
||||||
|
|
||||||
for test_case in test_cases {
|
for test_case in test_cases {
|
||||||
let bot_display_name = Some(bot_display_name.to_owned());
|
|
||||||
|
|
||||||
let result = super::determine_text_controller(
|
let result = super::determine_text_controller(
|
||||||
command_prefix,
|
command_prefix,
|
||||||
test_case.input,
|
test_case.input,
|
||||||
test_case.room_text_generation_prefix_requirement_type,
|
test_case.room_text_generation_prefix_requirement_type,
|
||||||
test_case.is_mentioning_bot,
|
test_case.is_mentioning_bot,
|
||||||
&bot_user_id,
|
|
||||||
&bot_display_name,
|
|
||||||
);
|
);
|
||||||
assert_eq!(result, test_case.expected, "Test case: {}", test_case.name);
|
assert_eq!(result, test_case.expected, "Test case: {}", test_case.name);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -35,8 +35,8 @@ pub async fn handle_image(
|
|||||||
};
|
};
|
||||||
|
|
||||||
let params = MatrixMessageProcessingParams::new(
|
let params = MatrixMessageProcessingParams::new(
|
||||||
bot.user_id().as_str().to_owned(),
|
bot.user_id().to_owned(),
|
||||||
message_context.combined_admin_and_user_regexes(),
|
Some(message_context.combined_admin_and_user_regexes()),
|
||||||
);
|
);
|
||||||
|
|
||||||
let conversation = create_llm_conversation_for_matrix_thread(
|
let conversation = create_llm_conversation_for_matrix_thread(
|
||||||
|
|||||||
@@ -46,6 +46,8 @@ mod tests {
|
|||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn test_build_prompt() {
|
fn test_build_prompt() {
|
||||||
|
let timestamp = chrono::Utc::now();
|
||||||
|
|
||||||
let test_cases = vec![
|
let test_cases = vec![
|
||||||
// Simple case
|
// Simple case
|
||||||
TestCase {
|
TestCase {
|
||||||
@@ -59,6 +61,7 @@ mod tests {
|
|||||||
messages: vec![Message {
|
messages: vec![Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: "Must be blue".to_owned(),
|
message_text: "Must be blue".to_owned(),
|
||||||
|
timestamp,
|
||||||
}],
|
}],
|
||||||
expected_prompt: "Generate a picture of a dog\nOther criteria:\n- Must be blue",
|
expected_prompt: "Generate a picture of a dog\nOther criteria:\n- Must be blue",
|
||||||
},
|
},
|
||||||
@@ -68,14 +71,17 @@ mod tests {
|
|||||||
messages: vec![Message {
|
messages: vec![Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: "Must be blue".to_owned(),
|
message_text: "Must be blue".to_owned(),
|
||||||
|
timestamp,
|
||||||
},
|
},
|
||||||
Message {
|
Message {
|
||||||
author: Author::Assistant,
|
author: Author::Assistant,
|
||||||
message_text: "Whatever".to_owned(),
|
message_text: "Whatever".to_owned(),
|
||||||
|
timestamp,
|
||||||
},
|
},
|
||||||
Message {
|
Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: "Must be 3-legged.\nMust be flying.".to_owned(),
|
message_text: "Must be 3-legged.\nMust be flying.".to_owned(),
|
||||||
|
timestamp,
|
||||||
}],
|
}],
|
||||||
expected_prompt: "Generate a picture of an elephant\nOther criteria:\n- Must be blue\n- Must be 3-legged.. Must be flying.",
|
expected_prompt: "Generate a picture of an elephant\nOther criteria:\n- Must be blue\n- Must be 3-legged.. Must be flying.",
|
||||||
},
|
},
|
||||||
@@ -85,18 +91,22 @@ mod tests {
|
|||||||
messages: vec![Message {
|
messages: vec![Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: "Must be blue".to_owned(),
|
message_text: "Must be blue".to_owned(),
|
||||||
|
timestamp,
|
||||||
},
|
},
|
||||||
Message {
|
Message {
|
||||||
author: Author::Assistant,
|
author: Author::Assistant,
|
||||||
message_text: "Whatever".to_owned(),
|
message_text: "Whatever".to_owned(),
|
||||||
|
timestamp,
|
||||||
},
|
},
|
||||||
Message {
|
Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: "Again".to_owned(),
|
message_text: "Again".to_owned(),
|
||||||
|
timestamp,
|
||||||
},
|
},
|
||||||
Message {
|
Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: "again".to_owned(),
|
message_text: "again".to_owned(),
|
||||||
|
timestamp,
|
||||||
}],
|
}],
|
||||||
expected_prompt: "Generate a picture of a grizzly bear\nOther criteria:\n- Must be blue",
|
expected_prompt: "Generate a picture of a grizzly bear\nOther criteria:\n- Must be blue",
|
||||||
},
|
},
|
||||||
|
|||||||
@@ -1,3 +1,5 @@
|
|||||||
|
use chrono::{DateTime, Utc};
|
||||||
|
|
||||||
#[derive(Debug, Clone, PartialEq)]
|
#[derive(Debug, Clone, PartialEq)]
|
||||||
pub enum Author {
|
pub enum Author {
|
||||||
Prompt,
|
Prompt,
|
||||||
@@ -9,6 +11,7 @@ pub enum Author {
|
|||||||
pub struct Message {
|
pub struct Message {
|
||||||
pub author: Author,
|
pub author: Author,
|
||||||
pub message_text: String,
|
pub message_text: String,
|
||||||
|
pub timestamp: DateTime<Utc>,
|
||||||
}
|
}
|
||||||
|
|
||||||
#[derive(Debug)]
|
#[derive(Debug)]
|
||||||
@@ -52,39 +55,59 @@ impl Conversation {
|
|||||||
messages: new_messages,
|
messages: new_messages,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
pub fn start_time(&self) -> Option<DateTime<Utc>> {
|
||||||
|
self.messages.first().map(|message| message.timestamp)
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
#[cfg(test)]
|
#[cfg(test)]
|
||||||
mod tests {
|
mod tests {
|
||||||
use super::*;
|
use super::*;
|
||||||
|
use chrono::{TimeZone, Utc};
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn combine_consecutive_messages() {
|
fn combine_consecutive_messages() {
|
||||||
|
let timestamp_1 = Utc.with_ymd_and_hms(2024, 9, 20, 18, 34, 15).unwrap();
|
||||||
|
|
||||||
|
let timestamp_2 = Utc.with_ymd_and_hms(2024, 9, 21, 18, 34, 15).unwrap();
|
||||||
|
|
||||||
|
let timestamp_3 = Utc.with_ymd_and_hms(2024, 9, 22, 18, 34, 15).unwrap();
|
||||||
|
|
||||||
let conversation = Conversation {
|
let conversation = Conversation {
|
||||||
messages: vec![
|
messages: vec![
|
||||||
|
// User's turn
|
||||||
Message {
|
Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: "Hello".to_string(),
|
message_text: "Hello".to_string(),
|
||||||
|
timestamp: timestamp_1,
|
||||||
},
|
},
|
||||||
Message {
|
Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: "How are you?".to_string(),
|
message_text: "How are you?".to_string(),
|
||||||
|
timestamp: timestamp_2,
|
||||||
},
|
},
|
||||||
Message {
|
Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: "I'm OK, btw.".to_string(),
|
message_text: "I'm OK, btw.".to_string(),
|
||||||
|
timestamp: timestamp_3,
|
||||||
},
|
},
|
||||||
|
// Assistant's turn
|
||||||
Message {
|
Message {
|
||||||
author: Author::Assistant,
|
author: Author::Assistant,
|
||||||
message_text: "Hi there!".to_string(),
|
message_text: "Hi there!".to_string(),
|
||||||
|
timestamp: timestamp_2,
|
||||||
},
|
},
|
||||||
Message {
|
Message {
|
||||||
author: Author::Assistant,
|
author: Author::Assistant,
|
||||||
message_text: "I'm doing well, thank you.".to_string(),
|
message_text: "I'm doing well, thank you.".to_string(),
|
||||||
|
timestamp: timestamp_3,
|
||||||
},
|
},
|
||||||
|
// User's turn
|
||||||
Message {
|
Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: "That's great!".to_string(),
|
message_text: "That's great!".to_string(),
|
||||||
|
timestamp: timestamp_3,
|
||||||
},
|
},
|
||||||
],
|
],
|
||||||
};
|
};
|
||||||
@@ -97,12 +120,17 @@ mod tests {
|
|||||||
conversation.messages[0].message_text,
|
conversation.messages[0].message_text,
|
||||||
"Hello\nHow are you?\nI'm OK, btw."
|
"Hello\nHow are you?\nI'm OK, btw."
|
||||||
);
|
);
|
||||||
|
assert_eq!(conversation.messages[0].timestamp, timestamp_1);
|
||||||
|
|
||||||
assert_eq!(conversation.messages[1].author, Author::Assistant);
|
assert_eq!(conversation.messages[1].author, Author::Assistant);
|
||||||
assert_eq!(
|
assert_eq!(
|
||||||
conversation.messages[1].message_text,
|
conversation.messages[1].message_text,
|
||||||
"Hi there!\nI'm doing well, thank you."
|
"Hi there!\nI'm doing well, thank you."
|
||||||
);
|
);
|
||||||
|
assert_eq!(conversation.messages[1].timestamp, timestamp_2);
|
||||||
|
|
||||||
assert_eq!(conversation.messages[2].author, Author::User);
|
assert_eq!(conversation.messages[2].author, Author::User);
|
||||||
assert_eq!(conversation.messages[2].message_text, "That's great!");
|
assert_eq!(conversation.messages[2].message_text, "That's great!");
|
||||||
|
assert_eq!(conversation.messages[2].timestamp, timestamp_3);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,3 +1,5 @@
|
|||||||
|
use mxlink::matrix_sdk::ruma::OwnedUserId;
|
||||||
|
|
||||||
use crate::utils::status::create_error_message_text;
|
use crate::utils::status::create_error_message_text;
|
||||||
use crate::utils::text_to_speech::create_transcribed_message_text;
|
use crate::utils::text_to_speech::create_transcribed_message_text;
|
||||||
|
|
||||||
@@ -5,15 +7,18 @@ use super::*;
|
|||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn test_messages_by_the_bot_are_identified_correctly() {
|
fn test_messages_by_the_bot_are_identified_correctly() {
|
||||||
let bot_user_id = "@bot:example.com";
|
let bot_user_id =
|
||||||
|
OwnedUserId::try_from("@bot:example.com").expect("Failed to parse bot user ID");
|
||||||
|
|
||||||
let matrix_message = super::super::matrix::MatrixMessage {
|
let matrix_message = super::super::matrix::MatrixMessage {
|
||||||
sender_id: bot_user_id.to_owned(),
|
sender_id: bot_user_id.to_owned(),
|
||||||
message_type: super::super::matrix::MatrixMessageType::Text,
|
message_type: super::super::matrix::MatrixMessageType::Text,
|
||||||
message_text: "Hello!".to_owned(),
|
message_text: "Hello!".to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
|
|
||||||
let llm_message = convert_matrix_message_to_llm_message(&matrix_message, bot_user_id).unwrap();
|
let llm_message = convert_matrix_message_to_llm_message(&matrix_message, &bot_user_id).unwrap();
|
||||||
|
|
||||||
assert_eq!(llm_message.author, Author::Assistant);
|
assert_eq!(llm_message.author, Author::Assistant);
|
||||||
assert_eq!(llm_message.message_text, "Hello!");
|
assert_eq!(llm_message.message_text, "Hello!");
|
||||||
@@ -22,7 +27,8 @@ fn test_messages_by_the_bot_are_identified_correctly() {
|
|||||||
#[test]
|
#[test]
|
||||||
fn test_notice_messages_by_bot_with_speech_to_text_prefix_are_cleaned_up_and_considered_sent_by_user(
|
fn test_notice_messages_by_bot_with_speech_to_text_prefix_are_cleaned_up_and_considered_sent_by_user(
|
||||||
) {
|
) {
|
||||||
let bot_user_id = "@bot:example.com";
|
let bot_user_id =
|
||||||
|
OwnedUserId::try_from("@bot:example.com").expect("Failed to parse bot user ID");
|
||||||
|
|
||||||
let source_message_text = "Hello!";
|
let source_message_text = "Hello!";
|
||||||
let message_text = create_transcribed_message_text(source_message_text);
|
let message_text = create_transcribed_message_text(source_message_text);
|
||||||
@@ -33,9 +39,11 @@ fn test_notice_messages_by_bot_with_speech_to_text_prefix_are_cleaned_up_and_con
|
|||||||
sender_id: bot_user_id.to_owned(),
|
sender_id: bot_user_id.to_owned(),
|
||||||
message_type: super::super::matrix::MatrixMessageType::Notice,
|
message_type: super::super::matrix::MatrixMessageType::Notice,
|
||||||
message_text,
|
message_text,
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
|
|
||||||
let llm_message = convert_matrix_message_to_llm_message(&matrix_message, bot_user_id).unwrap();
|
let llm_message = convert_matrix_message_to_llm_message(&matrix_message, &bot_user_id).unwrap();
|
||||||
|
|
||||||
assert_eq!(llm_message.author, Author::User);
|
assert_eq!(llm_message.author, Author::User);
|
||||||
assert_eq!(llm_message.message_text, source_message_text);
|
assert_eq!(llm_message.message_text, source_message_text);
|
||||||
@@ -43,7 +51,8 @@ fn test_notice_messages_by_bot_with_speech_to_text_prefix_are_cleaned_up_and_con
|
|||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn test_notice_error_messages_by_bot_are_ignored() {
|
fn test_notice_error_messages_by_bot_are_ignored() {
|
||||||
let bot_user_id = "@bot:example.com";
|
let bot_user_id =
|
||||||
|
OwnedUserId::try_from("@bot:example.com").expect("Failed to parse bot user ID");
|
||||||
|
|
||||||
let source_message_text = "Some error happened";
|
let source_message_text = "Some error happened";
|
||||||
let message_text = create_error_message_text(source_message_text);
|
let message_text = create_error_message_text(source_message_text);
|
||||||
@@ -54,9 +63,11 @@ fn test_notice_error_messages_by_bot_are_ignored() {
|
|||||||
sender_id: bot_user_id.to_owned(),
|
sender_id: bot_user_id.to_owned(),
|
||||||
message_type: super::super::matrix::MatrixMessageType::Notice,
|
message_type: super::super::matrix::MatrixMessageType::Notice,
|
||||||
message_text,
|
message_text,
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
|
|
||||||
let llm_message = convert_matrix_message_to_llm_message(&matrix_message, bot_user_id);
|
let llm_message = convert_matrix_message_to_llm_message(&matrix_message, &bot_user_id);
|
||||||
|
|
||||||
assert!(llm_message.is_none());
|
assert!(llm_message.is_none());
|
||||||
}
|
}
|
||||||
@@ -68,7 +79,8 @@ fn test_other_notice_messages_by_the_bot_are_ignored() {
|
|||||||
// (except for speech-to-text-created transcriptions - see `test_notice_messages_by_bot_with_speech_to_text_prefix_are_cleaned_up_and_considered_sent_by_user()`).
|
// (except for speech-to-text-created transcriptions - see `test_notice_messages_by_bot_with_speech_to_text_prefix_are_cleaned_up_and_considered_sent_by_user()`).
|
||||||
// This test is to make sure that we don't accidentally start accepting other notice messages.
|
// This test is to make sure that we don't accidentally start accepting other notice messages.
|
||||||
|
|
||||||
let bot_user_id = "@bot:example.com";
|
let bot_user_id =
|
||||||
|
OwnedUserId::try_from("@bot:example.com").expect("Failed to parse bot user ID");
|
||||||
|
|
||||||
let message_text = "Something something";
|
let message_text = "Something something";
|
||||||
|
|
||||||
@@ -76,9 +88,11 @@ fn test_other_notice_messages_by_the_bot_are_ignored() {
|
|||||||
sender_id: bot_user_id.to_owned(),
|
sender_id: bot_user_id.to_owned(),
|
||||||
message_type: super::super::matrix::MatrixMessageType::Notice,
|
message_type: super::super::matrix::MatrixMessageType::Notice,
|
||||||
message_text: message_text.to_owned(),
|
message_text: message_text.to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
|
|
||||||
let llm_message = convert_matrix_message_to_llm_message(&matrix_message, bot_user_id);
|
let llm_message = convert_matrix_message_to_llm_message(&matrix_message, &bot_user_id);
|
||||||
|
|
||||||
assert!(llm_message.is_none());
|
assert!(llm_message.is_none());
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -16,7 +16,7 @@ pub fn shorten_messages_list_to_context_size(
|
|||||||
model: &str,
|
model: &str,
|
||||||
prompt_message: &Option<Message>,
|
prompt_message: &Option<Message>,
|
||||||
mut messages: Vec<Message>,
|
mut messages: Vec<Message>,
|
||||||
max_response_tokens: u32,
|
max_response_tokens: Option<u32>,
|
||||||
max_context_tokens: u32,
|
max_context_tokens: u32,
|
||||||
) -> Vec<Message> {
|
) -> Vec<Message> {
|
||||||
// Loading the tokenization data is an expensive process, so
|
// Loading the tokenization data is an expensive process, so
|
||||||
@@ -26,7 +26,8 @@ pub fn shorten_messages_list_to_context_size(
|
|||||||
// We want to retain the prompt in all cases, so we always count it first.
|
// We want to retain the prompt in all cases, so we always count it first.
|
||||||
// We also always reserve enough tokens for the maximum response we expect.
|
// We also always reserve enough tokens for the maximum response we expect.
|
||||||
let mut current_context_length: u32 = if let Some(prompt_message) = prompt_message {
|
let mut current_context_length: u32 = if let Some(prompt_message) = prompt_message {
|
||||||
calculate_token_size_for_message(&bpe, model, prompt_message) + max_response_tokens
|
calculate_token_size_for_message(&bpe, model, prompt_message)
|
||||||
|
+ max_response_tokens.unwrap_or(0)
|
||||||
} else {
|
} else {
|
||||||
0
|
0
|
||||||
};
|
};
|
||||||
@@ -85,6 +86,7 @@ pub mod test {
|
|||||||
let message = super::Message {
|
let message = super::Message {
|
||||||
author: super::Author::User,
|
author: super::Author::User,
|
||||||
message_text: "Hello there!".to_owned(),
|
message_text: "Hello there!".to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
|
|
||||||
let tokens = super::calculate_token_size_for_message(&bpe, model, &message);
|
let tokens = super::calculate_token_size_for_message(&bpe, model, &message);
|
||||||
@@ -98,11 +100,12 @@ pub mod test {
|
|||||||
|
|
||||||
let bpe = super::get_bpe_for_model(model);
|
let bpe = super::get_bpe_for_model(model);
|
||||||
|
|
||||||
let max_response_tokens: u32 = 5;
|
let max_response_tokens: Option<u32> = Some(5);
|
||||||
|
|
||||||
let prompt = super::Message {
|
let prompt = super::Message {
|
||||||
author: super::Author::Prompt,
|
author: super::Author::Prompt,
|
||||||
message_text: "You are a bot!".to_owned(),
|
message_text: "You are a bot!".to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
let prompt_length = 10;
|
let prompt_length = 10;
|
||||||
|
|
||||||
@@ -116,6 +119,7 @@ pub mod test {
|
|||||||
let first = super::Message {
|
let first = super::Message {
|
||||||
author: super::Author::User,
|
author: super::Author::User,
|
||||||
message_text: "Hello there!".to_owned(),
|
message_text: "Hello there!".to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
let first_length = 8;
|
let first_length = 8;
|
||||||
|
|
||||||
@@ -129,6 +133,7 @@ pub mod test {
|
|||||||
let second = super::Message {
|
let second = super::Message {
|
||||||
author: super::Author::Assistant,
|
author: super::Author::Assistant,
|
||||||
message_text: "Hello!".to_owned(),
|
message_text: "Hello!".to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
let second_length = 7;
|
let second_length = 7;
|
||||||
|
|
||||||
@@ -143,6 +148,7 @@ pub mod test {
|
|||||||
author: super::Author::User,
|
author: super::Author::User,
|
||||||
message_text: "This is the 3rd message in this conversation. It shall be preserved."
|
message_text: "This is the 3rd message in this conversation. It shall be preserved."
|
||||||
.to_owned(),
|
.to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
let third_length = 21;
|
let third_length = 21;
|
||||||
|
|
||||||
@@ -156,6 +162,7 @@ pub mod test {
|
|||||||
let forth = super::Message {
|
let forth = super::Message {
|
||||||
author: super::Author::Assistant,
|
author: super::Author::Assistant,
|
||||||
message_text: "This is yet another message that shall be preserved.".to_owned(),
|
message_text: "This is yet another message that shall be preserved.".to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
let forth_length = 15;
|
let forth_length = 15;
|
||||||
|
|
||||||
@@ -173,7 +180,7 @@ pub mod test {
|
|||||||
&Some(prompt),
|
&Some(prompt),
|
||||||
conversation_messages,
|
conversation_messages,
|
||||||
max_response_tokens,
|
max_response_tokens,
|
||||||
prompt_length + max_response_tokens + forth_length + third_length,
|
prompt_length + max_response_tokens.unwrap_or(0) + forth_length + third_length,
|
||||||
);
|
);
|
||||||
|
|
||||||
assert_eq!(2, new_conversation_messages.len());
|
assert_eq!(2, new_conversation_messages.len());
|
||||||
@@ -195,11 +202,12 @@ pub mod test {
|
|||||||
|
|
||||||
let bpe = super::get_bpe_for_model(model);
|
let bpe = super::get_bpe_for_model(model);
|
||||||
|
|
||||||
let max_response_tokens: u32 = 5;
|
let max_response_tokens: Option<u32> = Some(5);
|
||||||
|
|
||||||
let prompt = super::Message {
|
let prompt = super::Message {
|
||||||
author: super::Author::User,
|
author: super::Author::User,
|
||||||
message_text: "あなたはボットです。".to_owned(),
|
message_text: "あなたはボットです。".to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
let prompt_length = 14;
|
let prompt_length = 14;
|
||||||
|
|
||||||
@@ -213,6 +221,7 @@ pub mod test {
|
|||||||
let first = super::Message {
|
let first = super::Message {
|
||||||
author: super::Author::User,
|
author: super::Author::User,
|
||||||
message_text: "こんにちは!".to_owned(),
|
message_text: "こんにちは!".to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
let first_length = 7;
|
let first_length = 7;
|
||||||
|
|
||||||
@@ -226,6 +235,7 @@ pub mod test {
|
|||||||
let second = super::Message {
|
let second = super::Message {
|
||||||
author: super::Author::Assistant,
|
author: super::Author::Assistant,
|
||||||
message_text: "こんにちは。今日は元気ですか。".to_owned(),
|
message_text: "こんにちは。今日は元気ですか。".to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
let second_length = 15;
|
let second_length = 15;
|
||||||
|
|
||||||
@@ -239,6 +249,7 @@ pub mod test {
|
|||||||
let third = super::Message {
|
let third = super::Message {
|
||||||
author: super::Author::User,
|
author: super::Author::User,
|
||||||
message_text: "これは第3のメッセージなので、保存されます。".to_owned(),
|
message_text: "これは第3のメッセージなので、保存されます。".to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
let third_length = 22;
|
let third_length = 22;
|
||||||
|
|
||||||
@@ -252,6 +263,7 @@ pub mod test {
|
|||||||
let forth = super::Message {
|
let forth = super::Message {
|
||||||
author: super::Author::Assistant,
|
author: super::Author::Assistant,
|
||||||
message_text: "これはもう一つの保存されますメッセージです。".to_owned(),
|
message_text: "これはもう一つの保存されますメッセージです。".to_owned(),
|
||||||
|
timestamp: chrono::Utc::now(),
|
||||||
};
|
};
|
||||||
let forth_length = 21;
|
let forth_length = 21;
|
||||||
|
|
||||||
@@ -269,7 +281,7 @@ pub mod test {
|
|||||||
&Some(prompt),
|
&Some(prompt),
|
||||||
conversation_messages,
|
conversation_messages,
|
||||||
max_response_tokens,
|
max_response_tokens,
|
||||||
prompt_length + max_response_tokens + forth_length + third_length,
|
prompt_length + max_response_tokens.unwrap_or(0) + forth_length + third_length,
|
||||||
);
|
);
|
||||||
|
|
||||||
assert_eq!(2, new_conversation_messages.len());
|
assert_eq!(2, new_conversation_messages.len());
|
||||||
|
|||||||
@@ -1,12 +1,14 @@
|
|||||||
|
use matrix_sdk::ruma::OwnedUserId;
|
||||||
|
|
||||||
use super::{Author, Message};
|
use super::{Author, Message};
|
||||||
use crate::conversation::matrix::{MatrixMessage, MatrixMessageType};
|
use crate::conversation::matrix::{MatrixMessage, MatrixMessageType};
|
||||||
use crate::utils::text_to_speech as text_to_speech_utils;
|
use crate::utils::text_to_speech as text_to_speech_utils;
|
||||||
|
|
||||||
pub fn convert_matrix_message_to_llm_message(
|
pub fn convert_matrix_message_to_llm_message(
|
||||||
matrix_message: &MatrixMessage,
|
matrix_message: &MatrixMessage,
|
||||||
bot_user_id: &str,
|
bot_user_id: &OwnedUserId,
|
||||||
) -> Option<Message> {
|
) -> Option<Message> {
|
||||||
if matrix_message.sender_id == bot_user_id {
|
if matrix_message.sender_id == bot_user_id.as_str() {
|
||||||
return convert_bot_message(matrix_message);
|
return convert_bot_message(matrix_message);
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -15,28 +17,43 @@ pub fn convert_matrix_message_to_llm_message(
|
|||||||
|
|
||||||
fn convert_bot_message(matrix_message: &MatrixMessage) -> Option<Message> {
|
fn convert_bot_message(matrix_message: &MatrixMessage) -> Option<Message> {
|
||||||
match matrix_message.message_type {
|
match matrix_message.message_type {
|
||||||
MatrixMessageType::Text => convert_bot_text_message(&matrix_message.message_text),
|
MatrixMessageType::Text => {
|
||||||
MatrixMessageType::Notice => convert_bot_notice_message(&matrix_message.message_text),
|
convert_bot_text_message(&matrix_message.message_text, &matrix_message.timestamp)
|
||||||
|
}
|
||||||
|
MatrixMessageType::Notice => {
|
||||||
|
convert_bot_notice_message(&matrix_message.message_text, &matrix_message.timestamp)
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn convert_bot_text_message(text: &str) -> Option<Message> {
|
fn convert_bot_text_message(
|
||||||
|
text: &str,
|
||||||
|
timestamp: &chrono::DateTime<chrono::Utc>,
|
||||||
|
) -> Option<Message> {
|
||||||
Some(Message {
|
Some(Message {
|
||||||
author: Author::Assistant,
|
author: Author::Assistant,
|
||||||
message_text: text.to_owned(),
|
message_text: text.to_owned(),
|
||||||
|
timestamp: timestamp.to_owned(),
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
fn convert_bot_notice_message(text: &str) -> Option<Message> {
|
fn convert_bot_notice_message(
|
||||||
|
text: &str,
|
||||||
|
timestamp: &chrono::DateTime<chrono::Utc>,
|
||||||
|
) -> Option<Message> {
|
||||||
// Notice messages sent by the bot are usually transcriptions of previous messages sent by the user.
|
// Notice messages sent by the bot are usually transcriptions of previous messages sent by the user.
|
||||||
// Such transcriptions are prefixed with an emoji and blockquoted.
|
// Such transcriptions are prefixed with an emoji and blockquoted.
|
||||||
// If we find a notice that doesn't match this pattern, we skip it.
|
// If we find a notice that doesn't match this pattern, we skip it.
|
||||||
|
//
|
||||||
|
// It should be noted that transcriptions are sometimes posted as regular notice messages which do not include
|
||||||
|
// the `> 🦻` formatting. This function will not handle these properly.
|
||||||
|
|
||||||
if let Some(text) = text_to_speech_utils::parse_transcribed_message_text(text) {
|
if let Some(text) = text_to_speech_utils::parse_transcribed_message_text(text) {
|
||||||
// This is a transcription message. We remove the prefix and consider it as a message sent by the user.
|
// This is a transcription message. We remove the prefix and consider it as a message sent by the user.
|
||||||
return Some(Message {
|
return Some(Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: text.to_owned(),
|
message_text: text.to_owned(),
|
||||||
|
timestamp: timestamp.to_owned(),
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -47,5 +64,6 @@ fn convert_user_message(matrix_message: &MatrixMessage) -> Option<Message> {
|
|||||||
Some(Message {
|
Some(Message {
|
||||||
author: Author::User,
|
author: Author::User,
|
||||||
message_text: matrix_message.message_text.clone(),
|
message_text: matrix_message.message_text.clone(),
|
||||||
|
timestamp: matrix_message.timestamp.to_owned(),
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,10 +1,15 @@
|
|||||||
|
use chrono::{DateTime, Utc};
|
||||||
use regex::Regex;
|
use regex::Regex;
|
||||||
|
|
||||||
|
use mxlink::matrix_sdk::ruma::OwnedUserId;
|
||||||
|
|
||||||
#[derive(Clone)]
|
#[derive(Clone)]
|
||||||
pub struct MatrixMessage {
|
pub struct MatrixMessage {
|
||||||
pub sender_id: String,
|
pub sender_id: OwnedUserId,
|
||||||
pub message_type: MatrixMessageType,
|
pub message_type: MatrixMessageType,
|
||||||
pub message_text: String,
|
pub message_text: String,
|
||||||
|
pub mentioned_users: Vec<OwnedUserId>,
|
||||||
|
pub timestamp: DateTime<Utc>,
|
||||||
}
|
}
|
||||||
|
|
||||||
#[derive(Clone)]
|
#[derive(Clone)]
|
||||||
@@ -13,26 +18,42 @@ pub enum MatrixMessageType {
|
|||||||
Notice,
|
Notice,
|
||||||
}
|
}
|
||||||
|
|
||||||
#[derive(Default, Clone)]
|
#[derive(Clone)]
|
||||||
pub struct MatrixMessageProcessingParams {
|
pub struct MatrixMessageProcessingParams {
|
||||||
pub(crate) bot_user_id: String,
|
pub(crate) bot_user_id: OwnedUserId,
|
||||||
pub(crate) allowed_users: Vec<Regex>,
|
|
||||||
|
|
||||||
// If non-empty, these prefixes will be stripped when processing the message
|
/// The prefixes that will be stripped when processing the messages in the context (thread or reply chain),
|
||||||
pub(crate) first_message_stripped_prefixes: Vec<String>,
|
/// which are found to be mentioning the bot user (`bot_user_id`).
|
||||||
|
pub(crate) bot_user_prefixes_to_strip: Vec<String>,
|
||||||
|
|
||||||
|
/// The prefixes that will be stripped when processing the 1st message in the context (thread or reply chain).
|
||||||
|
pub(crate) first_message_prefixes_to_strip: Vec<String>,
|
||||||
|
|
||||||
|
/// A list of users whose messages are allowed.
|
||||||
|
/// If None, all messages are allowed.
|
||||||
|
/// If Some, only messages from the allowed users (and the bot itself, `bot_user_id`) are allowed.
|
||||||
|
pub(crate) allowed_users: Option<Vec<Regex>>,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl MatrixMessageProcessingParams {
|
impl MatrixMessageProcessingParams {
|
||||||
pub fn new(bot_user_id: String, allowed_users: Vec<Regex>) -> Self {
|
pub fn new(bot_user_id: OwnedUserId, allowed_users: Option<Vec<Regex>>) -> Self {
|
||||||
Self {
|
Self {
|
||||||
bot_user_id,
|
bot_user_id,
|
||||||
|
bot_user_prefixes_to_strip: vec![],
|
||||||
|
|
||||||
|
first_message_prefixes_to_strip: vec![],
|
||||||
|
|
||||||
allowed_users,
|
allowed_users,
|
||||||
..Default::default()
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
pub fn with_first_message_stripped_prefixes(mut self, value: Vec<String>) -> Self {
|
pub fn with_bot_user_prefixes_to_strip(mut self, value: Vec<String>) -> Self {
|
||||||
self.first_message_stripped_prefixes = value;
|
self.bot_user_prefixes_to_strip = value;
|
||||||
|
self
|
||||||
|
}
|
||||||
|
|
||||||
|
pub fn with_first_message_prefixes_to_strip(mut self, value: Vec<String>) -> Self {
|
||||||
|
self.first_message_prefixes_to_strip = value;
|
||||||
self
|
self
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -5,9 +5,12 @@ use std::sync::Arc;
|
|||||||
|
|
||||||
use mxlink::matrix_sdk::ruma::{OwnedEventId, OwnedUserId};
|
use mxlink::matrix_sdk::ruma::{OwnedEventId, OwnedUserId};
|
||||||
use mxlink::matrix_sdk::{
|
use mxlink::matrix_sdk::{
|
||||||
|
deserialized_responses::TimelineEvent,
|
||||||
ruma::events::{
|
ruma::events::{
|
||||||
|
relation::Thread,
|
||||||
room::message::{
|
room::message::{
|
||||||
MessageType, OriginalSyncRoomMessageEvent, Relation, RoomMessageEventContent,
|
sanitize::remove_plain_reply_fallback, MessageType, OriginalSyncRoomMessageEvent,
|
||||||
|
Relation, RoomMessageEventContent,
|
||||||
},
|
},
|
||||||
AnyMessageLikeEvent, AnyMessageLikeEventContent, AnyTimelineEvent, MessageLikeEvent,
|
AnyMessageLikeEvent, AnyMessageLikeEventContent, AnyTimelineEvent, MessageLikeEvent,
|
||||||
},
|
},
|
||||||
@@ -16,7 +19,12 @@ use mxlink::matrix_sdk::{
|
|||||||
use mxlink::{MatrixLink, ThreadGetMessagesParams, ThreadInfo};
|
use mxlink::{MatrixLink, ThreadGetMessagesParams, ThreadInfo};
|
||||||
|
|
||||||
use super::{MatrixMessage, MatrixMessageProcessingParams, MatrixMessageType, RoomEventFetcher};
|
use super::{MatrixMessage, MatrixMessageProcessingParams, MatrixMessageType, RoomEventFetcher};
|
||||||
use crate::entity::{MessagePayload, ThreadContext, ThreadContextFirstMessage};
|
use crate::entity::{InteractionContext, InteractionTrigger, MessagePayload};
|
||||||
|
|
||||||
|
struct DetailedMessagePayload {
|
||||||
|
is_mentioning_bot: bool,
|
||||||
|
message_payload: MessagePayload,
|
||||||
|
}
|
||||||
|
|
||||||
pub async fn get_matrix_messages_in_thread(
|
pub async fn get_matrix_messages_in_thread(
|
||||||
matrix_link: MatrixLink,
|
matrix_link: MatrixLink,
|
||||||
@@ -42,23 +50,123 @@ pub async fn get_matrix_messages_in_thread(
|
|||||||
Ok(messages)
|
Ok(messages)
|
||||||
}
|
}
|
||||||
|
|
||||||
pub async fn process_matrix_messages_in_thread(
|
pub async fn get_matrix_messages_in_reply_chain(
|
||||||
|
event_fetcher: &Arc<RoomEventFetcher>,
|
||||||
|
room: &Room,
|
||||||
|
event_id: OwnedEventId,
|
||||||
|
) -> Result<Vec<MatrixMessage>, mxlink::matrix_sdk::Error> {
|
||||||
|
let messages_native =
|
||||||
|
get_matrix_messages_in_reply_chain_native(event_fetcher, room, event_id).await?;
|
||||||
|
|
||||||
|
let mut messages: Vec<MatrixMessage> = Vec::new();
|
||||||
|
|
||||||
|
for matrix_native_message in messages_native {
|
||||||
|
let Some(message) = convert_matrix_native_event_to_matrix_message(&matrix_native_message)
|
||||||
|
else {
|
||||||
|
continue;
|
||||||
|
};
|
||||||
|
|
||||||
|
messages.push(message);
|
||||||
|
}
|
||||||
|
|
||||||
|
Ok(messages)
|
||||||
|
}
|
||||||
|
|
||||||
|
async fn get_matrix_messages_in_reply_chain_native(
|
||||||
|
event_fetcher: &Arc<RoomEventFetcher>,
|
||||||
|
room: &Room,
|
||||||
|
event_id: OwnedEventId,
|
||||||
|
) -> Result<Vec<AnyMessageLikeEvent>, mxlink::matrix_sdk::Error> {
|
||||||
|
let mut next_event_id = Some(event_id.clone());
|
||||||
|
|
||||||
|
let mut messages: Vec<AnyMessageLikeEvent> = Vec::new();
|
||||||
|
let mut handled_event_ids: Vec<OwnedEventId> = Vec::new();
|
||||||
|
|
||||||
|
while let Some(next_event_id_in_loop) = next_event_id {
|
||||||
|
let event = event_fetcher
|
||||||
|
.fetch_event_in_room(&next_event_id_in_loop, room)
|
||||||
|
.await
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
if handled_event_ids.contains(&next_event_id_in_loop) {
|
||||||
|
tracing::warn!(
|
||||||
|
"Not following loop-causing event: {}",
|
||||||
|
next_event_id_in_loop
|
||||||
|
);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
handled_event_ids.push(next_event_id_in_loop.clone());
|
||||||
|
|
||||||
|
let event_deserialized = event.event.deserialize()?;
|
||||||
|
|
||||||
|
let AnyTimelineEvent::MessageLike(message_like_event) = event_deserialized else {
|
||||||
|
tracing::warn!(
|
||||||
|
"Not proceeding past non-MessageLike event: {:?}",
|
||||||
|
event_deserialized
|
||||||
|
);
|
||||||
|
break;
|
||||||
|
};
|
||||||
|
|
||||||
|
next_event_id = match message_like_event.clone() {
|
||||||
|
AnyMessageLikeEvent::RoomEncrypted(_) => None,
|
||||||
|
AnyMessageLikeEvent::RoomMessage(room_message) => {
|
||||||
|
if let MessageLikeEvent::Original(room_message_original) = room_message {
|
||||||
|
match room_message_original.content.relates_to {
|
||||||
|
Some(Relation::Reply { in_reply_to }) => Some(in_reply_to.event_id.clone()),
|
||||||
|
_ => None,
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
None
|
||||||
|
}
|
||||||
|
}
|
||||||
|
_ => None,
|
||||||
|
};
|
||||||
|
|
||||||
|
messages.push(message_like_event);
|
||||||
|
}
|
||||||
|
|
||||||
|
messages.reverse();
|
||||||
|
|
||||||
|
Ok(messages)
|
||||||
|
}
|
||||||
|
|
||||||
|
pub async fn process_matrix_messages(
|
||||||
messages: &[MatrixMessage],
|
messages: &[MatrixMessage],
|
||||||
params: &MatrixMessageProcessingParams,
|
params: &MatrixMessageProcessingParams,
|
||||||
) -> Vec<MatrixMessage> {
|
) -> Vec<MatrixMessage> {
|
||||||
let mut messages_filtered: Vec<MatrixMessage> = Vec::new();
|
let mut messages_filtered: Vec<MatrixMessage> = Vec::new();
|
||||||
|
|
||||||
for (i, message) in messages.iter().enumerate() {
|
for (i, message) in messages.iter().enumerate() {
|
||||||
if !is_message_from_allowed_sender(message, ¶ms.bot_user_id, ¶ms.allowed_users) {
|
if !is_message_from_allowed_sender(
|
||||||
|
message,
|
||||||
|
¶ms.bot_user_id,
|
||||||
|
params.allowed_users.as_deref(),
|
||||||
|
) {
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
|
|
||||||
let mut message = message.clone();
|
let mut message = message.clone();
|
||||||
|
|
||||||
if i == 0 && !params.first_message_stripped_prefixes.is_empty() {
|
if i == 0 && !params.first_message_prefixes_to_strip.is_empty() {
|
||||||
let mut message_text = message.message_text.clone();
|
let mut message_text = message.message_text.clone();
|
||||||
|
|
||||||
for prefix in ¶ms.first_message_stripped_prefixes {
|
for prefix in ¶ms.first_message_prefixes_to_strip {
|
||||||
|
if let Some(message_text_stripped) = message_text.strip_prefix(prefix) {
|
||||||
|
message_text = message_text_stripped.to_owned();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
message.message_text = message_text.trim().to_owned();
|
||||||
|
}
|
||||||
|
|
||||||
|
// We only strip `bot_user_prefixes_to_strip`-defined prefixes from messages that mention the bot user.
|
||||||
|
if !params.bot_user_prefixes_to_strip.is_empty()
|
||||||
|
&& message.mentioned_users.contains(¶ms.bot_user_id)
|
||||||
|
{
|
||||||
|
let mut message_text = message.message_text.clone();
|
||||||
|
|
||||||
|
for prefix in ¶ms.bot_user_prefixes_to_strip {
|
||||||
if let Some(message_text_stripped) = message_text.strip_prefix(prefix) {
|
if let Some(message_text_stripped) = message_text.strip_prefix(prefix) {
|
||||||
message_text = message_text_stripped.to_owned();
|
message_text = message_text_stripped.to_owned();
|
||||||
}
|
}
|
||||||
@@ -73,16 +181,25 @@ pub async fn process_matrix_messages_in_thread(
|
|||||||
messages_filtered
|
messages_filtered
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Tells if the given message is from an allowed sender.
|
||||||
|
///
|
||||||
|
/// If allowed_users is None, all messages are allowed.
|
||||||
|
/// If allowed_users is Some, only messages from the allowed users (and the `bot_user_id`) are allowed.
|
||||||
fn is_message_from_allowed_sender(
|
fn is_message_from_allowed_sender(
|
||||||
matrix_message: &MatrixMessage,
|
matrix_message: &MatrixMessage,
|
||||||
bot_user_id: &str,
|
bot_user_id: &OwnedUserId,
|
||||||
allowed_users: &[regex::Regex],
|
allowed_users: Option<&[regex::Regex]>,
|
||||||
) -> bool {
|
) -> bool {
|
||||||
if matrix_message.sender_id == bot_user_id {
|
if matrix_message.sender_id == *bot_user_id {
|
||||||
return true;
|
return true;
|
||||||
}
|
}
|
||||||
|
|
||||||
if mxidwc::match_user_id(&matrix_message.sender_id, allowed_users) {
|
if let Some(allowed_users) = allowed_users {
|
||||||
|
if mxidwc::match_user_id(matrix_message.sender_id.as_str(), allowed_users) {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
// No allowed users configured, so all messages are allowed
|
||||||
return true;
|
return true;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -108,30 +225,67 @@ pub fn convert_matrix_native_event_to_matrix_message(
|
|||||||
_ => return None,
|
_ => return None,
|
||||||
};
|
};
|
||||||
|
|
||||||
|
let is_reply = matches!(room_message.relates_to, Some(Relation::Reply { .. }));
|
||||||
|
|
||||||
|
let text = if is_reply {
|
||||||
|
// For regular replies, we need to strip the fallback-for-rich replies part.
|
||||||
|
// See: https://spec.matrix.org/v1.11/client-server-api/#fallbacks-for-rich-replies
|
||||||
|
remove_plain_reply_fallback(&text).to_owned()
|
||||||
|
} else {
|
||||||
|
text
|
||||||
|
};
|
||||||
|
|
||||||
|
let timestamp = chrono::DateTime::<chrono::Utc>::from(
|
||||||
|
matrix_native_event
|
||||||
|
.origin_server_ts()
|
||||||
|
.to_system_time()
|
||||||
|
.unwrap_or_else(std::time::SystemTime::now),
|
||||||
|
);
|
||||||
|
|
||||||
|
let mentioned_users = room_message
|
||||||
|
.mentions
|
||||||
|
.map(|m| m.user_ids.iter().map(|u| u.to_owned()).collect())
|
||||||
|
.unwrap_or(vec![]);
|
||||||
|
|
||||||
Some(MatrixMessage {
|
Some(MatrixMessage {
|
||||||
sender_id: matrix_native_event.sender().to_string(),
|
sender_id: matrix_native_event.sender().to_owned(),
|
||||||
message_type: if is_notice {
|
message_type: if is_notice {
|
||||||
MatrixMessageType::Notice
|
MatrixMessageType::Notice
|
||||||
} else {
|
} else {
|
||||||
MatrixMessageType::Text
|
MatrixMessageType::Text
|
||||||
},
|
},
|
||||||
message_text: text,
|
message_text: text,
|
||||||
|
mentioned_users,
|
||||||
|
timestamp,
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Determines the thread context (relationship within the thread + first thread message payload) for an incoming (new) room event.
|
/// Determines the interaction context for an incoming (new) room event.
|
||||||
/// This room event is assumed to be the "newest message" in the thread (or a top-level message).
|
///
|
||||||
/// If the given event is a regular reply (not a thread reply), this function will return `None`.
|
/// This context is created based on the "newest message" (`current_event`), which is:
|
||||||
/// If the given event is a top-level message, this function will consider this event as the start of the thread.
|
/// - either a top-level message, which may or may not be mentioning the bot
|
||||||
/// If the given event is a thread reply, this function will inspect the thread root event and will return the thread context.
|
/// - this function will inspect the event and will likely start a new threaded conversation
|
||||||
/// If the thread root event is not found, is redacted, or is of some unsupported MessagePayload type, this function will return `None`.
|
///
|
||||||
pub async fn determine_thread_context_for_room_event(
|
/// - or a thread reply
|
||||||
|
/// - this function will inspect the thread root event and will return the interaction context
|
||||||
|
/// - if the bot only reacts to prefixed messsages (or mentions), this function may ignore the given thread reply, unless it mentions the bot (which causes a synthetic "first message" to be produced)
|
||||||
|
/// - if the thread root event is not found, is redacted, or is of some unsupported MessagePayload type, this function will return `None`
|
||||||
|
///
|
||||||
|
/// - or an in-room (non-threaded) reply to a room message, which may or may not be mentioning the bot
|
||||||
|
/// - replies that do not mention the bot cause this function to return `None`
|
||||||
|
/// - other replies create a interaction context which points to a "first message" which is synthetic
|
||||||
|
#[tracing::instrument(name = "determine_interaction_context_for_room_event", skip_all, fields(room_id = room.room_id().as_str(), event_id = current_event.event_id.as_str()))]
|
||||||
|
pub async fn determine_interaction_context_for_room_event(
|
||||||
bot_user_id: &OwnedUserId,
|
bot_user_id: &OwnedUserId,
|
||||||
|
bot_display_name: &Option<String>,
|
||||||
room: &Room,
|
room: &Room,
|
||||||
current_event: &OriginalSyncRoomMessageEvent,
|
current_event: &OriginalSyncRoomMessageEvent,
|
||||||
current_event_payload: &MessagePayload,
|
current_event_payload: &MessagePayload,
|
||||||
event_fetcher: &Arc<RoomEventFetcher>,
|
event_fetcher: &Arc<RoomEventFetcher>,
|
||||||
) -> anyhow::Result<Option<ThreadContext>> {
|
) -> anyhow::Result<Option<InteractionContext>> {
|
||||||
|
let current_event_is_mentioning_bot =
|
||||||
|
is_event_mentioning_bot(¤t_event.content, bot_user_id, bot_display_name);
|
||||||
|
|
||||||
let Some(relation) = ¤t_event.content.relates_to else {
|
let Some(relation) = ¤t_event.content.relates_to else {
|
||||||
// This is a top-level message. We consider it the start of the thread.
|
// This is a top-level message. We consider it the start of the thread.
|
||||||
let thread_info = ThreadInfo::new(
|
let thread_info = ThreadInfo::new(
|
||||||
@@ -139,25 +293,75 @@ pub async fn determine_thread_context_for_room_event(
|
|||||||
current_event.event_id.clone(),
|
current_event.event_id.clone(),
|
||||||
);
|
);
|
||||||
|
|
||||||
let is_mentioning_bot = is_event_mentioning_bot(¤t_event.content, bot_user_id);
|
return Ok(Some(InteractionContext {
|
||||||
|
thread_info,
|
||||||
return Ok(Some(ThreadContext {
|
trigger: InteractionTrigger {
|
||||||
info: thread_info,
|
is_mentioning_bot: current_event_is_mentioning_bot,
|
||||||
first_message: ThreadContextFirstMessage {
|
|
||||||
is_mentioning_bot,
|
|
||||||
payload: current_event_payload.clone(),
|
payload: current_event_payload.clone(),
|
||||||
},
|
},
|
||||||
}));
|
}));
|
||||||
};
|
};
|
||||||
|
|
||||||
let Relation::Thread(thread) = relation else {
|
match relation {
|
||||||
// This is a reply or a replacement, etc. It's not a thread.
|
Relation::Thread(thread) => {
|
||||||
// We don't care about this.
|
determine_interaction_context_for_room_event_related_to_thread(
|
||||||
return Ok(None);
|
bot_user_id,
|
||||||
};
|
bot_display_name,
|
||||||
|
room,
|
||||||
|
current_event,
|
||||||
|
event_fetcher,
|
||||||
|
current_event_is_mentioning_bot,
|
||||||
|
thread,
|
||||||
|
)
|
||||||
|
.await
|
||||||
|
}
|
||||||
|
Relation::Reply { in_reply_to } => {
|
||||||
|
determine_interaction_context_for_room_event_related_to_reply(
|
||||||
|
current_event,
|
||||||
|
current_event_is_mentioning_bot,
|
||||||
|
in_reply_to.event_id.clone(),
|
||||||
|
)
|
||||||
|
.await
|
||||||
|
}
|
||||||
|
|
||||||
|
// This is a replacement or something else. It's not something we support.
|
||||||
|
_ => return Ok(None),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
async fn determine_interaction_context_for_room_event_related_to_thread(
|
||||||
|
bot_user_id: &OwnedUserId,
|
||||||
|
bot_display_name: &Option<String>,
|
||||||
|
room: &Room,
|
||||||
|
current_event: &OriginalSyncRoomMessageEvent,
|
||||||
|
event_fetcher: &Arc<RoomEventFetcher>,
|
||||||
|
current_event_is_mentioning_bot: bool,
|
||||||
|
thread: &Thread,
|
||||||
|
) -> anyhow::Result<Option<InteractionContext>> {
|
||||||
let thread_info = ThreadInfo::new(thread.event_id.clone(), current_event.event_id.clone());
|
let thread_info = ThreadInfo::new(thread.event_id.clone(), current_event.event_id.clone());
|
||||||
|
|
||||||
|
tracing::trace!(
|
||||||
|
?current_event_is_mentioning_bot,
|
||||||
|
is_thread_root_only = thread_info.is_thread_root_only(),
|
||||||
|
"Dealing with a thread reply",
|
||||||
|
);
|
||||||
|
|
||||||
|
if current_event_is_mentioning_bot && !thread_info.is_thread_root_only() {
|
||||||
|
// If the current event is a thread reply and is mentioning the bot,
|
||||||
|
// it's probably someone trying to involve us in the threaded conversation.
|
||||||
|
// See: https://github.com/etkecc/baibot/issues/15
|
||||||
|
//
|
||||||
|
// In such cases, we don't care what the thread root event is like or what the current event is like,
|
||||||
|
// we want text-generation to be triggered for this whole thread regardless.
|
||||||
|
return Ok(Some(InteractionContext {
|
||||||
|
thread_info,
|
||||||
|
trigger: InteractionTrigger {
|
||||||
|
is_mentioning_bot: true,
|
||||||
|
payload: MessagePayload::SynthethicChatCompletionTriggerInThread,
|
||||||
|
},
|
||||||
|
}));
|
||||||
|
}
|
||||||
|
|
||||||
let start_time = std::time::Instant::now();
|
let start_time = std::time::Instant::now();
|
||||||
|
|
||||||
let thread_start_timeline_event = event_fetcher
|
let thread_start_timeline_event = event_fetcher
|
||||||
@@ -183,34 +387,116 @@ pub async fn determine_thread_context_for_room_event(
|
|||||||
"Fetched thread start event"
|
"Fetched thread start event"
|
||||||
);
|
);
|
||||||
|
|
||||||
let thread_start_timeline_event_deserialized =
|
let thread_start_detailed_message_payload = timeline_event_to_detailed_message_payload(
|
||||||
match thread_start_timeline_event.event.deserialize() {
|
&thread.event_id,
|
||||||
|
thread_start_timeline_event,
|
||||||
|
thread_info.clone(),
|
||||||
|
bot_user_id,
|
||||||
|
bot_display_name,
|
||||||
|
)?;
|
||||||
|
|
||||||
|
let Some(detailed_message_payload) = thread_start_detailed_message_payload else {
|
||||||
|
return Ok(None);
|
||||||
|
};
|
||||||
|
|
||||||
|
Ok(Some(InteractionContext {
|
||||||
|
thread_info,
|
||||||
|
trigger: InteractionTrigger {
|
||||||
|
is_mentioning_bot: detailed_message_payload.is_mentioning_bot,
|
||||||
|
payload: detailed_message_payload.message_payload,
|
||||||
|
},
|
||||||
|
}))
|
||||||
|
}
|
||||||
|
|
||||||
|
async fn determine_interaction_context_for_room_event_related_to_reply(
|
||||||
|
current_event: &OriginalSyncRoomMessageEvent,
|
||||||
|
current_event_is_mentioning_bot: bool,
|
||||||
|
reply_to_event_id: OwnedEventId,
|
||||||
|
) -> anyhow::Result<Option<InteractionContext>> {
|
||||||
|
tracing::trace!(?current_event_is_mentioning_bot, "Dealing with a reply");
|
||||||
|
|
||||||
|
if !current_event_is_mentioning_bot {
|
||||||
|
// If the current event is not mentioning the bot, we don't care about it.
|
||||||
|
tracing::trace!("Ignoring reply event which does not mention the bot");
|
||||||
|
return Ok(None);
|
||||||
|
}
|
||||||
|
|
||||||
|
let thread_info = ThreadInfo::new(reply_to_event_id.clone(), current_event.event_id.clone());
|
||||||
|
|
||||||
|
Ok(Some(InteractionContext {
|
||||||
|
thread_info,
|
||||||
|
trigger: InteractionTrigger {
|
||||||
|
is_mentioning_bot: true,
|
||||||
|
payload: MessagePayload::SynthethicChatCompletionTriggerForReply,
|
||||||
|
},
|
||||||
|
}))
|
||||||
|
}
|
||||||
|
|
||||||
|
fn is_event_mentioning_bot(
|
||||||
|
event_content: &RoomMessageEventContent,
|
||||||
|
bot_user_id: &OwnedUserId,
|
||||||
|
bot_display_name: &Option<String>,
|
||||||
|
) -> bool {
|
||||||
|
if let Some(mentions) = &event_content.mentions {
|
||||||
|
mentions
|
||||||
|
.user_ids
|
||||||
|
.iter()
|
||||||
|
.any(|user_id| user_id == bot_user_id)
|
||||||
|
} else {
|
||||||
|
// For compatibility with clients that do not support the new Mentions specification
|
||||||
|
// (see https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions),
|
||||||
|
// we also do string matching here.
|
||||||
|
//
|
||||||
|
// As of 2024-10-03, at least Element iOS does not support the new Mentions specification
|
||||||
|
// and is still quite widespread.
|
||||||
|
//
|
||||||
|
// We may consider dropping this string-matching behavior altogether in the future,
|
||||||
|
// so improving this compatibility block is not a high priority.
|
||||||
|
if event_content.body().contains(bot_user_id.as_str()) {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
if let Some(bot_display_name) = bot_display_name {
|
||||||
|
return event_content.body().contains(bot_display_name);
|
||||||
|
}
|
||||||
|
|
||||||
|
false
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn timeline_event_to_detailed_message_payload(
|
||||||
|
timeline_event_id: &OwnedEventId,
|
||||||
|
timeline_event: TimelineEvent,
|
||||||
|
thread_info: ThreadInfo,
|
||||||
|
bot_user_id: &OwnedUserId,
|
||||||
|
bot_display_name: &Option<String>,
|
||||||
|
) -> anyhow::Result<Option<DetailedMessagePayload>> {
|
||||||
|
let timeline_event_deserialized = match timeline_event.event.deserialize() {
|
||||||
Ok(value) => value,
|
Ok(value) => value,
|
||||||
Err(err) => {
|
Err(err) => {
|
||||||
return Err(anyhow::format_err!(
|
return Err(anyhow::format_err!(
|
||||||
"Failed to deserialize thread start event {}: {:?}",
|
"Failed to deserialize timeline event {}: {:?}",
|
||||||
thread.event_id,
|
timeline_event_id,
|
||||||
err
|
err
|
||||||
));
|
));
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
let AnyTimelineEvent::MessageLike(thread_start_message_like_event) =
|
let AnyTimelineEvent::MessageLike(thread_start_message_like_event) =
|
||||||
thread_start_timeline_event_deserialized
|
timeline_event_deserialized
|
||||||
else {
|
else {
|
||||||
tracing::trace!(
|
tracing::trace!(
|
||||||
"Ignoring non-MessageLike thread start event: {:?}",
|
"Ignoring non-MessageLike timeline event: {:?}",
|
||||||
thread_start_timeline_event_deserialized
|
timeline_event_deserialized
|
||||||
);
|
);
|
||||||
return Ok(None);
|
return Ok(None);
|
||||||
};
|
};
|
||||||
|
|
||||||
let (thread_start_message_is_mentioning_bot, thread_start_message_payload) =
|
let (is_mentioning_bot, message_payload) = match thread_start_message_like_event {
|
||||||
match thread_start_message_like_event {
|
|
||||||
AnyMessageLikeEvent::RoomEncrypted(room_message) => {
|
AnyMessageLikeEvent::RoomEncrypted(room_message) => {
|
||||||
tracing::warn!(
|
tracing::warn!(
|
||||||
"Could not inspect thread start event {} because it failed to decrypt: {:?}",
|
"Could not inspect event {} because it failed to decrypt: {:?}",
|
||||||
thread.event_id.clone(),
|
timeline_event_id.clone(),
|
||||||
room_message
|
room_message
|
||||||
);
|
);
|
||||||
|
|
||||||
@@ -230,58 +516,69 @@ pub async fn determine_thread_context_for_room_event(
|
|||||||
let Ok(room_message_payload) = room_message_payload else {
|
let Ok(room_message_payload) = room_message_payload else {
|
||||||
tracing::debug!(
|
tracing::debug!(
|
||||||
msg_type = room_message_original.content.msgtype(),
|
msg_type = room_message_original.content.msgtype(),
|
||||||
"Ignoring thread start message of unknown type",
|
"Ignoring event message of unknown type",
|
||||||
);
|
);
|
||||||
return Ok(None);
|
return Ok(None);
|
||||||
};
|
};
|
||||||
|
|
||||||
let is_mentioning_bot =
|
let is_mentioning_bot = is_event_mentioning_bot(
|
||||||
is_event_mentioning_bot(&room_message_original.content, bot_user_id);
|
&room_message_original.content,
|
||||||
|
bot_user_id,
|
||||||
|
bot_display_name,
|
||||||
|
);
|
||||||
|
|
||||||
(is_mentioning_bot, room_message_payload)
|
(is_mentioning_bot, room_message_payload)
|
||||||
} else {
|
} else {
|
||||||
tracing::error!("Ignoring thread start message which appears to be redacted");
|
tracing::error!("Ignoring event message which appears to be redacted");
|
||||||
|
|
||||||
return Ok(None);
|
return Ok(None);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
other => {
|
other => {
|
||||||
tracing::trace!(
|
tracing::trace!("Ignoring unknown MessageLike event: {:?}", other);
|
||||||
"Ignoring unknown MessageLike thread start event: {:?}",
|
|
||||||
other
|
|
||||||
);
|
|
||||||
return Ok(None);
|
return Ok(None);
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
Ok(Some(ThreadContext {
|
Ok(Some(DetailedMessagePayload {
|
||||||
info: thread_info,
|
is_mentioning_bot,
|
||||||
first_message: ThreadContextFirstMessage {
|
message_payload,
|
||||||
is_mentioning_bot: thread_start_message_is_mentioning_bot,
|
|
||||||
payload: thread_start_message_payload,
|
|
||||||
},
|
|
||||||
}))
|
}))
|
||||||
}
|
}
|
||||||
|
|
||||||
fn is_event_mentioning_bot(
|
/// Creates a list of prefixes to strip from the beginning of message texts that mention the bot user.
|
||||||
event_content: &RoomMessageEventContent,
|
///
|
||||||
|
/// Different clients do mentions differently.
|
||||||
|
/// The body text containing the mention usually contains one of:
|
||||||
|
/// - the full user ID (includes a @ prefix by default)
|
||||||
|
/// - the localpart (with a @ prefix)
|
||||||
|
/// - the localpart (without a @ prefix)
|
||||||
|
/// - the display name (with a @ prefix)
|
||||||
|
/// - the display name (without a @ prefix)
|
||||||
|
///
|
||||||
|
/// Some add a `: ` suffix after the mention.
|
||||||
|
///
|
||||||
|
/// There's no guarantee that the mention is at the start even.
|
||||||
|
/// It being there is most common and we try to strip it from there
|
||||||
|
/// as best as we can.
|
||||||
|
pub fn create_list_of_bot_user_prefixes_to_strip(
|
||||||
bot_user_id: &OwnedUserId,
|
bot_user_id: &OwnedUserId,
|
||||||
) -> bool {
|
bot_display_name: &Option<String>,
|
||||||
if let Some(mentions) = &event_content.mentions {
|
) -> Vec<String> {
|
||||||
mentions
|
let bot_user_id_localpart = bot_user_id.localpart();
|
||||||
.user_ids
|
|
||||||
.iter()
|
let mut prefixes_to_strip = vec![
|
||||||
.any(|user_id| user_id == bot_user_id)
|
bot_user_id.as_str().to_owned(),
|
||||||
} else {
|
format!("@{}", bot_user_id_localpart),
|
||||||
// For compatibility with clients that do not support the new Mentions specification
|
bot_user_id_localpart.to_owned(),
|
||||||
// (see https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions),
|
];
|
||||||
// we also do string matching here.
|
|
||||||
//
|
if let Some(bot_display_name) = bot_display_name {
|
||||||
// It may be even better to match not only against the MXID, but also against the bot's
|
prefixes_to_strip.push(format!("@{}", bot_display_name));
|
||||||
// room-specific display name.
|
prefixes_to_strip.push(bot_display_name.to_owned());
|
||||||
//
|
|
||||||
// We may consider dropping this string-matching behavior altogether in the future,
|
|
||||||
// so improving this compatibility block is not a high priority.
|
|
||||||
event_content.body().contains(bot_user_id.as_str())
|
|
||||||
}
|
}
|
||||||
|
|
||||||
|
prefixes_to_strip.push(":".to_owned());
|
||||||
|
|
||||||
|
prefixes_to_strip
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,29 +1,42 @@
|
|||||||
|
use chrono::{TimeZone, Utc};
|
||||||
|
|
||||||
|
use mxlink::matrix_sdk::ruma::OwnedUserId;
|
||||||
|
|
||||||
use crate::conversation::matrix::{
|
use crate::conversation::matrix::{
|
||||||
MatrixMessage, MatrixMessageProcessingParams, MatrixMessageType,
|
MatrixMessage, MatrixMessageProcessingParams, MatrixMessageType,
|
||||||
};
|
};
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn is_message_from_allowed_sender() {
|
fn is_message_from_allowed_sender() {
|
||||||
let bot_user_id = "@bot:example.com";
|
let bot_user_id =
|
||||||
let allowed_user_id = "@user.someone:example.com";
|
OwnedUserId::try_from("@bot:example.com").expect("Failed to parse bot user ID");
|
||||||
let unallowed_user_id = "@another:example.com";
|
let allowed_user_id = OwnedUserId::try_from("@user.someone:example.com").unwrap();
|
||||||
|
let unallowed_user_id = OwnedUserId::try_from("@another:example.com").unwrap();
|
||||||
|
|
||||||
|
let timestamp = Utc.with_ymd_and_hms(2024, 9, 20, 18, 34, 15).unwrap();
|
||||||
|
|
||||||
let bot_message = MatrixMessage {
|
let bot_message = MatrixMessage {
|
||||||
sender_id: bot_user_id.to_owned(),
|
sender_id: bot_user_id.to_owned(),
|
||||||
message_type: MatrixMessageType::Text,
|
message_type: MatrixMessageType::Text,
|
||||||
message_text: "Hello!".to_owned(),
|
message_text: "Hello!".to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp,
|
||||||
};
|
};
|
||||||
|
|
||||||
let allowed_user_message = MatrixMessage {
|
let allowed_user_message = MatrixMessage {
|
||||||
sender_id: allowed_user_id.to_owned(),
|
sender_id: allowed_user_id.to_owned(),
|
||||||
message_type: MatrixMessageType::Text,
|
message_type: MatrixMessageType::Text,
|
||||||
message_text: "Hello!".to_owned(),
|
message_text: "Hello!".to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp,
|
||||||
};
|
};
|
||||||
|
|
||||||
let unallowed_user_message = MatrixMessage {
|
let unallowed_user_message = MatrixMessage {
|
||||||
sender_id: unallowed_user_id.to_owned(),
|
sender_id: unallowed_user_id.to_owned(),
|
||||||
message_type: MatrixMessageType::Text,
|
message_type: MatrixMessageType::Text,
|
||||||
message_text: "Hello!".to_owned(),
|
message_text: "Hello!".to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp,
|
||||||
};
|
};
|
||||||
|
|
||||||
let parsed_regex = match mxidwc::parse_pattern("@user.*:example.com") {
|
let parsed_regex = match mxidwc::parse_pattern("@user.*:example.com") {
|
||||||
@@ -36,65 +49,106 @@ fn is_message_from_allowed_sender() {
|
|||||||
let allowed_users = vec![parsed_regex];
|
let allowed_users = vec![parsed_regex];
|
||||||
|
|
||||||
assert!(
|
assert!(
|
||||||
super::is_message_from_allowed_sender(&bot_message, bot_user_id, &vec![]),
|
super::is_message_from_allowed_sender(&bot_message, &bot_user_id, Some(&allowed_users)),
|
||||||
"Bot message should be allowed"
|
"Bot message should be allowed"
|
||||||
);
|
);
|
||||||
|
|
||||||
assert!(
|
assert!(
|
||||||
super::is_message_from_allowed_sender(&allowed_user_message, bot_user_id, &allowed_users),
|
super::is_message_from_allowed_sender(
|
||||||
|
&allowed_user_message,
|
||||||
|
&bot_user_id,
|
||||||
|
Some(&allowed_users)
|
||||||
|
),
|
||||||
"Allowed user message should be allowed"
|
"Allowed user message should be allowed"
|
||||||
);
|
);
|
||||||
|
|
||||||
assert!(
|
assert!(
|
||||||
!super::is_message_from_allowed_sender(
|
!super::is_message_from_allowed_sender(
|
||||||
&unallowed_user_message,
|
&unallowed_user_message,
|
||||||
bot_user_id,
|
&bot_user_id,
|
||||||
&allowed_users
|
Some(&allowed_users),
|
||||||
),
|
),
|
||||||
"Unallowed user message should be ignored"
|
"Unallowed user message should be ignored"
|
||||||
);
|
);
|
||||||
|
|
||||||
|
assert!(
|
||||||
|
super::is_message_from_allowed_sender(&unallowed_user_message, &bot_user_id, None,),
|
||||||
|
"An empty list of allowed users lets everyone through"
|
||||||
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
#[tokio::test]
|
#[tokio::test]
|
||||||
async fn process_matrix_messages_in_thread() {
|
async fn process_matrix_messages() {
|
||||||
let bot_user_id = "@bot:example.com";
|
let bot_user_id =
|
||||||
let allowed_user_id = "@user.someone:example.com";
|
OwnedUserId::try_from("@bot:example.com").expect("Failed to parse bot user ID");
|
||||||
let unallowed_user_id = "@another:example.com";
|
let allowed_user_id = OwnedUserId::try_from("@user.someone:example.com").unwrap();
|
||||||
|
let unallowed_user_id = OwnedUserId::try_from("@another:example.com").unwrap();
|
||||||
|
|
||||||
|
let timestamp = Utc.with_ymd_and_hms(2024, 9, 20, 18, 34, 15).unwrap();
|
||||||
|
|
||||||
let allowed_user_message = MatrixMessage {
|
let allowed_user_message = MatrixMessage {
|
||||||
sender_id: allowed_user_id.to_owned(),
|
sender_id: allowed_user_id.to_owned(),
|
||||||
message_type: MatrixMessageType::Text,
|
message_type: MatrixMessageType::Text,
|
||||||
message_text: "Hello from the user!".to_owned(),
|
message_text: "Hello from the user!".to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp,
|
||||||
};
|
};
|
||||||
|
|
||||||
let allowed_user_message_with_prefix = MatrixMessage {
|
let allowed_user_message_with_prefix = MatrixMessage {
|
||||||
sender_id: allowed_user_id.to_owned(),
|
sender_id: allowed_user_id.to_owned(),
|
||||||
message_type: MatrixMessageType::Text,
|
message_type: MatrixMessageType::Text,
|
||||||
message_text: "!bai Hello from the user!".to_owned(),
|
message_text: "!bai Hello from the user!".to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp,
|
||||||
};
|
};
|
||||||
|
|
||||||
let allowed_user_message_with_prefix_no_space = MatrixMessage {
|
let allowed_user_message_with_prefix_no_space = MatrixMessage {
|
||||||
sender_id: allowed_user_id.to_owned(),
|
sender_id: allowed_user_id.to_owned(),
|
||||||
message_type: MatrixMessageType::Text,
|
message_type: MatrixMessageType::Text,
|
||||||
message_text: "!baiHello from the user!".to_owned(),
|
message_text: "!baiHello from the user!".to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp,
|
||||||
};
|
};
|
||||||
|
|
||||||
let allowed_user_message_with_prefix_full_width_space = MatrixMessage {
|
let allowed_user_message_with_prefix_full_width_space = MatrixMessage {
|
||||||
sender_id: allowed_user_id.to_owned(),
|
sender_id: allowed_user_id.to_owned(),
|
||||||
message_type: MatrixMessageType::Text,
|
message_type: MatrixMessageType::Text,
|
||||||
message_text: "!bai Hello from the user!".to_owned(),
|
message_text: "!bai Hello from the user!".to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp,
|
||||||
};
|
};
|
||||||
|
|
||||||
let bot_message = MatrixMessage {
|
let bot_message = MatrixMessage {
|
||||||
sender_id: bot_user_id.to_owned(),
|
sender_id: bot_user_id.to_owned(),
|
||||||
message_type: MatrixMessageType::Text,
|
message_type: MatrixMessageType::Text,
|
||||||
message_text: "Hello from the bot!".to_owned(),
|
message_text: "Hello from the bot!".to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp,
|
||||||
|
};
|
||||||
|
|
||||||
|
let allowed_user_message_with_bot_mention = MatrixMessage {
|
||||||
|
sender_id: allowed_user_id.to_owned(),
|
||||||
|
message_type: MatrixMessageType::Text,
|
||||||
|
message_text: "@baibot: Hello from the user!".to_owned(),
|
||||||
|
mentioned_users: vec![bot_user_id.to_owned()],
|
||||||
|
timestamp,
|
||||||
|
};
|
||||||
|
|
||||||
|
// The message text is the same as above - it mentions the bot, but the actually-mentioned user is another user.
|
||||||
|
let allowed_user_message_with_another_user_mention = MatrixMessage {
|
||||||
|
sender_id: allowed_user_id.to_owned(),
|
||||||
|
message_type: MatrixMessageType::Text,
|
||||||
|
message_text: allowed_user_message_with_bot_mention.message_text.clone(),
|
||||||
|
mentioned_users: vec![allowed_user_id.to_owned()],
|
||||||
|
timestamp,
|
||||||
};
|
};
|
||||||
|
|
||||||
let unallowed_user_message = MatrixMessage {
|
let unallowed_user_message = MatrixMessage {
|
||||||
sender_id: unallowed_user_id.to_owned(),
|
sender_id: unallowed_user_id.to_owned(),
|
||||||
message_type: MatrixMessageType::Text,
|
message_type: MatrixMessageType::Text,
|
||||||
message_text: "Hello from an unallowed user!".to_owned(),
|
message_text: "Hello from an unallowed user!".to_owned(),
|
||||||
|
mentioned_users: vec![],
|
||||||
|
timestamp,
|
||||||
};
|
};
|
||||||
|
|
||||||
let parsed_regex = match mxidwc::parse_pattern("@user.*:example.com") {
|
let parsed_regex = match mxidwc::parse_pattern("@user.*:example.com") {
|
||||||
@@ -106,12 +160,24 @@ async fn process_matrix_messages_in_thread() {
|
|||||||
|
|
||||||
let allowed_users = vec![parsed_regex];
|
let allowed_users = vec![parsed_regex];
|
||||||
|
|
||||||
let message_processing_params_basic =
|
let message_processing_params_basic = super::MatrixMessageProcessingParams::new(
|
||||||
super::MatrixMessageProcessingParams::new(bot_user_id.to_owned(), allowed_users.clone());
|
bot_user_id.to_owned(),
|
||||||
|
Some(allowed_users.clone()),
|
||||||
|
);
|
||||||
|
|
||||||
let message_processing_params_with_prefix_stripping =
|
let message_processing_params_with_prefix_stripping =
|
||||||
super::MatrixMessageProcessingParams::new(bot_user_id.to_owned(), allowed_users.clone())
|
super::MatrixMessageProcessingParams::new(
|
||||||
.with_first_message_stripped_prefixes(vec!["!bai".to_owned()]);
|
bot_user_id.to_owned(),
|
||||||
|
Some(allowed_users.clone()),
|
||||||
|
)
|
||||||
|
.with_first_message_prefixes_to_strip(vec!["!bai".to_owned()]);
|
||||||
|
|
||||||
|
let message_processing_params_with_bot_user_prefix_stripping =
|
||||||
|
super::MatrixMessageProcessingParams::new(
|
||||||
|
bot_user_id.to_owned(),
|
||||||
|
Some(allowed_users.clone()),
|
||||||
|
)
|
||||||
|
.with_bot_user_prefixes_to_strip(vec!["@baibot: ".to_owned(), "@baibot".to_owned()]);
|
||||||
|
|
||||||
struct TestCase {
|
struct TestCase {
|
||||||
name: String,
|
name: String,
|
||||||
@@ -195,10 +261,23 @@ async fn process_matrix_messages_in_thread() {
|
|||||||
"!bai Hello from the user!".to_owned(),
|
"!bai Hello from the user!".to_owned(),
|
||||||
],
|
],
|
||||||
},
|
},
|
||||||
|
TestCase {
|
||||||
|
name: "Messages that mention the bot user get the bot user prefix stripped"
|
||||||
|
.to_owned(),
|
||||||
|
messages: vec![
|
||||||
|
allowed_user_message_with_bot_mention.clone(),
|
||||||
|
allowed_user_message_with_another_user_mention.clone(),
|
||||||
|
],
|
||||||
|
message_processing_params: message_processing_params_with_bot_user_prefix_stripping.clone(),
|
||||||
|
expected_message_texts: vec![
|
||||||
|
"Hello from the user!".to_owned(),
|
||||||
|
"@baibot: Hello from the user!".to_owned(),
|
||||||
|
],
|
||||||
|
},
|
||||||
];
|
];
|
||||||
|
|
||||||
for test_case in test_cases {
|
for test_case in test_cases {
|
||||||
let processed_messages = super::process_matrix_messages_in_thread(
|
let processed_messages = super::process_matrix_messages(
|
||||||
&test_case.messages,
|
&test_case.messages,
|
||||||
&test_case.message_processing_params,
|
&test_case.message_processing_params,
|
||||||
)
|
)
|
||||||
@@ -216,3 +295,41 @@ async fn process_matrix_messages_in_thread() {
|
|||||||
);
|
);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn create_list_of_bot_user_prefixes_to_strip() {
|
||||||
|
let bot_user_id =
|
||||||
|
OwnedUserId::try_from("@baibot:example.com").expect("Failed to parse bot user ID");
|
||||||
|
|
||||||
|
// Test case 1: Bot user with no display name
|
||||||
|
let bot_display_name = None;
|
||||||
|
let prefixes =
|
||||||
|
super::create_list_of_bot_user_prefixes_to_strip(&bot_user_id, &bot_display_name);
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
prefixes,
|
||||||
|
vec![
|
||||||
|
"@baibot:example.com".to_string(),
|
||||||
|
"@baibot".to_string(),
|
||||||
|
"baibot".to_string(),
|
||||||
|
":".to_string()
|
||||||
|
]
|
||||||
|
);
|
||||||
|
|
||||||
|
// Test case 2: Bot user with display name
|
||||||
|
let bot_display_name = Some("Assistant".to_string());
|
||||||
|
let prefixes =
|
||||||
|
super::create_list_of_bot_user_prefixes_to_strip(&bot_user_id, &bot_display_name);
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
prefixes,
|
||||||
|
vec![
|
||||||
|
"@baibot:example.com".to_string(),
|
||||||
|
"@baibot".to_string(),
|
||||||
|
"baibot".to_string(),
|
||||||
|
"@Assistant".to_string(),
|
||||||
|
"Assistant".to_string(),
|
||||||
|
":".to_string()
|
||||||
|
]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|||||||
@@ -1,9 +1,14 @@
|
|||||||
|
use std::sync::Arc;
|
||||||
|
|
||||||
use mxlink::matrix_sdk::ruma::OwnedEventId;
|
use mxlink::matrix_sdk::ruma::OwnedEventId;
|
||||||
use mxlink::MatrixLink;
|
use mxlink::MatrixLink;
|
||||||
|
|
||||||
|
use crate::conversation::matrix::MatrixMessage;
|
||||||
|
|
||||||
use super::llm::{convert_matrix_message_to_llm_message, Conversation, Message};
|
use super::llm::{convert_matrix_message_to_llm_message, Conversation, Message};
|
||||||
use super::matrix::{
|
use super::matrix::{
|
||||||
get_matrix_messages_in_thread, process_matrix_messages_in_thread, MatrixMessageProcessingParams,
|
get_matrix_messages_in_reply_chain, get_matrix_messages_in_thread, process_matrix_messages,
|
||||||
|
MatrixMessageProcessingParams, RoomEventFetcher,
|
||||||
};
|
};
|
||||||
|
|
||||||
pub async fn create_llm_conversation_for_matrix_thread(
|
pub async fn create_llm_conversation_for_matrix_thread(
|
||||||
@@ -14,7 +19,33 @@ pub async fn create_llm_conversation_for_matrix_thread(
|
|||||||
) -> Result<Conversation, mxlink::matrix_sdk::Error> {
|
) -> Result<Conversation, mxlink::matrix_sdk::Error> {
|
||||||
let messages = get_matrix_messages_in_thread(matrix_link, room, thread_id).await?;
|
let messages = get_matrix_messages_in_thread(matrix_link, room, thread_id).await?;
|
||||||
|
|
||||||
let messages_filtered = process_matrix_messages_in_thread(&messages, params).await;
|
let llm_messages = filter_messages_and_convert_to_llm_messages(messages, params).await;
|
||||||
|
|
||||||
|
Ok(Conversation {
|
||||||
|
messages: llm_messages,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
pub async fn create_llm_conversation_for_matrix_reply_chain(
|
||||||
|
event_fetcher: &Arc<RoomEventFetcher>,
|
||||||
|
room: &mxlink::matrix_sdk::Room,
|
||||||
|
event_id: OwnedEventId,
|
||||||
|
params: &MatrixMessageProcessingParams,
|
||||||
|
) -> Result<Conversation, mxlink::matrix_sdk::Error> {
|
||||||
|
let messages = get_matrix_messages_in_reply_chain(event_fetcher, room, event_id).await?;
|
||||||
|
|
||||||
|
let llm_messages = filter_messages_and_convert_to_llm_messages(messages, params).await;
|
||||||
|
|
||||||
|
Ok(Conversation {
|
||||||
|
messages: llm_messages,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
async fn filter_messages_and_convert_to_llm_messages(
|
||||||
|
messages: Vec<MatrixMessage>,
|
||||||
|
params: &MatrixMessageProcessingParams,
|
||||||
|
) -> Vec<Message> {
|
||||||
|
let messages_filtered = process_matrix_messages(&messages, params).await;
|
||||||
|
|
||||||
let mut llm_messages: Vec<Message> = Vec::new();
|
let mut llm_messages: Vec<Message> = Vec::new();
|
||||||
|
|
||||||
@@ -28,7 +59,5 @@ pub async fn create_llm_conversation_for_matrix_thread(
|
|||||||
llm_messages.push(llm_message);
|
llm_messages.push(llm_message);
|
||||||
}
|
}
|
||||||
|
|
||||||
Ok(Conversation {
|
llm_messages
|
||||||
messages: llm_messages,
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -2,4 +2,6 @@ pub(crate) mod llm;
|
|||||||
pub(crate) mod matrix;
|
pub(crate) mod matrix;
|
||||||
mod matrix_llm_bridge;
|
mod matrix_llm_bridge;
|
||||||
|
|
||||||
pub(crate) use matrix_llm_bridge::create_llm_conversation_for_matrix_thread;
|
pub(crate) use matrix_llm_bridge::{
|
||||||
|
create_llm_conversation_for_matrix_reply_chain, create_llm_conversation_for_matrix_thread,
|
||||||
|
};
|
||||||
|
|||||||
13
src/entity/interaction_context.rs
Normal file
13
src/entity/interaction_context.rs
Normal file
@@ -0,0 +1,13 @@
|
|||||||
|
use mxlink::ThreadInfo;
|
||||||
|
|
||||||
|
use super::MessagePayload;
|
||||||
|
|
||||||
|
pub struct InteractionContext {
|
||||||
|
pub thread_info: ThreadInfo,
|
||||||
|
pub trigger: InteractionTrigger,
|
||||||
|
}
|
||||||
|
|
||||||
|
pub struct InteractionTrigger {
|
||||||
|
pub is_mentioning_bot: bool,
|
||||||
|
pub payload: MessagePayload,
|
||||||
|
}
|
||||||
@@ -15,6 +15,8 @@ pub struct MessageContext {
|
|||||||
admin_whitelist_regexes: Vec<regex::Regex>,
|
admin_whitelist_regexes: Vec<regex::Regex>,
|
||||||
trigger_event_info: TriggerEventInfo,
|
trigger_event_info: TriggerEventInfo,
|
||||||
thread_info: ThreadInfo,
|
thread_info: ThreadInfo,
|
||||||
|
|
||||||
|
bot_display_name: Option<String>,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl MessageContext {
|
impl MessageContext {
|
||||||
@@ -31,9 +33,20 @@ impl MessageContext {
|
|||||||
admin_whitelist_regexes,
|
admin_whitelist_regexes,
|
||||||
trigger_event_info,
|
trigger_event_info,
|
||||||
thread_info,
|
thread_info,
|
||||||
|
|
||||||
|
bot_display_name: None,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
pub fn with_bot_display_name(mut self, value: Option<String>) -> Self {
|
||||||
|
self.bot_display_name = value;
|
||||||
|
self
|
||||||
|
}
|
||||||
|
|
||||||
|
pub fn bot_display_name(&self) -> &Option<String> {
|
||||||
|
&self.bot_display_name
|
||||||
|
}
|
||||||
|
|
||||||
pub fn room(&self) -> &Room {
|
pub fn room(&self) -> &Room {
|
||||||
&self.room
|
&self.room
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -6,10 +6,28 @@ use mxlink::matrix_sdk::ruma::{OwnedEventId, OwnedUserId};
|
|||||||
use mxlink::ThreadInfo;
|
use mxlink::ThreadInfo;
|
||||||
|
|
||||||
/// MessagePayload is like matrix-sdk's MessageType, but represents only message types that the bot deals with and payloads are massaged a bit.
|
/// MessagePayload is like matrix-sdk's MessageType, but represents only message types that the bot deals with and payloads are massaged a bit.
|
||||||
|
///
|
||||||
|
/// This also includes a few synthetic events.
|
||||||
#[derive(Debug, Clone)]
|
#[derive(Debug, Clone)]
|
||||||
pub enum MessagePayload {
|
pub enum MessagePayload {
|
||||||
Text(TextMessageEventContent),
|
/// A synthetic message payload that indicates that the bot should produce a reply inside a thread.
|
||||||
|
/// This does not represent an actual message event, it's just a way to trigger a chat completion.
|
||||||
|
///
|
||||||
|
/// When this is invoked, the ThreadInfo contains the full thread details (which represents our context).
|
||||||
|
///
|
||||||
|
/// See: https://github.com/etkecc/baibot/issues/15
|
||||||
|
SynthethicChatCompletionTriggerInThread,
|
||||||
|
|
||||||
|
/// A synthetic message payload that indicates that the bot should produce a reply to a specific message.
|
||||||
|
/// This does not represent an actual message event, it's just a way to trigger a chat completion.
|
||||||
|
///
|
||||||
|
/// When this is invoked, the ThreadInfo would refer to the reply-message that triggered us.
|
||||||
|
/// We can follow the chain upward from it to get the full context.
|
||||||
|
///
|
||||||
|
/// See: https://github.com/etkecc/baibot/issues/15
|
||||||
|
SynthethicChatCompletionTriggerForReply,
|
||||||
|
|
||||||
|
Text(TextMessageEventContent),
|
||||||
Audio(AudioMessageEventContent),
|
Audio(AudioMessageEventContent),
|
||||||
|
|
||||||
Reaction {
|
Reaction {
|
||||||
|
|||||||
@@ -1,15 +1,15 @@
|
|||||||
pub mod catch_up_marker;
|
pub mod catch_up_marker;
|
||||||
pub mod cfg;
|
pub mod cfg;
|
||||||
pub mod globalconfig;
|
pub mod globalconfig;
|
||||||
|
mod interaction_context;
|
||||||
mod message_context;
|
mod message_context;
|
||||||
mod message_payload;
|
mod message_payload;
|
||||||
mod room_config_context;
|
mod room_config_context;
|
||||||
pub mod roomconfig;
|
pub mod roomconfig;
|
||||||
mod thread_context;
|
|
||||||
mod trigger_event_info;
|
mod trigger_event_info;
|
||||||
|
|
||||||
|
pub use interaction_context::{InteractionContext, InteractionTrigger};
|
||||||
pub use message_context::MessageContext;
|
pub use message_context::MessageContext;
|
||||||
pub use message_payload::MessagePayload;
|
pub use message_payload::MessagePayload;
|
||||||
pub use room_config_context::RoomConfigContext;
|
pub use room_config_context::RoomConfigContext;
|
||||||
pub use thread_context::{ThreadContext, ThreadContextFirstMessage};
|
|
||||||
pub use trigger_event_info::TriggerEventInfo;
|
pub use trigger_event_info::TriggerEventInfo;
|
||||||
|
|||||||
@@ -1,13 +0,0 @@
|
|||||||
use mxlink::ThreadInfo;
|
|
||||||
|
|
||||||
use super::MessagePayload;
|
|
||||||
|
|
||||||
pub struct ThreadContext {
|
|
||||||
pub info: ThreadInfo,
|
|
||||||
pub first_message: ThreadContextFirstMessage,
|
|
||||||
}
|
|
||||||
|
|
||||||
pub struct ThreadContextFirstMessage {
|
|
||||||
pub is_mentioning_bot: bool,
|
|
||||||
pub payload: MessagePayload,
|
|
||||||
}
|
|
||||||
@@ -5,12 +5,18 @@ use super::text::{block_quote, block_unquote};
|
|||||||
/// Creates a text message which is based on transcribed audio.
|
/// Creates a text message which is based on transcribed audio.
|
||||||
/// This text message is prefixed with an emoji and blockquoted, to indicate that it is a transcription.
|
/// This text message is prefixed with an emoji and blockquoted, to indicate that it is a transcription.
|
||||||
/// To reverse the process, use `parse_transcribed_message_text()`.
|
/// To reverse the process, use `parse_transcribed_message_text()`.
|
||||||
|
///
|
||||||
|
/// It should be noted that in certain cases (Transcribe-only mode), transcriptions are posted as regular notice messages which do not include
|
||||||
|
/// the `> 🦻` prefixing. That is, not every transcribed message will pass through here (intentionally).
|
||||||
pub fn create_transcribed_message_text(text: &str) -> String {
|
pub fn create_transcribed_message_text(text: &str) -> String {
|
||||||
block_quote(&format!("{} {}", AgentPurpose::SpeechToText.emoji(), text))
|
block_quote(&format!("{} {}", AgentPurpose::SpeechToText.emoji(), text))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Parses a transcribed message text, reversing the process done by `create_transcribed_message_text()`.
|
/// Parses a transcribed message text, reversing the process done by `create_transcribed_message_text()`.
|
||||||
/// If the provided text string does not match the expected format, None is returned.
|
/// If the provided text string does not match the expected format, None is returned.
|
||||||
|
///
|
||||||
|
/// It should be noted that in certain cases (Transcribe-only mode), transcriptions are posted as regular notice messages which do not include
|
||||||
|
/// the `> 🦻` prefixing. This function will not handle these properly.
|
||||||
pub fn parse_transcribed_message_text(text: &str) -> Option<String> {
|
pub fn parse_transcribed_message_text(text: &str) -> Option<String> {
|
||||||
if !text.starts_with("> ") {
|
if !text.starts_with("> ") {
|
||||||
return None;
|
return None;
|
||||||
|
|||||||
Reference in New Issue
Block a user