Compare commits

...

9 Commits

Author SHA1 Message Date
Slavi Pantaleev
3b7c28a55e Release 1.0.5 2024-09-14 10:47:56 +03:00
Slavi Pantaleev
3b25b92a81 Implement more fine-grained typing notices sending
The previous approach (implemented in dd1dd78312) was simple
(send typing notices for as long as the "controller" is running),
but this proved to be overly simplistic and unable to handle edge-cases:

- in multi-user rooms (or rooms with a prefix requirement), the bot
  used to send a typing notice while "working", but its work consisted
  of ignoring the message. So it then sent a "not typing" notice.
  This is wasteful and otherwise problematic - certain clients (like nheko)
  do not handle this "race" well.

- certain reactions (anything other than 🗣️ right now) are meant to be
  ignored. There's no point in doing the same "typing / not typing"
  dance

- there are other instances where the bot may do work, but doesn't (due
  to configuration or lack of capabilities)

This new more fine-grained implementation of typing notices aims to:

- only send a typing notice if actual "slow work" will be done

- avoid stopping & restarting typing notices (wasteful) if a chain of work is to
  be performed (processing voice messages and doing speech-to-text +
  text-generation + ...). Rather, maintaining typing notice sending
  throughout
2024-09-14 10:39:20 +03:00
Slavi Pantaleev
509f683365 Fix typo 2024-09-14 09:30:58 +03:00
Slavi Pantaleev
a986e29f51 Release 1.0.4 2024-09-13 21:46:15 +03:00
Slavi Pantaleev
dd1dd78312 Rework typing notifications
Previously, the bot only had rudimentary typing notification support.

It used to send a single notification when starting a long task
and did not bother with notifications anymore.
By default matrix-rust-sdk gives these notifications a validity of 4
seconds, so it would expire shortly. If the bot takes longer to respond,
you'd see the typing notification expire and wonder if a response is
coming.

Another edge case is the bot sending an answer quicker and the typing
notice still being on. Some clients (like element-web) seem to hide the
typing notice when a new message comes, so they don't experience this as
problematic.

The reworked typing notification system should be robust:

- typing notices are sent continuously, until the bot finishes doing
  work
- if the bot is performing multiple actions in a room (even for
  different people), typing notices would continue to be sent until the
  bot becomes idle
- as soon as the bot becomes idle, a "not typing anymore" notice is sent
  to clear the state
2024-09-13 21:40:24 +03:00
Slavi Pantaleev
f2b1115dc9 Populate CHANGELOG 2024-09-13 21:39:19 +03:00
Slavi Pantaleev
5742d88d45 Release 1.0.3 2024-09-13 12:33:35 +03:00
Slavi Pantaleev
1be035d94c Upgrade mxlink (1.0.0 -> 1.1.0)
This brings in a fix that allows for auto-recovery from errors
that occur during matrix-rust-sdk startup.
2024-09-13 11:38:07 +03:00
Slavi Pantaleev
601420d561 Fix broken links to "sample provider configs"
[skip ci]
2024-09-12 22:18:51 +03:00
10 changed files with 60 additions and 30 deletions

View File

@@ -1 +1,18 @@
There's nothing here yet.
# (2024-09-14) Version 1.0.5
Further [improves](https://github.com/etkecc/baibot/commit/3b25b92a81a05ebaf1c6dbabf675fbfbe6c9f418) the typing notification logic, so that it tolerates edge cases better.
# (2024-09-14) Version 1.0.4
[Improves](https://github.com/etkecc/baibot/commit/dd1dd78312e3db7f92b37fb3b4750fbe35de7115) the typing notification logic.
# (2024-09-13) Version 1.0.3
Contains [fixes](https://github.com/etkecc/rust-mxlink/commit/f339fc85e69aa7f614394ad303d1614cd307319c) for [some](https://github.com/etkecc/baibot/issues/1) startup failures caused by partial initialization (errors during startup).
# (2024-09-12) Version 1.0.0
Initial release. 🎉

6
Cargo.lock generated
View File

@@ -297,7 +297,7 @@ dependencies = [
[[package]]
name = "baibot"
version = "1.0.2"
version = "1.0.5"
dependencies = [
"anthropic-rs",
"anyhow",
@@ -2122,9 +2122,9 @@ dependencies = [
[[package]]
name = "mxlink"
version = "1.0.0"
version = "1.2.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "80cb97e517e3bc8e85b4787ed87037cd324526862ca9375c6d038699010bf6c2"
checksum = "f1aa513435471af3e1131dc18a1f4eb2e6fb42c11308f1be9b4f0316b3809e32"
dependencies = [
"base64 0.22.1",
"chacha20poly1305",

View File

@@ -7,7 +7,7 @@ license = "AGPL-3.0-or-later"
readme = "README.md"
keywords = ["matrix", "chat", "bot", "AI", "LLM"]
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
version = "1.0.2"
version = "1.0.5"
edition = "2021"
[lib]
@@ -22,7 +22,7 @@ base64 = "0.22.*"
# We'd rather not depend on this, but we cannot use the ruma-events EventContent macro without it.
matrix-sdk = { version = "0.7.1", default-features = false }
mxidwc = "1.0.*"
mxlink = "1.0.*"
mxlink = ">=1.2.1"
etke_openai_api_rust = "0.1.*"
quick_cache = "0.6.*"
regex = "1.10.*"

View File

@@ -49,7 +49,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- create a room-local agent: `!bai agent create-room-local anthropic my-anthropic-agent`
- create a global agent: `!bai agent create-global anthropic my-anthropic-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/anthropic.yml).
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/anthropic.yml).
### Groq
@@ -63,7 +63,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- create a room-local agent: `!bai agent create-room-local groq my-groq-agent`
- create a global agent: `!bai agent create-global groq my-groq-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/groq.yml).
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/groq.yml).
### LocalAI
@@ -77,7 +77,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- create a room-local agent: `!bai agent create-room-local localai my-localai-agent`
- create a global agent: `!bai agent create-global localai my-localai-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/localai.yml).
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/localai.yml).
### Mistral
@@ -91,7 +91,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- create a room-local agent: `!bai agent create-room-local mistral my-mistral-agent`
- create a global agent: `!bai agent create-global mistral my-mistral-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/mistral.yml).
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/mistral.yml).
### Ollama
@@ -105,7 +105,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- create a room-local agent: `!bai agent create-room-local ollama my-ollama-agent`
- create a global agent: `!bai agent create-global ollama my-ollama-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/ollama.yml).
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/ollama.yml).
### OpenAI
@@ -122,7 +122,7 @@ For services which are not fully compatible with the OpenAI API, consider using
- create a room-local agent: `!bai agent create-room-local openai my-openai-agent`
- create a global agent: `!bai agent create-global openai my-openai-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/openai.yml).
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/openai.yml).
### OpenAI Compatible
@@ -139,7 +139,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- create a room-local agent: `!bai agent create-room-local openai-compatible my-openai-compatible-agent`
- create a global agent: `!bai agent create-global openai-compatible my-openai-compatible-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/openai-compatible.yml).
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/openai-compatible.yml).
### OpenRouter
@@ -153,7 +153,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- create a room-local agent: `!bai agent create-room-local openrouter my-openrouter-agent`
- create a global agent: `!bai agent create-global openrouter my-openrouter-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/openrouter.yml).
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/openrouter.yml).
### Together AI
@@ -167,4 +167,4 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- create a room-local agent: `!bai agent create-room-local together-ai my-together-ai-agent`
- create a global agent: `!bai agent create-global together-ai my-together-ai-agent`
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/together-ai.yml).
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](./sample-provider-configs/together-ai.yml).

View File

@@ -9,6 +9,7 @@ use mxlink::matrix_sdk::Room;
use mxlink::{
InitConfig, LoginConfig, LoginCredentials, LoginEncryption, MatrixLink, PersistenceConfig,
TypingNoticeGuard,
};
use mxlink::helpers::account_data_config::{
@@ -210,6 +211,14 @@ impl Bot {
.await
}
pub(crate) async fn start_typing_notice(&self, room: &Room) -> TypingNoticeGuard {
self.inner
.matrix_link
.rooms()
.start_typing_notice(room)
.await
}
pub async fn start(&self) -> anyhow::Result<()> {
self.rooms().attach_event_handlers().await;
self.messaging().attach_event_handlers().await;

View File

@@ -100,8 +100,6 @@ pub async fn handle_room_local(
return Ok(());
};
message_context.room().typing_notice(true).await?;
if !try_to_ping_agent_or_complain(bot, message_context, &parsed_config.agent).await {
return Ok(());
}
@@ -215,8 +213,6 @@ pub async fn handle_global(
return Ok(());
};
message_context.room().typing_notice(true).await?;
if !try_to_ping_agent_or_complain(bot, message_context, &parsed_config.agent).await {
return Ok(());
}

View File

@@ -43,6 +43,8 @@ pub async fn handle(
) -> anyhow::Result<()> {
let mut original_message_is_audio = false;
let mut _typing_notice_guard: Option<mxlink::TypingNoticeGuard> = None;
let speech_to_text_flow_type = message_context
.room_config_context()
.speech_to_text_flow_type();
@@ -58,11 +60,11 @@ pub async fn handle(
return Ok(());
}
SpeechToTextFlowType::TranscribeAndGenerateText => {
tracing::debug!("Will be trascribing and possibly generating text..");
tracing::debug!("Will be transcribing and possibly generating text..");
MessageResponseType::InThread(message_context.thread_info().clone())
}
SpeechToTextFlowType::OnlyTranscribe => {
tracing::debug!("Will only be trascribing audio to text..");
tracing::debug!("Will only be transcribing audio to text..");
if message_context.thread_info().is_thread_root_only() {
MessageResponseType::Reply(message_context.thread_info().root_event_id.clone())
} else {
@@ -71,6 +73,10 @@ pub async fn handle(
}
};
if _typing_notice_guard.is_none() {
_typing_notice_guard = Some(bot.start_typing_notice(message_context.room()).await);
}
let Some(speech_to_text_created_event_id_result) =
handle_stage_speech_to_text(bot, message_context, audio_content, response_type).await
else {
@@ -96,6 +102,10 @@ pub async fn handle(
.room_config_context()
.should_auto_text_generate(original_message_is_audio)
{
if _typing_notice_guard.is_none() {
_typing_notice_guard = Some(bot.start_typing_notice(message_context.room()).await);
}
let speech_to_text_created_event_id_reaction_event_id =
if let Some(speech_to_text_created_event_id) = speech_to_text_created_event_id {
let reaction_event_response = bot
@@ -213,6 +223,10 @@ pub async fn handle(
match text_to_speech_stage_params {
Some(TextToSpeechParams::Perform(text_to_speech_eligible_payload, response_type)) => {
if _typing_notice_guard.is_none() {
_typing_notice_guard = Some(bot.start_typing_notice(message_context.room()).await);
}
let _tts_result = generate_and_send_tts_for_message(
bot,
matrix_link.clone(),
@@ -337,8 +351,6 @@ async fn handle_stage_text_generation(
)
.await?;
_ = message_context.room().typing_notice(true).await;
let prefixes_to_strip = match controller_type {
ChatCompletionControllerType::ViaText { prefixes_to_strip } => prefixes_to_strip.clone(),
ChatCompletionControllerType::ViaAudio => vec![],
@@ -498,8 +510,6 @@ async fn handle_stage_speech_to_text_actual_transcribing(
.get_media_content(&media_request, true)
.await?;
_ = message_context.room().typing_notice(true).await;
let span = tracing::debug_span!(
"speech_to_text_generation",
agent_id = agent.identifier().as_string()

View File

@@ -56,8 +56,6 @@ pub async fn handle_image(
original_prompt.to_owned()
};
message_context.room().typing_notice(true).await?;
let span = tracing::debug_span!(
"image_generation",
agent_id = agent.identifier().as_string()
@@ -139,7 +137,7 @@ pub async fn handle_sticker(
return Ok(());
};
message_context.room().typing_notice(true).await?;
let _typing_notice_guard = bot.start_typing_notice(message_context.room()).await;
let span = tracing::debug_span!(
"sticker_generation",

View File

@@ -52,6 +52,8 @@ pub(super) async fn handle(
return Ok(());
};
let _typing_notice_guard = bot.start_typing_notice(message_context.room()).await;
crate::controller::utils::text_to_speech::generate_and_send_tts_for_message(
bot,
matrix_link,

View File

@@ -18,8 +18,6 @@ pub async fn generate_and_send_tts_for_message(
text_message_event_id: &OwnedEventId,
text_content: &str,
) -> bool {
_ = message_context.room().typing_notice(true).await;
let reaction_event_response = bot
.reacting()
.react_no_fail(