Compare commits

..

13 Commits

Author SHA1 Message Date
Slavi Pantaleev
68a2fb161f Release 1.7.4 2025-06-10 16:38:03 +03:00
Slavi Pantaleev
3a3eb58d7b Update Cargo.lock-pinned dependencies 2025-06-10 16:37:13 +03:00
Slavi Pantaleev
74d988e650 Update mxlink (1.8.0 -> 1.8.1) 2025-06-10 16:33:44 +03:00
Slavi Pantaleev
ed8bedcd7e Release 1.7.3 2025-06-10 15:28:56 +03:00
Slavi Pantaleev
2842632969 Disable async-openai feature of tiktoken-rs crate
We don't make use of this feature, so it's not necessary.

Ref: af52be12fa/tiktoken-rs/Cargo.toml (L21)
2025-06-10 15:26:25 +03:00
Slavi Pantaleev
10a5bd2abb Update OpenAI model in sample config (gpt-4o -> gpt-4.1) 2025-06-10 15:20:00 +03:00
Slavi Pantaleev
5308b75f52 Update dependencies 2025-06-10 15:18:58 +03:00
Slavi Pantaleev
dad61e1270 Adjust docker run command on installation instructions to ensure /tmp is writable
Ref: https://github.com/etkecc/baibot/issues/43
2025-06-09 10:54:11 +03:00
Slavi Pantaleev
91986a129c Release 1.7.2 2025-05-11 23:20:58 +03:00
Slavi Pantaleev
264f683d6a Allow image_generation.size to be null for OpenAI and default it to that
The API spec for image creation and image editing says "string or null",
so we're allowing `null` now to trigger automatic selection.
2025-05-11 23:20:07 +03:00
Slavi Pantaleev
62f0f4fa0d Release 1.7.1 2025-05-11 22:20:26 +03:00
Slavi Pantaleev
69627abd74 Add image-editing feature documentation to the !bai usage command and adjust texts a bit 2025-05-11 22:19:15 +03:00
Slavi Pantaleev
d2660be33c Update sample config for OpenAI to use gpt-image-1, not dall-e-3 2025-05-10 12:30:16 +03:00
13 changed files with 732 additions and 594 deletions

View File

@@ -1,3 +1,19 @@
# (2025-06-10) Version 1.7.4
- (**Internal Improvement**) Dependency updates.
# (2025-06-10) Version 1.7.3
- (**Internal Improvement**) Dependency updates. This version is based on [mxlink](https://crates.io/crates/mxlink)@1.8.0 (which is based on the newly released [matrix-sdk](https://crates.io/crates/matrix-sdk)@[0.12.0](https://github.com/matrix-org/matrix-rust-sdk/releases/tag/matrix-sdk-0.12.0), which contains fixes for important security vulnerabilities)
# (2025-05-11) Version 1.7.2
- (**Bugfix**) Allow `image_generation.size` configuration value for OpenAI to be `null` to allow the model to choose the size automatically and default to that
# (2025-05-11) Version 1.7.1
- (**Bugfix**) Fix lack of documentation for the new [image-editing](./docs/features.md#-image-editing) feature in the `!bai usage` command's output
# (2025-05-10) Version 1.7.0
- (**Feature**) Add vision support to the OpenAI and Anthropic providers. You can now mix text and images in your conversations - fixes [issue #5](https://github.com/etkecc/baibot/issues/5)

1222
Cargo.lock generated

File diff suppressed because it is too large Load Diff

View File

@@ -7,7 +7,7 @@ license = "AGPL-3.0-or-later"
readme = "README.md"
keywords = ["matrix", "chat", "bot", "AI", "LLM"]
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
version = "1.7.0"
version = "1.7.4"
edition = "2024"
[lib]
@@ -22,17 +22,17 @@ base64 = "0.22.*"
chrono = { version = "0.4.*", default-features = false, features = ["std", "now"] }
# We'd rather not depend on this, but we cannot use the ruma-events EventContent macro without it.
# We add the `native-tls` feature, because of https://github.com/etkecc/rust-mxlink/issues/1
matrix-sdk = { version = "0.11.0", default-features = false, features = ["native-tls"] }
matrix-sdk = { version = "0.12.0", default-features = false, features = ["native-tls"] }
mxidwc = "1.0.*"
mxlink = ">=1.7.0"
mxlink = ">=1.8.1"
etke_openai_api_rust = "0.1.*"
quick_cache = "0.6.*"
regex = "1.11.*"
serde = { version = "1.0.*", features = ["derive"], default-features = false }
serde_json = "1.0.*"
serde_yaml = "0.9.*"
tempfile = "3.19.*"
tiktoken-rs = { version = "0.6.*", features = ["async-openai"] }
tempfile = "3.20.*"
tiktoken-rs = { version = "0.7.*", default-features = false }
tokio = { version = "1.45.*", features = ["rt", "rt-multi-thread", "macros"] }
tracing = "0.1.*"
tracing-subscriber = { version = "0.3.*", features = ["env-filter"] }

View File

@@ -53,6 +53,7 @@ CONTAINER_IMAGE_NAME=ghcr.io/etkecc/baibot:v1.0.0
--env BAIBOT_PERSISTENCE_DATA_DIR_PATH=/data \
--mount type=bind,src=/path/to/config.yml,dst=/app/config.yml,ro \
--mount type=bind,src=/path/to/data,dst=/data \
--tmpfs=/tmp:rw,noexec,nosuid,size=1024m \
$CONTAINER_IMAGE_NAME
```

View File

@@ -20,5 +20,5 @@ text_to_speech:
image_generation:
model_id: gpt-image-1
style: null
size: 1024x1024
size: null
quality: null

View File

@@ -16,5 +16,5 @@ text_to_speech:
image_generation:
model_id: gpt-image-1
style: null
size: 1024x1024
size: null
quality: null

View File

@@ -79,7 +79,7 @@ Simply send a command like `!bai image create A beautiful sunset over the ocean`
See a [🖼️ Screenshot of the Image Creation feature](./screenshots/image-creation.webp).
You can then, respond in the same message thread with:
You can then respond in the same message thread with:
- more messages, to add more criteria to your prompt.
- a message saying `again`, to generate one more image with the current prompt.
@@ -91,7 +91,7 @@ Simply send a command like `!bai image edit Turn the following image into an ani
See a [🖼️ Screenshot of the Image Editing feature (manipulating a single image)](./screenshots/image-editing-single-image.webp) and a [🖼️ Screenshot of the Image Editing feature (manipulating multiple images)](./screenshots/image-editing-multiple-images.webp).
You can then, respond in the same message thread with:
You can then respond in the same message thread with:
- more messages, to add more criteria to your prompt.
- one or more images, to provide the images that the bot will operate on.
@@ -101,7 +101,7 @@ You can then, respond in the same message thread with:
#### 🫵 Creating stickers
A variation of [creating images](#creating-images) is to create "sticker images".
A variation of [creating images](#creating-images) is creating "sticker images".
See a [🖼️ Screenshot of the Sticker Creation feature](./screenshots/sticker-generation.webp).

View File

@@ -76,7 +76,7 @@ agents:
# base_url: https://api.openai.com/v1
# api_key: ""
# text_generation:
# model_id: gpt-4o
# model_id: gpt-4.1
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
# temperature: 1.0
# max_response_tokens: 16384
@@ -91,10 +91,10 @@ agents:
# speed: 1.0
# response_format: opus
# image_generation:
# model_id: dall-e-3
# style: vivid
# size: 1024x1024
# quality: standard
# model_id: gpt-image-1
# style: null
# size: null
# quality: null
#
# - id: localai
# provider: localai

View File

@@ -153,7 +153,7 @@ pub struct ImageGenerationConfig {
pub style: Option<async_openai::types::ImageStyle>,
#[serde(default = "default_image_size")]
pub size: async_openai::types::ImageSize,
pub size: Option<async_openai::types::ImageSize>,
#[serde(default = "default_image_quality")]
pub quality: Option<async_openai::types::ImageQuality>,
@@ -186,8 +186,8 @@ fn default_image_style() -> Option<async_openai::types::ImageStyle> {
None
}
fn default_image_size() -> async_openai::types::ImageSize {
async_openai::types::ImageSize::S1024x1024
fn default_image_size() -> Option<async_openai::types::ImageSize> {
None
}
fn default_image_quality() -> Option<async_openai::types::ImageQuality> {

View File

@@ -284,11 +284,8 @@ impl ControllerTrait for Controller {
let size = params
.size_override
.map(|s| {
convert_string_to_enum::<async_openai::types::ImageSize>(&s)
.unwrap_or(image_generation_config.size)
})
.unwrap_or(image_generation_config.size);
.map(|s| convert_string_to_enum::<async_openai::types::ImageSize>(&s).unwrap())
.or(image_generation_config.size);
let response_format = match model.clone() {
ImageModel::DallE2 => Some(ImageResponseFormat::B64Json),
@@ -303,10 +300,7 @@ impl ControllerTrait for Controller {
let mut request_builder = CreateImageRequestArgs::default();
request_builder
.model(model)
.prompt(prompt.to_owned())
.size(size);
request_builder.model(model).prompt(prompt.to_owned());
if let Some(response_format) = response_format {
request_builder.response_format(response_format);
@@ -320,6 +314,10 @@ impl ControllerTrait for Controller {
request_builder.quality(quality.clone());
}
if let Some(size) = size {
request_builder.size(size);
}
let request = request_builder.build()?;
tracing::trace!(
@@ -382,9 +380,9 @@ impl ControllerTrait for Controller {
}
let dalle2_size = match image_generation_config.size {
async_openai::types::ImageSize::S256x256 => Some(DallE2ImageSize::S256x256),
async_openai::types::ImageSize::S512x512 => Some(DallE2ImageSize::S512x512),
async_openai::types::ImageSize::S1024x1024 => Some(DallE2ImageSize::S1024x1024),
Some(async_openai::types::ImageSize::S256x256) => Some(DallE2ImageSize::S256x256),
Some(async_openai::types::ImageSize::S512x512) => Some(DallE2ImageSize::S512x512),
Some(async_openai::types::ImageSize::S1024x1024) => Some(DallE2ImageSize::S1024x1024),
_ => None,
};

View File

@@ -224,9 +224,11 @@ impl TryInto<OpenAIImageGenerationConfig> for ImageGenerationConfig {
fn try_into(self) -> Result<OpenAIImageGenerationConfig, Self::Error> {
let size = if let Some(size) = &self.size {
convert_string_to_enum::<async_openai::types::ImageSize>(size)?
Some(convert_string_to_enum::<async_openai::types::ImageSize>(
size,
)?)
} else {
async_openai::types::ImageSize::S1024x1024
None
};
let style = if let Some(style) = &self.style {

View File

@@ -3,5 +3,5 @@ pub fn heading() -> &'static str {
}
pub fn intro() -> &'static str {
"The bot can perform various tasks, such as 💬 Text Generation, 🗣️ Text-to-Speech, 🦻 Speech-to-Text, 🖌️ Image Creation, and more."
"The bot can perform various tasks, such as 💬 Text Generation, 🗣️ Text-to-Speech, 🦻 Speech-to-Text, 🖌️ Image Generation, and more."
}

View File

@@ -34,20 +34,31 @@ By default, the bot will also perform 💬 Text Generation on the text. This is
If all your messages are in the same language, you can improve accuracy & latency by configuring the language via the **🦻 Speech-to-Text / 🔤 Language** setting.
### 🖌️ Image Creation
### Image Generation
#### Creating images
#### 🖌️ Creating images
Simply send a command like `%command_prefix% image create A beautiful sunset over the ocean` and the bot will start a threaded conversation and post an image based on your prompt.
You can then, respond in the same message thread with:
You can then respond in the same message thread with:
- more messages, to add more criteria to your prompt.
- a message saying `again`, to generate one more image with the current prompt.
#### Creating stickers
#### 🎨 Editing images
A variation of **creating images** is to create "sticker images".
Simply send a command like `%command_prefix% image edit Turn the following image into an anime-style drawing` and the bot will start a threaded conversation asking for more details.
You can then respond in the same message thread with:
- more messages, to add more criteria to your prompt.
- one or more images, to provide the images that the bot will operate on.
- a message saying `go`, to start the image generation process.
- a message saying `again`, to prompt the bot to generate one more image edit with the current prompt.
#### 🫵 Creating stickers
A variation of **creating images** is creating "sticker images".
To create a sticker, send a command like `%command_prefix% sticker A huge bowl of steaming ramen with a mountain of beansprouts on top`.