Compare commits
32 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
61d18b2e13 | ||
|
|
d831c08306 | ||
|
|
c70387b0c3 | ||
|
|
38516f2e17 | ||
|
|
5481b5a763 | ||
|
|
26bc437678 | ||
|
|
ec93f1ee2a | ||
|
|
b920b6e556 | ||
|
|
7136d34843 | ||
|
|
e0b4a40dd8 | ||
|
|
691aeeb1c7 | ||
|
|
257ffae9e7 | ||
|
|
f7bf3d7b60 | ||
|
|
3a88b0d656 | ||
|
|
08c689a889 | ||
|
|
ae8e878817 | ||
|
|
edbd72ece6 | ||
|
|
22906aa2d3 | ||
|
|
99bde53ef6 | ||
|
|
b3fd8e548f | ||
|
|
062fbbb8ef | ||
|
|
2801c78ad9 | ||
|
|
f4c698ad33 | ||
|
|
bd39001417 | ||
|
|
2692d0322e | ||
|
|
8eb70f0f2c | ||
|
|
5c0a7be7a2 | ||
|
|
0a8f9fc3e5 | ||
|
|
1ac3b2e060 | ||
|
|
a3ef9fd1bf | ||
|
|
df507eb201 | ||
|
|
4dcd9eff40 |
52
CHANGELOG.md
52
CHANGELOG.md
@@ -1,3 +1,55 @@
|
|||||||
|
# (2026-02-04) Version 1.14.0
|
||||||
|
|
||||||
|
- (**Feature**) The `openai` provider now uses OpenAI's [Responses API](https://platform.openai.com/docs/api-reference/responses) (instead of the older Chat Completions API), adding support for [🛠️ built-in tools](./docs/features.md#️-built-in-tools-openai-only) (`web_search` and `code_interpreter`). These tools are **disabled by default** and can be enabled via the `text_generation.tools` configuration (see the [sample configuration](https://github.com/etkecc/baibot/blob/c70387b0c38d8d0f30bba2179a2a21a3710dbeaf/docs/sample-provider-configs/openai.yml#L12-L15)). To enable tools on an existing agent, you need to [update the agent](./docs/agents.md#updating-agents) to re-create it with the `text_generation.tools` section added and enable the tools you need. Thanks to [Layla Manley](https://github.com/yeslayla) for the contribution in [#62](https://github.com/etkecc/baibot/pull/62)!
|
||||||
|
|
||||||
|
- (**Bugfix**) Fix sticker generation for newer GPT image models (`gpt-image-1`, `gpt-image-1-mini`, `gpt-image-1.5`) which don't support the previously hardcoded `256x256` size (minimum is `1024x1024`)
|
||||||
|
|
||||||
|
- (**Internal Improvement**) Dependency updates
|
||||||
|
|
||||||
|
|
||||||
|
# (2026-01-23) Version 1.13.0
|
||||||
|
|
||||||
|
- (**Improvement**) Extend auto-switching to support cheaper models (`gpt-image-1-mini`) for `gpt-image-1` and `gpt-image-1.5` when generating stickers ([e0b4a40](https://github.com/etkecc/baibot/commit/e0b4a40))
|
||||||
|
|
||||||
|
- (**Internal Improvement**) Upgrade Rust compiler (1.92.0 -> 1.93.0) ([691aeeb](https://github.com/etkecc/baibot/commit/691aeeb))
|
||||||
|
|
||||||
|
- (**Internal Improvement**) Dependency updates
|
||||||
|
|
||||||
|
|
||||||
|
# (2025-12-21) Version 1.12.0
|
||||||
|
|
||||||
|
- (**Improvement**) Upgrade [async-openai](https://crates.io/crates/async-openai) (0.31.1 -> 0.32.2) and add support for OpenAI's `gpt-image-1.5` model ([08c689a](https://github.com/etkecc/baibot/commit/08c689a), [f7bf3d7](https://github.com/etkecc/baibot/commit/f7bf3d7))
|
||||||
|
|
||||||
|
- (**Internal Improvement**) Dependency updates
|
||||||
|
|
||||||
|
|
||||||
|
# (2025-12-15) Version 1.11.0
|
||||||
|
|
||||||
|
- (**Feature**) Add support for custom avatars via file path and for keeping the already-set avatar (for those who wish to manage it by themselves via other means). See the [sample config](./etc/app/config.yml.dist) for details. ([062fbbb](https://github.com/etkecc/baibot/commit/062fbbb8ef9ad600db483a431c5c782402191023))
|
||||||
|
|
||||||
|
- (**Internal Improvement**) Dependency updates ([99bde53](https://github.com/etkecc/baibot/commit/99bde53ef648a5a9086a96778fde4a9dbc1ede58))
|
||||||
|
|
||||||
|
- (**Internal Improvement**) Documentation updates ([b3fd8e5](https://github.com/etkecc/baibot/commit/b3fd8e548f83fe46398ced4760d7e2bb7588c24d))
|
||||||
|
|
||||||
|
- (**Internal Improvement**) Upgrade Rust compiler (1.91.1 -> 1.92.0) ([22906aa](https://github.com/etkecc/baibot/commit/22906aa2d3cae51815fad2560a545eaa69c247b6))
|
||||||
|
|
||||||
|
|
||||||
|
# (2025-12-06) Version 1.10.0
|
||||||
|
|
||||||
|
- (**Internal Improvement**) Dependency updates. This version is based on [mxlink](https://crates.io/crates/mxlink)@1.11.0 (which is based on the newly released [matrix-sdk](https://crates.io/crates/matrix-sdk)@[0.16.0](https://github.com/matrix-org/matrix-rust-sdk/releases/tag/matrix-sdk-0.16.0).
|
||||||
|
|
||||||
|
# (2025-11-30) Version 1.9.0
|
||||||
|
|
||||||
|
- (**Internal Improvement**) Upgrade [async-openai](https://crates.io/crates/async-openai) from our own etkecc fork (0.28.1-patched) to the official upstream version 0.31.1. This upgrade required some code adaptations to the new module structure, etc. While tested, regressions are possible.
|
||||||
|
|
||||||
|
# (2025-11-28) Version 1.8.3
|
||||||
|
|
||||||
|
- (**Improvement**) Add support for the `BAIBOT_PERSISTENCE_SESSION_ENCRYPTION_KEY` environment variable for configuring `persistence.session_encryption_key`
|
||||||
|
|
||||||
|
- (**Improvement**) Add support for the `BAIBOT_USER_ENCRYPTION_RECOVERY_RESET_ALLOWED` environment variable for configuring `user.encryption.recovery_reset_allowed`
|
||||||
|
|
||||||
|
- (**Internal Improvement**) Dependency updates.
|
||||||
|
|
||||||
# (2025-11-20) Version 1.8.2
|
# (2025-11-20) Version 1.8.2
|
||||||
|
|
||||||
- (**Internal Improvement**) Dependency and compiler updates (Rust 1.89.0 -> 1.91.1).
|
- (**Internal Improvement**) Dependency and compiler updates (Rust 1.89.0 -> 1.91.1).
|
||||||
|
|||||||
1498
Cargo.lock
generated
1498
Cargo.lock
generated
File diff suppressed because it is too large
Load Diff
13
Cargo.toml
13
Cargo.toml
@@ -7,7 +7,7 @@ license = "AGPL-3.0-or-later"
|
|||||||
readme = "README.md"
|
readme = "README.md"
|
||||||
keywords = ["matrix", "chat", "bot", "AI", "LLM"]
|
keywords = ["matrix", "chat", "bot", "AI", "LLM"]
|
||||||
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
|
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
|
||||||
version = "1.8.2"
|
version = "1.14.0"
|
||||||
edition = "2024"
|
edition = "2024"
|
||||||
|
|
||||||
[lib]
|
[lib]
|
||||||
@@ -17,23 +17,24 @@ path = "src/lib.rs"
|
|||||||
[dependencies]
|
[dependencies]
|
||||||
anthropic = { git = "https://github.com/etkecc/anthropic-rs.git", branch = "fix-content-block-image" }
|
anthropic = { git = "https://github.com/etkecc/anthropic-rs.git", branch = "fix-content-block-image" }
|
||||||
anyhow = "1.0.*"
|
anyhow = "1.0.*"
|
||||||
async-openai = { git = "https://github.com/etkecc/async-openai", branch = "async-openai-v0.28.1-patched" }
|
async-openai = { version = "0.32.3", features = ["audio", "chat-completion", "image", "responses"] }
|
||||||
base64 = "0.22.*"
|
base64 = "0.22.*"
|
||||||
chrono = { version = "0.4.*", default-features = false, features = ["std", "now"] }
|
chrono = { version = "0.4.*", default-features = false, features = ["std", "now"] }
|
||||||
# We'd rather not depend on this, but we cannot use the ruma-events EventContent macro without it.
|
# We'd rather not depend on this, but we cannot use the ruma-events EventContent macro without it.
|
||||||
# We add the `native-tls` feature, because of https://github.com/etkecc/rust-mxlink/issues/1
|
# We add the `native-tls` feature, because of https://github.com/etkecc/rust-mxlink/issues/1
|
||||||
matrix-sdk = { version = "0.14.0", default-features = false, features = ["native-tls"] }
|
matrix-sdk = { version = "0.16.0", default-features = false, features = ["native-tls"] }
|
||||||
|
mime_guess = "2.0.*"
|
||||||
mxidwc = "1.0.*"
|
mxidwc = "1.0.*"
|
||||||
mxlink = ">=1.10.0"
|
mxlink = ">=1.11.0"
|
||||||
etke_openai_api_rust = "0.1.*"
|
etke_openai_api_rust = "0.1.*"
|
||||||
quick_cache = "0.6.*"
|
quick_cache = "0.6.*"
|
||||||
regex = "1.12.*"
|
regex = "1.12.*"
|
||||||
serde = { version = "1.0.*", features = ["derive"], default-features = false }
|
serde = { version = "1.0.*", features = ["derive"], default-features = false }
|
||||||
serde_json = "1.0.*"
|
serde_json = "1.0.*"
|
||||||
serde_yaml = "0.9.*"
|
serde_yaml = "0.9.*"
|
||||||
tempfile = "3.23.*"
|
tempfile = "3.24.*"
|
||||||
tiktoken-rs = { version = "0.9.*", default-features = false }
|
tiktoken-rs = { version = "0.9.*", default-features = false }
|
||||||
tokio = { version = "1.48.*", features = ["rt", "rt-multi-thread", "macros"] }
|
tokio = { version = "1.49.*", features = ["rt", "rt-multi-thread", "macros"] }
|
||||||
tracing = "0.1.*"
|
tracing = "0.1.*"
|
||||||
tracing-subscriber = { version = "0.3.*", features = ["env-filter"] }
|
tracing-subscriber = { version = "0.3.*", features = ["env-filter"] }
|
||||||
url = "2.5.*"
|
url = "2.5.*"
|
||||||
|
|||||||
@@ -4,7 +4,7 @@
|
|||||||
# #
|
# #
|
||||||
#######################################
|
#######################################
|
||||||
|
|
||||||
FROM docker.io/rust:1.91.1-slim-trixie AS build
|
FROM docker.io/rust:1.93.0-slim-trixie AS build
|
||||||
|
|
||||||
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev
|
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev
|
||||||
|
|
||||||
|
|||||||
@@ -4,7 +4,7 @@
|
|||||||
# #
|
# #
|
||||||
#######################################
|
#######################################
|
||||||
|
|
||||||
FROM docker.io/rust:1.91.1-slim-trixie AS build
|
FROM docker.io/rust:1.93.0-slim-trixie AS build
|
||||||
|
|
||||||
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev
|
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev
|
||||||
|
|
||||||
|
|||||||
@@ -17,7 +17,7 @@ It's influenced by [chaz](https://github.com/arcuru/chaz), but does **not** use
|
|||||||
|
|
||||||
- Supports **different use purposes** (depending on the [☁️ provider](./docs/providers.md) & model):
|
- Supports **different use purposes** (depending on the [☁️ provider](./docs/providers.md) & model):
|
||||||
|
|
||||||
- [💬 text-generation](./docs/features.md#-text-generation): communicating with you via text (though certain models may "see" images as well)
|
- [💬 text-generation](./docs/features.md#-text-generation): communicating with you via text (though certain models may "see" images as well). The [OpenAI provider](./docs/providers.md#openai) also supports [🛠️ built-in tools](./docs/features.md#️-built-in-tools-openai-only) (web search, code interpreter)
|
||||||
- [🦻 speech-to-text](./docs/features.md#-speech-to-text): turning your voice messages into text
|
- [🦻 speech-to-text](./docs/features.md#-speech-to-text): turning your voice messages into text
|
||||||
- [🗣️ text-to-speech](./docs/features.md#%EF%B8%8F-text-to-speech): turning bot or users text messages into voice messages
|
- [🗣️ text-to-speech](./docs/features.md#%EF%B8%8F-text-to-speech): turning bot or users text messages into voice messages
|
||||||
- [🖌️ image-generation](./docs/features.md#image-generation): creating and editing images based on instructions
|
- [🖌️ image-generation](./docs/features.md#image-generation): creating and editing images based on instructions
|
||||||
|
|||||||
@@ -43,7 +43,8 @@ Administrators cannot be changed without adjusting the bot's configuration on th
|
|||||||
|
|
||||||
Room-local agent managers are users privileged to **create their own [agents](./agents.md)** (see `!bai agent`) in rooms.
|
Room-local agent managers are users privileged to **create their own [agents](./agents.md)** (see `!bai agent`) in rooms.
|
||||||
|
|
||||||
**⚠️ WARNING**: Letting regular users create agents which contact arbitrary network services **may be a security issue**.
|
> [!WARNING]
|
||||||
|
> Letting regular users create agents which contact arbitrary network services **may be a security issue**.
|
||||||
|
|
||||||
The following commands are available:
|
The following commands are available:
|
||||||
- **Show** the currently allowed users: `!bai access room-local-agent-managers`
|
- **Show** the currently allowed users: `!bai access room-local-agent-managers`
|
||||||
|
|||||||
@@ -12,12 +12,15 @@ This file is created from the template found in [etc/app/config.yml.dist](../../
|
|||||||
|
|
||||||
Certain keys can be left unset, in which case [📝 hardcoded defaults](../../src/entity/cfg/defaults.rs) would be used.
|
Certain keys can be left unset, in which case [📝 hardcoded defaults](../../src/entity/cfg/defaults.rs) would be used.
|
||||||
|
|
||||||
Each configuration key found in the YAML configuration can be overridden by setting an environment variable (dots should be replaced with `_`). Example:
|
Some configuration keys found in the YAML configuration can be overridden by setting an environment variable (dots should be replaced with `_`). Example:
|
||||||
|
|
||||||
- to override `command_prefix`, set an environment variable `BAIBOT_COMMAND_PREFIX`
|
- to override `command_prefix`, set an environment variable `BAIBOT_COMMAND_PREFIX`
|
||||||
- to override `homeserver.server_name`, set an environment variable `BAIBOT_HOMESERVER_SERVER_NAME`
|
- to override `homeserver.server_name`, set an environment variable `BAIBOT_HOMESERVER_SERVER_NAME`
|
||||||
|
|
||||||
The static configuration contains an `initial_global_config` key, which is used to populate the bot's global configuration (stored as [dynamic configuration](#dynamic-configuration)) the first time the bot starts. Modifying this subsequently will not have any effect. After initial global configuration creation, it's expected to be managed dynamically via chat commands.
|
You can see the list of supported environment variables in the [🦀 src/entity/cfg/env.rs](../../src/entity/cfg/env.rs) file.
|
||||||
|
|
||||||
|
> [!WARNING]
|
||||||
|
> The static configuration contains an `initial_global_config` key, which is used to populate the bot's global configuration (stored as [dynamic configuration](#dynamic-configuration)) the first time the bot starts. Modifying this subsequently will not have any effect. After initial global configuration creation, it's expected to be managed dynamically via chat commands.
|
||||||
|
|
||||||
|
|
||||||
### Dynamic configuration
|
### Dynamic configuration
|
||||||
|
|||||||
@@ -40,6 +40,23 @@ You may also wish to see:
|
|||||||
- [📖 Usage / 💬 Text Generation](./usage.md#-text-generation) section for more details on how to use the bot for Text Generation in a room
|
- [📖 Usage / 💬 Text Generation](./usage.md#-text-generation) section for more details on how to use the bot for Text Generation in a room
|
||||||
|
|
||||||
|
|
||||||
|
#### 🛠️ Built-in Tools (OpenAI only)
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
The [OpenAI provider](./providers.md#openai) supports built-in tools that extend the model's capabilities:
|
||||||
|
|
||||||
|
- [🔍 Web Search](https://platform.openai.com/docs/guides/tools-web-search) (`web_search`): allows the model to search the web for up-to-date information. [🖼️ Screenshot](./screenshots/text-generation-tools-web-search.webp)
|
||||||
|
|
||||||
|
- [💻 Code Interpreter](https://platform.openai.com/docs/guides/tools-code-interpreter) (`code_interpreter`): allows the model to write and execute Python code in a sandbox
|
||||||
|
|
||||||
|
These tools are **disabled by default** and need to be explicitly enabled in the agent's `text_generation.tools` configuration. See the [OpenAI sample configuration](https://github.com/etkecc/baibot/blob/c70387b0c38d8d0f30bba2179a2a21a3710dbeaf/docs/sample-provider-configs/openai.yml#L12-L15) for reference.
|
||||||
|
|
||||||
|
To enable tools on an existing dynamically-created agent, you need to [update the agent](./agents.md#updating-agents) to re-create it with the `text_generation.tools` section added and enable the tools you need
|
||||||
|
|
||||||
|
💡 **Note**: These tools run on OpenAI's infrastructure and may incur additional costs. Web search results include citations that are incorporated into the response.
|
||||||
|
|
||||||
|
|
||||||
#### On-demand involvement
|
#### On-demand involvement
|
||||||
|
|
||||||
In the following 2 cases, it's useful to involve the bot in conversations on-demand:
|
In the following 2 cases, it's useful to involve the bot in conversations on-demand:
|
||||||
|
|||||||
@@ -23,7 +23,7 @@ The list of supported providers is below.
|
|||||||
|
|
||||||
### How to choose a provider
|
### How to choose a provider
|
||||||
|
|
||||||
If you're not sure which provider to start with, **we recommend [OpenAI](#openai)** as it's the most popular and has the **widest range of capabilities**: [💬 text-generation](./features.md#-text-generation) (no vision), [🖌️ image-generation](./features.md#️image-generation), [🦻 speech-to-text](./features.md#-speech-to-text), [🗣️ text-to-speech](./features.md#️-text-to-speech).
|
If you're not sure which provider to start with, **we recommend [OpenAI](#openai)** as it's the most popular and has the **widest range of capabilities**: [💬 text-generation](./features.md#-text-generation) (incl. vision, incl. [🛠️ tools](./features.md#️-built-in-tools-openai-only)), [🖌️ image-generation](./features.md#️image-generation), [🦻 speech-to-text](./features.md#-speech-to-text), [🗣️ text-to-speech](./features.md#️-text-to-speech).
|
||||||
|
|
||||||
You don't need to choose just one though. The bot supports [mixing & matching models](./features.md#-mixing--matching-models), so you can use multiple providers at the same time.
|
You don't need to choose just one though. The bot supports [mixing & matching models](./features.md#-mixing--matching-models), so you can use multiple providers at the same time.
|
||||||
|
|
||||||
@@ -47,7 +47,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
|
|||||||
|
|
||||||
- 🆔 Identifier: `anthropic`
|
- 🆔 Identifier: `anthropic`
|
||||||
- 🔗 Links: [🏠 Home page](https://www.anthropic.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Anthropic), [👤 Sign up](https://console.anthropic.com/), [📋 Models list](https://docs.anthropic.com/en/docs/about-claude/models)
|
- 🔗 Links: [🏠 Home page](https://www.anthropic.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Anthropic), [👤 Sign up](https://console.anthropic.com/), [📋 Models list](https://docs.anthropic.com/en/docs/about-claude/models)
|
||||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (incl. vision)
|
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (incl. vision, no tools)
|
||||||
- 🗲 Quick start:
|
- 🗲 Quick start:
|
||||||
- create a room-local agent: `!bai agent create-room-local anthropic my-anthropic-agent`
|
- create a room-local agent: `!bai agent create-room-local anthropic my-anthropic-agent`
|
||||||
- create a global agent: `!bai agent create-global anthropic my-anthropic-agent`
|
- create a global agent: `!bai agent create-global anthropic my-anthropic-agent`
|
||||||
@@ -61,7 +61,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
|
|||||||
|
|
||||||
- 🆔 Identifier: `groq`
|
- 🆔 Identifier: `groq`
|
||||||
- 🔗 Links: [🏠 Home page](https://groq.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Groq), [👤 Sign up](https://console.groq.com/login), [📋 Models list](https://console.groq.com/docs/models)
|
- 🔗 Links: [🏠 Home page](https://groq.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Groq), [👤 Sign up](https://console.groq.com/login), [📋 Models list](https://console.groq.com/docs/models)
|
||||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision), [🦻 speech-to-text](./features.md#-speech-to-text)
|
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools), [🦻 speech-to-text](./features.md#-speech-to-text)
|
||||||
- 🗲 Quick start:
|
- 🗲 Quick start:
|
||||||
- create a room-local agent: `!bai agent create-room-local groq my-groq-agent`
|
- create a room-local agent: `!bai agent create-room-local groq my-groq-agent`
|
||||||
- create a global agent: `!bai agent create-global groq my-groq-agent`
|
- create a global agent: `!bai agent create-global groq my-groq-agent`
|
||||||
@@ -75,7 +75,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
|
|||||||
|
|
||||||
- 🆔 Identifier: `localai`
|
- 🆔 Identifier: `localai`
|
||||||
- 🔗 Links: [🏠 Home page](https://localai.io/), [📋 Models list](https://localai.io/gallery.html)
|
- 🔗 Links: [🏠 Home page](https://localai.io/), [📋 Models list](https://localai.io/gallery.html)
|
||||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
||||||
- 🗲 Quick start:
|
- 🗲 Quick start:
|
||||||
- create a room-local agent: `!bai agent create-room-local localai my-localai-agent`
|
- create a room-local agent: `!bai agent create-room-local localai my-localai-agent`
|
||||||
- create a global agent: `!bai agent create-global localai my-localai-agent`
|
- create a global agent: `!bai agent create-global localai my-localai-agent`
|
||||||
@@ -89,7 +89,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
|
|||||||
|
|
||||||
- 🆔 Identifier: `mistral`
|
- 🆔 Identifier: `mistral`
|
||||||
- 🔗 Links: [🏠 Home page](https://mistral.ai/), [🌐 Wiki](https://en.wikipedia.org/wiki/Mistral_AI), [👤 Sign up](https://auth.mistral.ai/ui/registration), [📋 Models list](https://docs.mistral.ai/getting-started/models/)
|
- 🔗 Links: [🏠 Home page](https://mistral.ai/), [🌐 Wiki](https://en.wikipedia.org/wiki/Mistral_AI), [👤 Sign up](https://auth.mistral.ai/ui/registration), [📋 Models list](https://docs.mistral.ai/getting-started/models/)
|
||||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
|
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
|
||||||
- 🗲 Quick start:
|
- 🗲 Quick start:
|
||||||
- create a room-local agent: `!bai agent create-room-local mistral my-mistral-agent`
|
- create a room-local agent: `!bai agent create-room-local mistral my-mistral-agent`
|
||||||
- create a global agent: `!bai agent create-global mistral my-mistral-agent`
|
- create a global agent: `!bai agent create-global mistral my-mistral-agent`
|
||||||
@@ -103,7 +103,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
|
|||||||
|
|
||||||
- 🆔 Identifier: `ollama`
|
- 🆔 Identifier: `ollama`
|
||||||
- 🔗 Links: [🏠 Home page](https://ollama.com/), [📋 Models list](https://ollama.com/library)
|
- 🔗 Links: [🏠 Home page](https://ollama.com/), [📋 Models list](https://ollama.com/library)
|
||||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
|
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
|
||||||
- 🗲 Quick start:
|
- 🗲 Quick start:
|
||||||
- create a room-local agent: `!bai agent create-room-local ollama my-ollama-agent`
|
- create a room-local agent: `!bai agent create-room-local ollama my-ollama-agent`
|
||||||
- create a global agent: `!bai agent create-global ollama my-ollama-agent`
|
- create a global agent: `!bai agent create-global ollama my-ollama-agent`
|
||||||
@@ -120,7 +120,7 @@ For services which are not fully compatible with the OpenAI API, consider using
|
|||||||
|
|
||||||
- 🆔 Identifier: `openai`
|
- 🆔 Identifier: `openai`
|
||||||
- 🔗 Links: [🏠 Home page](https://openai.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/OpenAI), [👤 Sign up](https://platform.openai.com/signup), [📋 Models list](https://platform.openai.com/docs/models)
|
- 🔗 Links: [🏠 Home page](https://openai.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/OpenAI), [👤 Sign up](https://platform.openai.com/signup), [📋 Models list](https://platform.openai.com/docs/models)
|
||||||
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (incl. vision), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (incl. vision, incl. [🛠️ tools](./features.md#️-built-in-tools-openai-only)), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
||||||
- 🗲 Quick start:
|
- 🗲 Quick start:
|
||||||
- create a room-local agent: `!bai agent create-room-local openai my-openai-agent`
|
- create a room-local agent: `!bai agent create-room-local openai my-openai-agent`
|
||||||
- create a global agent: `!bai agent create-global openai my-openai-agent`
|
- create a global agent: `!bai agent create-global openai my-openai-agent`
|
||||||
@@ -137,7 +137,7 @@ Some of these popular services already have **shortcut** providers (leading to t
|
|||||||
This provider is just as featureful as the [OpenAI](#openai) provider, but is more compatible with services which do not fully adhere to the [OpenAI API spec](https://github.com/openai/openai-openapi/).
|
This provider is just as featureful as the [OpenAI](#openai) provider, but is more compatible with services which do not fully adhere to the [OpenAI API spec](https://github.com/openai/openai-openapi/).
|
||||||
|
|
||||||
- 🆔 Identifier: `openai-compatible`
|
- 🆔 Identifier: `openai-compatible`
|
||||||
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (no vision), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (no vision, no tools), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
||||||
- 🗲 Quick start:
|
- 🗲 Quick start:
|
||||||
- create a room-local agent: `!bai agent create-room-local openai-compatible my-openai-compatible-agent`
|
- create a room-local agent: `!bai agent create-room-local openai-compatible my-openai-compatible-agent`
|
||||||
- create a global agent: `!bai agent create-global openai-compatible my-openai-compatible-agent`
|
- create a global agent: `!bai agent create-global openai-compatible my-openai-compatible-agent`
|
||||||
@@ -151,7 +151,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
|
|||||||
|
|
||||||
- 🆔 Identifier: `openrouter`
|
- 🆔 Identifier: `openrouter`
|
||||||
- 🔗 Links: [🏠 Home page](https://openrouter.ai/), [👤 Sign up](https://openrouter.ai/), [📋 Models list](https://openrouter.ai/models)
|
- 🔗 Links: [🏠 Home page](https://openrouter.ai/), [👤 Sign up](https://openrouter.ai/), [📋 Models list](https://openrouter.ai/models)
|
||||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
|
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
|
||||||
- 🗲 Quick start:
|
- 🗲 Quick start:
|
||||||
- create a room-local agent: `!bai agent create-room-local openrouter my-openrouter-agent`
|
- create a room-local agent: `!bai agent create-room-local openrouter my-openrouter-agent`
|
||||||
- create a global agent: `!bai agent create-global openrouter my-openrouter-agent`
|
- create a global agent: `!bai agent create-global openrouter my-openrouter-agent`
|
||||||
@@ -165,7 +165,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
|
|||||||
|
|
||||||
- 🆔 Identifier: `together-ai`
|
- 🆔 Identifier: `together-ai`
|
||||||
- 🔗 Links: [🏠 Home page](https://www.together.ai/), [👤 Sign up](https://api.together.ai/signup), [📋 Models list](https://api.together.xyz/models)
|
- 🔗 Links: [🏠 Home page](https://www.together.ai/), [👤 Sign up](https://api.together.ai/signup), [📋 Models list](https://api.together.xyz/models)
|
||||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
|
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
|
||||||
- 🗲 Quick start:
|
- 🗲 Quick start:
|
||||||
- create a room-local agent: `!bai agent create-room-local together-ai my-together-ai-agent`
|
- create a room-local agent: `!bai agent create-room-local together-ai my-together-ai-agent`
|
||||||
- create a global agent: `!bai agent create-global together-ai my-together-ai-agent`
|
- create a global agent: `!bai agent create-global together-ai my-together-ai-agent`
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
base_url: https://api.openai.com/v1
|
base_url: https://api.openai.com/v1
|
||||||
api_key: YOUR_API_KEY_HERE
|
api_key: YOUR_API_KEY_HERE
|
||||||
text_generation:
|
text_generation:
|
||||||
model_id: gpt-5.1
|
model_id: gpt-5.2
|
||||||
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
# Reasoning models need to use `max_completion_tokens` instead of `max_response_tokens`.
|
# Reasoning models need to use `max_completion_tokens` instead of `max_response_tokens`.
|
||||||
@@ -9,6 +9,10 @@ text_generation:
|
|||||||
max_response_tokens: null
|
max_response_tokens: null
|
||||||
max_completion_tokens: 128000
|
max_completion_tokens: 128000
|
||||||
max_context_tokens: 400000
|
max_context_tokens: 400000
|
||||||
|
# Built-in tools
|
||||||
|
tools:
|
||||||
|
web_search: false
|
||||||
|
code_interpreter: false
|
||||||
speech_to_text:
|
speech_to_text:
|
||||||
model_id: whisper-1
|
model_id: whisper-1
|
||||||
text_to_speech:
|
text_to_speech:
|
||||||
@@ -17,7 +21,7 @@ text_to_speech:
|
|||||||
speed: 1.0
|
speed: 1.0
|
||||||
response_format: opus
|
response_format: opus
|
||||||
image_generation:
|
image_generation:
|
||||||
model_id: gpt-image-1
|
model_id: gpt-image-1.5
|
||||||
style: null
|
style: null
|
||||||
size: null
|
size: null
|
||||||
quality: null
|
quality: null
|
||||||
|
|||||||
BIN
docs/screenshots/text-generation-tools-web-search.webp
Normal file
BIN
docs/screenshots/text-generation-tools-web-search.webp
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 66 KiB |
@@ -11,6 +11,12 @@ user:
|
|||||||
# Leave empty to use the default (baibot).
|
# Leave empty to use the default (baibot).
|
||||||
name: baibot
|
name: baibot
|
||||||
|
|
||||||
|
# An optional path to an image file to be used as a custom avatar image.
|
||||||
|
# - null or empty string: use the default avatar
|
||||||
|
# - "keep": don't touch the avatar, keep whatever is already set
|
||||||
|
# - any other value: path to a custom avatar image file
|
||||||
|
avatar: null
|
||||||
|
|
||||||
encryption:
|
encryption:
|
||||||
# An optional passphrase to use for backing up and recovering the bot's encryption keys.
|
# An optional passphrase to use for backing up and recovering the bot's encryption keys.
|
||||||
# You can use any string here.
|
# You can use any string here.
|
||||||
@@ -76,7 +82,7 @@ agents:
|
|||||||
# base_url: https://api.openai.com/v1
|
# base_url: https://api.openai.com/v1
|
||||||
# api_key: ""
|
# api_key: ""
|
||||||
# text_generation:
|
# text_generation:
|
||||||
# model_id: gpt-5.1
|
# model_id: gpt-5.2
|
||||||
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
# prompt: "You are a brief, but helpful bot called {{ baibot_name }} powered by the {{ baibot_model_id }} model. The date/time of this conversation's start is: {{ baibot_conversation_start_time_utc }}."
|
||||||
# temperature: 1.0
|
# temperature: 1.0
|
||||||
# # Reasoning models need to use `max_completion_tokens` instead of `max_response_tokens`.
|
# # Reasoning models need to use `max_completion_tokens` instead of `max_response_tokens`.
|
||||||
@@ -84,6 +90,10 @@ agents:
|
|||||||
# max_response_tokens: null
|
# max_response_tokens: null
|
||||||
# max_completion_tokens: 128000
|
# max_completion_tokens: 128000
|
||||||
# max_context_tokens: 400000
|
# max_context_tokens: 400000
|
||||||
|
# # Built-in tools
|
||||||
|
# tools:
|
||||||
|
# web_search: false
|
||||||
|
# code_interpreter: false
|
||||||
# speech_to_text:
|
# speech_to_text:
|
||||||
# model_id: whisper-1
|
# model_id: whisper-1
|
||||||
# text_to_speech:
|
# text_to_speech:
|
||||||
@@ -92,7 +102,7 @@ agents:
|
|||||||
# speed: 1.0
|
# speed: 1.0
|
||||||
# response_format: opus
|
# response_format: opus
|
||||||
# image_generation:
|
# image_generation:
|
||||||
# model_id: gpt-image-1
|
# model_id: gpt-image-1.5
|
||||||
# style: null
|
# style: null
|
||||||
# size: null
|
# size: null
|
||||||
# quality: null
|
# quality: null
|
||||||
|
|||||||
@@ -14,7 +14,7 @@ services:
|
|||||||
- /etc/passwd:/etc/passwd:ro
|
- /etc/passwd:/etc/passwd:ro
|
||||||
|
|
||||||
synapse:
|
synapse:
|
||||||
image: ghcr.io/element-hq/synapse:v1.142.1
|
image: ghcr.io/element-hq/synapse:v1.146.0
|
||||||
user: "${UID}:${GID}"
|
user: "${UID}:${GID}"
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
entrypoint: python
|
entrypoint: python
|
||||||
@@ -27,7 +27,7 @@ services:
|
|||||||
- ./synapse/media-store:/media-store
|
- ./synapse/media-store:/media-store
|
||||||
|
|
||||||
element-web:
|
element-web:
|
||||||
image: ghcr.io/element-hq/element-web:v1.12.4
|
image: ghcr.io/element-hq/element-web:v1.12.9
|
||||||
user: "${UID}:${GID}"
|
user: "${UID}:${GID}"
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
environment:
|
environment:
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
services:
|
services:
|
||||||
ollama:
|
ollama:
|
||||||
image: docker.io/ollama/ollama:0.13.0
|
image: docker.io/ollama/ollama:0.15.4
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
ports:
|
ports:
|
||||||
- "${SERVICE_OLLAMA_BIND_PORT_HTTP}:11434"
|
- "${SERVICE_OLLAMA_BIND_PORT_HTTP}:11434"
|
||||||
|
|||||||
@@ -69,6 +69,7 @@ impl AgentProvider {
|
|||||||
models_list_url: Some("https://docs.anthropic.com/en/docs/about-claude/models"),
|
models_list_url: Some("https://docs.anthropic.com/en/docs/about-claude/models"),
|
||||||
supported_purposes: vec![AgentPurpose::TextGeneration],
|
supported_purposes: vec![AgentPurpose::TextGeneration],
|
||||||
text_generation_supports_vision: true,
|
text_generation_supports_vision: true,
|
||||||
|
text_generation_supports_tools: false,
|
||||||
},
|
},
|
||||||
Self::Groq => AgentProviderInfo {
|
Self::Groq => AgentProviderInfo {
|
||||||
id: Self::Groq.to_static_str(),
|
id: Self::Groq.to_static_str(),
|
||||||
@@ -80,11 +81,12 @@ impl AgentProvider {
|
|||||||
models_list_url: Some("https://console.groq.com/docs/models"),
|
models_list_url: Some("https://console.groq.com/docs/models"),
|
||||||
supported_purposes: vec![AgentPurpose::TextGeneration, AgentPurpose::SpeechToText],
|
supported_purposes: vec![AgentPurpose::TextGeneration, AgentPurpose::SpeechToText],
|
||||||
text_generation_supports_vision: false,
|
text_generation_supports_vision: false,
|
||||||
|
text_generation_supports_tools: false,
|
||||||
},
|
},
|
||||||
Self::LocalAI => AgentProviderInfo {
|
Self::LocalAI => AgentProviderInfo {
|
||||||
id: Self::LocalAI.to_static_str(),
|
id: Self::LocalAI.to_static_str(),
|
||||||
name: "LocalAI",
|
name: "LocalAI",
|
||||||
description: "LocalAI is the free, Open Source OpenAI alternative. LocalAI act as a drop-in replacement REST API that’s compatible with OpenAI API specifications for local inferencing. It allows you to run LLMs, generate images, audio (and not only) locally or on-prem with consumer grade hardware, supporting multiple model families and architectures.",
|
description: "LocalAI is the free, Open Source OpenAI alternative. LocalAI act as a drop-in replacement REST API that's compatible with OpenAI API specifications for local inferencing. It allows you to run LLMs, generate images, audio (and not only) locally or on-prem with consumer grade hardware, supporting multiple model families and architectures.",
|
||||||
homepage_url: Some("https://localai.io/"),
|
homepage_url: Some("https://localai.io/"),
|
||||||
wiki_url: None,
|
wiki_url: None,
|
||||||
sign_up_url: None,
|
sign_up_url: None,
|
||||||
@@ -95,6 +97,7 @@ impl AgentProvider {
|
|||||||
AgentPurpose::SpeechToText,
|
AgentPurpose::SpeechToText,
|
||||||
],
|
],
|
||||||
text_generation_supports_vision: false,
|
text_generation_supports_vision: false,
|
||||||
|
text_generation_supports_tools: false,
|
||||||
},
|
},
|
||||||
Self::Mistral => AgentProviderInfo {
|
Self::Mistral => AgentProviderInfo {
|
||||||
id: Self::Mistral.to_static_str(),
|
id: Self::Mistral.to_static_str(),
|
||||||
@@ -106,6 +109,7 @@ impl AgentProvider {
|
|||||||
models_list_url: Some("https://docs.mistral.ai/getting-started/models/"),
|
models_list_url: Some("https://docs.mistral.ai/getting-started/models/"),
|
||||||
supported_purposes: vec![AgentPurpose::TextGeneration],
|
supported_purposes: vec![AgentPurpose::TextGeneration],
|
||||||
text_generation_supports_vision: false,
|
text_generation_supports_vision: false,
|
||||||
|
text_generation_supports_tools: false,
|
||||||
},
|
},
|
||||||
Self::Ollama => AgentProviderInfo {
|
Self::Ollama => AgentProviderInfo {
|
||||||
id: Self::Ollama.to_static_str(),
|
id: Self::Ollama.to_static_str(),
|
||||||
@@ -117,6 +121,7 @@ impl AgentProvider {
|
|||||||
models_list_url: Some("https://ollama.com/library"),
|
models_list_url: Some("https://ollama.com/library"),
|
||||||
supported_purposes: vec![AgentPurpose::TextGeneration],
|
supported_purposes: vec![AgentPurpose::TextGeneration],
|
||||||
text_generation_supports_vision: false,
|
text_generation_supports_vision: false,
|
||||||
|
text_generation_supports_tools: false,
|
||||||
},
|
},
|
||||||
Self::OpenAI => AgentProviderInfo {
|
Self::OpenAI => AgentProviderInfo {
|
||||||
id: Self::OpenAI.to_static_str(),
|
id: Self::OpenAI.to_static_str(),
|
||||||
@@ -133,6 +138,7 @@ impl AgentProvider {
|
|||||||
AgentPurpose::SpeechToText,
|
AgentPurpose::SpeechToText,
|
||||||
],
|
],
|
||||||
text_generation_supports_vision: true,
|
text_generation_supports_vision: true,
|
||||||
|
text_generation_supports_tools: true,
|
||||||
},
|
},
|
||||||
Self::OpenAICompat => AgentProviderInfo {
|
Self::OpenAICompat => AgentProviderInfo {
|
||||||
id: Self::OpenAICompat.to_static_str(),
|
id: Self::OpenAICompat.to_static_str(),
|
||||||
@@ -149,6 +155,7 @@ impl AgentProvider {
|
|||||||
AgentPurpose::SpeechToText,
|
AgentPurpose::SpeechToText,
|
||||||
],
|
],
|
||||||
text_generation_supports_vision: false,
|
text_generation_supports_vision: false,
|
||||||
|
text_generation_supports_tools: false,
|
||||||
},
|
},
|
||||||
Self::OpenRouter => AgentProviderInfo {
|
Self::OpenRouter => AgentProviderInfo {
|
||||||
id: Self::OpenRouter.to_static_str(),
|
id: Self::OpenRouter.to_static_str(),
|
||||||
@@ -160,6 +167,7 @@ impl AgentProvider {
|
|||||||
models_list_url: Some("https://openrouter.ai/models"),
|
models_list_url: Some("https://openrouter.ai/models"),
|
||||||
supported_purposes: vec![AgentPurpose::TextGeneration],
|
supported_purposes: vec![AgentPurpose::TextGeneration],
|
||||||
text_generation_supports_vision: false,
|
text_generation_supports_vision: false,
|
||||||
|
text_generation_supports_tools: false,
|
||||||
},
|
},
|
||||||
Self::TogetherAI => AgentProviderInfo {
|
Self::TogetherAI => AgentProviderInfo {
|
||||||
id: Self::TogetherAI.to_static_str(),
|
id: Self::TogetherAI.to_static_str(),
|
||||||
@@ -171,6 +179,7 @@ impl AgentProvider {
|
|||||||
models_list_url: Some("https://api.together.xyz/models"),
|
models_list_url: Some("https://api.together.xyz/models"),
|
||||||
supported_purposes: vec![AgentPurpose::TextGeneration],
|
supported_purposes: vec![AgentPurpose::TextGeneration],
|
||||||
text_generation_supports_vision: false,
|
text_generation_supports_vision: false,
|
||||||
|
text_generation_supports_tools: false,
|
||||||
},
|
},
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -192,4 +201,5 @@ pub struct AgentProviderInfo {
|
|||||||
pub models_list_url: Option<&'static str>,
|
pub models_list_url: Option<&'static str>,
|
||||||
pub supported_purposes: Vec<AgentPurpose>,
|
pub supported_purposes: Vec<AgentPurpose>,
|
||||||
pub text_generation_supports_vision: bool,
|
pub text_generation_supports_vision: bool,
|
||||||
|
pub text_generation_supports_tools: bool,
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ use mxlink::mime;
|
|||||||
|
|
||||||
#[derive(Default)]
|
#[derive(Default)]
|
||||||
pub struct ImageGenerationParams {
|
pub struct ImageGenerationParams {
|
||||||
pub size_override: Option<String>,
|
pub smallest_size_possible: bool,
|
||||||
|
|
||||||
pub cheaper_model_switching_allowed: bool,
|
pub cheaper_model_switching_allowed: bool,
|
||||||
|
|
||||||
@@ -10,8 +10,8 @@ pub struct ImageGenerationParams {
|
|||||||
}
|
}
|
||||||
|
|
||||||
impl ImageGenerationParams {
|
impl ImageGenerationParams {
|
||||||
pub fn with_size_override(mut self, value: Option<String>) -> Self {
|
pub fn with_smallest_size_possible(mut self, value: bool) -> Self {
|
||||||
self.size_override = value;
|
self.smallest_size_possible = value;
|
||||||
self
|
self
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -56,12 +56,8 @@ impl ImageSource {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
impl From<ImageSource> for async_openai::types::ImageInput {
|
impl From<ImageSource> for async_openai::types::images::ImageInput {
|
||||||
fn from(value: ImageSource) -> Self {
|
fn from(value: ImageSource) -> Self {
|
||||||
async_openai::types::ImageInput::from_vec_u8(
|
async_openai::types::images::ImageInput::from_vec_u8(value.filename, value.bytes)
|
||||||
value.filename,
|
|
||||||
value.bytes,
|
|
||||||
value.mime_type.to_string(),
|
|
||||||
)
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
use serde::{Deserialize, Serialize};
|
use serde::{Deserialize, Serialize};
|
||||||
|
|
||||||
use super::OPENAI_IMAGE_MODEL_GPT_IMAGE_1;
|
use super::OPENAI_IMAGE_MODEL_GPT_IMAGE_1_DOT_5;
|
||||||
use crate::agent::{default_prompt, provider::ConfigTrait};
|
use crate::agent::{default_prompt, provider::ConfigTrait};
|
||||||
|
|
||||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||||
@@ -64,6 +64,9 @@ pub struct TextGenerationConfig {
|
|||||||
|
|
||||||
#[serde(default)]
|
#[serde(default)]
|
||||||
pub max_context_tokens: u32,
|
pub max_context_tokens: u32,
|
||||||
|
|
||||||
|
#[serde(default)]
|
||||||
|
pub tools: ToolsConfig,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl Default for TextGenerationConfig {
|
impl Default for TextGenerationConfig {
|
||||||
@@ -75,12 +78,22 @@ impl Default for TextGenerationConfig {
|
|||||||
max_response_tokens: None,
|
max_response_tokens: None,
|
||||||
max_completion_tokens: Some(128_000),
|
max_completion_tokens: Some(128_000),
|
||||||
max_context_tokens: 400_000,
|
max_context_tokens: 400_000,
|
||||||
|
tools: ToolsConfig::default(),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn default_text_model_id() -> String {
|
fn default_text_model_id() -> String {
|
||||||
"gpt-5.1".to_owned()
|
"gpt-5.2".to_owned()
|
||||||
|
}
|
||||||
|
|
||||||
|
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
|
||||||
|
pub struct ToolsConfig {
|
||||||
|
#[serde(default)]
|
||||||
|
pub web_search: bool,
|
||||||
|
|
||||||
|
#[serde(default)]
|
||||||
|
pub code_interpreter: bool,
|
||||||
}
|
}
|
||||||
|
|
||||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||||
@@ -104,16 +117,16 @@ fn default_speech_to_text_model_id() -> String {
|
|||||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||||
pub struct TextToSpeechConfig {
|
pub struct TextToSpeechConfig {
|
||||||
#[serde(default = "default_text_to_speech_model_id")]
|
#[serde(default = "default_text_to_speech_model_id")]
|
||||||
pub model_id: async_openai::types::SpeechModel,
|
pub model_id: async_openai::types::audio::SpeechModel,
|
||||||
|
|
||||||
#[serde(default = "default_text_to_speech_voice")]
|
#[serde(default = "default_text_to_speech_voice")]
|
||||||
pub voice: async_openai::types::Voice,
|
pub voice: async_openai::types::audio::Voice,
|
||||||
|
|
||||||
#[serde(default = "default_text_to_speech_speed")]
|
#[serde(default = "default_text_to_speech_speed")]
|
||||||
pub speed: f32,
|
pub speed: f32,
|
||||||
|
|
||||||
#[serde(default = "default_text_to_speech_response_format")]
|
#[serde(default = "default_text_to_speech_response_format")]
|
||||||
pub response_format: async_openai::types::SpeechResponseFormat,
|
pub response_format: async_openai::types::audio::SpeechResponseFormat,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl Default for TextToSpeechConfig {
|
impl Default for TextToSpeechConfig {
|
||||||
@@ -127,22 +140,22 @@ impl Default for TextToSpeechConfig {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn default_text_to_speech_model_id() -> async_openai::types::SpeechModel {
|
fn default_text_to_speech_model_id() -> async_openai::types::audio::SpeechModel {
|
||||||
async_openai::types::SpeechModel::Tts1Hd
|
async_openai::types::audio::SpeechModel::Tts1Hd
|
||||||
}
|
}
|
||||||
|
|
||||||
fn default_text_to_speech_voice() -> async_openai::types::Voice {
|
fn default_text_to_speech_voice() -> async_openai::types::audio::Voice {
|
||||||
async_openai::types::Voice::Onyx
|
async_openai::types::audio::Voice::Onyx
|
||||||
}
|
}
|
||||||
|
|
||||||
fn default_text_to_speech_speed() -> f32 {
|
fn default_text_to_speech_speed() -> f32 {
|
||||||
1.0
|
1.0
|
||||||
}
|
}
|
||||||
|
|
||||||
fn default_text_to_speech_response_format() -> async_openai::types::SpeechResponseFormat {
|
fn default_text_to_speech_response_format() -> async_openai::types::audio::SpeechResponseFormat {
|
||||||
// The API defaults to mp3, but we prefer Opus because it's smaller.
|
// The API defaults to mp3, but we prefer Opus because it's smaller.
|
||||||
// Our clients should all have support for it.
|
// Our clients should all have support for it.
|
||||||
async_openai::types::SpeechResponseFormat::Opus
|
async_openai::types::audio::SpeechResponseFormat::Opus
|
||||||
}
|
}
|
||||||
|
|
||||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||||
@@ -150,19 +163,19 @@ pub struct ImageGenerationConfig {
|
|||||||
pub model_id: String,
|
pub model_id: String,
|
||||||
|
|
||||||
#[serde(default = "default_image_style")]
|
#[serde(default = "default_image_style")]
|
||||||
pub style: Option<async_openai::types::ImageStyle>,
|
pub style: Option<async_openai::types::images::ImageStyle>,
|
||||||
|
|
||||||
#[serde(default = "default_image_size")]
|
#[serde(default = "default_image_size")]
|
||||||
pub size: Option<async_openai::types::ImageSize>,
|
pub size: Option<async_openai::types::images::ImageSize>,
|
||||||
|
|
||||||
#[serde(default = "default_image_quality")]
|
#[serde(default = "default_image_quality")]
|
||||||
pub quality: Option<async_openai::types::ImageQuality>,
|
pub quality: Option<async_openai::types::images::ImageQuality>,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl Default for ImageGenerationConfig {
|
impl Default for ImageGenerationConfig {
|
||||||
fn default() -> Self {
|
fn default() -> Self {
|
||||||
Self {
|
Self {
|
||||||
model_id: OPENAI_IMAGE_MODEL_GPT_IMAGE_1.to_owned(),
|
model_id: OPENAI_IMAGE_MODEL_GPT_IMAGE_1_DOT_5.to_owned(),
|
||||||
style: default_image_style(),
|
style: default_image_style(),
|
||||||
size: default_image_size(),
|
size: default_image_size(),
|
||||||
quality: default_image_quality(),
|
quality: default_image_quality(),
|
||||||
@@ -173,23 +186,28 @@ impl Default for ImageGenerationConfig {
|
|||||||
impl ImageGenerationConfig {
|
impl ImageGenerationConfig {
|
||||||
pub fn model_id_as_openai_image_model(
|
pub fn model_id_as_openai_image_model(
|
||||||
&self,
|
&self,
|
||||||
) -> Result<async_openai::types::ImageModel, String> {
|
) -> Result<async_openai::types::images::ImageModel, String> {
|
||||||
match self.model_id.as_str() {
|
match self.model_id.as_str() {
|
||||||
"dall-e-2" => Ok(async_openai::types::ImageModel::DallE2),
|
"dall-e-2" => Ok(async_openai::types::images::ImageModel::DallE2),
|
||||||
"dall-e-3" => Ok(async_openai::types::ImageModel::DallE3),
|
"dall-e-3" => Ok(async_openai::types::images::ImageModel::DallE3),
|
||||||
other => Ok(async_openai::types::ImageModel::Other(other.to_owned())),
|
"gpt-image-1" => Ok(async_openai::types::images::ImageModel::GptImage1),
|
||||||
|
"gpt-image-1.5" => Ok(async_openai::types::images::ImageModel::GptImage1dot5),
|
||||||
|
"gpt-image-1-mini" => Ok(async_openai::types::images::ImageModel::GptImage1Mini),
|
||||||
|
other => Ok(async_openai::types::images::ImageModel::Other(
|
||||||
|
other.to_owned(),
|
||||||
|
)),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn default_image_style() -> Option<async_openai::types::ImageStyle> {
|
fn default_image_style() -> Option<async_openai::types::images::ImageStyle> {
|
||||||
None
|
None
|
||||||
}
|
}
|
||||||
|
|
||||||
fn default_image_size() -> Option<async_openai::types::ImageSize> {
|
fn default_image_size() -> Option<async_openai::types::images::ImageSize> {
|
||||||
None
|
None
|
||||||
}
|
}
|
||||||
|
|
||||||
fn default_image_quality() -> Option<async_openai::types::ImageQuality> {
|
fn default_image_quality() -> Option<async_openai::types::images::ImageQuality> {
|
||||||
None
|
None
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -4,9 +4,15 @@ use async_openai::{
|
|||||||
Client as OpenAIClient,
|
Client as OpenAIClient,
|
||||||
config::OpenAIConfig,
|
config::OpenAIConfig,
|
||||||
types::{
|
types::{
|
||||||
ChatCompletionRequestMessage, CreateChatCompletionRequestArgs, CreateImageEditRequestArgs,
|
audio::{AudioInput, CreateSpeechRequestArgs, CreateTranscriptionRequestArgs},
|
||||||
CreateImageRequestArgs, CreateSpeechRequestArgs, CreateTranscriptionRequestArgs,
|
responses::{
|
||||||
DallE2ImageSize, Image, ImageModel, ImageResponseFormat,
|
CodeInterpreterContainerAuto, CodeInterpreterTool, CodeInterpreterToolContainer,
|
||||||
|
CreateResponseArgs, OutputItem, OutputMessageContent, Tool, WebSearchTool,
|
||||||
|
},
|
||||||
|
images::{
|
||||||
|
CreateImageEditRequestArgs, CreateImageRequestArgs,
|
||||||
|
Image, ImageInput, ImageModel, ImageResponseFormat,
|
||||||
|
},
|
||||||
},
|
},
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -25,12 +31,9 @@ use crate::{
|
|||||||
use crate::{
|
use crate::{
|
||||||
agent::{
|
agent::{
|
||||||
AgentPurpose,
|
AgentPurpose,
|
||||||
provider::{
|
provider::entity::{
|
||||||
entity::{
|
ImageEditResult, ImageGenerationResult, ImageSource, PingResult,
|
||||||
ImageEditResult, ImageGenerationResult, ImageSource, PingResult,
|
TextToSpeechParams, TextToSpeechResult,
|
||||||
TextToSpeechParams, TextToSpeechResult,
|
|
||||||
},
|
|
||||||
openai::utils::convert_string_to_enum,
|
|
||||||
},
|
},
|
||||||
},
|
},
|
||||||
strings,
|
strings,
|
||||||
@@ -38,8 +41,6 @@ use crate::{
|
|||||||
|
|
||||||
use super::config::Config;
|
use super::config::Config;
|
||||||
|
|
||||||
use super::OPENAI_IMAGE_MODEL_GPT_IMAGE_1;
|
|
||||||
|
|
||||||
#[derive(Debug, Clone)]
|
#[derive(Debug, Clone)]
|
||||||
pub struct Controller {
|
pub struct Controller {
|
||||||
config: Config,
|
config: Config,
|
||||||
@@ -128,28 +129,44 @@ impl ControllerTrait for Controller {
|
|||||||
conversation_messages.insert(0, prompt_message);
|
conversation_messages.insert(0, prompt_message);
|
||||||
}
|
}
|
||||||
|
|
||||||
let openai_conversation_messages: Vec<ChatCompletionRequestMessage> =
|
let input = super::utils::convert_llm_messages_to_openai_response_input(conversation_messages);
|
||||||
super::utils::convert_llm_messages_to_openai_messages(conversation_messages);
|
|
||||||
|
|
||||||
let messages_count = openai_conversation_messages.len();
|
let messages_count = match &input {
|
||||||
|
async_openai::types::responses::InputParam::Items(items) => items.len(),
|
||||||
|
_ => 1,
|
||||||
|
};
|
||||||
|
|
||||||
let temperature = params
|
let temperature = params
|
||||||
.temperature_override
|
.temperature_override
|
||||||
.unwrap_or(text_generation_config.temperature);
|
.unwrap_or(text_generation_config.temperature);
|
||||||
|
|
||||||
let mut request_builder = CreateChatCompletionRequestArgs::default();
|
let mut request_builder = CreateResponseArgs::default();
|
||||||
|
|
||||||
request_builder
|
request_builder
|
||||||
.model(&text_generation_config.model_id)
|
.model(&text_generation_config.model_id)
|
||||||
.temperature(temperature)
|
.temperature(temperature)
|
||||||
.messages(openai_conversation_messages);
|
.input(input);
|
||||||
|
|
||||||
if let Some(max_response_tokens) = text_generation_config.max_response_tokens {
|
let mut tools = Vec::new();
|
||||||
request_builder.max_tokens(max_response_tokens);
|
if text_generation_config.tools.web_search {
|
||||||
|
tools.push(Tool::WebSearch(WebSearchTool::default()));
|
||||||
|
}
|
||||||
|
if text_generation_config.tools.code_interpreter {
|
||||||
|
tools.push(Tool::CodeInterpreter(CodeInterpreterTool {
|
||||||
|
container: CodeInterpreterToolContainer::Auto(
|
||||||
|
CodeInterpreterContainerAuto::default(),
|
||||||
|
),
|
||||||
|
}));
|
||||||
}
|
}
|
||||||
|
|
||||||
if let Some(max_completion_tokens) = text_generation_config.max_completion_tokens {
|
if !tools.is_empty() {
|
||||||
request_builder.max_completion_tokens(max_completion_tokens);
|
request_builder.tools(tools);
|
||||||
|
}
|
||||||
|
|
||||||
|
if let Some(max_response_tokens) = text_generation_config.max_response_tokens {
|
||||||
|
request_builder.max_output_tokens(max_response_tokens);
|
||||||
|
} else if let Some(max_completion_tokens) = text_generation_config.max_completion_tokens {
|
||||||
|
request_builder.max_output_tokens(max_completion_tokens);
|
||||||
}
|
}
|
||||||
|
|
||||||
let request = request_builder.build()?;
|
let request = request_builder.build()?;
|
||||||
@@ -159,33 +176,31 @@ impl ControllerTrait for Controller {
|
|||||||
model = format!("{:?}", request.model),
|
model = format!("{:?}", request.model),
|
||||||
?messages_count,
|
?messages_count,
|
||||||
request = request_as_json,
|
request = request_as_json,
|
||||||
"Sending OpenAI chat completion API request"
|
"Sending OpenAI response API request"
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
let response = self.client.chat().create(request).await?;
|
let response = self.client.responses().create(request).await?;
|
||||||
|
|
||||||
tracing::trace!(
|
tracing::trace!(
|
||||||
?response,
|
?response,
|
||||||
"Got response from the OpenAI chat completion API"
|
"Got response from the OpenAI response API"
|
||||||
);
|
);
|
||||||
|
|
||||||
// We only request 1 result, so there should only be 1 choice.
|
for item in response.output {
|
||||||
if let Some(choice) = response.choices.into_iter().next() {
|
if let OutputItem::Message(message) = item {
|
||||||
match choice.message.content {
|
for content in message.content {
|
||||||
Some(text) => {
|
if let OutputMessageContent::OutputText(text_content) = content {
|
||||||
return Ok(TextGenerationResult { text });
|
return Ok(TextGenerationResult {
|
||||||
}
|
text: text_content.text,
|
||||||
None => {
|
});
|
||||||
return Err(anyhow::anyhow!(
|
}
|
||||||
"No content was found in the response choice from the OpenAI chat completion API"
|
|
||||||
));
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
Err(anyhow::anyhow!(
|
Err(anyhow::anyhow!(
|
||||||
"No response messages choices were returned from the OpenAI chat completion API"
|
"No response messages choices were returned from the OpenAI response API"
|
||||||
))
|
))
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -209,11 +224,7 @@ impl ControllerTrait for Controller {
|
|||||||
|
|
||||||
let request = CreateTranscriptionRequestArgs::default()
|
let request = CreateTranscriptionRequestArgs::default()
|
||||||
.model(&speech_to_text_config.model_id)
|
.model(&speech_to_text_config.model_id)
|
||||||
.file(async_openai::types::AudioInput::from_vec_u8(
|
.file(AudioInput::from_vec_u8(filename, media))
|
||||||
filename,
|
|
||||||
media,
|
|
||||||
mime_type.to_string(),
|
|
||||||
))
|
|
||||||
.language(language.clone())
|
.language(language.clone())
|
||||||
.build()?;
|
.build()?;
|
||||||
|
|
||||||
@@ -223,7 +234,7 @@ impl ControllerTrait for Controller {
|
|||||||
"Sending OpenAI speech-to-text API request"
|
"Sending OpenAI speech-to-text API request"
|
||||||
);
|
);
|
||||||
|
|
||||||
let response = self.client.audio().transcribe(request).await?;
|
let response = self.client.audio().transcription().create(request).await?;
|
||||||
|
|
||||||
tracing::trace!(
|
tracing::trace!(
|
||||||
?response,
|
?response,
|
||||||
@@ -255,10 +266,13 @@ impl ControllerTrait for Controller {
|
|||||||
let model = if params.cheaper_model_switching_allowed {
|
let model = if params.cheaper_model_switching_allowed {
|
||||||
// Switch to a cheaper model
|
// Switch to a cheaper model
|
||||||
match original_model {
|
match original_model {
|
||||||
async_openai::types::ImageModel::DallE2 => async_openai::types::ImageModel::DallE2,
|
ImageModel::DallE2 => ImageModel::DallE2,
|
||||||
async_openai::types::ImageModel::DallE3 => async_openai::types::ImageModel::DallE2,
|
ImageModel::DallE3 => ImageModel::DallE2,
|
||||||
async_openai::types::ImageModel::Other(_) => {
|
ImageModel::GptImage1 => ImageModel::GptImage1Mini,
|
||||||
async_openai::types::ImageModel::DallE2
|
ImageModel::GptImage1dot5 => ImageModel::GptImage1Mini,
|
||||||
|
ImageModel::GptImage1Mini => ImageModel::GptImage1Mini,
|
||||||
|
ImageModel::Other(_) => {
|
||||||
|
ImageModel::DallE2
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
} else {
|
} else {
|
||||||
@@ -269,11 +283,24 @@ impl ControllerTrait for Controller {
|
|||||||
// Switch to a cheaper quality
|
// Switch to a cheaper quality
|
||||||
match &image_generation_config.quality {
|
match &image_generation_config.quality {
|
||||||
Some(quality) => match quality {
|
Some(quality) => match quality {
|
||||||
async_openai::types::ImageQuality::Standard => {
|
async_openai::types::images::ImageQuality::Standard => {
|
||||||
Some(async_openai::types::ImageQuality::Standard)
|
Some(async_openai::types::images::ImageQuality::Standard)
|
||||||
}
|
}
|
||||||
async_openai::types::ImageQuality::HD => {
|
async_openai::types::images::ImageQuality::HD => {
|
||||||
Some(async_openai::types::ImageQuality::Standard)
|
Some(async_openai::types::images::ImageQuality::Standard)
|
||||||
|
}
|
||||||
|
// New quality levels - keep as-is or downgrade to Standard
|
||||||
|
async_openai::types::images::ImageQuality::High => {
|
||||||
|
Some(async_openai::types::images::ImageQuality::Standard)
|
||||||
|
}
|
||||||
|
async_openai::types::images::ImageQuality::Medium => {
|
||||||
|
Some(async_openai::types::images::ImageQuality::Medium)
|
||||||
|
}
|
||||||
|
async_openai::types::images::ImageQuality::Low => {
|
||||||
|
Some(async_openai::types::images::ImageQuality::Low)
|
||||||
|
}
|
||||||
|
async_openai::types::images::ImageQuality::Auto => {
|
||||||
|
Some(async_openai::types::images::ImageQuality::Auto)
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
None => None,
|
None => None,
|
||||||
@@ -282,20 +309,21 @@ impl ControllerTrait for Controller {
|
|||||||
image_generation_config.quality.clone()
|
image_generation_config.quality.clone()
|
||||||
};
|
};
|
||||||
|
|
||||||
let size = params
|
let size = if params.smallest_size_possible {
|
||||||
.size_override
|
Some(get_sticker_size(&model))
|
||||||
.map(|s| convert_string_to_enum::<async_openai::types::ImageSize>(&s).unwrap())
|
} else {
|
||||||
.or(image_generation_config.size);
|
image_generation_config.size
|
||||||
|
};
|
||||||
|
|
||||||
let response_format = match model.clone() {
|
let response_format = match model.clone() {
|
||||||
ImageModel::DallE2 => Some(ImageResponseFormat::B64Json),
|
ImageModel::DallE2 => Some(ImageResponseFormat::B64Json),
|
||||||
ImageModel::DallE3 => Some(ImageResponseFormat::B64Json),
|
ImageModel::DallE3 => Some(ImageResponseFormat::B64Json),
|
||||||
ImageModel::Other(model_str) => match model_str.as_str() {
|
// gpt-image-1 only outputs base64 and we don't need to specify the response format.
|
||||||
// gpt-image-1 only outputs base64 and we don't need to specify the response format.
|
// In fact, specifying the response format results in an error.
|
||||||
// In fact, specifying the response format results in an error.
|
ImageModel::GptImage1 => None,
|
||||||
OPENAI_IMAGE_MODEL_GPT_IMAGE_1 => None,
|
ImageModel::GptImage1Mini => None,
|
||||||
_ => Some(ImageResponseFormat::B64Json),
|
ImageModel::GptImage1dot5 => None,
|
||||||
},
|
ImageModel::Other(_) => Some(ImageResponseFormat::B64Json),
|
||||||
};
|
};
|
||||||
|
|
||||||
let mut request_builder = CreateImageRequestArgs::default();
|
let mut request_builder = CreateImageRequestArgs::default();
|
||||||
@@ -329,15 +357,15 @@ impl ControllerTrait for Controller {
|
|||||||
"Sending OpenAI image generation API request"
|
"Sending OpenAI image generation API request"
|
||||||
);
|
);
|
||||||
|
|
||||||
let response = self.client.images().create(request).await?;
|
let response = self.client.images().generate(request).await?;
|
||||||
|
|
||||||
if let Some(image) = response.data.into_iter().next() {
|
if let Some(image) = response.data.into_iter().next() {
|
||||||
match image.deref() {
|
match image.deref() {
|
||||||
async_openai::types::Image::B64Json {
|
Image::B64Json {
|
||||||
b64_json,
|
b64_json,
|
||||||
revised_prompt,
|
revised_prompt,
|
||||||
} => {
|
} => {
|
||||||
let bytes = base64_decode(b64_json)?;
|
let bytes = base64_decode(b64_json.as_ref())?;
|
||||||
|
|
||||||
return Ok(ImageGenerationResult {
|
return Ok(ImageGenerationResult {
|
||||||
bytes,
|
bytes,
|
||||||
@@ -374,15 +402,15 @@ impl ControllerTrait for Controller {
|
|||||||
return Err(anyhow::anyhow!("No image sources provided"));
|
return Err(anyhow::anyhow!("No image sources provided"));
|
||||||
}
|
}
|
||||||
|
|
||||||
let mut image_inputs = Vec::new();
|
let mut image_inputs: Vec<ImageInput> = Vec::new();
|
||||||
for image in images {
|
for image in images {
|
||||||
image_inputs.push(image.into());
|
image_inputs.push(image.into());
|
||||||
}
|
}
|
||||||
|
|
||||||
let dalle2_size = match image_generation_config.size {
|
let dalle2_size = match image_generation_config.size {
|
||||||
Some(async_openai::types::ImageSize::S256x256) => Some(DallE2ImageSize::S256x256),
|
Some(async_openai::types::images::ImageSize::S256x256) => Some(async_openai::types::images::ImageSize::S256x256),
|
||||||
Some(async_openai::types::ImageSize::S512x512) => Some(DallE2ImageSize::S512x512),
|
Some(async_openai::types::images::ImageSize::S512x512) => Some(async_openai::types::images::ImageSize::S512x512),
|
||||||
Some(async_openai::types::ImageSize::S1024x1024) => Some(DallE2ImageSize::S1024x1024),
|
Some(async_openai::types::images::ImageSize::S1024x1024) => Some(async_openai::types::images::ImageSize::S1024x1024),
|
||||||
_ => None,
|
_ => None,
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -391,16 +419,18 @@ impl ControllerTrait for Controller {
|
|||||||
.map_err(|err| anyhow::anyhow!(err))?;
|
.map_err(|err| anyhow::anyhow!(err))?;
|
||||||
|
|
||||||
let response_format = match model.clone() {
|
let response_format = match model.clone() {
|
||||||
async_openai::types::ImageModel::DallE2 => {
|
ImageModel::DallE2 => {
|
||||||
Some(async_openai::types::ImageResponseFormat::B64Json)
|
Some(ImageResponseFormat::B64Json)
|
||||||
}
|
}
|
||||||
async_openai::types::ImageModel::DallE3 => {
|
ImageModel::DallE3 => {
|
||||||
Some(async_openai::types::ImageResponseFormat::B64Json)
|
Some(ImageResponseFormat::B64Json)
|
||||||
}
|
}
|
||||||
async_openai::types::ImageModel::Other(model_str) => match model_str.as_str() {
|
// gpt-image-1 only outputs base64 and we don't need to specify the response format.
|
||||||
OPENAI_IMAGE_MODEL_GPT_IMAGE_1 => None,
|
// In fact, specifying the response format results in an error.
|
||||||
_ => Some(async_openai::types::ImageResponseFormat::B64Json),
|
ImageModel::GptImage1 => None,
|
||||||
},
|
ImageModel::GptImage1Mini => None,
|
||||||
|
ImageModel::GptImage1dot5 => None,
|
||||||
|
ImageModel::Other(_) => Some(ImageResponseFormat::B64Json),
|
||||||
};
|
};
|
||||||
|
|
||||||
let mut request_builder = CreateImageEditRequestArgs::default();
|
let mut request_builder = CreateImageEditRequestArgs::default();
|
||||||
@@ -429,12 +459,12 @@ impl ControllerTrait for Controller {
|
|||||||
"Sending OpenAI image edit API request"
|
"Sending OpenAI image edit API request"
|
||||||
);
|
);
|
||||||
|
|
||||||
let response = self.client.images().create_edit(request).await?;
|
let response = self.client.images().edit(request).await?;
|
||||||
|
|
||||||
if let Some(image_data) = response.data.into_iter().next() {
|
if let Some(image_data) = response.data.into_iter().next() {
|
||||||
match image_data.deref() {
|
match image_data.deref() {
|
||||||
Image::B64Json { b64_json, .. } => {
|
Image::B64Json { b64_json, .. } => {
|
||||||
let bytes = base64_decode(b64_json)?;
|
let bytes = base64_decode(b64_json.as_ref())?;
|
||||||
return Ok(ImageEditResult {
|
return Ok(ImageEditResult {
|
||||||
bytes,
|
bytes,
|
||||||
mime_type: mxlink::mime::IMAGE_PNG,
|
mime_type: mxlink::mime::IMAGE_PNG,
|
||||||
@@ -471,7 +501,7 @@ impl ControllerTrait for Controller {
|
|||||||
|
|
||||||
let voice = if let Some(voice_string) = params.voice_override {
|
let voice = if let Some(voice_string) = params.voice_override {
|
||||||
// This is a hacky way to construct a Voice enum from the string we have.
|
// This is a hacky way to construct a Voice enum from the string we have.
|
||||||
let voice: serde_json::Result<async_openai::types::Voice> =
|
let voice: serde_json::Result<async_openai::types::audio::Voice> =
|
||||||
serde_json::from_str(&format!("\"{}\"", voice_string));
|
serde_json::from_str(&format!("\"{}\"", voice_string));
|
||||||
match voice {
|
match voice {
|
||||||
Ok(voice) => voice,
|
Ok(voice) => voice,
|
||||||
@@ -511,7 +541,7 @@ impl ControllerTrait for Controller {
|
|||||||
"Sending OpenAI text-to-speech API request"
|
"Sending OpenAI text-to-speech API request"
|
||||||
);
|
);
|
||||||
|
|
||||||
let result = self.client.audio().speech(request).await?;
|
let result = self.client.audio().speech().create(request).await?;
|
||||||
|
|
||||||
Ok(TextToSpeechResult {
|
Ok(TextToSpeechResult {
|
||||||
bytes: result.bytes.into(),
|
bytes: result.bytes.into(),
|
||||||
@@ -570,15 +600,15 @@ impl ControllerTrait for Controller {
|
|||||||
}
|
}
|
||||||
|
|
||||||
fn response_format_to_mime_type(
|
fn response_format_to_mime_type(
|
||||||
response_format: &async_openai::types::SpeechResponseFormat,
|
response_format: &async_openai::types::audio::SpeechResponseFormat,
|
||||||
) -> Option<mxlink::mime::Mime> {
|
) -> Option<mxlink::mime::Mime> {
|
||||||
let content_type = match response_format {
|
let content_type = match response_format {
|
||||||
async_openai::types::SpeechResponseFormat::Mp3 => "audio/mp3".to_owned(),
|
async_openai::types::audio::SpeechResponseFormat::Mp3 => "audio/mp3".to_owned(),
|
||||||
async_openai::types::SpeechResponseFormat::Wav => "audio/wav".to_owned(),
|
async_openai::types::audio::SpeechResponseFormat::Wav => "audio/wav".to_owned(),
|
||||||
async_openai::types::SpeechResponseFormat::Opus => "audio/ogg".to_owned(),
|
async_openai::types::audio::SpeechResponseFormat::Opus => "audio/ogg".to_owned(),
|
||||||
async_openai::types::SpeechResponseFormat::Aac => "audio/aac".to_owned(),
|
async_openai::types::audio::SpeechResponseFormat::Aac => "audio/aac".to_owned(),
|
||||||
async_openai::types::SpeechResponseFormat::Flac => "audio/flac".to_owned(),
|
async_openai::types::audio::SpeechResponseFormat::Flac => "audio/flac".to_owned(),
|
||||||
async_openai::types::SpeechResponseFormat::Pcm => "audio/L8".to_owned(),
|
async_openai::types::audio::SpeechResponseFormat::Pcm => "audio/L8".to_owned(),
|
||||||
};
|
};
|
||||||
|
|
||||||
match content_type.parse() {
|
match content_type.parse() {
|
||||||
@@ -606,3 +636,17 @@ fn audio_mime_type_to_file_name(mime_type: &mxlink::mime::Mime) -> Option<String
|
|||||||
|
|
||||||
Some(format!("audio.{}", file_extension))
|
Some(format!("audio.{}", file_extension))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Returns the smallest supported size for stickers based on what the image model supports.
|
||||||
|
fn get_sticker_size(model: &ImageModel) -> async_openai::types::images::ImageSize {
|
||||||
|
use async_openai::types::images::ImageSize;
|
||||||
|
|
||||||
|
match model {
|
||||||
|
ImageModel::DallE2 => ImageSize::S256x256,
|
||||||
|
ImageModel::DallE3 => ImageSize::S1024x1024,
|
||||||
|
ImageModel::GptImage1 => ImageSize::S1024x1024,
|
||||||
|
ImageModel::GptImage1Mini => ImageSize::S1024x1024,
|
||||||
|
ImageModel::GptImage1dot5 => ImageSize::S1024x1024,
|
||||||
|
ImageModel::Other(_) => ImageSize::S1024x1024,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
@@ -16,7 +16,7 @@ use super::super::AgentInstantiationResult;
|
|||||||
use super::ConfigTrait;
|
use super::ConfigTrait;
|
||||||
use super::controller::ControllerType;
|
use super::controller::ControllerType;
|
||||||
|
|
||||||
pub const OPENAI_IMAGE_MODEL_GPT_IMAGE_1: &str = "gpt-image-1";
|
pub const OPENAI_IMAGE_MODEL_GPT_IMAGE_1_DOT_5: &str = "gpt-image-1.5";
|
||||||
|
|
||||||
pub fn create_controller_from_yaml_value_config(
|
pub fn create_controller_from_yaml_value_config(
|
||||||
agent_id: &str,
|
agent_id: &str,
|
||||||
|
|||||||
@@ -1,8 +1,6 @@
|
|||||||
use async_openai::types::{
|
use async_openai::types::responses::{
|
||||||
ChatCompletionRequestAssistantMessageArgs, ChatCompletionRequestMessage,
|
EasyInputContent, EasyInputMessage, ImageDetail, InputContent, InputImageContent, InputItem,
|
||||||
ChatCompletionRequestMessageContentPartImage, ChatCompletionRequestSystemMessageArgs,
|
InputParam, MessageType, Role,
|
||||||
ChatCompletionRequestUserMessageArgs, ChatCompletionRequestUserMessageContent,
|
|
||||||
ChatCompletionRequestUserMessageContentPart, ImageUrlArgs,
|
|
||||||
};
|
};
|
||||||
|
|
||||||
use crate::conversation::llm::{
|
use crate::conversation::llm::{
|
||||||
@@ -10,93 +8,41 @@ use crate::conversation::llm::{
|
|||||||
};
|
};
|
||||||
use crate::utils::base64::base64_encode;
|
use crate::utils::base64::base64_encode;
|
||||||
|
|
||||||
pub fn convert_llm_messages_to_openai_messages(
|
pub fn convert_llm_messages_to_openai_response_input(
|
||||||
conversation_messages: Vec<LLMMessage>,
|
conversation_messages: Vec<LLMMessage>,
|
||||||
) -> Vec<ChatCompletionRequestMessage> {
|
) -> InputParam {
|
||||||
let mut openai_conversation_messages: Vec<ChatCompletionRequestMessage> =
|
let mut items = Vec::with_capacity(conversation_messages.len());
|
||||||
Vec::with_capacity(conversation_messages.len());
|
|
||||||
|
|
||||||
for message in conversation_messages {
|
for message in conversation_messages {
|
||||||
let openai_message = convert_llm_message_to_openai_message(message);
|
let role = match message.author {
|
||||||
if let Some(openai_message) = openai_message {
|
LLMAuthor::Prompt => Role::System,
|
||||||
openai_conversation_messages.push(openai_message);
|
LLMAuthor::Assistant => Role::Assistant,
|
||||||
}
|
LLMAuthor::User => Role::User,
|
||||||
}
|
};
|
||||||
|
|
||||||
openai_conversation_messages
|
let content = match message.content {
|
||||||
}
|
LLMMessageContent::Text(text) => EasyInputContent::Text(text),
|
||||||
|
LLMMessageContent::Image(image_details) => {
|
||||||
|
let image_url = format!(
|
||||||
|
"data:{};base64,{}",
|
||||||
|
image_details.mime,
|
||||||
|
base64_encode(&image_details.data)
|
||||||
|
);
|
||||||
|
|
||||||
fn convert_llm_message_to_openai_message(
|
EasyInputContent::ContentList(vec![InputContent::InputImage(InputImageContent {
|
||||||
llm_message: LLMMessage,
|
image_url: Some(image_url),
|
||||||
) -> Option<ChatCompletionRequestMessage> {
|
detail: ImageDetail::Auto,
|
||||||
match &llm_message.content {
|
file_id: None,
|
||||||
LLMMessageContent::Text(text) => Some(match llm_message.author {
|
})])
|
||||||
LLMAuthor::Prompt => ChatCompletionRequestSystemMessageArgs::default()
|
|
||||||
.content(text.clone())
|
|
||||||
.build()
|
|
||||||
.expect("Failed building OpenAI system message")
|
|
||||||
.into(),
|
|
||||||
LLMAuthor::Assistant => ChatCompletionRequestAssistantMessageArgs::default()
|
|
||||||
.content(text.clone())
|
|
||||||
.build()
|
|
||||||
.expect("Failed building OpenAI assistant message")
|
|
||||||
.into(),
|
|
||||||
LLMAuthor::User => ChatCompletionRequestUserMessageArgs::default()
|
|
||||||
.content(text.clone())
|
|
||||||
.build()
|
|
||||||
.expect("Failed building OpenAI user message")
|
|
||||||
.into(),
|
|
||||||
}),
|
|
||||||
LLMMessageContent::Image(image_details) => {
|
|
||||||
let image_url = format!(
|
|
||||||
"data:{};base64,{}",
|
|
||||||
image_details.mime,
|
|
||||||
base64_encode(&image_details.data)
|
|
||||||
);
|
|
||||||
|
|
||||||
let part = ChatCompletionRequestUserMessageContentPart::ImageUrl(
|
|
||||||
ChatCompletionRequestMessageContentPartImage {
|
|
||||||
image_url: ImageUrlArgs::default()
|
|
||||||
.url(image_url)
|
|
||||||
.build()
|
|
||||||
.expect("Failed building OpenAI image url"),
|
|
||||||
},
|
|
||||||
);
|
|
||||||
|
|
||||||
let message_content = ChatCompletionRequestUserMessageContent::Array(vec![part]);
|
|
||||||
|
|
||||||
match llm_message.author {
|
|
||||||
LLMAuthor::User => Some(
|
|
||||||
ChatCompletionRequestUserMessageArgs::default()
|
|
||||||
.content(message_content)
|
|
||||||
.build()
|
|
||||||
.expect("Failed building OpenAI user message")
|
|
||||||
.into(),
|
|
||||||
),
|
|
||||||
_ => {
|
|
||||||
tracing::warn!(
|
|
||||||
"OpenAI API does not support image content for messages authored by {:?}. This message part will be skipped.",
|
|
||||||
llm_message.author
|
|
||||||
);
|
|
||||||
None
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
}
|
};
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
pub(super) fn convert_string_to_enum<T>(value: &str) -> Result<T, String>
|
items.push(InputItem::EasyMessage(EasyInputMessage {
|
||||||
where
|
r#type: MessageType::Message,
|
||||||
T: serde::de::DeserializeOwned,
|
role,
|
||||||
{
|
content,
|
||||||
// This is a hacky way to construct an enum from the string we have.
|
}));
|
||||||
let enum_result: serde_json::Result<T> = serde_json::from_str(&format!("\"{}\"", value));
|
|
||||||
match enum_result {
|
|
||||||
Ok(enum_result) => Ok(enum_result),
|
|
||||||
Err(err) => {
|
|
||||||
tracing::debug!(?err, "Failed to parse into enum");
|
|
||||||
|
|
||||||
Err(format!("The value ({}) is not supported.", value))
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
|
InputParam::Items(items)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -95,6 +95,7 @@ impl TryInto<OpenAITextGenerationConfig> for TextGenerationConfig {
|
|||||||
max_response_tokens: self.max_response_tokens,
|
max_response_tokens: self.max_response_tokens,
|
||||||
max_completion_tokens: None,
|
max_completion_tokens: None,
|
||||||
max_context_tokens: self.max_context_tokens,
|
max_context_tokens: self.max_context_tokens,
|
||||||
|
tools: Default::default(),
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -161,11 +162,11 @@ impl TryInto<OpenAITextToSpeechConfig> for TextToSpeechConfig {
|
|||||||
type Error = String;
|
type Error = String;
|
||||||
|
|
||||||
fn try_into(self) -> Result<OpenAITextToSpeechConfig, Self::Error> {
|
fn try_into(self) -> Result<OpenAITextToSpeechConfig, Self::Error> {
|
||||||
let model_id = convert_string_to_enum::<async_openai::types::SpeechModel>(&self.model_id)?;
|
let model_id = convert_string_to_enum::<async_openai::types::audio::SpeechModel>(&self.model_id)?;
|
||||||
|
|
||||||
let voice = convert_string_to_enum::<async_openai::types::Voice>(&self.voice)?;
|
let voice = convert_string_to_enum::<async_openai::types::audio::Voice>(&self.voice)?;
|
||||||
|
|
||||||
let response_format = convert_string_to_enum::<async_openai::types::SpeechResponseFormat>(
|
let response_format = convert_string_to_enum::<async_openai::types::audio::SpeechResponseFormat>(
|
||||||
&self.response_format,
|
&self.response_format,
|
||||||
)?;
|
)?;
|
||||||
|
|
||||||
@@ -224,7 +225,7 @@ impl TryInto<OpenAIImageGenerationConfig> for ImageGenerationConfig {
|
|||||||
|
|
||||||
fn try_into(self) -> Result<OpenAIImageGenerationConfig, Self::Error> {
|
fn try_into(self) -> Result<OpenAIImageGenerationConfig, Self::Error> {
|
||||||
let size = if let Some(size) = &self.size {
|
let size = if let Some(size) = &self.size {
|
||||||
Some(convert_string_to_enum::<async_openai::types::ImageSize>(
|
Some(convert_string_to_enum::<async_openai::types::images::ImageSize>(
|
||||||
size,
|
size,
|
||||||
)?)
|
)?)
|
||||||
} else {
|
} else {
|
||||||
@@ -232,7 +233,7 @@ impl TryInto<OpenAIImageGenerationConfig> for ImageGenerationConfig {
|
|||||||
};
|
};
|
||||||
|
|
||||||
let style = if let Some(style) = &self.style {
|
let style = if let Some(style) = &self.style {
|
||||||
Some(convert_string_to_enum::<async_openai::types::ImageStyle>(
|
Some(convert_string_to_enum::<async_openai::types::images::ImageStyle>(
|
||||||
style,
|
style,
|
||||||
)?)
|
)?)
|
||||||
} else {
|
} else {
|
||||||
@@ -240,7 +241,7 @@ impl TryInto<OpenAIImageGenerationConfig> for ImageGenerationConfig {
|
|||||||
};
|
};
|
||||||
|
|
||||||
let quality = if let Some(quality) = &self.quality {
|
let quality = if let Some(quality) = &self.quality {
|
||||||
Some(convert_string_to_enum::<async_openai::types::ImageQuality>(
|
Some(convert_string_to_enum::<async_openai::types::images::ImageQuality>(
|
||||||
quality,
|
quality,
|
||||||
)?)
|
)?)
|
||||||
} else {
|
} else {
|
||||||
|
|||||||
@@ -3,6 +3,8 @@ use etke_openai_api_rust::chat::{ChatApi, ChatBody};
|
|||||||
use etke_openai_api_rust::images::{ImagesApi, ImagesBody};
|
use etke_openai_api_rust::images::{ImagesApi, ImagesBody};
|
||||||
use etke_openai_api_rust::{Auth, Message, OpenAI};
|
use etke_openai_api_rust::{Auth, Message, OpenAI};
|
||||||
|
|
||||||
|
const SMALLEST_IMAGE_SIZE: &str = "256x256";
|
||||||
|
|
||||||
use super::super::ControllerTrait;
|
use super::super::ControllerTrait;
|
||||||
use crate::utils::base64::base64_decode;
|
use crate::utils::base64::base64_decode;
|
||||||
use crate::{
|
use crate::{
|
||||||
@@ -303,9 +305,11 @@ impl ControllerTrait for Controller {
|
|||||||
// when they span multiple lines.
|
// when they span multiple lines.
|
||||||
let prompt = prompt.replace("\n", " ");
|
let prompt = prompt.replace("\n", " ");
|
||||||
|
|
||||||
let size: Option<String> = params
|
let size: Option<String> = if params.smallest_size_possible {
|
||||||
.size_override
|
Some(SMALLEST_IMAGE_SIZE.to_owned())
|
||||||
.or_else(|| image_generation_config.size.clone());
|
} else {
|
||||||
|
image_generation_config.size.clone()
|
||||||
|
};
|
||||||
|
|
||||||
let request = ImagesBody {
|
let request = ImagesBody {
|
||||||
model: Some(image_generation_config.model_id.to_owned()),
|
model: Some(image_generation_config.model_id.to_owned()),
|
||||||
|
|||||||
@@ -1,3 +1,4 @@
|
|||||||
|
use std::fs;
|
||||||
use std::sync::Arc;
|
use std::sync::Arc;
|
||||||
use std::{future::Future, pin::Pin};
|
use std::{future::Future, pin::Pin};
|
||||||
|
|
||||||
@@ -18,12 +19,13 @@ use mxlink::helpers::account_data_config::{
|
|||||||
RoomConfigManager as AccountDataRoomConfigManager,
|
RoomConfigManager as AccountDataRoomConfigManager,
|
||||||
};
|
};
|
||||||
use mxlink::helpers::encryption::Manager as EncryptionManager;
|
use mxlink::helpers::encryption::Manager as EncryptionManager;
|
||||||
|
use mxlink::mime::Mime;
|
||||||
|
|
||||||
use crate::agent::Manager as AgentManager;
|
use crate::agent::Manager as AgentManager;
|
||||||
use crate::entity::catch_up_marker::{
|
use crate::entity::catch_up_marker::{
|
||||||
CatchUpMarker, CatchUpMarkerManager, DelayedCatchUpMarkerManager,
|
CatchUpMarker, CatchUpMarkerManager, DelayedCatchUpMarkerManager,
|
||||||
};
|
};
|
||||||
use crate::entity::cfg::Config;
|
use crate::entity::cfg::{Avatar, Config};
|
||||||
use crate::entity::globalconfig::{GlobalConfig, GlobalConfigurationManager};
|
use crate::entity::globalconfig::{GlobalConfig, GlobalConfigurationManager};
|
||||||
use crate::entity::roomconfig::{RoomConfig, RoomConfigurationManager};
|
use crate::entity::roomconfig::{RoomConfig, RoomConfigurationManager};
|
||||||
|
|
||||||
@@ -316,34 +318,72 @@ impl Bot {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
let should_update_avatar = match ¤t_avatar_url {
|
let desired_avatar: Option<(Vec<u8>, Mime)> = match &self.inner.config.user.avatar {
|
||||||
Some(avatar_url) => {
|
Avatar::Keep => {
|
||||||
let request = MediaRequestParameters {
|
tracing::info!("Avatar configured to keep current, skipping avatar management");
|
||||||
source: MediaSource::Plain(avatar_url.to_owned()),
|
None
|
||||||
format: MediaFormat::File,
|
}
|
||||||
};
|
Avatar::Default => {
|
||||||
|
tracing::info!("Avatar configured to use default");
|
||||||
let content = media
|
Some((
|
||||||
.get_media_content(&request, true)
|
LOGO_BYTES.to_vec(),
|
||||||
.await
|
LOGO_MIME_TYPE
|
||||||
.map_err(|e| anyhow::anyhow!("Failed fetching existing avatar: {:?}", e))?;
|
.parse()
|
||||||
|
.expect("Failed parsing mime type for logo"),
|
||||||
content.as_slice() != LOGO_BYTES
|
))
|
||||||
|
}
|
||||||
|
Avatar::Custom(avatar_path) => {
|
||||||
|
tracing::info!(?avatar_path, "Avatar configured to use custom path");
|
||||||
|
let bytes = fs::read(avatar_path).map_err(|e| {
|
||||||
|
anyhow::anyhow!("Failed reading avatar from {:?}: {:?}", avatar_path, e)
|
||||||
|
})?;
|
||||||
|
let mime = mime_guess::from_path(avatar_path).first_or_octet_stream();
|
||||||
|
tracing::debug!(?mime, bytes_len = bytes.len(), "Loaded custom avatar");
|
||||||
|
Some((bytes, mime))
|
||||||
}
|
}
|
||||||
None => true,
|
|
||||||
};
|
};
|
||||||
|
|
||||||
if should_update_avatar {
|
if let Some((desired_bytes, mime_type)) = desired_avatar {
|
||||||
tracing::info!("Updating avatar..");
|
let should_update_avatar = match ¤t_avatar_url {
|
||||||
|
Some(avatar_url) => {
|
||||||
|
tracing::debug!(?avatar_url, "Fetching current avatar to compare");
|
||||||
|
let request = MediaRequestParameters {
|
||||||
|
source: MediaSource::Plain(avatar_url.to_owned()),
|
||||||
|
format: MediaFormat::File,
|
||||||
|
};
|
||||||
|
|
||||||
let mime_type = LOGO_MIME_TYPE
|
let content = media
|
||||||
.parse()
|
.get_media_content(&request, true)
|
||||||
.expect("Failed parsing mime type for logo");
|
.await
|
||||||
|
.map_err(|e| anyhow::anyhow!("Failed fetching existing avatar: {:?}", e))?;
|
||||||
|
|
||||||
account
|
let needs_update = content.as_slice() != desired_bytes;
|
||||||
.upload_avatar(&mime_type, LOGO_BYTES.to_vec())
|
|
||||||
.await
|
tracing::debug!(
|
||||||
.map_err(|e| anyhow::anyhow!("Failed uploading avatar: {:?}", e))?;
|
current_bytes_len = content.len(),
|
||||||
|
desired_bytes_len = desired_bytes.len(),
|
||||||
|
?needs_update,
|
||||||
|
"Compared current and desired avatar"
|
||||||
|
);
|
||||||
|
|
||||||
|
needs_update
|
||||||
|
}
|
||||||
|
None => {
|
||||||
|
tracing::debug!("No current avatar set, will upload");
|
||||||
|
true
|
||||||
|
}
|
||||||
|
};
|
||||||
|
|
||||||
|
if should_update_avatar {
|
||||||
|
tracing::info!("Updating avatar..");
|
||||||
|
account
|
||||||
|
.upload_avatar(&mime_type, desired_bytes)
|
||||||
|
.await
|
||||||
|
.map_err(|e| anyhow::anyhow!("Failed uploading avatar: {:?}", e))?;
|
||||||
|
tracing::info!("Avatar updated successfully");
|
||||||
|
} else {
|
||||||
|
tracing::debug!("Avatar already up to date, skipping upload");
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
Ok(())
|
Ok(())
|
||||||
|
|||||||
@@ -5,7 +5,7 @@ use anyhow::anyhow;
|
|||||||
|
|
||||||
use crate::agent::AgentPurpose;
|
use crate::agent::AgentPurpose;
|
||||||
|
|
||||||
pub use crate::entity::cfg::{Config, defaults as cfg_defaults, env as cfg_env};
|
pub use crate::entity::cfg::{Avatar, Config, defaults as cfg_defaults, env as cfg_env};
|
||||||
|
|
||||||
pub fn load() -> anyhow::Result<Config> {
|
pub fn load() -> anyhow::Result<Config> {
|
||||||
let config_file_path = env::var(cfg_env::BAIBOT_CONFIG_FILE_PATH)
|
let config_file_path = env::var(cfg_env::BAIBOT_CONFIG_FILE_PATH)
|
||||||
@@ -33,7 +33,13 @@ pub fn load() -> anyhow::Result<Config> {
|
|||||||
cfg_env::BAIBOT_USER_ENCRYPTION_RECOVERY_PASSPHRASE => {
|
cfg_env::BAIBOT_USER_ENCRYPTION_RECOVERY_PASSPHRASE => {
|
||||||
config.user.encryption.recovery_passphrase = Some(value);
|
config.user.encryption.recovery_passphrase = Some(value);
|
||||||
}
|
}
|
||||||
|
cfg_env::BAIBOT_USER_ENCRYPTION_RECOVERY_RESET_ALLOWED => {
|
||||||
|
config.user.encryption.recovery_reset_allowed = value.parse::<bool>()?;
|
||||||
|
}
|
||||||
cfg_env::BAIBOT_USER_NAME => config.user.name = value,
|
cfg_env::BAIBOT_USER_NAME => config.user.name = value,
|
||||||
|
cfg_env::BAIBOT_USER_AVATAR => {
|
||||||
|
config.user.avatar = Avatar::from_string(value);
|
||||||
|
}
|
||||||
cfg_env::BAIBOT_COMMAND_PREFIX => config.command_prefix = value,
|
cfg_env::BAIBOT_COMMAND_PREFIX => config.command_prefix = value,
|
||||||
cfg_env::BAIBOT_ROOM_POST_JOIN_SELF_INTRODUCTION_ENABLED => {
|
cfg_env::BAIBOT_ROOM_POST_JOIN_SELF_INTRODUCTION_ENABLED => {
|
||||||
config.room.post_join_self_introduction_enabled = value.parse::<bool>()?;
|
config.room.post_join_self_introduction_enabled = value.parse::<bool>()?;
|
||||||
@@ -51,6 +57,9 @@ pub fn load() -> anyhow::Result<Config> {
|
|||||||
cfg_env::BAIBOT_PERSISTENCE_DATA_DIR_PATH => {
|
cfg_env::BAIBOT_PERSISTENCE_DATA_DIR_PATH => {
|
||||||
config.persistence.data_dir_path = Some(value);
|
config.persistence.data_dir_path = Some(value);
|
||||||
}
|
}
|
||||||
|
cfg_env::BAIBOT_PERSISTENCE_SESSION_ENCRYPTION_KEY => {
|
||||||
|
config.persistence.session_encryption_key = Some(value);
|
||||||
|
}
|
||||||
cfg_env::BAIBOT_PERSISTENCE_CONFIG_ENCRYPTION_KEY => {
|
cfg_env::BAIBOT_PERSISTENCE_CONFIG_ENCRYPTION_KEY => {
|
||||||
config.persistence.config_encryption_key = Some(value);
|
config.persistence.config_encryption_key = Some(value);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -12,9 +12,6 @@ use crate::strings;
|
|||||||
use crate::utils::mime::get_file_extension;
|
use crate::utils::mime::get_file_extension;
|
||||||
use crate::{Bot, entity::MessageContext};
|
use crate::{Bot, entity::MessageContext};
|
||||||
|
|
||||||
// We may make this configurable (per room, etc.) in the future, but for now it's hardcoded.
|
|
||||||
const STICKER_SIZE: &str = "256x256";
|
|
||||||
|
|
||||||
pub async fn handle_image(
|
pub async fn handle_image(
|
||||||
bot: &Bot,
|
bot: &Bot,
|
||||||
matrix_link: MatrixLink,
|
matrix_link: MatrixLink,
|
||||||
@@ -177,7 +174,7 @@ pub async fn handle_sticker(
|
|||||||
);
|
);
|
||||||
|
|
||||||
let params = ImageGenerationParams::default()
|
let params = ImageGenerationParams::default()
|
||||||
.with_size_override(Some(STICKER_SIZE.to_owned()))
|
.with_smallest_size_possible(true)
|
||||||
.with_cheaper_model_switching_allowed(true)
|
.with_cheaper_model_switching_allowed(true)
|
||||||
.with_cheaper_quality_switching_allowed(true);
|
.with_cheaper_quality_switching_allowed(true);
|
||||||
|
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
use std::path::PathBuf;
|
use std::path::PathBuf;
|
||||||
|
|
||||||
use mxlink::helpers::encryption::EncryptionKey;
|
use mxlink::helpers::encryption::EncryptionKey;
|
||||||
use serde::{Deserialize, Serialize};
|
use serde::{Deserialize, Deserializer, Serialize};
|
||||||
|
|
||||||
use crate::{
|
use crate::{
|
||||||
agent::{AgentDefinition, AgentPurpose, PublicIdentifier},
|
agent::{AgentDefinition, AgentPurpose, PublicIdentifier},
|
||||||
@@ -83,6 +83,52 @@ impl ConfigHomeserver {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Configuration for the bot's avatar.
|
||||||
|
///
|
||||||
|
/// - `Default`: Use the built-in default avatar (null, empty string, or missing in config)
|
||||||
|
/// - `Keep`: Don't touch the avatar, keep whatever is already set ("keep" in config)
|
||||||
|
/// - `Custom(String)`: Use a custom avatar from the specified file path
|
||||||
|
#[derive(Debug, Clone, PartialEq, Serialize)]
|
||||||
|
pub enum Avatar {
|
||||||
|
/// Use the built-in default avatar
|
||||||
|
Default,
|
||||||
|
/// Keep the current avatar, don't change it
|
||||||
|
Keep,
|
||||||
|
/// Use a custom avatar from the specified file path
|
||||||
|
Custom(String),
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Default for Avatar {
|
||||||
|
fn default() -> Self {
|
||||||
|
Avatar::Default
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
impl<'de> Deserialize<'de> for Avatar {
|
||||||
|
fn deserialize<D>(deserializer: D) -> Result<Self, D::Error>
|
||||||
|
where
|
||||||
|
D: Deserializer<'de>,
|
||||||
|
{
|
||||||
|
let value: Option<String> = Option::deserialize(deserializer)?;
|
||||||
|
Ok(match value {
|
||||||
|
None => Avatar::Default,
|
||||||
|
Some(s) => Avatar::from_string(s),
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Avatar {
|
||||||
|
pub fn from_string(value: String) -> Self {
|
||||||
|
if value.is_empty() {
|
||||||
|
Avatar::Default
|
||||||
|
} else if value.eq_ignore_ascii_case("keep") {
|
||||||
|
Avatar::Keep
|
||||||
|
} else {
|
||||||
|
Avatar::Custom(value)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
#[derive(Debug, Serialize, Deserialize)]
|
#[derive(Debug, Serialize, Deserialize)]
|
||||||
pub struct ConfigUser {
|
pub struct ConfigUser {
|
||||||
pub mxid_localpart: String,
|
pub mxid_localpart: String,
|
||||||
@@ -93,6 +139,9 @@ pub struct ConfigUser {
|
|||||||
|
|
||||||
#[serde(default)]
|
#[serde(default)]
|
||||||
pub encryption: ConfigUserEncryption,
|
pub encryption: ConfigUserEncryption,
|
||||||
|
|
||||||
|
#[serde(default)]
|
||||||
|
pub avatar: Avatar,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl ConfigUser {
|
impl ConfigUser {
|
||||||
|
|||||||
@@ -6,8 +6,11 @@ pub const BAIBOT_HOMESERVER_URL: &str = "BAIBOT_HOMESERVER_URL";
|
|||||||
pub const BAIBOT_USER_MXID_LOCALPART: &str = "BAIBOT_USER_MXID_LOCALPART";
|
pub const BAIBOT_USER_MXID_LOCALPART: &str = "BAIBOT_USER_MXID_LOCALPART";
|
||||||
pub const BAIBOT_USER_PASSWORD: &str = "BAIBOT_USER_PASSWORD";
|
pub const BAIBOT_USER_PASSWORD: &str = "BAIBOT_USER_PASSWORD";
|
||||||
pub const BAIBOT_USER_NAME: &str = "BAIBOT_USER_NAME";
|
pub const BAIBOT_USER_NAME: &str = "BAIBOT_USER_NAME";
|
||||||
|
pub const BAIBOT_USER_AVATAR: &str = "BAIBOT_USER_AVATAR";
|
||||||
pub const BAIBOT_USER_ENCRYPTION_RECOVERY_PASSPHRASE: &str =
|
pub const BAIBOT_USER_ENCRYPTION_RECOVERY_PASSPHRASE: &str =
|
||||||
"BAIBOT_USER_ENCRYPTION_RECOVERY_PASSPHRASE";
|
"BAIBOT_USER_ENCRYPTION_RECOVERY_PASSPHRASE";
|
||||||
|
pub const BAIBOT_USER_ENCRYPTION_RECOVERY_RESET_ALLOWED: &str =
|
||||||
|
"BAIBOT_USER_ENCRYPTION_RECOVERY_RESET_ALLOWED";
|
||||||
|
|
||||||
pub const BAIBOT_COMMAND_PREFIX: &str = "BAIBOT_COMMAND_PREFIX";
|
pub const BAIBOT_COMMAND_PREFIX: &str = "BAIBOT_COMMAND_PREFIX";
|
||||||
|
|
||||||
|
|||||||
@@ -2,4 +2,4 @@ mod config;
|
|||||||
pub mod defaults;
|
pub mod defaults;
|
||||||
pub mod env;
|
pub mod env;
|
||||||
|
|
||||||
pub use config::Config;
|
pub use config::{Avatar, Config};
|
||||||
|
|||||||
@@ -109,11 +109,21 @@ pub fn help_provider_details(id: &str, info: &AgentProviderInfo) -> String {
|
|||||||
let mut purpose_line = format!("{} {}", purpose.emoji(), purpose.as_str());
|
let mut purpose_line = format!("{} {}", purpose.emoji(), purpose.as_str());
|
||||||
|
|
||||||
if let AgentPurpose::TextGeneration = purpose {
|
if let AgentPurpose::TextGeneration = purpose {
|
||||||
|
let mut extras = vec![];
|
||||||
|
|
||||||
if info.text_generation_supports_vision {
|
if info.text_generation_supports_vision {
|
||||||
purpose_line = format!("{} ({})", purpose_line, "incl. vision");
|
extras.push("incl. vision");
|
||||||
} else {
|
} else {
|
||||||
purpose_line = format!("{} ({})", purpose_line, "no vision");
|
extras.push("no vision");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
if info.text_generation_supports_tools {
|
||||||
|
extras.push("incl. tools");
|
||||||
|
} else {
|
||||||
|
extras.push("no tools");
|
||||||
|
}
|
||||||
|
|
||||||
|
purpose_line = format!("{} ({})", purpose_line, extras.join(", "));
|
||||||
}
|
}
|
||||||
|
|
||||||
capabilities.push(purpose_line);
|
capabilities.push(purpose_line);
|
||||||
|
|||||||
@@ -64,7 +64,7 @@ To create a sticker, send a command like `%command_prefix% sticker A huge bowl o
|
|||||||
|
|
||||||
The difference from **creating images** is that the bot will:
|
The difference from **creating images** is that the bot will:
|
||||||
|
|
||||||
- create a smaller-resolution image (`256x256`) - smaller/quicker, but still good enough for a sticker
|
- create a smaller-resolution image (as small as the model allows) - smaller/quicker, but still good enough for a sticker
|
||||||
- potentially switch to a different (cheaper or otherwise more suitable) model, if available
|
- potentially switch to a different (cheaper or otherwise more suitable) model, if available
|
||||||
- post the image directly to the room (as a reply to your message), without starting a threaded conversation
|
- post the image directly to the room (as a reply to your message), without starting a threaded conversation
|
||||||
"#;
|
"#;
|
||||||
|
|||||||
Reference in New Issue
Block a user