Compare commits

..

16 Commits

Author SHA1 Message Date
Slavi Pantaleev
61d18b2e13 Release 1.14.0 2026-02-04 03:19:30 +02:00
Slavi Pantaleev
d831c08306 Add support for OpenAI built-in tools (web_search, code_interpreter)
Document the new tools feature in README, features.md, and providers.md.
Update the `!bai providers` command output to show vision/tools support
consistently for all providers.

Based on #62 by @yeslayla which migrated the OpenAI provider to the
Responses API.

See: https://github.com/etkecc/baibot/pull/62
2026-02-04 03:17:19 +02:00
Slavi Pantaleev
c70387b0c3 Fix sticker generation for newer GPT image models
Sticker generation was failing when using newer GPT image models
(gpt-image-1, gpt-image-1-mini, gpt-image-1.5). The issue occurred
because stickers requested 256x256 size, but these models only support
1024x1024, 1536x1024, 1024x1536, and auto.

To reproduce, send `!bai sticker Something` to an agent configured
with a GPT image model. The error was:

  invalid_request_error: Invalid value: '256x256'. Supported values
  are: '1024x1024', '1024x1536', '1536x1024', and 'auto'. (param: size)
  (code: invalid_value)

The fix replaces the hardcoded 256x256 size override with a
`smallest_size_possible` flag, letting each provider determine the
appropriate sticker size based on the model being used.

The `openai_compat` provider still defaults to requesting 256x256 in all cases
(regardless of model name).
2026-02-04 02:39:51 +02:00
Slavi Pantaleev
38516f2e17 Update services 2026-02-04 02:28:12 +02:00
Slavi Pantaleev
5481b5a763 Update dependencies 2026-02-04 01:54:53 +02:00
Slavi Pantaleev
26bc437678 Add tools config to OpenAI example in config.yml.dist 2026-02-04 01:41:27 +02:00
Layla
ec93f1ee2a Implement OpenAI's response API and add support for built-in tools (web search & code interpreter). 2026-02-04 01:41:27 +02:00
Slavi Pantaleev
b920b6e556 Release 1.13.0 2026-01-23 00:33:29 +02:00
Slavi Pantaleev
7136d34843 Update dependencies 2026-01-23 00:23:25 +02:00
Slavi Pantaleev
e0b4a40dd8 Configure switching to cheaper models (gpt-image-1-mini) for gpt-image-1 & gpt-image-1.5 2026-01-22 22:31:51 +02:00
Slavi Pantaleev
691aeeb1c7 Upgrade Rust (1.92.0 -> 1.93.0) 2026-01-22 22:23:18 +02:00
Slavi Pantaleev
257ffae9e7 Release 1.12.0 2025-12-21 23:27:08 +02:00
Slavi Pantaleev
f7bf3d7b60 Upgrade async-openai (0.32.1 -> 0.32.2)
This brings in an important bugfix related to image generation.
Ref:
- https://github.com/64bit/async-openai/issues/507
- https://github.com/64bit/async-openai/pull/508
2025-12-21 23:21:03 +02:00
Slavi Pantaleev
3a88b0d656 Make gpt-image-1.5 the default image model for OpenAI 2025-12-21 12:53:44 +02:00
Slavi Pantaleev
08c689a889 Upgrade async-openai (0.31.1 -> 0.32.1) and adapt, adding support for gpt-image-1.5 2025-12-21 12:23:25 +02:00
Slavi Pantaleev
ae8e878817 Update services 2025-12-21 11:33:08 +02:00
24 changed files with 517 additions and 448 deletions

View File

@@ -1,3 +1,28 @@
# (2026-02-04) Version 1.14.0
- (**Feature**) The `openai` provider now uses OpenAI's [Responses API](https://platform.openai.com/docs/api-reference/responses) (instead of the older Chat Completions API), adding support for [🛠️ built-in tools](./docs/features.md#️-built-in-tools-openai-only) (`web_search` and `code_interpreter`). These tools are **disabled by default** and can be enabled via the `text_generation.tools` configuration (see the [sample configuration](https://github.com/etkecc/baibot/blob/c70387b0c38d8d0f30bba2179a2a21a3710dbeaf/docs/sample-provider-configs/openai.yml#L12-L15)). To enable tools on an existing agent, you need to [update the agent](./docs/agents.md#updating-agents) to re-create it with the `text_generation.tools` section added and enable the tools you need. Thanks to [Layla Manley](https://github.com/yeslayla) for the contribution in [#62](https://github.com/etkecc/baibot/pull/62)!
- (**Bugfix**) Fix sticker generation for newer GPT image models (`gpt-image-1`, `gpt-image-1-mini`, `gpt-image-1.5`) which don't support the previously hardcoded `256x256` size (minimum is `1024x1024`)
- (**Internal Improvement**) Dependency updates
# (2026-01-23) Version 1.13.0
- (**Improvement**) Extend auto-switching to support cheaper models (`gpt-image-1-mini`) for `gpt-image-1` and `gpt-image-1.5` when generating stickers ([e0b4a40](https://github.com/etkecc/baibot/commit/e0b4a40))
- (**Internal Improvement**) Upgrade Rust compiler (1.92.0 -> 1.93.0) ([691aeeb](https://github.com/etkecc/baibot/commit/691aeeb))
- (**Internal Improvement**) Dependency updates
# (2025-12-21) Version 1.12.0
- (**Improvement**) Upgrade [async-openai](https://crates.io/crates/async-openai) (0.31.1 -> 0.32.2) and add support for OpenAI's `gpt-image-1.5` model ([08c689a](https://github.com/etkecc/baibot/commit/08c689a), [f7bf3d7](https://github.com/etkecc/baibot/commit/f7bf3d7))
- (**Internal Improvement**) Dependency updates
# (2025-12-15) Version 1.11.0
- (**Feature**) Add support for custom avatars via file path and for keeping the already-set avatar (for those who wish to manage it by themselves via other means). See the [sample config](./etc/app/config.yml.dist) for details. ([062fbbb](https://github.com/etkecc/baibot/commit/062fbbb8ef9ad600db483a431c5c782402191023))

572
Cargo.lock generated

File diff suppressed because it is too large Load Diff

View File

@@ -7,7 +7,7 @@ license = "AGPL-3.0-or-later"
readme = "README.md"
keywords = ["matrix", "chat", "bot", "AI", "LLM"]
include = ["/etc/assets/baibot-torso-768.png", "/src", "/README.md", "/CHANGELOG.md", "/LICENSE"]
version = "1.11.0"
version = "1.14.0"
edition = "2024"
[lib]
@@ -17,7 +17,7 @@ path = "src/lib.rs"
[dependencies]
anthropic = { git = "https://github.com/etkecc/anthropic-rs.git", branch = "fix-content-block-image" }
anyhow = "1.0.*"
async-openai = { version = "0.31.1", features = ["audio", "chat-completion", "image"] }
async-openai = { version = "0.32.3", features = ["audio", "chat-completion", "image", "responses"] }
base64 = "0.22.*"
chrono = { version = "0.4.*", default-features = false, features = ["std", "now"] }
# We'd rather not depend on this, but we cannot use the ruma-events EventContent macro without it.
@@ -32,9 +32,9 @@ regex = "1.12.*"
serde = { version = "1.0.*", features = ["derive"], default-features = false }
serde_json = "1.0.*"
serde_yaml = "0.9.*"
tempfile = "3.23.*"
tempfile = "3.24.*"
tiktoken-rs = { version = "0.9.*", default-features = false }
tokio = { version = "1.48.*", features = ["rt", "rt-multi-thread", "macros"] }
tokio = { version = "1.49.*", features = ["rt", "rt-multi-thread", "macros"] }
tracing = "0.1.*"
tracing-subscriber = { version = "0.3.*", features = ["env-filter"] }
url = "2.5.*"

View File

@@ -4,7 +4,7 @@
# #
#######################################
FROM docker.io/rust:1.92.0-slim-trixie AS build
FROM docker.io/rust:1.93.0-slim-trixie AS build
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev

View File

@@ -4,7 +4,7 @@
# #
#######################################
FROM docker.io/rust:1.92.0-slim-trixie AS build
FROM docker.io/rust:1.93.0-slim-trixie AS build
RUN apt-get update && apt-get install -y build-essential pkg-config libssl-dev libsqlite3-dev

View File

@@ -17,7 +17,7 @@ It's influenced by [chaz](https://github.com/arcuru/chaz), but does **not** use
- Supports **different use purposes** (depending on the [☁️ provider](./docs/providers.md) & model):
- [💬 text-generation](./docs/features.md#-text-generation): communicating with you via text (though certain models may "see" images as well)
- [💬 text-generation](./docs/features.md#-text-generation): communicating with you via text (though certain models may "see" images as well). The [OpenAI provider](./docs/providers.md#openai) also supports [🛠️ built-in tools](./docs/features.md#️-built-in-tools-openai-only) (web search, code interpreter)
- [🦻 speech-to-text](./docs/features.md#-speech-to-text): turning your voice messages into text
- [🗣️ text-to-speech](./docs/features.md#%EF%B8%8F-text-to-speech): turning bot or users text messages into voice messages
- [🖌️ image-generation](./docs/features.md#image-generation): creating and editing images based on instructions

View File

@@ -40,6 +40,23 @@ You may also wish to see:
- [📖 Usage / 💬 Text Generation](./usage.md#-text-generation) section for more details on how to use the bot for Text Generation in a room
#### 🛠️ Built-in Tools (OpenAI only)
The [OpenAI provider](./providers.md#openai) supports built-in tools that extend the model's capabilities:
- [🔍 Web Search](https://platform.openai.com/docs/guides/tools-web-search) (`web_search`): allows the model to search the web for up-to-date information. [🖼️ Screenshot](./screenshots/text-generation-tools-web-search.webp)
- [💻 Code Interpreter](https://platform.openai.com/docs/guides/tools-code-interpreter) (`code_interpreter`): allows the model to write and execute Python code in a sandbox
These tools are **disabled by default** and need to be explicitly enabled in the agent's `text_generation.tools` configuration. See the [OpenAI sample configuration](https://github.com/etkecc/baibot/blob/c70387b0c38d8d0f30bba2179a2a21a3710dbeaf/docs/sample-provider-configs/openai.yml#L12-L15) for reference.
To enable tools on an existing dynamically-created agent, you need to [update the agent](./agents.md#updating-agents) to re-create it with the `text_generation.tools` section added and enable the tools you need
💡 **Note**: These tools run on OpenAI's infrastructure and may incur additional costs. Web search results include citations that are incorporated into the response.
#### On-demand involvement
In the following 2 cases, it's useful to involve the bot in conversations on-demand:

View File

@@ -23,7 +23,7 @@ The list of supported providers is below.
### How to choose a provider
If you're not sure which provider to start with, **we recommend [OpenAI](#openai)** as it's the most popular and has the **widest range of capabilities**: [💬 text-generation](./features.md#-text-generation) (no vision), [🖌️ image-generation](./features.md#️image-generation), [🦻 speech-to-text](./features.md#-speech-to-text), [🗣️ text-to-speech](./features.md#️-text-to-speech).
If you're not sure which provider to start with, **we recommend [OpenAI](#openai)** as it's the most popular and has the **widest range of capabilities**: [💬 text-generation](./features.md#-text-generation) (incl. vision, incl. [🛠️ tools](./features.md#️-built-in-tools-openai-only)), [🖌️ image-generation](./features.md#️image-generation), [🦻 speech-to-text](./features.md#-speech-to-text), [🗣️ text-to-speech](./features.md#️-text-to-speech).
You don't need to choose just one though. The bot supports [mixing & matching models](./features.md#-mixing--matching-models), so you can use multiple providers at the same time.
@@ -47,7 +47,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `anthropic`
- 🔗 Links: [🏠 Home page](https://www.anthropic.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Anthropic), [👤 Sign up](https://console.anthropic.com/), [📋 Models list](https://docs.anthropic.com/en/docs/about-claude/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (incl. vision)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (incl. vision, no tools)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local anthropic my-anthropic-agent`
- create a global agent: `!bai agent create-global anthropic my-anthropic-agent`
@@ -61,7 +61,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `groq`
- 🔗 Links: [🏠 Home page](https://groq.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Groq), [👤 Sign up](https://console.groq.com/login), [📋 Models list](https://console.groq.com/docs/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local groq my-groq-agent`
- create a global agent: `!bai agent create-global groq my-groq-agent`
@@ -75,7 +75,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `localai`
- 🔗 Links: [🏠 Home page](https://localai.io/), [📋 Models list](https://localai.io/gallery.html)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local localai my-localai-agent`
- create a global agent: `!bai agent create-global localai my-localai-agent`
@@ -89,7 +89,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `mistral`
- 🔗 Links: [🏠 Home page](https://mistral.ai/), [🌐 Wiki](https://en.wikipedia.org/wiki/Mistral_AI), [👤 Sign up](https://auth.mistral.ai/ui/registration), [📋 Models list](https://docs.mistral.ai/getting-started/models/)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local mistral my-mistral-agent`
- create a global agent: `!bai agent create-global mistral my-mistral-agent`
@@ -103,7 +103,7 @@ You don't need to choose just one though. The bot supports [mixing & matching mo
- 🆔 Identifier: `ollama`
- 🔗 Links: [🏠 Home page](https://ollama.com/), [📋 Models list](https://ollama.com/library)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local ollama my-ollama-agent`
- create a global agent: `!bai agent create-global ollama my-ollama-agent`
@@ -120,7 +120,7 @@ For services which are not fully compatible with the OpenAI API, consider using
- 🆔 Identifier: `openai`
- 🔗 Links: [🏠 Home page](https://openai.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/OpenAI), [👤 Sign up](https://platform.openai.com/signup), [📋 Models list](https://platform.openai.com/docs/models)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (incl. vision), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (incl. vision, incl. [🛠️ tools](./features.md#️-built-in-tools-openai-only)), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local openai my-openai-agent`
- create a global agent: `!bai agent create-global openai my-openai-agent`
@@ -137,7 +137,7 @@ Some of these popular services already have **shortcut** providers (leading to t
This provider is just as featureful as the [OpenAI](#openai) provider, but is more compatible with services which do not fully adhere to the [OpenAI API spec](https://github.com/openai/openai-openapi/).
- 🆔 Identifier: `openai-compatible`
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (no vision), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-creation), [💬 text-generation](./features.md#-text-generation) (no vision, no tools), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local openai-compatible my-openai-compatible-agent`
- create a global agent: `!bai agent create-global openai-compatible my-openai-compatible-agent`
@@ -151,7 +151,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- 🆔 Identifier: `openrouter`
- 🔗 Links: [🏠 Home page](https://openrouter.ai/), [👤 Sign up](https://openrouter.ai/), [📋 Models list](https://openrouter.ai/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local openrouter my-openrouter-agent`
- create a global agent: `!bai agent create-global openrouter my-openrouter-agent`
@@ -165,7 +165,7 @@ This provider is just as featureful as the [OpenAI](#openai) provider, but is mo
- 🆔 Identifier: `together-ai`
- 🔗 Links: [🏠 Home page](https://www.together.ai/), [👤 Sign up](https://api.together.ai/signup), [📋 Models list](https://api.together.xyz/models)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision)
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation) (no vision, no tools)
- 🗲 Quick start:
- create a room-local agent: `!bai agent create-room-local together-ai my-together-ai-agent`
- create a global agent: `!bai agent create-global together-ai my-together-ai-agent`

View File

@@ -9,6 +9,10 @@ text_generation:
max_response_tokens: null
max_completion_tokens: 128000
max_context_tokens: 400000
# Built-in tools
tools:
web_search: false
code_interpreter: false
speech_to_text:
model_id: whisper-1
text_to_speech:
@@ -17,7 +21,7 @@ text_to_speech:
speed: 1.0
response_format: opus
image_generation:
model_id: gpt-image-1
model_id: gpt-image-1.5
style: null
size: null
quality: null

Binary file not shown.

After

Width:  |  Height:  |  Size: 66 KiB

View File

@@ -90,6 +90,10 @@ agents:
# max_response_tokens: null
# max_completion_tokens: 128000
# max_context_tokens: 400000
# # Built-in tools
# tools:
# web_search: false
# code_interpreter: false
# speech_to_text:
# model_id: whisper-1
# text_to_speech:
@@ -98,7 +102,7 @@ agents:
# speed: 1.0
# response_format: opus
# image_generation:
# model_id: gpt-image-1
# model_id: gpt-image-1.5
# style: null
# size: null
# quality: null

View File

@@ -14,7 +14,7 @@ services:
- /etc/passwd:/etc/passwd:ro
synapse:
image: ghcr.io/element-hq/synapse:v1.144.0
image: ghcr.io/element-hq/synapse:v1.146.0
user: "${UID}:${GID}"
restart: unless-stopped
entrypoint: python
@@ -27,7 +27,7 @@ services:
- ./synapse/media-store:/media-store
element-web:
image: ghcr.io/element-hq/element-web:v1.12.6
image: ghcr.io/element-hq/element-web:v1.12.9
user: "${UID}:${GID}"
restart: unless-stopped
environment:

View File

@@ -1,6 +1,6 @@
services:
ollama:
image: docker.io/ollama/ollama:0.13.3
image: docker.io/ollama/ollama:0.15.4
restart: unless-stopped
ports:
- "${SERVICE_OLLAMA_BIND_PORT_HTTP}:11434"

View File

@@ -69,6 +69,7 @@ impl AgentProvider {
models_list_url: Some("https://docs.anthropic.com/en/docs/about-claude/models"),
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: true,
text_generation_supports_tools: false,
},
Self::Groq => AgentProviderInfo {
id: Self::Groq.to_static_str(),
@@ -80,11 +81,12 @@ impl AgentProvider {
models_list_url: Some("https://console.groq.com/docs/models"),
supported_purposes: vec![AgentPurpose::TextGeneration, AgentPurpose::SpeechToText],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::LocalAI => AgentProviderInfo {
id: Self::LocalAI.to_static_str(),
name: "LocalAI",
description: "LocalAI is the free, Open Source OpenAI alternative. LocalAI act as a drop-in replacement REST API that’s compatible with OpenAI API specifications for local inferencing. It allows you to run LLMs, generate images, audio (and not only) locally or on-prem with consumer grade hardware, supporting multiple model families and architectures.",
description: "LocalAI is the free, Open Source OpenAI alternative. LocalAI act as a drop-in replacement REST API that's compatible with OpenAI API specifications for local inferencing. It allows you to run LLMs, generate images, audio (and not only) locally or on-prem with consumer grade hardware, supporting multiple model families and architectures.",
homepage_url: Some("https://localai.io/"),
wiki_url: None,
sign_up_url: None,
@@ -95,6 +97,7 @@ impl AgentProvider {
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::Mistral => AgentProviderInfo {
id: Self::Mistral.to_static_str(),
@@ -106,6 +109,7 @@ impl AgentProvider {
models_list_url: Some("https://docs.mistral.ai/getting-started/models/"),
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::Ollama => AgentProviderInfo {
id: Self::Ollama.to_static_str(),
@@ -117,6 +121,7 @@ impl AgentProvider {
models_list_url: Some("https://ollama.com/library"),
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::OpenAI => AgentProviderInfo {
id: Self::OpenAI.to_static_str(),
@@ -133,6 +138,7 @@ impl AgentProvider {
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: true,
text_generation_supports_tools: true,
},
Self::OpenAICompat => AgentProviderInfo {
id: Self::OpenAICompat.to_static_str(),
@@ -149,6 +155,7 @@ impl AgentProvider {
AgentPurpose::SpeechToText,
],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::OpenRouter => AgentProviderInfo {
id: Self::OpenRouter.to_static_str(),
@@ -160,6 +167,7 @@ impl AgentProvider {
models_list_url: Some("https://openrouter.ai/models"),
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
Self::TogetherAI => AgentProviderInfo {
id: Self::TogetherAI.to_static_str(),
@@ -171,6 +179,7 @@ impl AgentProvider {
models_list_url: Some("https://api.together.xyz/models"),
supported_purposes: vec![AgentPurpose::TextGeneration],
text_generation_supports_vision: false,
text_generation_supports_tools: false,
},
}
}
@@ -192,4 +201,5 @@ pub struct AgentProviderInfo {
pub models_list_url: Option<&'static str>,
pub supported_purposes: Vec<AgentPurpose>,
pub text_generation_supports_vision: bool,
pub text_generation_supports_tools: bool,
}

View File

@@ -2,7 +2,7 @@ use mxlink::mime;
#[derive(Default)]
pub struct ImageGenerationParams {
pub size_override: Option<String>,
pub smallest_size_possible: bool,
pub cheaper_model_switching_allowed: bool,
@@ -10,8 +10,8 @@ pub struct ImageGenerationParams {
}
impl ImageGenerationParams {
pub fn with_size_override(mut self, value: Option<String>) -> Self {
self.size_override = value;
pub fn with_smallest_size_possible(mut self, value: bool) -> Self {
self.smallest_size_possible = value;
self
}

View File

@@ -1,6 +1,6 @@
use serde::{Deserialize, Serialize};
use super::OPENAI_IMAGE_MODEL_GPT_IMAGE_1;
use super::OPENAI_IMAGE_MODEL_GPT_IMAGE_1_DOT_5;
use crate::agent::{default_prompt, provider::ConfigTrait};
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -64,6 +64,9 @@ pub struct TextGenerationConfig {
#[serde(default)]
pub max_context_tokens: u32,
#[serde(default)]
pub tools: ToolsConfig,
}
impl Default for TextGenerationConfig {
@@ -75,6 +78,7 @@ impl Default for TextGenerationConfig {
max_response_tokens: None,
max_completion_tokens: Some(128_000),
max_context_tokens: 400_000,
tools: ToolsConfig::default(),
}
}
}
@@ -83,6 +87,15 @@ fn default_text_model_id() -> String {
"gpt-5.2".to_owned()
}
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
pub struct ToolsConfig {
#[serde(default)]
pub web_search: bool,
#[serde(default)]
pub code_interpreter: bool,
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct SpeechToTextConfig {
#[serde(default = "default_speech_to_text_model_id")]
@@ -162,7 +175,7 @@ pub struct ImageGenerationConfig {
impl Default for ImageGenerationConfig {
fn default() -> Self {
Self {
model_id: OPENAI_IMAGE_MODEL_GPT_IMAGE_1.to_owned(),
model_id: OPENAI_IMAGE_MODEL_GPT_IMAGE_1_DOT_5.to_owned(),
style: default_image_style(),
size: default_image_size(),
quality: default_image_quality(),
@@ -178,8 +191,11 @@ impl ImageGenerationConfig {
"dall-e-2" => Ok(async_openai::types::images::ImageModel::DallE2),
"dall-e-3" => Ok(async_openai::types::images::ImageModel::DallE3),
"gpt-image-1" => Ok(async_openai::types::images::ImageModel::GptImage1),
"gpt-image-1.5" => Ok(async_openai::types::images::ImageModel::GptImage1dot5),
"gpt-image-1-mini" => Ok(async_openai::types::images::ImageModel::GptImage1Mini),
other => Ok(async_openai::types::images::ImageModel::Other(other.to_owned())),
other => Ok(async_openai::types::images::ImageModel::Other(
other.to_owned(),
)),
}
}
}

View File

@@ -5,7 +5,10 @@ use async_openai::{
config::OpenAIConfig,
types::{
audio::{AudioInput, CreateSpeechRequestArgs, CreateTranscriptionRequestArgs},
chat::{ChatCompletionRequestMessage, CreateChatCompletionRequestArgs},
responses::{
CodeInterpreterContainerAuto, CodeInterpreterTool, CodeInterpreterToolContainer,
CreateResponseArgs, OutputItem, OutputMessageContent, Tool, WebSearchTool,
},
images::{
CreateImageEditRequestArgs, CreateImageRequestArgs,
Image, ImageInput, ImageModel, ImageResponseFormat,
@@ -28,12 +31,9 @@ use crate::{
use crate::{
agent::{
AgentPurpose,
provider::{
entity::{
ImageEditResult, ImageGenerationResult, ImageSource, PingResult,
TextToSpeechParams, TextToSpeechResult,
},
openai::utils::convert_string_to_enum,
provider::entity::{
ImageEditResult, ImageGenerationResult, ImageSource, PingResult,
TextToSpeechParams, TextToSpeechResult,
},
},
strings,
@@ -129,28 +129,44 @@ impl ControllerTrait for Controller {
conversation_messages.insert(0, prompt_message);
}
let openai_conversation_messages: Vec<ChatCompletionRequestMessage> =
super::utils::convert_llm_messages_to_openai_messages(conversation_messages);
let input = super::utils::convert_llm_messages_to_openai_response_input(conversation_messages);
let messages_count = openai_conversation_messages.len();
let messages_count = match &input {
async_openai::types::responses::InputParam::Items(items) => items.len(),
_ => 1,
};
let temperature = params
.temperature_override
.unwrap_or(text_generation_config.temperature);
let mut request_builder = CreateChatCompletionRequestArgs::default();
let mut request_builder = CreateResponseArgs::default();
request_builder
.model(&text_generation_config.model_id)
.temperature(temperature)
.messages(openai_conversation_messages);
.input(input);
if let Some(max_response_tokens) = text_generation_config.max_response_tokens {
request_builder.max_tokens(max_response_tokens);
let mut tools = Vec::new();
if text_generation_config.tools.web_search {
tools.push(Tool::WebSearch(WebSearchTool::default()));
}
if text_generation_config.tools.code_interpreter {
tools.push(Tool::CodeInterpreter(CodeInterpreterTool {
container: CodeInterpreterToolContainer::Auto(
CodeInterpreterContainerAuto::default(),
),
}));
}
if let Some(max_completion_tokens) = text_generation_config.max_completion_tokens {
request_builder.max_completion_tokens(max_completion_tokens);
if !tools.is_empty() {
request_builder.tools(tools);
}
if let Some(max_response_tokens) = text_generation_config.max_response_tokens {
request_builder.max_output_tokens(max_response_tokens);
} else if let Some(max_completion_tokens) = text_generation_config.max_completion_tokens {
request_builder.max_output_tokens(max_completion_tokens);
}
let request = request_builder.build()?;
@@ -160,33 +176,31 @@ impl ControllerTrait for Controller {
model = format!("{:?}", request.model),
?messages_count,
request = request_as_json,
"Sending OpenAI chat completion API request"
"Sending OpenAI response API request"
);
}
let response = self.client.chat().create(request).await?;
let response = self.client.responses().create(request).await?;
tracing::trace!(
?response,
"Got response from the OpenAI chat completion API"
"Got response from the OpenAI response API"
);
// We only request 1 result, so there should only be 1 choice.
if let Some(choice) = response.choices.into_iter().next() {
match choice.message.content {
Some(text) => {
return Ok(TextGenerationResult { text });
}
None => {
return Err(anyhow::anyhow!(
"No content was found in the response choice from the OpenAI chat completion API"
));
for item in response.output {
if let OutputItem::Message(message) = item {
for content in message.content {
if let OutputMessageContent::OutputText(text_content) = content {
return Ok(TextGenerationResult {
text: text_content.text,
});
}
}
}
}
Err(anyhow::anyhow!(
"No response messages choices were returned from the OpenAI chat completion API"
"No response messages choices were returned from the OpenAI response API"
))
}
@@ -254,10 +268,12 @@ impl ControllerTrait for Controller {
match original_model {
ImageModel::DallE2 => ImageModel::DallE2,
ImageModel::DallE3 => ImageModel::DallE2,
ImageModel::GptImage1 => ImageModel::GptImage1Mini,
ImageModel::GptImage1dot5 => ImageModel::GptImage1Mini,
ImageModel::GptImage1Mini => ImageModel::GptImage1Mini,
ImageModel::Other(_) => {
ImageModel::DallE2
}
_ => original_model.clone(),
}
} else {
original_model
@@ -293,10 +309,11 @@ impl ControllerTrait for Controller {
image_generation_config.quality.clone()
};
let size = params
.size_override
.map(|s| convert_string_to_enum::<async_openai::types::images::ImageSize>(&s).unwrap())
.or(image_generation_config.size);
let size = if params.smallest_size_possible {
Some(get_sticker_size(&model))
} else {
image_generation_config.size
};
let response_format = match model.clone() {
ImageModel::DallE2 => Some(ImageResponseFormat::B64Json),
@@ -305,6 +322,7 @@ impl ControllerTrait for Controller {
// In fact, specifying the response format results in an error.
ImageModel::GptImage1 => None,
ImageModel::GptImage1Mini => None,
ImageModel::GptImage1dot5 => None,
ImageModel::Other(_) => Some(ImageResponseFormat::B64Json),
};
@@ -411,6 +429,7 @@ impl ControllerTrait for Controller {
// In fact, specifying the response format results in an error.
ImageModel::GptImage1 => None,
ImageModel::GptImage1Mini => None,
ImageModel::GptImage1dot5 => None,
ImageModel::Other(_) => Some(ImageResponseFormat::B64Json),
};
@@ -617,3 +636,17 @@ fn audio_mime_type_to_file_name(mime_type: &mxlink::mime::Mime) -> Option<String
Some(format!("audio.{}", file_extension))
}
/// Returns the smallest supported size for stickers based on what the image model supports.
fn get_sticker_size(model: &ImageModel) -> async_openai::types::images::ImageSize {
use async_openai::types::images::ImageSize;
match model {
ImageModel::DallE2 => ImageSize::S256x256,
ImageModel::DallE3 => ImageSize::S1024x1024,
ImageModel::GptImage1 => ImageSize::S1024x1024,
ImageModel::GptImage1Mini => ImageSize::S1024x1024,
ImageModel::GptImage1dot5 => ImageSize::S1024x1024,
ImageModel::Other(_) => ImageSize::S1024x1024,
}
}

View File

@@ -16,7 +16,7 @@ use super::super::AgentInstantiationResult;
use super::ConfigTrait;
use super::controller::ControllerType;
pub const OPENAI_IMAGE_MODEL_GPT_IMAGE_1: &str = "gpt-image-1";
pub const OPENAI_IMAGE_MODEL_GPT_IMAGE_1_DOT_5: &str = "gpt-image-1.5";
pub fn create_controller_from_yaml_value_config(
agent_id: &str,

View File

@@ -1,12 +1,6 @@
use async_openai::types::{
chat::{
ChatCompletionRequestAssistantMessageArgs, ChatCompletionRequestMessage,
ChatCompletionRequestMessageContentPartImage,
ChatCompletionRequestSystemMessageArgs,
ChatCompletionRequestUserMessageArgs, ChatCompletionRequestUserMessageContent,
ChatCompletionRequestUserMessageContentPart,
ImageUrlArgs,
},
use async_openai::types::responses::{
EasyInputContent, EasyInputMessage, ImageDetail, InputContent, InputImageContent, InputItem,
InputParam, MessageType, Role,
};
use crate::conversation::llm::{
@@ -14,93 +8,41 @@ use crate::conversation::llm::{
};
use crate::utils::base64::base64_encode;
pub fn convert_llm_messages_to_openai_messages(
pub fn convert_llm_messages_to_openai_response_input(
conversation_messages: Vec<LLMMessage>,
) -> Vec<ChatCompletionRequestMessage> {
let mut openai_conversation_messages: Vec<ChatCompletionRequestMessage> =
Vec::with_capacity(conversation_messages.len());
) -> InputParam {
let mut items = Vec::with_capacity(conversation_messages.len());
for message in conversation_messages {
let openai_message = convert_llm_message_to_openai_message(message);
if let Some(openai_message) = openai_message {
openai_conversation_messages.push(openai_message);
}
}
let role = match message.author {
LLMAuthor::Prompt => Role::System,
LLMAuthor::Assistant => Role::Assistant,
LLMAuthor::User => Role::User,
};
openai_conversation_messages
}
let content = match message.content {
LLMMessageContent::Text(text) => EasyInputContent::Text(text),
LLMMessageContent::Image(image_details) => {
let image_url = format!(
"data:{};base64,{}",
image_details.mime,
base64_encode(&image_details.data)
);
fn convert_llm_message_to_openai_message(
llm_message: LLMMessage,
) -> Option<ChatCompletionRequestMessage> {
match &llm_message.content {
LLMMessageContent::Text(text) => Some(match llm_message.author {
LLMAuthor::Prompt => ChatCompletionRequestSystemMessageArgs::default()
.content(text.clone())
.build()
.expect("Failed building OpenAI system message")
.into(),
LLMAuthor::Assistant => ChatCompletionRequestAssistantMessageArgs::default()
.content(text.clone())
.build()
.expect("Failed building OpenAI assistant message")
.into(),
LLMAuthor::User => ChatCompletionRequestUserMessageArgs::default()
.content(text.clone())
.build()
.expect("Failed building OpenAI user message")
.into(),
}),
LLMMessageContent::Image(image_details) => {
let image_url = format!(
"data:{};base64,{}",
image_details.mime,
base64_encode(&image_details.data)
);
let part = ChatCompletionRequestUserMessageContentPart::ImageUrl(
ChatCompletionRequestMessageContentPartImage {
image_url: ImageUrlArgs::default()
.url(image_url)
.build()
.expect("Failed building OpenAI image url"),
},
);
let message_content = ChatCompletionRequestUserMessageContent::Array(vec![part]);
match llm_message.author {
LLMAuthor::User => Some(
ChatCompletionRequestUserMessageArgs::default()
.content(message_content)
.build()
.expect("Failed building OpenAI user message")
.into(),
),
_ => {
tracing::warn!(
"OpenAI API does not support image content for messages authored by {:?}. This message part will be skipped.",
llm_message.author
);
None
}
EasyInputContent::ContentList(vec![InputContent::InputImage(InputImageContent {
image_url: Some(image_url),
detail: ImageDetail::Auto,
file_id: None,
})])
}
}
}
}
};
pub(super) fn convert_string_to_enum<T>(value: &str) -> Result<T, String>
where
T: serde::de::DeserializeOwned,
{
// This is a hacky way to construct an enum from the string we have.
let enum_result: serde_json::Result<T> = serde_json::from_str(&format!("\"{}\"", value));
match enum_result {
Ok(enum_result) => Ok(enum_result),
Err(err) => {
tracing::debug!(?err, "Failed to parse into enum");
Err(format!("The value ({}) is not supported.", value))
}
items.push(InputItem::EasyMessage(EasyInputMessage {
r#type: MessageType::Message,
role,
content,
}));
}
InputParam::Items(items)
}

View File

@@ -95,6 +95,7 @@ impl TryInto<OpenAITextGenerationConfig> for TextGenerationConfig {
max_response_tokens: self.max_response_tokens,
max_completion_tokens: None,
max_context_tokens: self.max_context_tokens,
tools: Default::default(),
})
}
}

View File

@@ -3,6 +3,8 @@ use etke_openai_api_rust::chat::{ChatApi, ChatBody};
use etke_openai_api_rust::images::{ImagesApi, ImagesBody};
use etke_openai_api_rust::{Auth, Message, OpenAI};
const SMALLEST_IMAGE_SIZE: &str = "256x256";
use super::super::ControllerTrait;
use crate::utils::base64::base64_decode;
use crate::{
@@ -303,9 +305,11 @@ impl ControllerTrait for Controller {
// when they span multiple lines.
let prompt = prompt.replace("\n", " ");
let size: Option<String> = params
.size_override
.or_else(|| image_generation_config.size.clone());
let size: Option<String> = if params.smallest_size_possible {
Some(SMALLEST_IMAGE_SIZE.to_owned())
} else {
image_generation_config.size.clone()
};
let request = ImagesBody {
model: Some(image_generation_config.model_id.to_owned()),

View File

@@ -12,9 +12,6 @@ use crate::strings;
use crate::utils::mime::get_file_extension;
use crate::{Bot, entity::MessageContext};
// We may make this configurable (per room, etc.) in the future, but for now it's hardcoded.
const STICKER_SIZE: &str = "256x256";
pub async fn handle_image(
bot: &Bot,
matrix_link: MatrixLink,
@@ -177,7 +174,7 @@ pub async fn handle_sticker(
);
let params = ImageGenerationParams::default()
.with_size_override(Some(STICKER_SIZE.to_owned()))
.with_smallest_size_possible(true)
.with_cheaper_model_switching_allowed(true)
.with_cheaper_quality_switching_allowed(true);

View File

@@ -109,11 +109,21 @@ pub fn help_provider_details(id: &str, info: &AgentProviderInfo) -> String {
let mut purpose_line = format!("{} {}", purpose.emoji(), purpose.as_str());
if let AgentPurpose::TextGeneration = purpose {
let mut extras = vec![];
if info.text_generation_supports_vision {
purpose_line = format!("{} ({})", purpose_line, "incl. vision");
extras.push("incl. vision");
} else {
purpose_line = format!("{} ({})", purpose_line, "no vision");
extras.push("no vision");
}
if info.text_generation_supports_tools {
extras.push("incl. tools");
} else {
extras.push("no tools");
}
purpose_line = format!("{} ({})", purpose_line, extras.join(", "));
}
capabilities.push(purpose_line);

View File

@@ -64,7 +64,7 @@ To create a sticker, send a command like `%command_prefix% sticker A huge bowl o
The difference from **creating images** is that the bot will:
- create a smaller-resolution image (`256x256`) - smaller/quicker, but still good enough for a sticker
- create a smaller-resolution image (as small as the model allows) - smaller/quicker, but still good enough for a sticker
- potentially switch to a different (cheaper or otherwise more suitable) model, if available
- post the image directly to the room (as a reply to your message), without starting a threaded conversation
"#;