Commit Graph

72 Commits

Author SHA1 Message Date
Slavi Pantaleev
a865a26093 Update tiktoken-rs to 0.11, adding support for newer GPT models
Supersedes https://github.com/etkecc/baibot/pull/116

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 10:52:22 +03:00
kschwank
2d659964a7 Add sender context mode for text generation (#104)
Add a per-room/global `sender_context_mode` setting that optionally
prefixes conversation messages with sender metadata before sending them
to the model provider.

This helps models distinguish between participants in multi-user rooms.
                                                                                                                                                                                                                                                             
Three modes are supported:
- `disabled` (default, no change)
- `matrix_user_id` (prefixes with `[sender=@user:server]`)
- `matrix_user_id_and_timestamp` (adds `send_at`; Example: `[sender=@user:server sent_at=<ISO 8601>]`)
                                                                                                                                                                                                                                                             
Sender context is applied to user and assistant text messages only,
skipping system prompts and non-text content.

Mixed-sender merged turns (something we intentionally do for Anthropic)
have their `sender_id` cleared to avoid misattribution.
2026-03-25 20:14:34 +02:00
Slavi Pantaleev
d9b5524c97 Fix OpenAI response input for async-openai 0.34
async-openai 0.34 adds a required phase field to EasyInputMessage,
which broke our Responses API request construction.

Set phase to None because baibot does not currently model assistant commentary vs final-answer turns,
so omitting phase preserves the previous behavior while keeping the request compatible with the new crate.
2026-03-24 12:51:56 +02:00
Slavi Pantaleev
12b938d2d1 Skip files with application/octet-stream MIME type
Files with unrecognized MIME types (application/octet-stream) are
not supported by any LLM provider and would cause errors that
permanently break the conversation thread. Instead, represent them
as a text message describing the attachment so the LLM is still
aware a file was sent without the thread becoming unusable.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 12:22:38 +02:00
Slavi Pantaleev
f3d1b32ad7 Use mime_guess for file extension MIME type detection
Replace the hand-maintained extension-to-MIME mapping with the
mime_guess crate, which was already in the dependency tree.
This covers hundreds of file extensions out of the box.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 12:22:38 +02:00
Slavi Pantaleev
3d3bd3c9f9 Add support for file attachments (m.file) in conversations
Files sent as m.file Matrix messages are now downloaded, MIME-detected,
and forwarded to LLM providers alongside the conversation context,
similar to how m.image is already handled.

- OpenAI provider: sends files inline as base64 data URLs
- Anthropic provider: skips files with a warning (library limitation)
- OpenAI-compat provider: skips files with a warning (library limitation)

Controller routing respects the existing prefix requirement setting.
MIME detection expanded to cover PDF, text, code, and document formats.
Docs updated to reflect file support and known limitations.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 12:22:38 +02:00
Slavi Pantaleev
6aa3d70d57 Bump default OpenAI text-generation model (gpt-5.2 -> gpt-5.4) 2026-03-11 09:53:02 +02:00
Taylor Southwick
4852d1fe92 Add support for access tokens using MAS (#83)
* Add support for access tokens using MAS

* use 1.13.0

* Update dependencies

* Harden auth credential selection in matrix link init

Use the same non-empty access-token criterion for auth mode selection and bind the token directly from the branch condition.
Return explicit configuration errors for missing or empty `device_id`/`password` instead of panicking, so invalid auth config fails gracefully.

* Centralize and harden user auth config handling

Move authentication-mode resolution into typed config parsing with ConfigUserAuth,
so downstream login setup consumes validated credentials instead of re-checking raw optional fields.

Enforce explicit password-vs-token selection, validate token/device/user-id requirements in one place,
and normalize empty auth env overrides to unset values for consistent behavior across YAML and environment input.

* Add auth config unit tests

Move auth_config tests into a dedicated cfg test module file to keep production config code compact while preserving behavior coverage. The tests cover password/token mode selection, missing/both auth method rejection, missing device_id, and empty-value handling.

* Use conventional mxlink version requirement

Replace the unconventional wildcard lower-bound expression with a standard semver lower bound for readability and tooling consistency.

---------

Co-authored-by: Slavi Pantaleev <slavi@devture.com>
2026-03-07 10:26:40 +02:00
Slavi Pantaleev
faf92cac09 Switch from deprecated serde_yaml to serde_yaml_ng
serde_yaml is deprecated and unmaintained. serde_yaml_ng is the community
fork with a compatible API, so this is a straightforward rename across the
codebase.
2026-02-10 14:34:57 +02:00
Slavi Pantaleev
a82e9a1d1f Add prek pre-commit hooks via mise, fix formatting and clippy warnings
- Add mise.toml (prek 0.3.2) and .pre-commit-config.yaml with hooks for
  trailing whitespace, end-of-file, YAML check, merge conflicts, large files,
  cargo fmt, cargo clippy (-D warnings), and unit tests
- Add prek/mise recipes to justfile
- Run cargo fmt to fix formatting issues
- Fix all clippy warnings: collapse nested if statements, derive Default for Avatar
2026-02-10 14:33:19 +02:00
Slavi Pantaleev
d831c08306 Add support for OpenAI built-in tools (web_search, code_interpreter)
Document the new tools feature in README, features.md, and providers.md.
Update the `!bai providers` command output to show vision/tools support
consistently for all providers.

Based on #62 by @yeslayla which migrated the OpenAI provider to the
Responses API.

See: https://github.com/etkecc/baibot/pull/62
2026-02-04 03:17:19 +02:00
Slavi Pantaleev
c70387b0c3 Fix sticker generation for newer GPT image models
Sticker generation was failing when using newer GPT image models
(gpt-image-1, gpt-image-1-mini, gpt-image-1.5). The issue occurred
because stickers requested 256x256 size, but these models only support
1024x1024, 1536x1024, 1024x1536, and auto.

To reproduce, send `!bai sticker Something` to an agent configured
with a GPT image model. The error was:

  invalid_request_error: Invalid value: '256x256'. Supported values
  are: '1024x1024', '1024x1536', '1536x1024', and 'auto'. (param: size)
  (code: invalid_value)

The fix replaces the hardcoded 256x256 size override with a
`smallest_size_possible` flag, letting each provider determine the
appropriate sticker size based on the model being used.

The `openai_compat` provider still defaults to requesting 256x256 in all cases
(regardless of model name).
2026-02-04 02:39:51 +02:00
Layla
ec93f1ee2a Implement OpenAI's response API and add support for built-in tools (web search & code interpreter). 2026-02-04 01:41:27 +02:00
Slavi Pantaleev
e0b4a40dd8 Configure switching to cheaper models (gpt-image-1-mini) for gpt-image-1 & gpt-image-1.5 2026-01-22 22:31:51 +02:00
Slavi Pantaleev
3a88b0d656 Make gpt-image-1.5 the default image model for OpenAI 2025-12-21 12:53:44 +02:00
Slavi Pantaleev
08c689a889 Upgrade async-openai (0.31.1 -> 0.32.1) and adapt, adding support for gpt-image-1.5 2025-12-21 12:23:25 +02:00
Slavi Pantaleev
062fbbb8ef Add support for custom avatars (via file path) and for not touching the already-set avatar
This is based on the work done in https://github.com/etkecc/baibot/pull/60 by https://github.com/Fmstrat (Ben Curtis),
with various changes on top to make the code more idiomatic and flexible.

This commit squashes the following patches (newest first):

- Improve handling of `user.avatar` configuration (null & empty string being the same now) and add support for a special `keep` value
- Minor import reordering
- Simplify avatar configuration (`user.avatar.source` -> `user.avatar`)
- Combine `logo_bytes` and `mime_type` determination logic and do not fall back to default avatar if reading the custom avatar file fails
- Switch from deprecated `mime_guess::guess_mime_type(avatar_path)` to `mime_guess::from_path(avatar_path).first_or_octet_stream()`
- Relax `mime_guess` constraint and order alphabetically
- Use `mime` from `mxlink`
- Add mime-type support and switch to user.avatar.source
- (Original work by Fmstrat) Add support for custom avatars

Co-authored-by: Fmstrat <nospam@nowsci.com>
2025-12-15 08:07:24 +02:00
Slavi Pantaleev
2801c78ad9 Bump default OpenAI text-generation model (gpt-5.1 -> gpt-5.2) 2025-12-12 16:05:16 +02:00
Slavi Pantaleev
8eb70f0f2c Upgrade async-openai from our own etkecc fork to upstream's 0.31.1
Switches `async-openai` from our own etkecc fork (0.28.1-patched) to the
official crates.io version 0.31.1.

We adapt to async-openai's types reorganization and making use of crate
features to only enable what we need.
2025-11-30 11:48:56 +02:00
Slavi Pantaleev
df507eb201 Make use of the BAIBOT_USER_ENCRYPTION_RECOVERY_RESET_ALLOWED environment variable for configuring config.user.encryption.recovery_reset_allowed 2025-11-28 14:18:35 +02:00
Slavi Pantaleev
4dcd9eff40 Make use of the BAIBOT_PERSISTENCE_SESSION_ENCRYPTION_KEY environment variable for configuring config.persistence.session_encryption_key
Seems like we already had a constant defined, but weren't making use of it.
2025-11-28 14:17:23 +02:00
Slavi Pantaleev
da97361e1b Bump default OpenAI text-generation model (gpt-5 -> gpt-5.1) 2025-11-20 05:25:20 +02:00
Slavi Pantaleev
624b9de35b Upgrade mxlink (1.9.0 -> 1.10.0) and matrix-sdk (0.13.0 -> 0.14.0) 2025-09-08 14:26:37 +03:00
Slavi Pantaleev
941bf7ca42 Update sample configs for OpenAI (gpt-5) to specify max_completion_tokens, not max_response_tokens
Fixup for b43f61f5ff
2025-09-08 10:37:00 +03:00
Slavi Pantaleev
b43f61f5ff Change default OpenAI model (gpt-4.1 -> gpt-5)
Ref: https://openai.com/index/introducing-gpt-5/
2025-08-08 07:22:19 +03:00
Slavi Pantaleev
264f683d6a Allow image_generation.size to be null for OpenAI and default it to that
The API spec for image creation and image editing says "string or null",
so we're allowing `null` now to trigger automatic selection.
2025-05-11 23:20:07 +03:00
Slavi Pantaleev
69627abd74 Add image-editing feature documentation to the !bai usage command and adjust texts a bit 2025-05-11 22:19:15 +03:00
Slavi Pantaleev
a84135ff32 fmt 2025-05-10 11:47:50 +03:00
Slavi Pantaleev
231528a0d8 Document which providers support vision 2025-05-10 11:47:00 +03:00
Slavi Pantaleev
2f9c3dfce0 Use patched anthropic-rs library to add Vision support to text conversations
Related to: https://github.com/AbdelStark/anthropic-rs/pull/11
2025-05-10 10:25:15 +03:00
Slavi Pantaleev
de958208b2 Use patched async-openai library to work around a few upstream issues
Related to:

- CreateImageEditRequest forces an application/octet-stream content type
  for images (https://github.com/64bit/async-openai/issues/364)

- CreateImageEditRequest only deals with a single image
  (https://github.com/64bit/async-openai/issues/363)

This is a continuation of 8f86289373
and fixes the Image Editing feature for OpenAI.
2025-05-10 10:00:36 +03:00
Slavi Pantaleev
ac4f2080ce Make some improvements as suggested by clippy 2025-05-10 09:33:44 +03:00
Slavi Pantaleev
3ffa50b7b9 Use ImageInput|AudioInput::from_vec_u8 helper 2025-05-10 09:29:06 +03:00
Slavi Pantaleev
8f86289373 Initial work on Vision support in text conversations and Image Editing
This is a huge patch which does some major refactoring like:

- renaming "Image Generation" to "Image Creation" in most places,
  to better match its new command (`!bai image create`)

- relocating image creation command (`!bai image` -> `!bai image create`),
  so it wouldn't conflict with the new image editing command (`!bai image edit`)

- introducing a new image editing command (`!bai image edit`), which
  is meant to work only with the OpenAI provider, but doesn't fully work yet
  due to https://github.com/64bit/async-openai/issues/364, though a next patch will fix it

- adding support for reading images off of Matrix conversations and forwarding them to
  text conversations. Works for OpenAI, but not for Anthropic yet
  (requires custom patches) and not for OpenAI-Compat (no support for
  images there)

- relocating some utils around (base64, mime)
2025-05-10 09:18:01 +03:00
Slavi Pantaleev
e0dcc39a72 Default OpenAI image-generation model to gpt-image-1 (previously dall-e-3) 2025-05-03 09:42:46 +03:00
Slavi Pantaleev
c94376109c Avoid passing response_format to OpenAI's image generation API for the gpt-image-1 model
The API reference for `response_format` says:

> This parameter isn't supported for gpt-image-1 which will always return base64-encoded images.

Related to https://github.com/etkecc/baibot/issues/40
2025-05-03 09:42:44 +03:00
Slavi Pantaleev
256ed05662 Make style, quality and internal response_format image generation parameters optional
Some OpenAI models (like `gpt-image-1`) either don't support these or
only support specific other values.

Related to https://github.com/etkecc/baibot/issues/40
2025-05-03 09:42:39 +03:00
Slavi Pantaleev
8222681e27 Default OpenAI text-generation model to gpt-4.1 (previously gpt-4o) 2025-05-03 09:42:37 +03:00
Slavi Pantaleev
f304b93c68 Improve in-room error reporting details when image generation fails
What previously was a generic error message like:

> ⚠️ Error: An error occurred while processing your message. Please try again.

.. now becomes a much more helpful error message like:

> ⚠️ Error: There was a problem performing image-generation via the room-local/my-openai-agent agent:
>
> invalid_request_error: Invalid value: 'standard'. Supported values are: 'low', 'medium', 'high', and 'auto'. (param: quality) (code: invalid_value)

Related to https://github.com/etkecc/baibot/issues/40
2025-05-03 09:42:28 +03:00
Slavi Pantaleev
59e2746578 Remove explicit lifetime to fix clippy-reported warning 2025-02-27 09:58:58 +02:00
Slavi Pantaleev
47d8edea70 Add support for configuring max_completion_tokens for OpenAI
Related to db9422740c
2025-02-27 09:58:58 +02:00
Slavi Pantaleev
692d61b239 Replace Anthropic library (anthropic-rs -> anthropic) and switch default recommended model (claude-3-5-sonnet-20240620 -> claude-3-7-sonnet-20250219)
Fixes https://github.com/etkecc/baibot/issues/22

Ultimate related to `anthropic-rs` hardcoding models as enumeration variants in the code
and not updating them. See:
- https://github.com/roushou/mesh/issues/1
- https://github.com/roushou/mesh/pull/2

https://github.com/cortesi/misanthropy was also considered as an
alternative, but it did not allow configuring the base API URL like our
old Anthropic library (`anthropic-rs`) and like our new choice (`anthropic`).
We'd rather not lose support for this, so we're going with the `anthropic` library.
2025-02-27 09:44:49 +02:00
Slavi Pantaleev
406141cd7d fmt 2025-02-27 07:46:16 +02:00
Slavi Pantaleev
c051da2f4a Add config setting controlling if a self-introduction message is posted after joining a room
Fixes https://github.com/etkecc/baibot/issues/32
2025-02-26 20:51:24 +02:00
Slavi Pantaleev
06b2b6d776 Use progress indicator emoji (⏳), not 🦻 to indicate that speech-to-text is happening
🦻 is used for another purpose - to denote that a message is one coming
from speech-to-text, by:

- having the bot react to its own speech-to-text transcription message
  with the 🦻 emoji when it's posted in a non-thread

- having the bot prefix its speech-to-text transcription messages with
  `> 🦻` when it's posted in a thread

⏳ is already used as a progress indicator for other features, so it
makes sense to use it for indicating that speech-to-text is happening for a given audio message as well.
While 🦻 was an even more descriptive illustration of what's actually happening to the audio message
("it's being heard by the bot"), us using the 🦻 emoji for different things didn't seem good.
2025-02-26 20:51:24 +02:00
Slavi Pantaleev
a1bd292752 Add support for making Text-To-Speech send regular text messages instead of notices
When speech-to-text/flow-type = `only_transcribe`, the bot will now send
text messages by default, not notices.

While notice messages may be less desirable with other bots in the room,
it's probably a better default for most people who enable "transcribe-only" mode.

This is an improvement related to https://github.com/etkecc/baibot/issues/14
2025-02-26 20:51:08 +02:00
Slavi Pantaleev
ec1879d212 Populate image/audio attachment body with a filename, not with text
Various clients (including newer versions of Element Web), do not like
it when the `body` field of the attachment is not a file name.

For images, a preview may not be shown and downloading the attachment
may suggest that the whole long text is used as a filename (which is odd).

There is value (improved accessibility, etc.)
in adding better descriptions (especially to generated images),
but given that it's currently problematic, I'm getting rid of it.
It's better and safer if we stick to using filenames.
2025-01-24 11:35:08 +02:00
Slavi Pantaleev
39a184e5d0 Adapt to mxlink 1.4.0 (matrix-sdk 0.8.0) 2024-11-19 20:57:35 +02:00
Slavi Pantaleev
9d166e35ba Add missing typing notices sending functionality while generating images 2024-11-19 20:46:23 +02:00
Slavi Pantaleev
d9a045a5e4 Make fallback user mentions support also match against the bot's room-specific username
It seems like Element iOS benefits from this.
2024-10-03 16:28:49 +03:00