This is a huge patch which does some major refactoring like:
- renaming "Image Generation" to "Image Creation" in most places,
to better match its new command (`!bai image create`)
- relocating image creation command (`!bai image` -> `!bai image create`),
so it wouldn't conflict with the new image editing command (`!bai image edit`)
- introducing a new image editing command (`!bai image edit`), which
is meant to work only with the OpenAI provider, but doesn't fully work yet
due to https://github.com/64bit/async-openai/issues/364, though a next patch will fix it
- adding support for reading images off of Matrix conversations and forwarding them to
text conversations. Works for OpenAI, but not for Anthropic yet
(requires custom patches) and not for OpenAI-Compat (no support for
images there)
- relocating some utils around (base64, mime)
The API reference for `response_format` says:
> This parameter isn't supported for gpt-image-1 which will always return base64-encoded images.
Related to https://github.com/etkecc/baibot/issues/40
What previously was a generic error message like:
> ⚠️ Error: An error occurred while processing your message. Please try again.
.. now becomes a much more helpful error message like:
> ⚠️ Error: There was a problem performing image-generation via the room-local/my-openai-agent agent:
>
> invalid_request_error: Invalid value: 'standard'. Supported values are: 'low', 'medium', 'high', and 'auto'. (param: quality) (code: invalid_value)
Related to https://github.com/etkecc/baibot/issues/40
🦻 is used for another purpose - to denote that a message is one coming
from speech-to-text, by:
- having the bot react to its own speech-to-text transcription message
with the 🦻 emoji when it's posted in a non-thread
- having the bot prefix its speech-to-text transcription messages with
`> 🦻` when it's posted in a thread
⏳ is already used as a progress indicator for other features, so it
makes sense to use it for indicating that speech-to-text is happening for a given audio message as well.
While 🦻 was an even more descriptive illustration of what's actually happening to the audio message
("it's being heard by the bot"), us using the 🦻 emoji for different things didn't seem good.
When speech-to-text/flow-type = `only_transcribe`, the bot will now send
text messages by default, not notices.
While notice messages may be less desirable with other bots in the room,
it's probably a better default for most people who enable "transcribe-only" mode.
This is an improvement related to https://github.com/etkecc/baibot/issues/14
Various clients (including newer versions of Element Web), do not like
it when the `body` field of the attachment is not a file name.
For images, a preview may not be shown and downloading the attachment
may suggest that the whole long text is used as a filename (which is odd).
There is value (improved accessibility, etc.)
in adding better descriptions (especially to generated images),
but given that it's currently problematic, I'm getting rid of it.
It's better and safer if we stick to using filenames.
This patch introduces a new `baibot_conversation_start_time_utc`
variable which indicates the time the conversation got started.
Using `baibot_now_utc` is still possible, but given that the current
time is a moving target, its use is in conflict with prompt caching.
Because the new `baibot_conversation_start_time_utc` prompt variable
is a more reasonable default, we're now using it in all sample configs.
The other prerequisite seems to be not using a `prompt` (`prompt: null`),
but we already supported this.
It'd be nice to add an optional `max_completion_tokens` parameter as
well, for the benefit of the o1 models, but this is not yet supported by
async-openai.
Possibly tracked here: https://github.com/64bit/async-openai/issues/272
Since 2024-10-02, `gpt-4o` is actually the same as `gpt-4o-2024-08-06`.
We previously used `gpt-4o-2024-08-06`, because it was pointing to a
much better (longer context) model. Since they're both the same now,
we'd better stick to the unpinned model and make it easier for future
users to get upgrades.
Fallback support was intentionally removed in 9908512968,
because it was deemed OK to do so.
It turns out that Element iOS still doesn't properly do user mentions
(and likely never will, until Element X replaces it), so we can't just
drop the fallback user mentions logic without affecting all these
clients. It's possible that the Element Android is no better (unverified claim).
This actually fixes 2 issues.
Fixes https://github.com/etkecc/baibot/issues/14
Fixes https://github.com/etkecc/baibot/issues/17
When people enable transcribe-only mode and the bot replies outside of a
thread, messages will no longer look like this: `> 🦻 Transcribed text`.
Instead, they will:
- look like this: `Transcribed text`
- get an emoji reaction (🦻) sent by the bot itself,
to indicate that the message is a transcription
---------------------------------------
As https://github.com/etkecc/baibot/issues/14 discusses,
the `> 🦻` prefixing of messages also served the purpose of indicating
to the bot that this is not its own message, but rather something it
"heard" from a user.
Given that out-of-thread replies no longer include this, they could be
mistaken for bot messages.
Because transcribed messages are posted as notice messages, we can
easily tell them apart from regular text-generated messages by the bot
itself, so we can (and do) treat them differently.
Thankfully, the bot does not yet support building a text-generation
conversation from arbitrary messages (something discussed in
https://github.com/etkecc/baibot/issues/15), so these out-of-thread
replies having the wrong owner are not an issue for now.
If we do land support for this, we'll probably need to make the bot inspect such notice messages
posted by it, inspect their reactons and attribute them properly (🦻 -> user message).
Fixes https://github.com/etkecc/baibot/issues/10
This also includes them in the default prompts (for newly-created agents),
so that people can get a better experience out of the box.
The previous approach (implemented in dd1dd78312) was simple
(send typing notices for as long as the "controller" is running),
but this proved to be overly simplistic and unable to handle edge-cases:
- in multi-user rooms (or rooms with a prefix requirement), the bot
used to send a typing notice while "working", but its work consisted
of ignoring the message. So it then sent a "not typing" notice.
This is wasteful and otherwise problematic - certain clients (like nheko)
do not handle this "race" well.
- certain reactions (anything other than 🗣️ right now) are meant to be
ignored. There's no point in doing the same "typing / not typing"
dance
- there are other instances where the bot may do work, but doesn't (due
to configuration or lack of capabilities)
This new more fine-grained implementation of typing notices aims to:
- only send a typing notice if actual "slow work" will be done
- avoid stopping & restarting typing notices (wasteful) if a chain of work is to
be performed (processing voice messages and doing speech-to-text +
text-generation + ...). Rather, maintaining typing notice sending
throughout
Previously, the bot only had rudimentary typing notification support.
It used to send a single notification when starting a long task
and did not bother with notifications anymore.
By default matrix-rust-sdk gives these notifications a validity of 4
seconds, so it would expire shortly. If the bot takes longer to respond,
you'd see the typing notification expire and wonder if a response is
coming.
Another edge case is the bot sending an answer quicker and the typing
notice still being on. Some clients (like element-web) seem to hide the
typing notice when a new message comes, so they don't experience this as
problematic.
The reworked typing notification system should be robust:
- typing notices are sent continuously, until the bot finishes doing
work
- if the bot is performing multiple actions in a room (even for
different people), typing notices would continue to be sent until the
bot becomes idle
- as soon as the bot becomes idle, a "not typing anymore" notice is sent
to clear the state