Files
baibot-withmcp/docs/usage.md
Slavi Pantaleev 8f86289373 Initial work on Vision support in text conversations and Image Editing
This is a huge patch which does some major refactoring like:

- renaming "Image Generation" to "Image Creation" in most places,
  to better match its new command (`!bai image create`)

- relocating image creation command (`!bai image` -> `!bai image create`),
  so it wouldn't conflict with the new image editing command (`!bai image edit`)

- introducing a new image editing command (`!bai image edit`), which
  is meant to work only with the OpenAI provider, but doesn't fully work yet
  due to https://github.com/64bit/async-openai/issues/364, though a next patch will fix it

- adding support for reading images off of Matrix conversations and forwarding them to
  text conversations. Works for OpenAI, but not for Anthropic yet
  (requires custom patches) and not for OpenAI-Compat (no support for
  images there)

- relocating some utils around (base64, mime)
2025-05-10 09:18:01 +03:00

6.7 KiB

📖 Usage

This document covers how to use the bot in a room.

The 🌟 Features page also includes details about how each feature works and can be configured.

💬 Text Generation

This is related to the 💬 Text Generation feature.

If there's a text-generation handler agent configured, the bot may respond to messages sent in the room.

See screenshots of:

Whether the bot responds depends on:

Room messages start a threaded conversation where you can continue back-and-forth communication with the bot. Using on-demand involvement, you can can also mention the bot to provoke it to get involved in any conversation thread or reply chain.

Unless you've enabled the ♻️ Context Management feature, all messages will be sent to the agent's API each time. If the context management feature is enabled, older messages may be dropped.

🗣️ Text-to-Speech

This is related to the 🗣️ Text-to-Speech feature.

If there's a text-to-speech handler agent configured, the bot may convert text messages sent to the room to audio (voice).

See:

By default, the bot:

🦻 Speech-to-Text

This is related to the 🦻 Speech-to-Text feature.

If there's a speech-to-text handler agent configured, the bot may transcribe voice messages sent to the room to text.

See a 🖼️ Screenshot of the default flow for Speech-to-Text and Text-Generation.

The speech-to-text feature triggers automatically by default, but can be adjusted via the 🦻 Speech-to-Text / 🪄 Flow Type setting.

If all your messages are in the same language, you can improve accuracy & latency by configuring the language (see 🦻 Speech-to-Text / 🔤 Language).

Image Generation

This feature is not configurable at the moment. The configuration (size, quality, style) specified at the 🤖 agent level will be used.

Capabilities depend on the ☁️ provider and model used.

🖌️ Creating images

Simply send a command like !bai image create A beautiful sunset over the ocean and the bot will start a threaded conversation and post an image based on your prompt.

See a 🖼️ Screenshot of the Image Creation feature.

You can then, respond in the same message thread with:

  • more messages, to add more criteria to your prompt.
  • a message saying again, to generate one more image with the current prompt.

🎨 Editing images

Simply send a command like !bai image edit Turn the following image into an anime-style drawing and the bot will start a threaded conversation asking for more details.

See a 🖼️ Screenshot of the Image Editing feature.

You can then, respond in the same message thread with:

  • more messages, to add more criteria to your prompt.
  • one or more images, to provide the images that the bot will operate on.
  • a message saying go, to start the image generation process.
  • a message saying again, to prompt the bot to generate one more image edit with the current prompt.

🫵 Creating stickers

A variation of creating images is to create "sticker images".

See a 🖼️ Screenshot of the Sticker Creation feature.

To create a sticker, send a command like !bai sticker A huge ramen bowl with lots of chashu and a mountain of beansprouts on top.

The difference from creating images is that the bot will:

  • generate a smaller-resolution image (currently hardcoded to 256x256) - smaller/quicker, but still good enough for a sticker
  • potentially switch to a different (cheaper or otherwise more suitable) model, if available
  • post the image directly to the room (as a reply to your message), without starting a threaded conversation

Some models (like OpenAI's Dall-E-3) can only generate larger images (1024x1024, etc., for a higher charge), so we switching to a smaller/cheaper model (like Dall-E-2) is a way to generate a sticker cheaply.