Initial commit
9
docs/README.md
Normal file
@@ -0,0 +1,9 @@
|
||||
# Table of Contents
|
||||
|
||||
- [🔒 Access](./access.md)
|
||||
- [🤖 Agents](./agents.md) and [☁️ Providers](./providers.md)
|
||||
- [🛠️ Configuration](./configuration/README.md)
|
||||
- [🌟 Features](./features.md)
|
||||
- [📖 Usage](./usage.md)
|
||||
- [🚀 Installation](./installation.md)
|
||||
- [💻 Development](./development.md)
|
||||
46
docs/access.md
Normal file
@@ -0,0 +1,46 @@
|
||||
## 🔒 Access
|
||||
|
||||
This bot employs access control to decide who can use its services and manage its configuration.
|
||||
|
||||
|
||||
### 👋 Joining rooms
|
||||
|
||||
The bot automatically joins rooms when invited by someone considered a bot [user](#-users).
|
||||
|
||||
|
||||
### 👥 Users
|
||||
|
||||
The bot will ignore messages (and room invitations) from unallowed users.
|
||||
|
||||
Users can **use all the bot's [features](./features.md)** ([💬 Text Generation](./features.md#-text-generation), [🦻 Speech-to-Text](./features.md#-speech-to-text), etc.), but **cannot manage the bot's configuration**.
|
||||
|
||||
The bot can be used by users that match some [dynamically](./configuration/README.md#dynamic-configuration) configured [Matrix user id](https://spec.matrix.org/v1.11/#users) patterns.
|
||||
|
||||
The following commands are available:
|
||||
- **Show** the currently allowed users: `!bai access users`
|
||||
- **Set** the list of allowed users: `!bai access set-users SPACE_SEPARATED_PATTERNS`
|
||||
|
||||
Example patterns: `@*:example.com @*:another.com @someone:company.org`
|
||||
|
||||
|
||||
### 👮♂️ Administrators
|
||||
|
||||
Administrators can **manage the bot's configuration and access control**.
|
||||
|
||||
The bot can be administrated by users that match some [statically](./configuration/README.md#static-configuration) configured [Matrix user id](https://spec.matrix.org/v1.11/#users) patterns.
|
||||
|
||||
Administrators cannot be changed without adjusting the bot's configuration on the server.
|
||||
|
||||
|
||||
### 💼 Room-local agent managers
|
||||
|
||||
Room-local agent managers are users privileged to **create their own [agents](./agents.md)** (see `!bai agent`) in rooms.
|
||||
Letting regular users create agents which contact arbitrary network services **may be a security issue**.
|
||||
|
||||
No room-local agent manager patterns are configured, so new agents can only be created by administrators.
|
||||
|
||||
The following commands are available:
|
||||
- **Show** the currently allowed users: `!bai access room-local-agent-managers`
|
||||
- **Set** the list of allowed users: `!bai access set-room-local-agent-managers SPACE_SEPARATED_PATTERNS`
|
||||
|
||||
Example patterns: `@*:synapse.127.0.0.1.nip.io @*:another.com @someone:company.org`
|
||||
61
docs/agents.md
Normal file
@@ -0,0 +1,61 @@
|
||||
## 🤖 Agents
|
||||
|
||||
An agent is an instantiation and configuration of some [☁️ provider](./providers.md).
|
||||
It can support different capabilities (text-generation, speech-to-text, etc.) depending on the provider used and on the configuration of the agent.
|
||||
|
||||
Agents can be set as **[🤝 handlers](./configuration/handlers.md) for various purposes** (text-generation, speech-to-text, etc.) globally or in specific rooms. Send a `!bai config status` command to see the current configuration.
|
||||
|
||||
Agents can be **defined [statically](./configuration/README.md#static-configuration)** (in the server configuration) **or dynamically** (via commands sent to the bot).
|
||||
|
||||
When [creating agents](#creating-agents) dynamically, you can do it **per-room or globally**.
|
||||
Globally-defined agents can be used by any authorized bot user in any room, while room-local agents can only be used in the room where they were defined.
|
||||
|
||||
Agent configuration (like all other configuration) is stored in the Matrix Account Data of the bot user and is **potentially encrypted** (if enabled in the configuration), so that your configuration data is safe even on untrusted homeservers.
|
||||
|
||||
|
||||
### Listing agents
|
||||
|
||||
To **list** all available agents: `!bai agent list`
|
||||
|
||||
|
||||
#### Creating agents
|
||||
|
||||
See a [🖼️ Screenshot of the agent creation process](./screenshots/agent-creation.webp).
|
||||
|
||||
To **create** a new agent, you need to specify the [provider](./providers.md) and an agent id of your choosing.
|
||||
|
||||
- **Create** a new agent:
|
||||
- (Accessible in **this room only**) `!bai agent create-room-local PROVIDER_ID AGENT_ID`
|
||||
- (Accessible in **all rooms**) `!bai agent create-global PROVIDER AGENT_ID`
|
||||
- Example: `!bai agent create-room-local openai my-openai-agent`
|
||||
|
||||
The `AGENT_ID` is a unique identifier for the agent. It can be any string which **doesn't contain spaces and `/`**.
|
||||
|
||||
Depending on where the agent is defined (within a room, globally, or [statically](./configuration/README.md#static-configuration)), this id will get a prefix (e.g. `room-local/`, `global/` or `static/`). The combined id (prefix + agent id) makes the **full agent identifier** (refered to as `FULL_AGENT_IDENTIFIER` in commands below).
|
||||
|
||||
When creating an agent, you will be given some sample [YAML](https://en.wikipedia.org/wiki/YAML) configuration which you can use to customize the agent's behavior.
|
||||
|
||||
This configuration varies depending on the [☁️ provider](./providers.md) used and the capabilities of the agent. Based on the configuration keys you pass, certain features will be enabled or disabled. For example, if you skip the `image_generation` key for an [OpenAI](./providers.md#openai) agent, it won't be able to generate images (see [🖌️ Image Generation](./features.md#-image-generation)).
|
||||
|
||||
After making your modifications to the sample YAML, you submit it back to the bot and the new agent will be created.
|
||||
|
||||
**To make use of the agent**, you need to [🤝 configure it as a handler for a given purpose](./configuration/handlers.md).
|
||||
|
||||
|
||||
### Showing agent details
|
||||
|
||||
To **show** full details for a given agent: `!bai agent details FULL_AGENT_IDENTIFIER`
|
||||
|
||||
This command requires a full agent identifier (e.g. `room-local/agent-id`).
|
||||
|
||||
|
||||
### Deleting agents
|
||||
|
||||
To **delete** an agent: `!bai agent delete FULL_AGENT_IDENTIFIER`
|
||||
|
||||
This command requires a full agent identifier (e.g. `room-local/agent-id`).
|
||||
|
||||
|
||||
### Updating agents
|
||||
|
||||
To **update** a given agent's configuration: show the agent's [details](#showing-agent-details) (current configuration), then [delete](#deleting-agents) it and finally [re-create](#creating-agents) it.
|
||||
48
docs/configuration/README.md
Normal file
@@ -0,0 +1,48 @@
|
||||
## 🛠️ Configuration
|
||||
|
||||
The bot's behavior is controlled by a combination of [static](#static-configuration) and [dynamic](#dynamic-configuration) configuration.
|
||||
|
||||
|
||||
### Static configuration
|
||||
|
||||
The bot can be configured using a [YAML](https://en.wikipedia.org/wiki/YAML) configuration file as well as [environment variables](https://en.wikipedia.org/wiki/Environment_variable).
|
||||
|
||||
When running the bot locally (during [🧑💻 development](../development.md)), the bot's configuration is read from the `var/app/config.yml` file.
|
||||
This file is created from the template found in [etc/app/config.yml.dist](../../etc/app/config.yml.dist).
|
||||
|
||||
Certain keys can be left unset, in which case [📝 hardcoded defaults](../../src/entity/cfg/defaults.rs) would be used.
|
||||
|
||||
Each configuration key found in the YAML configuration can be overridden by setting an environment variable (dots should be replaced with `_`). Example:
|
||||
|
||||
- to override `command_prefix`, set an environment variable `BAIBOT_COMMAND_PREFIX`
|
||||
- to override `homeserver.server_name`, set an environment variable `BAIBOT_HOMESERVER_SERVER_NAME`
|
||||
|
||||
The static configuration contains an `initial_global_config` key, which is used to populate the bot's global configuration (stored as [dynamic configuration](#dynamic-configuration)) the first time the bot starts. Modifying this subsequently will not have any effect. After initial global configuration creation, it's expected to be managed dynamically via chat commands.
|
||||
|
||||
|
||||
### Dynamic configuration
|
||||
|
||||
Besides the bot's [static configuration](#static-configuration), **the bot can also be configured dynamically at runtime (via chat messages)**.
|
||||
|
||||
This includes changes to [🔒 Access](../access.md), [🤖 Agents](../agents.md) and [🛠️ Room Settings](#room-settings).
|
||||
|
||||
|
||||
#### Room Settings
|
||||
|
||||
Room Settings come from 3 different levels with priority in the following order (higher to lower):
|
||||
|
||||
- 📍 per-room (`!bai config room ..` commands)
|
||||
- 🌐 globally (`!bai config global ..` commands)
|
||||
- 📝 as [hardcoded defaults](../../src/entity/cfg/defaults.rs)
|
||||
|
||||
You can adjust the following settings per room and/or globally:
|
||||
|
||||
- [💬 Text Generation](text-generation.md)
|
||||
- [🦻 Speech-to-Text](speech-to-text.md)
|
||||
- [🗣️ Text-to-Speech](text-to-speech.md)
|
||||
- [🖌️ Image Generation](image-generation.md)
|
||||
- [🤝 Handlers](handlers.md)
|
||||
|
||||
Refer to the bot's help messages (as a response to a `!bai config` help command) for the most up-to-date information on what Room Settings can be configured.
|
||||
|
||||
You can **get an overview of the configuration affecting the current room** (a mix of hardcoded defaults, agent defaults, global and room-level settings) by sending a `!bai config status` command to the room.
|
||||
32
docs/configuration/handlers.md
Normal file
@@ -0,0 +1,32 @@
|
||||
## 🤝 Handlers
|
||||
|
||||
### Introduction
|
||||
|
||||
You can use **different models in different rooms** (e.g. [OpenAI](../providers.md#openai) GPT-4o alongside [Llama](https://en.wikipedia.org/wiki/Llama_(language_model)) running on [Groq](../providers.md#groq), etc.)
|
||||
|
||||
You can also use **different models within the same room** (e.g. [💬 text-generation](#-text-generation) handled by one [agent](./agents.md), [🦻 speech-to-text](#-speech-to-text) handled by another, [🗣️ text-to-speech](#️-text-to-speech) by a 3rd, etc.)
|
||||
|
||||
The bot supports the following use-purposes:
|
||||
|
||||
- [💬 text-generation](../features.md#-text-generation): communicating with you via text
|
||||
- [🦻 speech-to-text](../features.md#-speech-to-text): turning your voice messages into text
|
||||
- [🗣️ text-to-speech](../features.md#️-text-to-speech): turning bot or users text messages into voice messages
|
||||
- [🖌️ image-generation](../features.md#-image-generation): generating images based on instructions
|
||||
|
||||
In a given room, each different purpose can be served by a different [provider](../providers.md) and model. This combination of provider and model configuration is called an [🤖 agent](../agents.md). Each purpose can be served by a different **handler** agent.
|
||||
|
||||
See a [🖼️ Screenshot of an example room configuration](./screenshots/config-status-handlers.webp).
|
||||
|
||||
|
||||
### Configuring
|
||||
|
||||
Handlers can be configured [dynamically](./README.md#dynamic-configuration):
|
||||
|
||||
- either per-room (e.g. `!bai config room set-handler text-generation room-local/openai-gpt-4o`)
|
||||
- or globally (e.g. `!bai config global set-handler text-generation global/openai-gpt-4o`)
|
||||
|
||||
The per-room configuration takes priority over the global configuration.
|
||||
|
||||
There's also a `catch-all` purpose that can be used as a fallback handler for messages that don't match any other handler.
|
||||
|
||||
💡 It's a good idea to globally-configure a powerful agent as a catch-all handler, so that the bot can always handle messages of any kind. You can then override individual handlers per room or globally.
|
||||
9
docs/configuration/image-generation.md
Normal file
@@ -0,0 +1,9 @@
|
||||
|
||||
## 🖌️ Image Generation
|
||||
|
||||
The Image Generation feature is not configurable at this moment.
|
||||
|
||||
You may also wish to see:
|
||||
|
||||
- [🌟 Features / 🖌️ Image Generation](../features.md#-image-generation) for a higher-level introduction to the Image Generation features
|
||||
- [📖 Usage / 🖌️ Image Generation](../usage.md#-image-generation) section for more details on how to use the bot for Image Generation in a room
|
||||
39
docs/configuration/speech-to-text.md
Normal file
@@ -0,0 +1,39 @@
|
||||
## 🦻 Speech-to-Text
|
||||
|
||||
Below are some configuration settings related to Speech-to-Text.
|
||||
|
||||
You may also wish to see:
|
||||
|
||||
- [🌟 Features / 🦻 Speech-to-Text](../features.md#-speech-to-text) for a higher-level introduction to the Speech-to-Text features
|
||||
- [📖 Usage / 🦻 Speech-to-Text](../usage.md#-speech-to-text) section for more details on how to use the bot for Speech-to-Text in a room
|
||||
|
||||
|
||||
### 🪄 Flow Type
|
||||
|
||||
Controls how voice messages sent by [👥 user](../access.md#-users) are handled.
|
||||
|
||||
The following configuration values are recognized:
|
||||
|
||||
- (default) `transcribe_and_generate_text`: the bot will turn [👥 user](../access.md#-users) voice messages into text and then generate text messages via [💬 Text Generation](../features.md#-text-generation). This is the default setting to allow for [Seamless voice interaction](../features.md#seamless-voice-interaction).
|
||||
|
||||
- `ignore`: the bot will ignore all audio messages
|
||||
|
||||
- `only_transcribe`: the bot will turn [👥 user](../access.md#-users) voice messages into text, but will **not** proceed with [💬 Text Generation](../features.md#-text-generation). Switching to this may be useful in some cases, as in [Transcribe-only mode](../features.md#transcribe-only-mode).
|
||||
|
||||
Example: `!bai config room speech-to-text set-flow-type ignore` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
||||
|
||||
|
||||
### 🔤 Language
|
||||
|
||||
Lets you specify the language of the input voice messages, to avoid using auto-detection.
|
||||
Supplying the input language using a 2-letter code (e.g. `ja`) as per [ISO-639-1](https://en.wikipedia.org/wiki/List_of_ISO_639-1_codes) may improve accuracy & latency.
|
||||
|
||||

|
||||
|
||||
In the above example screenshot, even without a language specified, the voice was understood correctly as [Bulgarian](https://en.wikipedia.org/wiki/Bulgarian_language), but was produced in latin, not [Cyrillic](https://en.wikipedia.org/wiki/Cyrillic_script), which is wrong.
|
||||
|
||||
If different [👥 user](../access.md#-users) are using different languages, do not specify a language.
|
||||
|
||||
💡 Certain models (like [OpenAI](../providers.md#openai)'s Whisper) may perform auto-translation if you specify a language, but you're speaking another one. You may abuse this side-effect for performing voice-to-text translation, but be aware that not all models behave this way.
|
||||
|
||||
Example (setting it to Japanese): `!bai config room speech-to-text set-language ja` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
||||
75
docs/configuration/text-generation.md
Normal file
@@ -0,0 +1,75 @@
|
||||
|
||||
## 💬 Text Generation
|
||||
|
||||
Below are some [🛠️ dynamic configuration settings](./README.md#dynamic-configuration) related to Text Generation.
|
||||
|
||||
You may also wish to see:
|
||||
|
||||
- [🌟 Features / 💬 Text Generation](../features.md#-text-generation) for a higher-level introduction to the Text Generation features
|
||||
- [📖 Usage / 💬 Text Generation](../usage.md#-text-generation) section for more details on how to use the bot for Text Generation in a room
|
||||
|
||||
|
||||
### 🗟 Prefix Requirement Type
|
||||
|
||||
In Direct Message rooms with the bot (1:1 rooms), it most usually makes sense for the bot to respond to **all** of your messages, as shown on this [🖼️ screenshot](../screenshots/text-generation.webp).
|
||||
|
||||
In group rooms (with multiple users), it may be more appropriate for the bot to only respond to messages that are **prefixed** with the command prefix (e.g. `!bai`), so that other chat exchange in the room will not trigger it. Such a setup is shown on this [🖼️ screenshot](../screenshots/text-generation-prefix-requirement.webp).
|
||||
|
||||
There are exceptions to these rules, and you can configure the bot to respond only to prefixed messages in a 1:1 room, or to respond to all messages even in a multi-user group room.
|
||||
|
||||
To support such use-cases, the bot has a `text-generation prefix-requirement-type` setting, which can be set to:
|
||||
|
||||
- (default) `no`: indicates that the bot would not require a prefix and would respond to all messages
|
||||
|
||||
- `command_prefix`: indicates that the bot would require that messages be prefixed with the command prefix (e.g. `!bai`) and would ignore all messages that are not prefixed
|
||||
|
||||
By default, the bot is **auto-configured (upon joining a new room)** to use the `no` setting in rooms that only include 2 users (you and the bot), and `command_prefix` in rooms with more than 2 users. To prevent surprises, the bot will **not** adjust this setting subsequently. You can manually adjust it via `!bai config room set-prefix-requirement-type VALUE`.
|
||||
|
||||
Example: `!bai config room text-generation set-prefix-requirement-type command_prefix` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
||||
|
||||
Regardless of this configuration, **the bot will also respond to messages which directly [mention](https://spec.matrix.org/latest/client-server-api/#user-and-room-mentions) the bot** (e.g. `@baibot`), even if they are not prefixed. An example of this can be seen on this [🖼️ screenshot](../screenshots/text-generation-prefix-requirement.webp).
|
||||
|
||||
|
||||
### 🪄 Auto Usage
|
||||
|
||||
Text generation is enabled by default (the `text-generation auto-usage` setting being set to `always`), but can be set to:
|
||||
|
||||
- (default) `always`: generate text for all messages (also see [🗟 Prefix Requirement Type](#-prefix-requirement-type))
|
||||
|
||||
- `never`: never generate text for messages
|
||||
|
||||
- `only_for_voice`: only generate text when the original user message was a voice message, later transcribed via [🦻 Speech-to-Text](../features.md#-speech-to-text)
|
||||
|
||||
- `only_for_text`: only generate text when original user message was a text message
|
||||
|
||||
Example: `!bai config room text-generation set-auto-usage only_for_voice` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
||||
|
||||
|
||||
### ♻️ Context Management
|
||||
|
||||
The bot also supports ♻️ **context management**, which automatically adjusts the message history length, etc.
|
||||
|
||||
This feature relies on [tokenization](https://en.wikipedia.org/wiki/Large_language_model#Tokenization) performed by the [tiktoken-rs](https://github.com/zurawiki/tiktoken-rs) library which is [poorly well-maintained](https://github.com/zurawiki/tiktoken-rs/issues/50) and only works well for [OpenAI](../providers.md#openai) models.
|
||||
|
||||
This setting is **disabled by default**, but can be enabled via `!bai config room context-management-enabled true` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings)).
|
||||
|
||||
|
||||
### ⌨️ Prompt Override
|
||||
|
||||
You can override the [system prompt](https://huggingface.co/docs/transformers/en/tasks/prompting) configured at the [🤖 agent](../agents.md) level.
|
||||
|
||||
Example (multi-line is supported):
|
||||
|
||||
```
|
||||
!bai config room text-generation set-prompt-override You're a UI/UX expert. Everything you say needs to consider design and usability.
|
||||
|
||||
Where appropriate, you'll mention best practices and common pitfalls.
|
||||
```
|
||||
|
||||
A prompt override can also be set globally, see [🛠️ Room Settings](./README.md#room-settings).
|
||||
|
||||
### 🌡️ Temperature Override
|
||||
|
||||
You can override the [temperature](https://blogs.novita.ai/what-are-large-language-model-settings-temperature-top-p-and-max-tokens/#what-is-llm-temperature) (randomness / creativity) parameter configured at the [🤖 agent](../agents.md) level.
|
||||
|
||||
Example: `!bai config room text-generation set-temperature-override 3.5` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
||||
63
docs/configuration/text-to-speech.md
Normal file
@@ -0,0 +1,63 @@
|
||||
|
||||
## 🗣️ Text-to-Speech
|
||||
|
||||
Below are some configuration settings related to Text-to-Speech.
|
||||
|
||||
You may also wish to see:
|
||||
|
||||
- [🌟 Features / 🗣️ Text-to-Speech](../features.md#-text-generation) for a higher-level introduction to the Text-to-Speech features
|
||||
- [📖 Usage / 🗣️ Text-to-Speech](../usage.md#-text-generation) section for more details on how to use the bot for Text-to-Speech in a room
|
||||
|
||||
|
||||
### 🪄 Bot Messages Flow Type
|
||||
|
||||
Controls how automatic text-to-speech functions for **messages sent by the bot**.
|
||||
|
||||
The following configuration values are recognized:
|
||||
|
||||
- (default) `on_demand_for_voice`: the bot will turn its own text messages into audio (voice) messages only after an allowed [👥 user](../access.md#-users) **reacts** to a bot's message with 🗣️. To make it easier for users to react without having to hunt for this emoji, the bot will automatically add a 🗣️ reaction to its own messages which are in response to a user audio (voice) message.
|
||||
|
||||
- `on_demand_always`: the bot will turn its own text messages into audio (voice) messages only after an allowed [👥 user](../access.md#-users) **reacts** to a bot's message with 🗣️. To make it easier for users to react without having to hunt for this emoji, the bot will automatically add a 🗣️ reaction to **all of its own messages**.
|
||||
|
||||
- `only_for_voice`: the bot will turn its own text messages into audio (voice) messages only if the original user message was a voice message. This is to allow for [Seamless voice interaction](../features.md#seamless-voice-interaction), where you can speak to the bot and then hear its responses
|
||||
|
||||
- `never`: the bot will never turn its own text messags into audio (voice) messages
|
||||
|
||||
- `always`: the bot will turn all its text messages into audio (voice) messages. This also allows for [Seamless voice interaction](../features.md#seamless-voice-interaction).
|
||||
|
||||
Example: `!bai config room text-to-speech set-bot-msgs-flow-type never` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
||||
|
||||
|
||||
### 🪄 User Messages Flow Type
|
||||
|
||||
Controls how automatic text-to-speech functions for **messages sent by [👥 users](../access.md#-users)**.
|
||||
|
||||
**Only works when automatic text-generation is disabled** (see [💬 Text Generation / 🪄 Auto Usage](./text-generation.md#-auto-usage)).
|
||||
|
||||
The following configuration values are recognized:
|
||||
|
||||
- (default) `never`: the bot will never turn [👥 user](../access.md#-users) text messages into audio (voice) messages
|
||||
|
||||
- `on_demand`: the bot will turn [👥 user](../access.md#-users) text messages into audio (voice) messages if the text message receives a 🗣️ reaction
|
||||
|
||||
- `always`: the bot will turn all [👥 user](../access.md#-users) text messages into audio (voice) messages. This is to allow for [Text-to-Speech-only mode](../features.md/#text-to-speech-only-mode).
|
||||
|
||||
Example: `!bai config room text-to-speech set-user-msgs-flow-type always` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
||||
|
||||
|
||||
### 🗲 Speed override
|
||||
|
||||
The speed override setting lets you speed up/down speech relative to the default speed configured at the [🤖 agent](../agents.md) level (usually `1.0`).
|
||||
|
||||
Values typically range from `0.25` to `4.0`, but may vary depending on the selected model.
|
||||
|
||||
Example: `!bai config room text-to-speech set-speed-override 1.5` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
||||
|
||||
|
||||
### 👫 Voice override
|
||||
|
||||
The voice override setting lets you change the voice being used by the text-to-speech model configured at the [🤖 agent](../agents.md) level (usually `onyx` when using [OpenAI](../providers.md#openai)).
|
||||
|
||||
Possible values (e.g. `onyx`) depend on the model you're using. For example, for [OpenAI](../providers.md#openai)'s Whisper model, [these voices](https://platform.openai.com/docs/guides/text-to-speech/voice-options) are available.
|
||||
|
||||
Example: `!bai config room text-to-speech set-voice-override nova` (this can also be set globally, see [🛠️ Room Settings](./README.md#room-settings))
|
||||
134
docs/development.md
Normal file
@@ -0,0 +1,134 @@
|
||||
## 🧑💻 Development
|
||||
|
||||
This documentation page contains information about **running the bot locally for development purposes**.
|
||||
This can also **helpful for quickly testing the bot in a containerized environment, with all dependency services included**.
|
||||
|
||||
For running the bot against your Matrix server, see the [🚀 Installation](./installation.md) documentation.
|
||||
|
||||
This bot is built in [🦀 Rust](https://www.rust-lang.org/) and uses the [mxlink](https://github.com/etkecc/rust-mxlink) library (built on top of [matrix-rust-sdk](https://github.com/matrix-org/matrix-rust-sdk)).
|
||||
|
||||
For local development, we run all dependency services in [🐋 Docker](https://www.docker.com/) containers via [docker-compose](https://docs.docker.com/compose/).
|
||||
|
||||
|
||||
### Prerequisites
|
||||
|
||||
- [🐋 Docker](https://www.docker.com/) and [docker-compose](https://docs.docker.com/compose/)
|
||||
- [Just](https://github.com/casey/just)
|
||||
- (Optional) [🦀 Rust](https://www.rust-lang.org/) - for compiling and running outside of a container
|
||||
- (Optional) an API key for some Large Language Model [☁️ provider](./providers.md) (e.g. [OpenAI](./providers.md#openai)), though we recommend using [LocalAI](#localai) or [Ollama](#ollama) for local development
|
||||
|
||||
|
||||
### Getting started guide
|
||||
|
||||
Developing [locally](#running-locally) is possible, but requires a [Rust](https://www.rust-lang.org/) toolchain.
|
||||
If this dependency is problematic for you, consider [🐋 running in a container](#running-in-a-container).
|
||||
|
||||
In any case, you will need [🐋 Docker](https://www.docker.com/) as [dependency services](../etc/services/) run there.
|
||||
|
||||
|
||||
#### Running locally
|
||||
|
||||
1. Start the core dependency services (Postgres, Synapse, Element Web): `just services-start`
|
||||
2. (Only the first time around) Prepare initial app configuration in `var/app/local/config.yml`: `just app-local-prepare`
|
||||
3. (Only the first time around) [Prepare your configuration file](#prepare-your-configuration-file)
|
||||
4. (Only the first time around) Prepare initial default Matrix user accounts (`admin` and `baibot`): `just users-prepare`
|
||||
5. (Optional) Start additional services depending on which [agent provider you've chosen](#choosing-an-agent-provider):
|
||||
- for [LocalAI](#localai):
|
||||
- Start services: `just localai-start`
|
||||
- Wait a while for LocalAI to start up. It has a lot of models to download. Monitor progress using `just localai-tail-logs`
|
||||
- When ready, you'll be able to reach LocalAI's web interface at http://localai.127.0.0.1.nip.io:42027/ (not that you really need it)
|
||||
- for [Ollama](#ollama):
|
||||
- Start services: `just ollama-start`
|
||||
- (Only the first time around) Pull the model configured in `agents.static_definitions` in the configuration file: `just ollama-pull-model gemma2:2b`
|
||||
6. Start the bot: `just run-locally`
|
||||
7. Go to http://element.127.0.0.1.nip.io:42025/ and login with `admin` / `admin`
|
||||
8. Create a new room and invite `@baibot:synapse.127.0.0.1.nip.io`
|
||||
9. When done, stop the bot (`Ctrl` + `C`)
|
||||
10. Stop the core dependency services: `just services-stop`
|
||||
11. (Optional) Stop additional services:
|
||||
- for [LocalAI](#localai): `just localai-stop`
|
||||
- for [Ollama](#ollama): `just ollama-stop`
|
||||
|
||||
|
||||
#### Running in a container
|
||||
|
||||
You can avoid having a [Rust](https://www.rust-lang.org/) toolchain installed locally and build/run this in a container.
|
||||
|
||||
1. Start the core dependency services (Postgres, Synapse, Element Web): `just services-start`
|
||||
2. (Only the first time around) Prepare initial app configuration in `var/app/container/config.yml`: `just app-container-prepare`
|
||||
3. (Only the first time around) [Prepare your configuration file](#prepare-your-configuration-file)
|
||||
4. (Only the first time around) Prepare initial default Matrix user accounts (`admin` and `baibot`): `just users-prepare`
|
||||
5. (Optional) Start additional services depending on which [agent provider you've chosen](#choosing-an-agent-provider):
|
||||
- for [LocalAI](#localai):
|
||||
- Start services: `just localai-start`
|
||||
- Wait a while for LocalAI to start up. It has a lot of models to download. Monitor progress using `just localai-tail-logs`
|
||||
- When ready, you'll be able to reach LocalAI's web interface at http://localai.127.0.0.1.nip.io:42027/ (not that you really need it)
|
||||
- for [Ollama](#ollama):
|
||||
- Start services: `just ollama-start`
|
||||
- (Only the first time around) Pull the model configured in `agents.static_definitions` in the configuration file: `just ollama-pull-model gemma2:2b`
|
||||
6. Start the bot: `just run-in-container`
|
||||
7. Go to http://element.127.0.0.1.nip.io:42025/ and login with `admin` / `admin`
|
||||
8. Create a new room and invite `@baibot:synapse.127.0.0.1.nip.io`
|
||||
9. When done, stop the bot (`Ctrl` + `C`)
|
||||
10. Stop the dependency services: `just services-stop`
|
||||
11. (Optional) Stop additional services:
|
||||
- for [LocalAI](#localai): `just localai-stop`
|
||||
- for [Ollama](#ollama): `just ollama-stop`
|
||||
|
||||
|
||||
#### Prepare your configuration file
|
||||
|
||||
This is about editing your configuration. The initial configuration is created based on `etc/app/config.yml.dist` when you run `just app-local-prepare` or `just app-container-prepare`.
|
||||
|
||||
Depending on whether you run locally or in a container, your configuration lives in a different file (`var/app/local/config.yml` and `var/app/container/config.yml`, respectively).
|
||||
|
||||
Before starting the bot, you may wish to adjust this configuration.
|
||||
|
||||
|
||||
##### Choosing an agent provider
|
||||
|
||||
You can create [🤖 agents](./agents.md) either [statically](./configuration/README.md#static-configuration) or [dynamically](./configuration/README.md#dynamic-configuration) using any of the supported [☁️ providers](./providers.md).
|
||||
|
||||
For getting started most quickly (and locally), we recommend using [LocalAI](#localai) or [Ollama](#ollama). These services are already configured to run as [local services via docker-compose](../etc/services/).
|
||||
|
||||
**Ollama is most lightweight** (~2GB for the container image + ~1.6GB for the model), but supports only [💬 text-generation](./features.md#-text-generation).
|
||||
|
||||
**LocalAI requires 4x more disk space** (~6GB for the container image + ~12GB for the models), but supports [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text) and [🖼️ image-generation](./features.md#️-image-generation).
|
||||
|
||||
**OpenAI supports all of these capabilities** as well and does not require powerful hardware or lots of disk space. However, it requires signup and an API key.
|
||||
|
||||
For local testing, **we recommend LocalAI**, because it runs fully locally and supports more features than Ollama.
|
||||
|
||||
###### LocalAI
|
||||
|
||||
[LocalAI](./providers.md#localai) supports all [🌟 features](./features.md) of the bot.
|
||||
|
||||
If you decided to go with [LocalAI](./providers.md#localai):
|
||||
|
||||
- enable the `localai` entry in the `agents.static_definitions` list in the configuration file
|
||||
- adjust the `initial_global_config.handler.catch_all` setting in the configuration file (`null` -> `static/localai`)
|
||||
|
||||
By default, we configure LocalAI to use the [All-In-One images](https://localai.io/basics/container/#all-in-one-images) running on the CPU.
|
||||
Performance is not great, but it should work reasonably well on good hardware.
|
||||
|
||||
If you'd like to use GPU acceleration, you may adjust the `SERVICE_LOCALAI_IMAGE_NAME` variable in [var/services/env](../var/services/env) (this file is automatically prepared for you based on [etc/services/env.dist](../etc/services/env.dist)) to use [other available LocalAI All-In-One images](https://localai.io/basics/container/#available-aio-images).
|
||||
|
||||
###### Ollama
|
||||
|
||||
[Ollama](./providers.md#ollama) only supports [💬 text-generation](./features.md#-text-generation).
|
||||
|
||||
If you decided to go with [Ollama](./providers.md#ollama):
|
||||
|
||||
- enable the `ollama` entry in the `agents.static_definitions` list in the configuration file
|
||||
- adjust the `initial_global_config.handler.catch_all` setting in the configuration file (`null` -> `static/ollama`)
|
||||
|
||||
The [gemma2:2b](https://ollama.com/library/gemma2:2b) model was chosen as a default, because it's smallest/lightest and should run well under [Ollama](./providers.md#ollama) on most machines.
|
||||
|
||||
###### OpenAI
|
||||
|
||||
[OpenAI](./providers.md#openai) supports all [🌟 features](./features.md) of the bot.
|
||||
|
||||
If you decided to go with [OpenAI](./providers.md#openai):
|
||||
|
||||
- enable the `openai` entry in the `agents.static_definitions` list in the configuration file
|
||||
- adjust the `initial_global_config.handler.catch_all` setting in the configuration file (`null` -> `static/openai`)
|
||||
150
docs/features.md
Normal file
@@ -0,0 +1,150 @@
|
||||
## 🌟 Features
|
||||
|
||||
### 🎨 Mixing & matching models
|
||||
|
||||
You can use **different models in different rooms** (e.g. [OpenAI](./providers.md#openai) GPT-4o alongside [Llama](https://en.wikipedia.org/wiki/Llama_(language_model)) running on [Groq](./providers.md#groq), etc.)
|
||||
|
||||
You can also use **different models within the same room** (e.g. [💬 text-generation](#-text-generation) handled by one [🤖 agent](./agents.md), [🦻 speech-to-text](#-speech-to-text) handled by another, [🗣️ text-to-speech](#️-text-to-speech) by a 3rd, etc.)
|
||||
|
||||
The bot supports the following use-purposes:
|
||||
|
||||
- [💬 text-generation](#-text-generation): communicating with you via text
|
||||
- [🦻 speech-to-text](#-speech-to-text): turning your voice messages into text
|
||||
- [🗣️ text-to-speech](#️-text-to-speech): turning bot or users text messages into voice messages
|
||||
- [🖌️ image-generation](#-image-generation): generating images based on instructions
|
||||
|
||||
In a given room, each different purpose can be served by a different [☁️ provider](./providers.md) and model. This combination of provider and model configuration is called an [🤖 agent](./agents.md). Each purpose can be served by a different **handler** agent.
|
||||
|
||||
See a [🖼️ Screenshot of an example room configuration](./screenshots/config-status-handlers.webp).
|
||||
|
||||
For more information about configuring handlers, see the [🤝 Handlers / Configuring](./configuration/handlers.md#configuring) documentation section.
|
||||
|
||||
|
||||
### 💬 Text Generation
|
||||
|
||||
Text Generation is the bot's ability to **respond to users' text messages with text**.
|
||||
|
||||

|
||||
|
||||
In multi-user (group) rooms, to avoid disturbing the normal conversation between people, the bot is auto-configured to only respond to messages starting with the command prefix (`!bai`) or direct mentions via the [💬 Text Generation / 🗟 Prefix Requirement Type](./configuration/text-generation.md#-prefix-requirement-type) setting.
|
||||
|
||||
A few other features (like [🗣️ Text-to-Speech](#️-text-to-speech) and [🦻 Speech-to-Text](#-speech-to-text)) combine well with Text Generation, so you **don't necessarily need to communicate with the bot via text** (with [Seamless voice interaction](#seamless-voice-interaction), you can communicate only with voice).
|
||||
|
||||
You may also wish to see:
|
||||
|
||||
- [🛠️ Configuration / 💬 Text Generation](./configuration/README.md#-text-generation) for configuration options related to Text Generation
|
||||
- [📖 Usage / 💬 Text Generation](./usage.md#-text-generation) section for more details on how to use the bot for Text Generation in a room
|
||||
|
||||
|
||||
### 🗣️ Text-to-Speech
|
||||
|
||||
Text-to-Speech is the bot's ability to **turn text messages into voice messages**.
|
||||
|
||||
It can be performed **on the bot's own text messages** (responses to yours due to [💬 Text Generation](#-text-generation)) and/or **on your own text messages**.
|
||||
|
||||
Text-to-Speech can be enabled to be done automatically or on-demand (only after reacting to a message with 🗣️), and is configurable for different message types ([🪄 Bot Messages Flow Type](./configuration/README.md#-bot-messages-flow-type) vs [🪄 User Messages Flow Type](./configuration/README.md#-user-messages-flow-type)).
|
||||
|
||||
By default, the bot **doesn't** perform text-to-speech. It can be configured for [Seamless voice interaction](#seamless-voice-interaction), where you can **speak to the bot** (instead of typing) and then **hear its responses**.
|
||||
|
||||
Another use-case is to have the bot operate in [Text-to-Speech-only mode](#text-to-speech-only-mode).
|
||||
|
||||
- [🛠️ Configuration / 🗣️ Text-to-Speech](./configuration/README.md#-text-to-speech) for configuration options related to Text-to-Speech
|
||||
- [📖 Usage / 🗣️ Text-to-Speech](./usage.md#-text-to-speech) section for more details on how to use the bot for Text-to-Speech in a room
|
||||
|
||||
|
||||
#### Text-to-Speech-only mode
|
||||
|
||||
You may wish to have the bot **automatically turn your text messages into voice messages**, but **without** doing [💬 Text Generation](#-text-generation).
|
||||
|
||||

|
||||
|
||||
This could be useful in a room with others, where you'd like to post text messages and have people in the room consume them more easily (by listening to audio).
|
||||
|
||||
To allow for this use-case, you can:
|
||||
|
||||
- disable [💬 Text Generation](#-text-generation) (via [💬 Text Generation / 🪄 Auto Usage](./configuration/text-generation.md#-auto-usage) setting): `!bai config room text-generation set-auto-usage never`
|
||||
|
||||
- enable [🗣️ Text-to-Speech](#️-text-to-speech) for user messages (via [🗣️ Text-to-Speech / 🪄 User Messages Flow Type](./configuration/text-to-speech.md#-user-messages-flow-type)): `!bai config room text-to-speech set-user-msgs-flow-type always` (or `on_demand`)
|
||||
|
||||
|
||||
### 🦻 Speech-to-Text
|
||||
|
||||
Speech-to-Text is the bot's ability to **turn voice messages into text**.
|
||||
|
||||

|
||||
|
||||
The default flow is shown in the screenshot above: your voice messages are transcribed to text and [💬 Text Generation](#-text-generation) is performed. By default, the bot offers [🗣️ Text-to-Speech](#️-text-to-speech) for its answers via a 🗣️ emoji. You can click it to trigger text-to-speech on-demand.
|
||||
|
||||
You may also configure the bot for [Seamless voice interaction](#seamless-voice-interaction) or [Transcribe-only mode](#transcribe-only-mode), etc.
|
||||
|
||||
|
||||
#### Seamless voice interaction
|
||||
|
||||
The bot can perform seamless voice interaction (🗣️-to-🗣️), allowing you to **speak to the bot** (instead of typing) and then **hear its responses**.
|
||||
|
||||

|
||||
|
||||
The flow is like this:
|
||||
|
||||
1. 👤 You sending a voice message
|
||||
2. 🤖 The bot:
|
||||
- (default) first turning your **voice message into text** ([🦻 Speech-to-Text](#-speech-to-text)) and posting it as a reply. This lets you you see what the bot heard.
|
||||
- (default) then **answering in text** ([💬 Text Generation](#-text-generation)). This lets you read/skim text, if you so prefer.
|
||||
- (can be enabled) finally **turning the answer's text into a voice message** ([🗣️ Text-to-Speech](#️-text-to-speech))
|
||||
3. 👤 You continuing the conversation via text or voice messages
|
||||
|
||||
⚠️ Certain clients (like [Element](https://element.io/)) only support sending voice messages as top-level room messages, not as thread replies. Until this client limitation is fixed, Element users can only send the 1st message as a voice message - subsequent replies in the same conversation thread will need to be sent as text messages.
|
||||
|
||||
By default, the last part of the aforementioned flow is **not enabled**, because we assume **a saner default is to reply with text and merely *offer* text-to-speech to those who want it**. Offering is done by the bot reacting to its own message with 🗣️, and letting you click this emoji to trigger text-to-speech on-demand.
|
||||
|
||||
To enable automatic text-to-speech for the bot's messages, set the [🗣️ Text-to-Speech / 🪄 Bot Messages Flow Type](./configuration/text-to-speech.md#-bot-messages-flow-type) setting to `only_for_voice` or `always` (e.g. `!bai config room text-to-speech set-bot-msgs-flow-type only_for_voice`).
|
||||
|
||||
|
||||
#### Transcribe-only mode
|
||||
|
||||
If you'd like to have the bot **only turn voice messages into text** (without generating text messages or voice messages), you can configure the bot for that.
|
||||
|
||||

|
||||
|
||||
To operate in this mode, you can:
|
||||
|
||||
- disable [💬 Text Generation](#-text-generation) (via [💬 Text Generation / 🪄 Auto Usage](./configuration/text-generation.md#-auto-usage) setting): `!bai config room text-generation set-auto-usage never`
|
||||
|
||||
- adjust the [🦻 Speech-to-Text / 🪄 Flow Type](./configuration/speech-to-text.md#-flow-type) setting to make the bot only transcribe (without doing [💬 Text Generation](#-text-generation)): `!bai config room speech-to-text set-flow-type only_transcribe`
|
||||
|
||||
|
||||
### 🖌️ Image Generation
|
||||
|
||||
Image generation is the bot's ability to **generate images** based on text prompts.
|
||||
|
||||
See a [🖼️ Screenshot of the Image Generation feature](./screenshots/image-generation.webp).
|
||||
|
||||
You may also wish to see:
|
||||
|
||||
- [🛠️ Configuration / 🖌️ Image Generation](./configuration/README.md#-image-generation) for configuration options related to Image Generation
|
||||
- [📖 Usage / 🖌️ Image Generation](./usage.md#-image-generation) section for more details on how to use the bot for Image Generation in a room
|
||||
- [🫵 Sticker Generation](#-sticker-generation) - a special case of Image Generation
|
||||
|
||||
|
||||
### 🫵 Sticker Generation
|
||||
|
||||
Sticker generation is the bot's ability to **generate sticker** images based on text prompts. It's a special case of [🖌️ Image Generation](#️-image-generation).
|
||||
|
||||
See a [🖼️ Screenshot of the Sticker Generation feature](./screenshots/sticker-generation.webp).
|
||||
|
||||
See [📖 Usage / 🖌️ Image Generation / Generating Stickers](./usage.md#generating-stickers) for details.
|
||||
|
||||
|
||||
### 🔒 Encryption
|
||||
|
||||
#### Message exchange
|
||||
|
||||
The bot works in both **unencrypted and encrypted Matrix rooms**.
|
||||
|
||||
If configured, the bot can make use of **Matrix's Secure Storage (Recovery) feature**, so that it can restore its encryption keys even its local database gets lost.
|
||||
|
||||
#### Configuration
|
||||
|
||||
The bot also stores its [🛠️ configuration](./configuration/README.md) (both 📍 per-room and 🌐globally) in Matrix Account Data, which is **generally stored as plain-text in the server**.
|
||||
|
||||
To overcome this Matrix limitation, the bot can **optionally encrypt the configuration data** before storing it in Account Data. This allows for the bot to be used securely even against untrusted servers, without leaking sensitive configuration data to them.
|
||||
93
docs/installation.md
Normal file
@@ -0,0 +1,93 @@
|
||||
## 🚀 Installation
|
||||
|
||||
☁️ The easiest way to use the bot is to **get a managed Matrix server from [etke.cc](https://etke.cc/)** and order baibot via the [order form](https://etke.cc/order/). Existing customers can request the inclusion of this additional service by [contacting support](https://etke.cc/contacts/).
|
||||
|
||||
💻 If you're managing your Matrix server with the help of the [matrix-docker-ansible-deploy](https://github.com/spantaleev/matrix-docker-ansible-deploy) Ansible playbook, you can easily **install the bot via the Ansible playbook**. See the playbook's [Setting up baibot](https://github.com/spantaleev/matrix-docker-ansible-deploy/blob/master/docs/configuring-playbook-bot-baibot.md) documentation page.
|
||||
|
||||
🐋 In other cases, we **recommend using our [prebuilt container images](https://github.com/etkecc/baibot/pkgs/container/baibot) and [running in a container](#-running-in-a-container)**. You can also [build a container image](#building-a-container-image) yourself.
|
||||
|
||||
🔨 If containers are not your thing, you can [build a binary](#-building-a-binary) yourself and [run it](#-running-a-binary).
|
||||
|
||||
🗲 For a quick experiment, you can refer to the [🧑💻 development documentation](./development.md) which contains information on how to build and run the bot (and its various dependency services) locally.
|
||||
|
||||
|
||||
### 🐋 Building a container image
|
||||
|
||||
We provide prebuilt container images for the `amd64` and `arm64` architectures, so **you don't necessarily need to build images yourself** and can jump to [Running in a container](#-running-in-a-container).
|
||||
|
||||
If you nevertheless wish to build a container image yourself, you can do so by running `just build-container-image`.
|
||||
This will build and tag your container image as `localhost/baibot:latest`.
|
||||
|
||||
|
||||
### 🐋 Running in a container
|
||||
|
||||
We recommend using a **tagged-release** (e.g. `v1.0.0`, not `latest`) of our [prebuilt container images](https://github.com/etkecc/baibot/pkgs/container/baibot), but you can also [build a container image](#-building-a-container-image) yourself.
|
||||
|
||||
You should:
|
||||
|
||||
- [🛠️ prepare a configuration file](#-preparing-a-configuration-file) (e.g. `cp etc/app/config.yml.dist /path/to/config.yml` & edit it)
|
||||
- prepare a data directory (`mkdir /path/to/data`)
|
||||
|
||||
The example below uses [🐋 Docker](https://www.docker.com/) to run the container, but other container runtimes like [Podman](https://podman.io/) should work as well.
|
||||
|
||||
```sh
|
||||
# Adjust the version tag to point to the latest available tagged version.
|
||||
# If building your own container image name, adjust to something like `localhost/baibot:latest`.
|
||||
CONTAINER_IMAGE_NAME=ghcr.io/etkecc/baibot:v1.0.0
|
||||
|
||||
/usr/bin/env docker run \
|
||||
-it \
|
||||
--rm \
|
||||
--name=baibot \
|
||||
--user=$(id -u):$(id -g) \
|
||||
--cap-drop=ALL \
|
||||
--read-only \
|
||||
--env BAIBOT_PERSISTENCE_DATA_DIR_PATH=/data \
|
||||
--mount type=bind,src=/path/to/config.yml,dst=/app/config.yml,ro \
|
||||
--mount type=bind,src=/path/to/data,dst=/data \
|
||||
$CONTAINER_IMAGE_NAME
|
||||
```
|
||||
|
||||
💡 If you've defined the `persistence.data_dir_path` setting in the `config.yml` file, you can skip the `BAIBOT_PERSISTENCE_DATA_DIR_PATH` environment variable.
|
||||
|
||||
|
||||
### 🔨 Building a binary
|
||||
|
||||
To build a binary, you need a [🦀 Rust](https://www.rust-lang.org/) toolchain.
|
||||
|
||||
Consult the [Dockerfile](../Dockerfile) file to learn what some of the build dependencies are (e.g. `libssl-dev`, `libsqlite3-dev`, etc., on Debian-based distros).
|
||||
|
||||
You can build a binary from the current project's source code:
|
||||
|
||||
- in `debug` mode via: `just build-debug`, yielding a binary in `target/debug/baibot`
|
||||
- (recommended) in `release` mode via: `just build-release`, yielding a binary in `target/release/baibot`
|
||||
|
||||
💡 Unless you're [🧑💻 developing](./development.md), you probably wish to build in release mode, as that provides a much smaller and more optimized binary.
|
||||
|
||||
📦 You can also install from the [baibot](https://crates.io/crates/baibot) crate published to [crates.io](https://crates.io) with the help of the [cargo](https://doc.rust-lang.org/cargo/) package manager by running: `cargo install baibot`.
|
||||
|
||||
|
||||
### 🖥️ Running a binary
|
||||
|
||||
Once you've [🔨 built a binary](#-building-a-binary) and [🛠️ prepared a configuration file](#-preparing-a-configuration-file), you can run it.
|
||||
|
||||
Consult the [Dockerfile](../Dockerfile) file to learn what some of the runtime dependencies are (e.g. `ca-certificates`, `sqlite3`, etc., on Debian-based distros).
|
||||
|
||||
You can run the binary like this:
|
||||
|
||||
```sh
|
||||
BAIBOT_CONFIG_FILE_PATH=/path/to/config.yml \
|
||||
BAIBOT_PERSISTENCE_DATA_DIR_PATH=/path/to/data \
|
||||
./target/release/baibot
|
||||
```
|
||||
|
||||
💡 If you've defined the `persistence.data_dir_path` setting in the `config.yml` file, you can skip the `BAIBOT_PERSISTENCE_DATA_DIR_PATH` environment variable.
|
||||
|
||||
💡 If your `config.yml` file is in your working directory (which may be different than the directory the binary lives in), you can skip the `BAIBOT_CONFIG_FILE_PATH` environment variable.
|
||||
|
||||
|
||||
### 🛠️ Preparing a configuration file
|
||||
|
||||
For an introduction to the configuration file, see the [🛠️ Configuration](./configuration/README.md) page.
|
||||
|
||||
Generally, you need to copy the configuration file template ([etc/app/config.yml.dist](../etc/app/config.yml.dist)) and make modifications as needed.
|
||||
170
docs/providers.md
Normal file
@@ -0,0 +1,170 @@
|
||||
## ☁️ Providers
|
||||
|
||||
[🤖 Agents](./agents.md) are powered by a provider. The provider could be a **local service** or a **cloud service**.
|
||||
|
||||
The list of supported providers is below.
|
||||
|
||||
|
||||
### Table of contents
|
||||
|
||||
- [How to choose a provider](#how-to-choose-a-provider)
|
||||
- [How to use a provider](#how-to-use-a-provider)
|
||||
- [Supported providers](#supported-providers)
|
||||
- [Anthropic](#anthropic)
|
||||
- [Groq](#groq)
|
||||
- [LocalAI](#localai)
|
||||
- [Mistral](#mistral)
|
||||
- [Ollama](#ollama)
|
||||
- [OpenAI](#openai)
|
||||
- [OpenAI Compatible](#openai-compatible)
|
||||
- [OpenRouter](#openrouter)
|
||||
- [Together AI](#together-ai)
|
||||
|
||||
|
||||
### How to choose a provider
|
||||
|
||||
If you're not sure which provider to start with, we **recommend [OpenAI](#openai)** as it's the most popular and has the **widest range of capabilities**: [💬 text-generation](./features.md#-text-generation), [🖌️ image-generation](./features.md#️-image-generation), [🦻 speech-to-text](./features.md#-speech-to-text), [🗣️ text-to-speech](./features.md#️-text-to-speech).
|
||||
|
||||
You don't need to choose just one though. The bot supports [mixing & matching models](./features.md#-mixing--matching-models), so you can use multiple providers at the same time.
|
||||
|
||||
|
||||
### How to use a provider
|
||||
|
||||
- sign up for it
|
||||
- obtain an API key
|
||||
- [create a new agent](./agents.md#creating-agents)
|
||||
- set it as a handler for some types of messages (see [Mixing & matching models](./features.md#-mixing--matching-models)) for a specific room or globally
|
||||
|
||||
|
||||
### Supported providers
|
||||
|
||||
### Anthropic
|
||||
|
||||
[Anthropic](https://www.anthropic.com/) is an American AI company founded by former OpenAI engineers and providing powerful language models.
|
||||
|
||||
- 🆔 Identifier: `anthropic`
|
||||
- 🔗 Links: [🏠 Home page](https://www.anthropic.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Anthropic), [👤 Sign up](https://console.anthropic.com/), [📋 Models list](https://docs.anthropic.com/en/docs/about-claude/models)
|
||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
|
||||
- 🗲 Quick start:
|
||||
- create a room-local agent: `!bai agent create-room-local anthropic my-anthropic-agent`
|
||||
- create a global agent: `!bai agent create-global anthropic my-anthropic-agent`
|
||||
|
||||
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/anthropic.yml).
|
||||
|
||||
|
||||
### Groq
|
||||
|
||||
[Groq](https://groq.com/) is an American company developing optimized Language Processing Units (LPU) and offering cloud service which runs various models (built by others) with very high performance.
|
||||
|
||||
- 🆔 Identifier: `groq`
|
||||
- 🔗 Links: [🏠 Home page](https://groq.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/Groq), [👤 Sign up](https://console.groq.com/login), [📋 Models list](https://console.groq.com/docs/models)
|
||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation), [🦻 speech-to-text](./features.md#-speech-to-text)
|
||||
- 🗲 Quick start:
|
||||
- create a room-local agent: `!bai agent create-room-local groq my-groq-agent`
|
||||
- create a global agent: `!bai agent create-global groq my-groq-agent`
|
||||
|
||||
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/groq.yml).
|
||||
|
||||
|
||||
### LocalAI
|
||||
|
||||
[LocalAI](https://localai.io/) is the free, Open Source OpenAI alternative. LocalAI act as a drop-in replacement REST API that’s compatible with OpenAI API specifications for local inferencing. It allows you to run LLMs, generate images, audio (and not only) locally or on-prem with consumer grade hardware, supporting multiple model families and architectures.
|
||||
|
||||
- 🆔 Identifier: `localai`
|
||||
- 🔗 Links: [🏠 Home page](https://localai.io/), [📋 Models list](https://localai.io/gallery.html)
|
||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
||||
- 🗲 Quick start:
|
||||
- create a room-local agent: `!bai agent create-room-local localai my-localai-agent`
|
||||
- create a global agent: `!bai agent create-global localai my-localai-agent`
|
||||
|
||||
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/localai.yml).
|
||||
|
||||
|
||||
### Mistral
|
||||
|
||||
[Mistral AI](https://mistral.ai/) is a research lab based in Europe (France) which produces their own language models.
|
||||
|
||||
- 🆔 Identifier: `mistral`
|
||||
- 🔗 Links: [🏠 Home page](https://mistral.ai/), [🌐 Wiki](https://en.wikipedia.org/wiki/Mistral_AI), [👤 Sign up](https://auth.mistral.ai/ui/registration), [📋 Models list](https://docs.mistral.ai/getting-started/models/)
|
||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
|
||||
- 🗲 Quick start:
|
||||
- create a room-local agent: `!bai agent create-room-local mistral my-mistral-agent`
|
||||
- create a global agent: `!bai agent create-global mistral my-mistral-agent`
|
||||
|
||||
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/mistral.yml).
|
||||
|
||||
|
||||
### Ollama
|
||||
|
||||
[Ollama](https://ollama.com/) lets you run various models in a [self-hosted](https://github.com/ollama/ollama?tab=readme-ov-file#ollama) way. This is more advanced and requires powerful hardware for running some of the better models, but ensures your data stays with you.
|
||||
|
||||
- 🆔 Identifier: `ollama`
|
||||
- 🔗 Links: [🏠 Home page](https://ollama.com/), [📋 Models list](https://ollama.com/library)
|
||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
|
||||
- 🗲 Quick start:
|
||||
- create a room-local agent: `!bai agent create-room-local ollama my-ollama-agent`
|
||||
- create a global agent: `!bai agent create-global ollama my-ollama-agent`
|
||||
|
||||
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/ollama.yml).
|
||||
|
||||
|
||||
### OpenAI
|
||||
|
||||
[OpenAI](https://openai.com/) is an American AI company providing powerful language models.
|
||||
|
||||
Use this provider either with the OpenAI API or with other OpenAI-compatible API services which **fully** adhere to the [OpenAI API spec](https://github.com/openai/openai-openapi/).
|
||||
For services which are not fully compatible with the OpenAI API, consider using the [OpenAI Compatible](#openai-compatible) provider.
|
||||
|
||||
- 🆔 Identifier: `openai`
|
||||
- 🔗 Links: [🏠 Home page](https://openai.com/), [🌐 Wiki](https://en.wikipedia.org/wiki/OpenAI), [👤 Sign up](https://platform.openai.com/signup), [📋 Models list](https://platform.openai.com/docs/models)
|
||||
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-generation), [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
||||
- 🗲 Quick start:
|
||||
- create a room-local agent: `!bai agent create-room-local openai my-openai-agent`
|
||||
- create a global agent: `!bai agent create-global openai my-openai-agent`
|
||||
|
||||
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/openai.yml).
|
||||
|
||||
|
||||
### OpenAI Compatible
|
||||
|
||||
This provider allows you to use OpenAI-compatible API services like [OpenRouter](https://openrouter.ai/), [Together AI](https://www.together.ai/), etc.
|
||||
|
||||
Some of these popular services already have **shortcut** providers (leading to this one behind the scenes) - this make it easier to get started.
|
||||
|
||||
This provider is just as featureful as the [OpenAI](#openai) provider, but is more compatible with services which do not fully adhere to the [OpenAI API spec](https://github.com/openai/openai-openapi/).
|
||||
|
||||
- 🆔 Identifier: `openai-compatible`
|
||||
- 🌟 Capabilities: [🖌️ image-generation](./features.md#️-image-generation), [💬 text-generation](./features.md#-text-generation), [🗣️ text-to-speech](./features.md#️-text-to-speech), [🦻 speech-to-text](./features.md#-speech-to-text)
|
||||
- 🗲 Quick start:
|
||||
- create a room-local agent: `!bai agent create-room-local openai-compatible my-openai-compatible-agent`
|
||||
- create a global agent: `!bai agent create-global openai-compatible my-openai-compatible-agent`
|
||||
|
||||
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/openai-compatible.yml).
|
||||
|
||||
|
||||
### OpenRouter
|
||||
|
||||
[OpenRouter](https://openrouter.ai/) is a unified interface for LLMs. The platform scouts for the lowest prices and best latencies/throughputs across dozens of providers, and lets you choose how to [prioritize](https://openrouter.ai/docs/provider-routing) them.
|
||||
|
||||
- 🆔 Identifier: `openrouter`
|
||||
- 🔗 Links: [🏠 Home page](https://openrouter.ai/), [👤 Sign up](https://openrouter.ai/), [📋 Models list](https://openrouter.ai/models)
|
||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
|
||||
- 🗲 Quick start:
|
||||
- create a room-local agent: `!bai agent create-room-local openrouter my-openrouter-agent`
|
||||
- create a global agent: `!bai agent create-global openrouter my-openrouter-agent`
|
||||
|
||||
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/openrouter.yml).
|
||||
|
||||
|
||||
### Together AI
|
||||
|
||||
[Together AI](https://www.together.ai/) makes it easy to run or [fine-tune](https://docs.together.ai/docs/fine-tuning-overview) leading open source models with only a few lines of code.
|
||||
|
||||
- 🆔 Identifier: `together-ai`
|
||||
- 🔗 Links: [🏠 Home page](https://www.together.ai/), [👤 Sign up](https://api.together.ai/signup), [📋 Models list](https://api.together.xyz/models)
|
||||
- 🌟 Capabilities: [💬 text-generation](./features.md#-text-generation)
|
||||
- 🗲 Quick start:
|
||||
- create a room-local agent: `!bai agent create-room-local together-ai my-together-ai-agent`
|
||||
- create a global agent: `!bai agent create-global together-ai my-together-ai-agent`
|
||||
|
||||
💡 When creating an agent, the bot will show you an up-to-date sample configuration for this provider which looks [like this](../sample-provider-configs/together-ai.yml).
|
||||
8
docs/sample-provider-configs/anthropic.yml
Normal file
@@ -0,0 +1,8 @@
|
||||
base_url: https://api.anthropic.com/v1
|
||||
api_key: YOUR_API_KEY_HERE
|
||||
text_generation:
|
||||
model_id: claude-3-5-sonnet-20240620
|
||||
prompt: You are a brief, but helpful bot.
|
||||
temperature: 1.0
|
||||
max_response_tokens: 8192
|
||||
max_context_tokens: 204800
|
||||
10
docs/sample-provider-configs/groq.yml
Normal file
@@ -0,0 +1,10 @@
|
||||
base_url: https://api.groq.com/openai/v1
|
||||
api_key: YOUR_API_KEY_HERE
|
||||
text_generation:
|
||||
model_id: llama3-70b-8192
|
||||
prompt: You are a brief, but helpful bot.
|
||||
temperature: 1.0
|
||||
max_response_tokens: 4096
|
||||
max_context_tokens: 131072
|
||||
speech_to_text:
|
||||
model_id: whisper-large-v3
|
||||
20
docs/sample-provider-configs/localai.yml
Normal file
@@ -0,0 +1,20 @@
|
||||
base_url: http://my-localai-self-hosted-service:8080/v1
|
||||
api_key: YOUR_API_KEY_HERE
|
||||
text_generation:
|
||||
model_id: gpt-4
|
||||
prompt: You are a brief, but helpful bot.
|
||||
temperature: 1.0
|
||||
max_response_tokens: 4096
|
||||
max_context_tokens: 128000
|
||||
speech_to_text:
|
||||
model_id: whisper-1
|
||||
text_to_speech:
|
||||
model_id: tts-1
|
||||
voice: onyx
|
||||
speed: 1.0
|
||||
response_format: opus
|
||||
image_generation:
|
||||
model_id: stablediffusion
|
||||
style: vivid
|
||||
size: 1024x1024
|
||||
quality: standard
|
||||
8
docs/sample-provider-configs/mistral.yml
Normal file
@@ -0,0 +1,8 @@
|
||||
base_url: https://api.mistral.ai/v1
|
||||
api_key: YOUR_API_KEY_HERE
|
||||
text_generation:
|
||||
model_id: mistral-large-latest
|
||||
prompt: You are a brief, but helpful bot.
|
||||
temperature: 1.0
|
||||
max_response_tokens: 4096
|
||||
max_context_tokens: 128000
|
||||
8
docs/sample-provider-configs/ollama.yml
Normal file
@@ -0,0 +1,8 @@
|
||||
base_url: http://my-ollama-self-hosted-service:11434/v1
|
||||
api_key: YOUR_API_KEY_HERE
|
||||
text_generation:
|
||||
model_id: gemma2:2b
|
||||
prompt: You are a brief, but helpful bot.
|
||||
temperature: 1.0
|
||||
max_response_tokens: 4096
|
||||
max_context_tokens: 128000
|
||||
10
docs/sample-provider-configs/openai-compatible.yml
Normal file
@@ -0,0 +1,10 @@
|
||||
base_url: ''
|
||||
api_key: YOUR_API_KEY_HERE
|
||||
text_generation:
|
||||
model_id: some-model
|
||||
prompt: You are a brief, but helpful bot.
|
||||
temperature: 1.0
|
||||
max_response_tokens: 4096
|
||||
max_context_tokens: 128000
|
||||
speech_to_text:
|
||||
model_id: whisper-1
|
||||
20
docs/sample-provider-configs/openai.yml
Normal file
@@ -0,0 +1,20 @@
|
||||
base_url: https://api.openai.com/v1
|
||||
api_key: YOUR_API_KEY_HERE
|
||||
text_generation:
|
||||
model_id: gpt-4o-2024-08-06
|
||||
prompt: You are a brief, but helpful bot.
|
||||
temperature: 1.0
|
||||
max_response_tokens: 16384
|
||||
max_context_tokens: 128000
|
||||
speech_to_text:
|
||||
model_id: whisper-1
|
||||
text_to_speech:
|
||||
model_id: tts-1-hd
|
||||
voice: onyx
|
||||
speed: 1.0
|
||||
response_format: opus
|
||||
image_generation:
|
||||
model_id: dall-e-3
|
||||
style: vivid
|
||||
size: 1024x1024
|
||||
quality: standard
|
||||
8
docs/sample-provider-configs/openrouter.yml
Normal file
@@ -0,0 +1,8 @@
|
||||
base_url: https://openrouter.ai/api/v1
|
||||
api_key: YOUR_API_KEY_HERE
|
||||
text_generation:
|
||||
model_id: mattshumer/reflection-70b:free
|
||||
prompt: You are a brief, but helpful bot.
|
||||
temperature: 1.0
|
||||
max_response_tokens: 2048
|
||||
max_context_tokens: 8192
|
||||
8
docs/sample-provider-configs/together-ai.yml
Normal file
@@ -0,0 +1,8 @@
|
||||
base_url: https://api.together.xyz/v1
|
||||
api_key: YOUR_API_KEY_HERE
|
||||
text_generation:
|
||||
model_id: meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo
|
||||
prompt: You are a brief, but helpful bot.
|
||||
temperature: 1.0
|
||||
max_response_tokens: 2048
|
||||
max_context_tokens: 8192
|
||||
BIN
docs/screenshots/agent-creation.webp
Normal file
|
After Width: | Height: | Size: 89 KiB |
BIN
docs/screenshots/config-status-handlers.webp
Normal file
|
After Width: | Height: | Size: 65 KiB |
BIN
docs/screenshots/image-generation.webp
Normal file
|
After Width: | Height: | Size: 684 KiB |
BIN
docs/screenshots/introduction-and-general-usage.webp
Normal file
|
After Width: | Height: | Size: 52 KiB |
BIN
docs/screenshots/speech-to-text-default-flow.webp
Normal file
|
After Width: | Height: | Size: 27 KiB |
BIN
docs/screenshots/speech-to-text-language.webp
Normal file
|
After Width: | Height: | Size: 37 KiB |
BIN
docs/screenshots/speech-to-text-transcribe-only-mode.webp
Normal file
|
After Width: | Height: | Size: 13 KiB |
BIN
docs/screenshots/sticker-generation.webp
Normal file
|
After Width: | Height: | Size: 130 KiB |
BIN
docs/screenshots/text-generation-prefix-requirement.webp
Normal file
|
After Width: | Height: | Size: 158 KiB |
BIN
docs/screenshots/text-generation.webp
Normal file
|
After Width: | Height: | Size: 11 KiB |
BIN
docs/screenshots/text-to-speech-only-mode.webp
Normal file
|
After Width: | Height: | Size: 16 KiB |
BIN
docs/screenshots/text-to-speech-seamless-voice-interaction.webp
Normal file
|
After Width: | Height: | Size: 23 KiB |
99
docs/usage.md
Normal file
@@ -0,0 +1,99 @@
|
||||
## 📖 Usage
|
||||
|
||||
This document covers how to use the bot in a room.
|
||||
|
||||
The [🌟 Features](./features.md) page also includes details about how each feature works and can be configured.
|
||||
|
||||
|
||||
### 💬 Text Generation
|
||||
|
||||
This is related to the [💬 Text Generation](./features.md#-text-generation) feature.
|
||||
|
||||
If there's a text-generation handler agent configured, the bot **may** respond to messages sent in the room.
|
||||
|
||||
🖼️ See screenshots of:
|
||||
|
||||
- the [default Text Generation flow](./screenshots/text-generation.webp) for 1:1 rooms
|
||||
- the [Text Generation flow in multi-user rooms](./screenshots/text-generation-prefix-requirement.webp) (where the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting is auto-configured to "required")
|
||||
|
||||
Whether the bot responds depends on:
|
||||
|
||||
- ([🔒 access](./access.md)) whether you're a whitelisted bot [👥 user](./access.md#-users)
|
||||
|
||||
- [🛠️ configuration](./configuration/README.md) whether there's a configured `text-generation` handler agent (or a `catch-all` handler agent). See [Mixing & matching models](./features.md#-mixing--matching-models)
|
||||
|
||||
- (🎨 agent capabilities) whether the configured `text-generation` (or `catch-all`) handler agent actually supports text-generation. The provider may lack support for this feature or it may be disabled in the [🤖 agents](./agents.md) configuration
|
||||
|
||||
- (the [🗟 Prefix Requirement](./configuration/text-generation.md#-prefix-requirement-type) setting) whether a prefix (e.g. `!bai`) is required in front of messages sent to the room. For multi-user rooms, this setting defaults to "required"
|
||||
|
||||
Room messages start a threaded conversation where you can continue back-and-forth communication with the bot.
|
||||
|
||||
Unless you've enabled the [♻️ Context Management](./features.md#️-context-management) feature, all messages will be sent to the agent's API each time. If the context management feature is enabled, older messages may be dropped.
|
||||
|
||||
|
||||
### 🗣️ Text-to-Speech
|
||||
|
||||
This is related to the [🗣️ Text-to-Speech](./features.md#️-text-to-speech) feature.
|
||||
|
||||
If there's a text-to-speech handler agent configured, the bot **may** convert text messages sent to the room to audio (voice).
|
||||
|
||||
See:
|
||||
|
||||
- a [🖼️ screenshot](./screenshots/text-to-speech-only-mode.webp) of the bot's [Text-to-Speech-only](./features.md#text-to-speech-only-mode) mode
|
||||
|
||||
- a [🖼️ screenshot](./screenshots/text-to-speech-seamless-voice-interaction.webp) of the bot's [Seamless voice interaction](./features.md#seamless-voice-interaction) mode
|
||||
|
||||
By default, the bot:
|
||||
|
||||
- will offer tex-to-speech for its own messages which are a response to voice message from your, as part of the [Seamless voice interaction](./features.md#seamless-voice-interaction) feature. This can be adjusted via the [🗣️ Text-to-Speech / 🪄 Bot Messages Flow Type](./configuration/text-to-speech.md#-bot-messages-flow-type) setting.
|
||||
|
||||
- does not turn your own text messages to audio (voice). If you'd like for the bot to operate in such a mode, use the [🗣️ Text-to-Speech / 🪄 User Messages Flow Type](./configuration/text-to-speech.md#-user-messages-flow-type) setting (see [Text-to-Speech-only mode](./features.md#text-to-speech-only-mode)).
|
||||
|
||||
|
||||
### 🦻 Speech-to-Text
|
||||
|
||||
This is related to the [🦻 Speech-to-Text](./features.md#-speech-to-text) feature.
|
||||
|
||||
If there's a speech-to-text handler agent configured, the bot **may** transcribe voice messages sent to the room to text.
|
||||
|
||||
See a [🖼️ Screenshot of the default flow for Speech-to-Text and Text-Generation](./screenshots/speech-to-text-default-flow.webp).
|
||||
|
||||
The speech-to-text feature triggers automatically by default, but can be adjusted via the [🦻 Speech-to-Text / 🪄 Flow Type](./features.md#-speech-to-text-flow-type) setting.
|
||||
|
||||
If all your messages are in the same language, you can improve accuracy & latency by configuring the language (see [🦻 Speech-to-Text / 🔤 Language](./configuration/speech-to-text.md#-language)).
|
||||
|
||||
|
||||
### 🖌️ Image Generation
|
||||
|
||||
This is related to the [🖌️ Image Generation](./features.md#️-image-generation) feature.
|
||||
|
||||
This feature is not configurable at the moment. The configuration (size, quality, style) specified at the [🤖 agent](./agents.md) level will be used.
|
||||
|
||||
|
||||
#### Generating images
|
||||
|
||||
Simply send a command like `!bai image A beautiful sunset over the ocean` and the bot will start a threaded conversation and post an image based on your prompt.
|
||||
|
||||
See a [🖼️ Screenshot of the Image Generation feature](./screenshots/image-generation.webp).
|
||||
|
||||
You can then, respond in the same message thread with:
|
||||
|
||||
- more messages, to add more criteria to your prompt.
|
||||
- a message saying `again`, to generate one more image with the current prompt.
|
||||
|
||||
|
||||
#### Generating stickers
|
||||
|
||||
A variation of [generating images](#generating-images) is to generate "sticker images".
|
||||
|
||||
See a [🖼️ Screenshot of the Sticker Generation feature](./screenshots/sticker-generation.webp).
|
||||
|
||||
To generate a sticker, send a command like `!bai sticker A huge ramen bowl with lots of chashu and a mountain of beansprouts on top`.
|
||||
|
||||
The difference from [generating images](#generating-images) is that the bot will:
|
||||
|
||||
- generate a smaller-resolution image (currently hardcoded to `256x256`) - smaller/quicker, but still good enough for a sticker
|
||||
- potentially switch to a different (cheaper or otherwise more suitable) model, if available
|
||||
- post the image directly to the room (as a reply to your message), without starting a threaded conversation
|
||||
|
||||
Some models (like [OpenAI](./providers.md#openai)'s [Dall-E-3](https://openai.com/index/dall-e-3/)) can only generate larger images (`1024x1024`, etc., for a higher charge), so we switching to a smaller/cheaper model (like [Dall-E-2](https://openai.com/index/dall-e-2/)) is a way to generate a sticker cheaply.
|
||||