Multi-Agent Fluid Conversation Manager is a platform for creating conversational agents trained on your own content, integrated with multiple LLMs and written in
Elixir, with Phoenix and OTP. It brings together my back-end and AI experience, and it is a simpler version of the chatbot management and deployment system I built for Elife.
The architecture is
event-driven: everything that happens in the system (an answered message, a model call, training progress) becomes an event on Phoenix PubSub. Persistence, token accounting, webhooks and a live WebSocket channel are independent subscribers, and a failure in one of them does not affect the others.
Each conversation runs in its own
supervised process: messages from the same user are handled in order, different users in parallel, the history stays in memory, and inactivity notices are the process's own timers, with no cron. If the server goes down mid-training, the coordinator rebuilds its queue from the database when it comes back up.
Answering and training are
composable pipelines, declared as a list of steps. Each step can be conditional, retried, time-boxed in a supervised task or run in parallel with others (language detection and semantic search run at the same time), and steps can be replaced, inserted or removed at runtime or through configuration.
Models are
swappable: each one is addressed as provider:model (OpenAI, Anthropic, Gemini or local models through Ollama), and switching an agent's model is a single PATCH, even mid-conversation. Each provider can have its own rate limiting, with retries that respect the API's limits.
RAG runs semantic search over embeddings in
PostgreSQL with pgvector, and agents can call
tools, local ones or from any MCP server (over stdio or HTTP). Conversations can be handed off to a human agent when integrated with an external system, and the test suite runs without a database or network, using in-memory storage and a fake model provider.
Technologies:
Elixir, Phoenix, PostgreSQL + pgvector, Ollama and MCP.
Real recordings with local models (qwen2.5 and embeddinggemma through Ollama, on CPU); waits for the model were shortened.