Hrida.ai Essentials
Hrida.ai Essentials
You have installed Hrida.ai, connected a provider, and had your first conversation. This page covers the six things that turn a basic chat UI into a setup that works well day-to-day. None of them are required, but most users end up reaching for all of them within the first week.
Work through in order, or jump to the section you need:
If you are setting up Hrida.ai for multiple users, also read the Deployment & Scaling guide. It covers infrastructure decisions (PostgreSQL, Redis, external vector databases, shared storage) that are separate from the feature-level essentials on this page. The two guides are additive — work through the essentials here for day-to-day usage, and the scaling guide for multi-user infrastructure.
Without tools, an LLM can only generate text. With tools, it can look things up, run code, and take actions on your behalf. You attach a Tool to a chat (or to a model), and the model decides during its response when to call it and with what arguments. Hrida.ai executes the call, feeds the result back, and the model continues with that new information.
Tool calling is what turns a chat model into an agent. Getting this set up unlocks most of the advanced features covered in the rest of this page.
There are two things worth understanding early on.
Hrida.ai connects a model to its tools using Native Mode: the model receives tool definitions as a structured part of its API request and returns structured tool calls in its response. It's faster, more reliable, preserves the KV cache, and is required for all of Hrida.ai's built-in system tools (Memory, Notes, Knowledge, Web Search, Image Gen, Code Interpreter). All new tool-calling features are built for Native Mode.
The model receives tool definitions as a structured part of its API request and returns structured tool calls in its response. Faster, more reliable, preserves KV cache, and required for all of Hrida.ai's built-in system tools (Memory, Notes, Knowledge, Web Search, Image Gen, Code Interpreter). All new tool-calling features are built for Native Mode.
Every major current model supports it (OpenAI, Anthropic, Gemini, Llama 3.1+, Qwen 2.5+, DeepSeek, GLM, and others). You can set it once for your entire instance:
Admin Panel > Settings > Models → click the Settings button at the top right of the models list → set Function Calling = Native → save. Every current and future model inherits the setting.
- Per model: Admin Panel > Settings > Models > [your model] > Model Parameters > Function Calling = Native
- Per chat: In the chat's Chat Controls (right sidebar)
If a specific model has trouble with Native Mode, the fix is to pick a stronger model for tool-calling workloads — not to fall back to the legacy, unsupported Default option. See the Tools reference if you're migrating an older deployment off it.
Many of the tools people look for are already built into Hrida.ai and just need to be turned on: web search, code run, image generation, memory, and knowledge-base retrieval are all available without installing any plugins. Once enabled, these appear automatically as system tools when using Native Mode.
Most of these need a small amount of setup (choosing a provider, adding an API key, or enabling a toggle). Setup guides for the most popular ones:
For anything not built in, the Zlabs Innovation is worth browsing. A few categories to give a sense of what is available:
- Observability / cost tracking: Langfuse, OpenLit, Portkey. Log every chat turn, token usage, and latency to your own stack.
- Smart-home / automation: Home Assistant tools that let the model control devices, routines, and scenes.
- Research: arXiv, PubMed, Semantic Scholar, Wolfram Alpha. Structured results with real citations.
- Databases / APIs: read-only SQL against your own database, or calls to your internal API.
- Domain-specific: weather, stocks, crypto, shipping tracking, recipes, and many more.
Tools appear in the + menu in the chat input. The model only sees the tools you have enabled for that conversation.
The Zlabs Innovation hosts thousands of third-party plugins. Before writing anything yourself, browse what is already there: sort by popularity, filter by category, and skim a few pages. Most of the time, someone has already built what you need (or something close enough to fork).
Hrida.ai ships with a lot out of the box, but its real power is that it is designed to be extended. Many of the advanced capabilities people show in demos (auto-translation, token/cost tracking, custom post-processing, niche provider integrations) are plugins built on top of the platform. Understanding the plugin landscape is the single biggest unlock for a new user.
There are two plugin families: Tools and Functions.
Tools give the model abilities it can call during a response:
| Source | What it does | Examples |
|---|---|---|
| Built-in | System tools that ship with Hrida.ai. Enable in the admin panel, no install needed. | Web Search, Code Interpreter, Image Generation, Memory, Notes, Knowledge retrieval |
| Custom — Tool | Code you write yourself or install from the Hrida.ai site. Manage in Workspace > Tools. | Langfuse / OpenLit observability, Home Assistant, arXiv / PubMed lookups, Wolfram Alpha, Jira / Linear, SQL queries |
| Custom — Tool server | External services connected via MCP or OpenAPI. Configure in Admin Panel > Settings > Tools. | Your own microservices, third-party APIs, existing MCP servers |
Functions run at the platform level and modify how Hrida.ai itself behaves. There are three types:
| Type | What it does | Examples |
|---|---|---|
| Pipes | Add a new "model" to the model picker, backed by custom code | Model-routing (cheap vs. expensive based on prompt), multi-step agent loops, custom LLM backends |
| Filters | Modify every request and/or response as it passes through, automatically on every chat turn | Context trimming, PII scrubbing, token / cost counting, Langfuse tracing, response reformatting |
| Actions | Add a button under each message that runs custom code when the user clicks it | "Regenerate follow-ups", "Translate reply", "Pin message", "Save to Knowledge" |
Both Tools and Functions can be browsed and installed from the Zlabs Innovation, which hosts thousands of third-party plugins. You can also write your own from scratch in the admin panel.
The Zlabs Innovation hosts the one-click catalog for both Tools and Functions. Pick one, click "Get", paste it into the admin panel, enable it, and configure its valves (the plugin's settings).
Whenever you think "it would be nice if Hrida.ai did X," it almost certainly already does via a plugin. There are thousands of plugins already written, and the one you need is usually already there. Even if nothing matches exactly, the closest hit is usually only about 20 lines off from what you want and you can fork it from the admin panel.
By default, background tasks (titles, tags, autocomplete) use your main chat model. Setting a dedicated task model is the easiest way to improve speed and reduce unnecessary API costs.
Every time Hrida.ai needs a short piece of "thinking" for a UI feature (writing a chat title for the sidebar, generating tags, suggesting follow-up questions, powering the autocomplete in the prompt box) it calls a Task Model. By default that task model is whatever main model you are currently chatting with, which means:
- Your expensive flagship model gets invoked every time you open a new chat just to write "Groceries list."
- On a slow local model, every keystroke feels laggy because autocomplete is waiting on a 30B-parameter model.
- A reasoning model (o1, r1, Claude with extended thinking) spends five seconds thinking before producing a three-word title.
These run in the background, so they are easy to overlook. A dedicated task model is a small change that makes a noticeable difference.
Fix: In Admin Panel > Settings > Interface, set a dedicated Task Model. There are two fields, because the right choice depends on what your main chat model is:
Task Model (External): Set to a fast, cheap, non-reasoning cloud model like gpt-5-nano, gemini-2.5-flash-lite, or llama-3.1-8b-instant.
Task Model (Local): Set to a tiny local model like qwen3:1b, gemma3:1b, or llama3.2:3b.
The main chat experience does not change. The background chores just stop dragging.
While you are in the Interface settings, you can also disable these chores entirely if you are on a low-spec machine or simply do not want them. Each one has both an admin toggle in the same page and an environment variable:
| Chore | Admin toggle (Settings > Interface) | Env var |
|---|---|---|
| Autocomplete (fires on every keystroke) | Autocomplete Generation | ENABLE_AUTOCOMPLETE_GENERATION=False |
| Follow-up suggestions | Follow-up Generation | ENABLE_FOLLOW_UP_GENERATION=False |
| Chat title generation | Title Generation | ENABLE_TITLE_GENERATION=False |
| Tag generation | Tags Generation | ENABLE_TAGS_GENERATION=False |
Autocomplete is the single biggest "make it snappy" toggle on weak hardware. It fires on every keystroke, so a slow task model turns the whole prompt box into molasses. Disable it first if the UI feels sluggish.
After enough back-and-forth you will eventually see:
The prompt is too long: 207601, model maximum context length: 202751
This error comes from your model provider, not from Hrida.ai. Every time you send a message, the entire conversation (system prompt, all previous turns, attached files, tool call results, and your new message) is sent as the "prompt." When the sum exceeds the model's context window, the provider rejects the request.
Hrida.ai intentionally does not ship a built-in trimmer, because:
- Every model uses a different tokenizer (GPT, Claude, Gemini, GLM, Llama all differ).
- Every model has a different context window (8k to 1M+).
- Every deployment wants a different policy (trim by tokens, by turns, by message count, drop attachments first, summarize older messages, etc.).
There is no single correct answer. The supported approach is to install a filter Function that trims the conversation on your terms.
filters for most common policies already exist and can be installed with one click. If none fits, the code is short enough to copy and adapt. See the full guide including a minimal "newest N turns" filter: Troubleshooting: Context Window / Prompt Too Long.
RAG (Retrieval-Augmented Generation) is the feature that lets you say "Here's a 400-page PDF, answer my questions about it" without the model having to read the whole thing every turn. Hrida.ai splits your documents into chunks, embeds them as vectors, stores them in a vector database, and at chat time retrieves just the relevant pieces to pass to the model.
Two ways to use it, in order of simplicity:
# shortcut in the input), or bind it to a model in Workspace > Models so that model always has it available.The defaults are reasonable for getting started. When you outgrow them, there are three knobs that matter most:
- Embedding engine. The default (SentenceTransformers
all-MiniLM-L6-v2) runs locally on CPU and consumes roughly 500 MB of RAM per worker. For any multi-user deployment, point at an external embeddings API (OpenAI, or Ollama withnomic-embed-text) viaRAG_EMBEDDING_ENGINE. - Content extraction engine. The default uses
pypdf, which leaks memory during heavy ingestion. For anything beyond casual use, switch to Tika or Docling viaCONTENT_EXTRACTION_ENGINE. - Vector database. The default ChromaDB (local SQLite-backed) does not tolerate multi-worker deployments. At scale, switch to PGVector — it is the only vector database officially supported and maintained by the Hrida.ai team. Milvus, Qdrant, and MariaDB Vector are also available but are community-maintained: they may break on upgrades and fixes depend on community contributions. See the env-configuration reference for setup and users disclaimers on each provider.
None of these matter for "a single user with a handful of PDFs." All of them start mattering the moment you have 100 documents or 10 concurrent users.
If you just want RAG to work well out of the box, these settings are a solid general-purpose starting point. They are not fine-tuned for every use case, but they will produce noticeably better results than the defaults for most document types.
Set these in Admin Panel > Settings > Documents:
| Setting | Default | Recommended value | Why |
|---|---|---|---|
| Text Splitter | character | token | Token-based splitting produces more consistent chunk sizes across document types |
| Markdown Header Splitting | On | On | Respects document structure by splitting at headings, keeping sections coherent |
| Chunk Size | 1000 | 2000 | Larger chunks preserve more surrounding context per retrieval hit |
| Chunk Overlap | 100 | 200 | More overlap means less chance of cutting a key sentence in half |
| Top K | 3 | 15 | Retrieves more candidate chunks, giving the model a wider pool of relevant context. If you are working with local models that have constrained context sizes, lower this to 5 to avoid filling the context window with retrieved chunks |
| Embedding Model | all-MiniLM-L6-v2 (local CPU) | External (OpenAI or Ollama) | The default works for a single user but consumes ~500 MB RAM per worker. For any multi-user setup, use an external embedding API instead |
The default SentenceTransformers model runs locally on CPU and is fine for a single user getting started. For anything beyond that, point at an external embeddings API: set RAG_EMBEDDING_ENGINE=openai with an OpenAI API key, or RAG_EMBEDDING_ENGINE=ollama with any Ollama embedding model (e.g., nomic-embed-text). This offloads the work and frees significant RAM.
If "run Python" is too restrictive and you want the model to actually work on your machine (clone repos, install packages, run test suites, spin up a local preview of a website, iterate on a data report against a real CSV), that is what Hrida Terminal is for. It connects a real shell (sandboxed in a Docker container by default, or bare-metal if you want) as a tool the model can call the same way it calls any other tool. In-chat file browser, live web previews, and skill definitions are included.
This is the biggest "aha" feature once you get past basic chat. It turns Hrida.ai from a chat UI into a place where the model actually builds things for you. If Native Mode is on and you have given the model a capable terminal, ask it to build you a small app or run an analysis on a folder of files and watch it go.
You do not need all of the above on day one. A reasonable order for a new install:
Everything else (enterprise SSO, multi-replica HA, Redis scaling, observability) is in Advanced Topics and Troubleshooting when and if you need it.
When something goes wrong, start here:
| Having problems with… | Read this |
|---|---|
| Connection refused, 401 errors, CORS failures, WebSocket disconnects | Connection Errors |
| "Prompt is too long" or context window exceeded | Context Window / Prompt Too Long |
| RAG not returning relevant results, uploads failing, knowledge base issues | RAG Troubleshooting |
| Web search not working or returning poor results | Web Search Troubleshooting |
| Image generation errors or provider setup | Image Generation Troubleshooting |
| Speech-to-text, text-to-speech, or audio playback | Audio Troubleshooting |
| SSO, OAuth, or LDAP login issues | SSO & OAuth Troubleshooting |
| High memory usage, slow responses, or worker crashes | Performance & RAM · Scaling Guide |
| Login loops, config drift, or database locks in multi-replica setups | Scaling & HA Troubleshooting · Scaling Guide |
| Locked out of admin account | Reset Admin Password |
| TLS certificate errors with custom/internal CAs | Custom CA Store |
| Alembic migration errors or manual schema fixes | Database Migration |
This page is the condensed version. The full docs go much deeper. If you did not find what you needed: