Skip to main content

Local inference

Run models entirely on your own hardware with the built-in llama.cpp engine. Your data never leaves the machine. Managed from Settings → Inference Runtime.

Experimental

The bundled inference runtime is marked experimental in the app.

Setting up the engine​

Settings → Inference Runtime → Install / Download fetches a prebuilt llama-server binary matching your platform and hardware. You choose:

SettingOptions
Versionlatest or a pinned llama.cpp release.
VariantmacOS: Auto / Metal (default). Windows: CPU, CUDA 12.4, CUDA 13.1, Vulkan. Linux: CPU, Vulkan, ROCm.

The running-instance panel shows the build number, binary path, and lets you re-download.

Running the engine​

ControlEffect
Start / Stop / RestartManage the llama-server process. Status: setting-up → starting → running (or failed / stopped).
PortWhich port the inference server binds.
Extra argsAdditional llama-server command-line flags.
Check for updates / UpdateUpdate the llama.cpp binary.
UninstallRemove the engine and its binary.

Engine logs and a PTY view are available in-app. Once running, the local engine appears as a model provider inside the connected Hrida.ai UI.

Models (Hugging Face)​

Settings → Models manages a local GGUF model cache at <data>/models/<repo>/<file>.

ActionDetail
Search Hugging FaceSearch repos by name; optionally pass a HF token for gated/private repos.
Browse repo filesList the GGUF files in a repo with sizes.
DownloadStreams the file with a live progress bar; cancellable mid-download.
Downloaded modelsLists cached models with size and date; delete individually.
Models directoryShown in the UI; click to open it in the file manager.

Hardware guidance​

A 7B model needs roughly 8 GB RAM, 13B roughly 16 GB. On lighter machines, connect to a remote server and let it do inference instead.

Hrida.ai is proprietary software of Zlabs Innovation. See the license for terms. © 2026 Zlabs Innovation.