Skip to content

Local AI

inbuxa can use a language model in two places:

  • The spam filter asks it for an opinion of each incoming message. The opinion is one signal among many, and its weight is capped.
  • Explain, in the console, asks it to put a delivery failure, a spam verdict, a log line or a setting into plain words, for an administrator.

Both are off on a new install, and nothing is sent anywhere until you set a model up. The model is one you run, on the mail server or on your own network. Nothing in inbuxa points at a hosted service, and the console warns you if the address you give isn't on your own network.

What each use sends to the model, and what it never sends, is on the Security page.

Running a model

Anything that serves an OpenAI-style chat completions API works. The two tested are llama.cpp's server and Ollama. The console suggests their usual addresses:

Server Address
llama.cpp http://127.0.0.1:8080/v1/chat/completions
Ollama http://127.0.0.1:11434/v1/chat/completions

The recommended model is Qwen3 4B Instruct 2507 (Apache-2.0), quantized to Q4_K_M, which is about 2.5 GB on disk and takes about 4 GB of memory running. No GPU is needed. What it costs in time, measured on CPU cores:

Cores for the model Typical message Slowest 1 in 20
4 6.3 s 11.0 s
2 10.8 s 20.4 s

These are the numbers for the smallest setup that works well: a 4B model on a few CPU cores, with no GPU. Treat them as a floor, not a forecast. More or faster cores bring them down, and a GPU brings them down a long way, so measure on your own hardware. On a small virtual server with slower cores, expect them to be higher. Four cores is the recommended minimum. On two, about one message in twelve runs past the 20-second limit and goes unclassified, which is safe but loses the signal. The default limits below are set for a small CPU-only machine.

The model's time is spent while the message is already being received, and never holds it up. What it costs is CPU: give the model its own cores, or it competes with the mail server for them. See Sizing.

Keep the model on the same machine

Point the address at loopback. Message text then never leaves the mail server. The model's port only has to answer the mail server, so bind it to loopback too.

On a cluster

The setting is one for the whole cluster, and so is the model's address: there's no per-node address. Loopback is what makes that work. With 127.0.0.1, each node asks the model on its own machine, so run one model on every node.

A node with no model, or whose model is down, just adds no AI signal. Its mail is filtered and delivered as normal. So does a message that arrives when Requests in flight at once are all taken: it isn't queued for the model, it's scored without it. The limits on calls in flight and the pause after failures are counted per node, so a cluster of three nodes can have three times Requests in flight at once running.

Turning it on

In the console, go to Settings › Spam filter › Local AI. It says Off and "The spam filter doesn't use a language model. Nothing is sent anywhere until you set one up."

The console's Local AI page, showing Off, with the Limits card below

  1. Choose Set up local AI spam filtering, then Guided.
  2. Step 1 says what will happen: only a message's subject and text are sent, with no addresses, headers or attachments. The opinion adds at most 2 points by default. A slow or missing model never holds mail up.
  3. Step 2 asks for the Model address, the Model name the server knows it by, and a Name in inbuxa. An address that isn't loopback or a private network gets a warning, not a refusal.
  4. Step 3 shows the Instructions for the model. The default was measured against real mail; change it only with a reason. The server adds its own framing around it, so a message is treated as data, never as instructions.
  5. Choose Turn on.

Step 2 of the guided setup: the model address, model name and name in inbuxa

The page then says On, and which model it asks at which address.

The Local AI page with the classifier on, the model it asks, and the Limits and Explanations cards

Running the setup again reuses a model with the same name rather than making a second one.

Without the guide

Manual goes to the two forms the guide fills in:

  • Settings › System › AI models: the model, with its address, name and timeout.
  • Settings › Spam filter › LLM classifier: the classifier, which is Disabled until set to Enabled with a model and a prompt. The categories and confidence levels it accepts are set here too.

Turning it off

Turn off on the Local AI page switches the classifier off. The model stays configured, so turning it back on is one step. Explain keeps working while a model exists; see below.

If you stop the model instead, and leave the classifier on, every message fails the model call quickly and is filtered without it. Turning the classifier off stops it trying.

Limits

The Limits card on the Local AI page keeps the model's influence small and its load bounded. Changes take effect with the next message. A value equal to the default is stored as "use the default", so it follows the default if that changes in a later release.

Limit Default What it does
Most the model can add to a score 2 points The cap on the model's word alone
Most the model can take off a score 1 point The cap on the model vouching for a message
Longest the spam filter waits for the model 20 s After that, the message is filtered without it
Requests in flight at once 4 Per node, shared by the spam filter and Explain. Explain only takes one while at least one stays free for mail
Most message text sent 2048 bytes More measured no better, and made small machines slow
Pause after repeated failures 60 s A dead model stops being asked for a while
Calls per account per hour from its own Sieve scripts 60 See Sieve

Reset to defaults puts all of them back.

The limits are a server-level setting. A tenant administrator can't see or change them.

What people see

When the model gave an opinion, the webmail shows it:

  • in a message's details, beside the spam filter's own working;
  • in a banner above a message that is in Junk, headed Language model's opinion.

Both show the model's category and confidence, and its short reason, and say "One of several signals the spam filter weighed". The reason is the model's own words, shown as plain text. There's no setting for it: people see it when the model gave one.

A promotional message in Junk Mail, with the language model's opinion above it: Unsolicited, High, and its reason

The opinion comes from the message's X-Spam-LLM header, which the server adds when the model's answer produced a tag, such as LLM_UNSOLICITED_HIGH. Those tags are scored like any other spam tag, under Settings › Spam filter › Scores.

Explain

Explain puts something in the console into plain words. It appears:

  • beside each recipient of a queued message that failed, temporarily or permanently;
  • under the result of classifying a message;
  • in each row of Logs, and beside each event of a trace, stored or live;
  • beside a field's help, for any setting that isn't a secret.

Explain open beside the DNS Management setting: the answer prepared for this release, and a reminder that it can be wrong

The answer opens in a side panel and appears as the model writes it, so the first words show within a few seconds. Answers are kept short: three or four sentences. It is the model's reading, and it can be wrong. Check it against the data it explains.

Not every answer needs the model:

  • Asked before. Each node remembers the answers it gave, in memory, for up to 24 hours: up to 1,000 of them. The same question, from any server administrator, is answered at once without asking the model again. Remembered answers are never written to disk, and a restart forgets them.
  • Prepared with the release. Each release ships explanations of the settings at their default values, written beforehand with the recommended model. A setting you haven't changed is explained at once. Once you change it, or a release changes its description, the model is asked instead.

The line under the answer says which: the model and node that answered just now, an answer given earlier on that node, or the release it was prepared for.

The console never sends text you type. It names what to explain, and the server builds the question from the stored data and the setting's documentation. Message bodies, subjects, attachment names, passwords and secrets are never part of it, and a secret setting can't be explained at all.

Who can use it

Explain needs the permission Ask the local model to explain things in the console (sysAiExplain). Only the superuser role has it by default. A tenant administrator is refused even with it.

It shows only when a model exists. That can be a model with the spam classifier off: Explain doesn't need the classifier.

Its settings

The Explanations card on the Local AI page:

Setting Default What it does
Offer explanations on Off hides Explain everywhere
Model Same as the spam filter Or a different model under Settings › System › AI models
Per administrator, per hour 30 Explanations each administrator can ask for
Longest wait 45 s After that, the panel says so

Mail comes first: an explanation only takes a request slot while at least one stays free for the spam filter. When the model is busy with incoming mail, paused after failures, or missing, the panel says which.

Sieve

A filter can ask the model a question itself, with llm_prompt(model, prompt, temperature). It returns the model's answer for the script to test.

A person's own scripts can call it only with the permission Interact with AI models (interactAi). The default user, tenant administrator and superuser roles have it. Calls are limited by Calls per account per hour from its own Sieve scripts. Whatever a script sends is what the model sees, so a script that sends message text sends it to your model.