Skip to content

Sizing

This page helps you choose the shape of an install before you build it: how many machines, how big, what faces the internet, and what doesn't.

These are estimates

The numbers here are extrapolated from a small set of measurements, listed at the end, not from every kind of hardware and mail. Your cores, disks and mail will differ. Use them to pick a starting size, then watch the dashboards and adjust.

What actually drives the size

In rough order of how often each one is the limit:

  1. Local AI, if you turn it on. The model costs far more CPU per message than everything else the server does. When it's on, it usually sets the size. See below.
  2. Disk. Mail kept is mail stored. Plan for the mailboxes you expect in a few years, not the ones you have.
  3. Staying up when a machine fails. That's what a cluster is for. It is rarely needed for capacity.
  4. Memory for people connected at once. Every open webmail tab keeps a push connection to the server, and every mail app on a phone or desktop keeps an IMAP session.
  5. Mail volume. The server itself handles far more mail per core than most installs ever see.

Estimates

The mail server

Resource Estimate From
CPU for mail Roughly 100–250 messages a second per core, received, spam-filtered and delivered 250 a second per core measured on fast cores; halved for typical virtual servers
Memory, idle 150–250 MB Measured on running servers
Memory per open webmail tab About 0.25 MB 1,000 JMAP push connections measured
Memory per mail app's IMAP session About 0.3 MB, or nothing with legacy protocols off 1,000 sessions measured
CPU for reading mail Roughly 1,500 "open a folder, list it, open a message" a second per core, over JMAP or IMAP Measured on fast cores
Memory under a heavy burst of incoming mail Up to about 1 GB more Measured at 900 messages a second
Disk per message Its size plus 5–50%: about 50–70 KB for a typical 48 KB message RocksDB and PostgreSQL measured

To turn that into a plan:

  • Mail volume. 1,000 people receiving 100 messages a day each is 100,000 a day: about 1 a second on average, and perhaps 10 a second at a busy moment. That is a small fraction of one core. Volume starts to matter in the tens of millions of messages a day.
  • Connections. Count open webmail tabs and mail apps, not people. 2,000 people with the webmail open is about 2,000 push connections, or about 0.5 GB of memory. If each also has a mail app on their phone, add 2,000 IMAP sessions, about 0.6 GB. With legacy protocols off, there are none of those.
  • Reading. Even a busy person opens a folder or a message every few seconds. A single core keeps up with thousands of people doing that at once, so reading mail is rarely what sets the size.
  • Disk. People × their mailbox size, plus a quarter for indexes and headroom. 500 people at 5 GB each is about 3 TB. With a cluster's blob store keeping three copies, that's three times as much raw disk.

The webmail and the console

Memory Grows with
Webmail 30–40 MB each People reading mail at the same time, not mail stored
Console 10–15 MB Almost nothing: a few administrators

Neither needs a machine of its own for capacity. They're separated for security, not size.

Local AI

The model reads each incoming message's text before giving its opinion. On CPU, that's the expensive part.

Measured Estimate
CPU per message About 10 s on 4 cores: 40–60 core-seconds
Messages an hour, 4 cores for the model About 250–350
Memory About 4 GB per model

So, with a CPU-only model:

Incoming messages a day Cores for the model, estimate
Up to about 5,000 4
About 10,000 8
About 50,000 30–40, or a GPU

These numbers are for the smallest setup that works well: a small model (4B), on a few CPU cores, with no GPU. They are a floor, not a forecast. Faster cores bring them down, and a GPU brings them down a long way, so measure on your own hardware before you size for it.

Two things make this safer than it looks:

  • Too little capacity never delays mail. A message that finds the model busy is scored without it. What you lose is the AI signal on those messages, not mail.
  • Give the model its own cores. Limit the model service to a set number of cores, as the packaged service does, so it can't starve the mail server on the same machine.

On a cluster, run one model per node, next to each node's mail server. Each node classifies the mail it receives.

Choosing a shape

One machine

Everything on one server: mail server, webmail, console, a reverse proxy, and the model if you use it. Storage is the built-in RocksDB.

Fits: up to a few hundred people, where a few hours of downtime to restore from backup is acceptable.

Without AI With AI
CPU 2 cores 6 cores, 4 of them for the model
Memory 4 GB 8–10 GB
Disk Mailboxes plus a quarter, on SSD The same, plus 3 GB for the model

Two machines

The mail server on one, the webmail, console and proxy on the other. This buys no capacity: it keeps the web front ends off the machine that holds the mail. See Where each service goes.

Fits: the same sizes as one machine, when you'd rather a bug in a web front end couldn't reach the mail.

A cluster

Three nodes, each running the mail server, with shared storage: PostgreSQL for the data store, an S3-compatible store for mail bodies, and a coordinator such as NATS. A webmail on each node, or on its own machines behind a load balancer.

Fits: anyone who needs mail to keep flowing when a machine goes down, for maintenance or failure. That's the reason to cluster. One node already handles more mail than most organizations send.

Why three: PostgreSQL's failover, the coordinator, and a replicated blob store all need a majority to agree on who is in charge. Two machines can't tell a failed partner from a broken link between them; three can.

Per node, estimate:

Without AI With AI
CPU 4 cores 8 cores or more, 4+ for the model
Memory 16 GB, much of it for PostgreSQL 24 GB
Disk Each node's share of mail, times the blob store's copies, plus the database The same

inbuxa's own service runs this shape: three nodes of 6 to 8 cores and 64 GB, shared with other services, each with the mail server, PostgreSQL (one leader, with standbys), NATS, an S3-compatible blob store and its own model, over a private network between them. That is well above what its mail needs, which is part of how the estimates on this page were checked.

More nodes

Add nodes when a single node's CPU is the limit, which in practice means AI, or when you want mail received in more places. Use roles to decide what each node does, and read replicas and sharded stores when the database becomes the limit.

What faces the internet

Whatever the shape, put these where they belong:

Reachable from What
The internet Port 25 on the mail server, which receives mail and has to answer anyone. Port 443 on the reverse proxy in front of the webmail
The internet, only if you use them 465 or 587 for sending from mail apps; 993 for IMAP; 995 for POP3; 4190 for ManageSieve. Turn off what nobody uses; see Protocols
Your own network or VPN The console. Administrators only
The machines themselves, never the internet PostgreSQL, the coordinator, the blob store, the model's port, Prometheus scraping, the server's setup and recovery ports
Between cluster nodes, privately The storage and coordinator traffic above. Run it over a private network or an encrypted tunnel between the nodes, not across the open internet

The server's API, which the webmail and console use, needs to reach them, and only them; see Security.

One sender, many messages

By default, the server accepts at most 5 messages a second from any one IP address, and 25 an hour from one sender domain to one recipient. That protects you from floods. It also slows down your own systems if they send in bulk from one address, such as an application server's notifications. Raise the limits for those addresses under Settings › Mail flow › Incoming rate limits, or with the limits guide, rather than turning them off.

Containers, bare metal, or both

Every part runs either way:

  • All containers. The server, console and webmail are published as images for linux/amd64 and linux/arm64. Simplest to set up and to update; see Installing.
  • Bare metal. The server is a single binary, and runs under systemd like any other daemon. Worth it when you want the mail server outside a container runtime, or in a network namespace of its own.
  • A mix. The mail server and the model on the host, the webmail and console in containers. That's what inbuxa's own service does: the part that holds the mail runs directly on the machine, and the web front ends, which are updated most often, run in containers.

For PostgreSQL, the coordinator and the blob store, use whatever you already know how to run and back up. The server doesn't mind where they are, as long as every node can reach them.

The measurements

Everything above was estimated from these. All on 2026-09-26, with inbuxa 2026.9.25.1, except the JMAP rows, which used 2026.9.26.

What Where Result
Incoming mail, 48 KB messages, one per connection, spam filter on 1, 2 and 4 cores (container CPU limits) of an AMD Ryzen AI 7 350, RocksDB on NVMe 247, 527 and 840–940 messages a second, every one delivered
IMAP sessions held open 4 cores, same machine 992 sessions: about 294 MB more memory
IMAP reading: open a folder, list 50 messages, open one 4 cores, 32 and 128 clients 1,400–1,600 a second, using about 2.5 cores
JMAP push connections held open, as the webmail keeps one per tab 4 cores, same machine, 2026.9.26 992 connections: about 247 MB more memory
JMAP reading: the same folder, list and message, as two JMAP requests 4 cores, 32 clients, 2026.9.26 About 3,300 a second, using about 2.2 cores; at 128 clients the load generator, not the server, was the limit
Disk used RocksDB, 95,000 messages of 48 KB 4.9 GB: about 51 KB each
Disk used A production PostgreSQL cluster About 48 KB per message of mail body, plus about 22 KB each in the database with its search index
Mail server memory, idle Three production nodes 120–185 MB each
Webmail and console memory Three production nodes 31–36 MB and 8–15 MB
Local AI, per message Production, AMD Ryzen 5 3600, 4 cores for the model About 10 s: 8.7–11 s reading, 1.3–1.5 s answering
Local AI, per message Calibration, fast cores, 4 and 2 cores 6.3 and 10.8 s typical
Local AI memory Production About 4.4 GB per model

The load tests ran from a single machine, with the server's rate limits switched off so that they measured the server rather than the limits: the per-IP limits on incoming mail, and the limit of 1,000 requests a minute per account over HTTP. Each test client made far more requests than a person does, so leave those limits on in production.