Cloud
Token billing. Ready today.
Call the same OpenAI-compatible API over the internet. Pay per million tokens. Competitive rates, no credit buckets, no usage theater.
Midium Cloud does not retain prompt data or generated response data.


Cloud
Call the same OpenAI-compatible API over the internet. Pay per million tokens. Competitive rates, no credit buckets, no usage theater.
Midium Cloud does not retain prompt data or generated response data.
Local
Download the desktop app for out-of-the-box production-grade inference on hardware you own. Same models. Same API. Unlimited on the box.
Midium Local is 100% air-gapped — built for compliance and data-sensitive work. Nothing leaves the machine. Vaults are a +$100/mo add-on: no vault count or usage restrictions.
Tool calling
Everyone serves the same open models. Midium’s reliability layer constrains generation so native OpenAI tool_calls stay well-formed — even when the raw model would fumble them. Same contract on cloud and local. Measured on Gemma 4 26B A4B against Ollama: identical 4-bit weights, Apple Silicon, BFCL v4’s official checker, 1,000 cases, temp 0.
Parallel calls are the reliability moat. On the much smaller E4B with thinking on, Ollama edges Midium — the flagship 26B is where the layer shows. Courier vs Ollama: Tool-Calling, Measured.
Vaults
Persistent knowledge bases, attached over MCP. Retrieval is semantic and exact. Agents write back as they work.
Cloud
A read is an MCP retrieve, search, or get. A write is an insert, update, or delete that mutates the index.
Local
Add Vaults to a Mini or Studio license. Same MCP surface, air-gapped, unlimited vaults, no usage caps. The company brain stays on a machine you own.
Prompts, files, outputs, and local vaults stay on your machine. The runtime never phones home with customer data. Air-gapped by architecture, not by policy.
Hosted inference is billed by the token. We do not retain prompts or generated responses.
Vault contents persist until you delete them — that is the product. We do not train on them. Local vaults never leave the machine.
Library
Cloud prices are per million tokens. Input is prefill. Output is decode. Local includes the full library on either a Mac Mini license ($150/mo) or a Mac Studio license ($300/mo).
| Model | Class | Cloud input / M | Cloud output / M | Local |
|---|---|---|---|---|
| Gemma 4 E2B | Lite | $0.01 | $0.05 | Included |
| Gemma 4 E4B | Lite | $0.014 | $0.07 | Included |
| Gemma 4 26B A4B | Balanced | $0.028 | $0.14 | Included |
| Qwen3.8 27B | Balanced | $0.032 | $0.16 | Included |
| Gemma 4 31B | Balanced | $0.054 | $0.27 | Included |
| Laguna XS 2.1 | Balanced | $0.052 | $0.26 | Included |
| Laguna S 2.1 | Frontier | $0.184 | $0.92 | Included |
| Inkling Small | Frontier | $0.30 | $1.50 | Included |
Same catalog on Midium Cloud and Midium Local. OpenAI-compatible API.