Midium

Production inference. Cloud or local.

Competitive token billing in the cloud. The same production-grade models on a machine you own. Native tool calls, measured. We do not retain prompts or generated responses.

Cloud

Token billing. Ready today.

Call the same OpenAI-compatible API over the internet. Pay per million tokens. Competitive rates, no credit buckets, no usage theater.

Midium Cloud does not retain prompt data or generated response data.

Local

Air-gapped. On your Mac.

  • Mac Mini$150/mo
  • Mac Studio$300/mo
  • Vaults add-on+$100/mo

Download the desktop app for out-of-the-box production-grade inference on hardware you own. Same models. Same API. Unlimited on the box.

Midium Local is 100% air-gapped — built for compliance and data-sensitive work. Nothing leaves the machine. Vaults are a +$100/mo add-on: no vault count or usage restrictions.

Tool calling

The engine that makes agents reliable.

Everyone serves the same open models. Midium’s reliability layer constrains generation so native OpenAI tool_calls stay well-formed — even when the raw model would fumble them. Same contract on cloud and local. Measured on Gemma 4 26B A4B against Ollama: identical 4-bit weights, Apple Silicon, BFCL v4’s official checker, 1,000 cases, temp 0.

BFCL tool-calling accuracy
0.902
vs Ollama’s best 0.874
Parallel-call points
+13.5
Matched no-think · 0.830 vs 0.695
Decode speed
+37%
123.8 vs 90.1 tok/s
Sooner to first token
2.3×
0.12s vs 0.28s

Parallel calls are the reliability moat. On the much smaller E4B with thinking on, Ollama edges Midium — the flagship 26B is where the layer shows. Courier vs Ollama: Tool-Calling, Measured.

Vaults

The database for your agents.

Persistent knowledge bases, attached over MCP. Retrieval is semantic and exact. Agents write back as they work.

Cloud

Unlimited vaults. Pay for traffic.

  • Reads / 1k$5
  • Writes / 1k$25

A read is an MCP retrieve, search, or get. A write is an insert, update, or delete that mutates the index.

Local

+$100/mo. No restrictions.

Add Vaults to a Mini or Studio license. Same MCP surface, air-gapped, unlimited vaults, no usage caps. The company brain stays on a machine you own.

  • MCP-accessible — attach a vault the way you attach any tool.
  • Auto-writeback — agents modify the vault as they work.
  • Semantic and deterministic retrieval in one engine.

Local stays private

Prompts, files, outputs, and local vaults stay on your machine. The runtime never phones home with customer data. Air-gapped by architecture, not by policy.

Cloud keeps nothing

Hosted inference is billed by the token. We do not retain prompts or generated responses.

Vaults persist

Vault contents persist until you delete them — that is the product. We do not train on them. Local vaults never leave the machine.

Library

Models and costs

Cloud prices are per million tokens. Input is prefill. Output is decode. Local includes the full library on either a Mac Mini license ($150/mo) or a Mac Studio license ($300/mo).

Midium model library and cloud token prices
ModelClassCloud input / MCloud output / MLocal
Gemma 4 E2BLite$0.01$0.05Included
Gemma 4 E4BLite$0.014$0.07Included
Gemma 4 26B A4BBalanced$0.028$0.14Included
Qwen3.8 27BBalanced$0.032$0.16Included
Gemma 4 31BBalanced$0.054$0.27Included
Laguna XS 2.1Balanced$0.052$0.26Included
Laguna S 2.1Frontier$0.184$0.92Included
Inkling SmallFrontier$0.30$1.50Included

Same catalog on Midium Cloud and Midium Local. OpenAI-compatible API.

Same experience both ways.

Point your client at the cloud or at localhost. The models, the native tool calls, the vaults, and the privacy posture are designed to match.