Arête logo: a pale curved triangle on a deep violet field

Arête

Small, fast language models built to run on the hardware you already own.

aretē (Greek): excellence; living up to your full potential.

The lineup

Six sizes, from a ThinkPad to a home datacenter. Tell us what you're running and we'll show what fits.

I have

Parameters
Design
Size at Q4_K_M
Runs on
Status

File sizes are estimates at Q4_K_M, about 0.6 GB per billion parameters. "Fits" leaves ~20% headroom for context.

How we build

We'd rather you run a model than admire its parameter count.

Speed first

Every model gets its own multi-token prediction head, trained in from the start. Mid-size models use mixture-of-experts so only about a billion parameters work on each token.

Quantization is the target

Every tier is designed to be run at Q4_K_M. That's the version we test, tune, and size against real GPU memory, not round numbers.

Open, on a delay

The newest flagship lives behind our API. Once two newer generations ship, its weights go public, so older models are always open.

The water cycle

Every name comes from one system, so you can tell what something is before you open it.

Language models

Named for the sky, from frost to storm. Bigger model, bigger weather.

Rime · Wisp · Cirrus · Stratus · Nimbus · Maelstrom

Datasets

Named for moving water: what the models learn from.

Undertow

Encoders

Named for light in the atmosphere: models that see and hear.

Halo · Aurora · Peal · Corona