EMPERO/00 — INDEX INDEPENDENT AI RESEARCH LAB · OPEN BY DEFAULT · BUILT IN GERMANY 2026
New · Qwen3.8-9B / 4B / 2B — distilled

Small models,
trained in the open.

We are an independent AI research lab in Germany, building language models efficient enough to run on hardware you own. Latest: a full-parameter distill of Qwen3.8-2.4T into 9B / 4B / 2B, and a GDN-aware 27B GGUF at 11.7 GiB. The open-weight Qwythos family has passed one million downloads on Hugging Face.

1M+
Model downloads · Hugging Face
262,144
Token context window · Qwen3.8 native
25
Open models · Hugging Face
01 Models Qwen3.8 distill · Ridge · Qwythos
Latest release Qwen3.8 2.4T teacher 262k context Apache-2.0

Qwen3.8-9B / 4B / 2B — one teacher, three students.

A full-parameter SFT of Qwen3.8-2.4T-A95B into the Qwen3.5 9B, 4B and 2B architectures. Same curriculum, three capacities: ~70k / 45k / 30k curated teacher traces of dense chain-of-thought — math, code, reasoning, instruction following, tool use. Every answer opens with a <think> block learned from the teacher, not from self-generated rollouts. Native function calling, 262,144-token context, day-one GGUF.

MMLU CoT · vs Qwen3.5 base
9B54.6 → 75.1
4B35.4 → 55.3
2B28.3 → 54.8
2B · GSM8K33 → 64
Native context262k
02 Abacus Where the models go to work
Product Native Rust macOS · Linux · Windows

A coding agent that runs on your model.

A fast terminal agent in one native binary, pointed at any OpenAI-compatible endpoint — a frontier API, or a Qwythos on the machine in front of you. Abacus is where our models get used in anger: proof that small, owned models handle real codebases.

  • Approval-gated by design — every mutation is shown first as a semantic per-file diff.
  • PLAN / BUILD workflow — the model declares read-only or mutating intent before acting.
  • Long-horizon sessions — persistent goals, parallel git-worktree subagents, context compaction.
  • Extensible — Agent Skills, declarative plugins and MCP servers.
03 Research The pipeline behind the models
R.01 — Post-training

FTPO — fix the token, not the model

Our Final-Token Preference Optimization method identifies the exact token where a failure mode begins and trains the model to prefer coherent alternatives at that one position. ~2,000 preference tuples took Qwythos looping from 6.7% to 0.0% with no capability loss.

R.02 — Data

rethink + SFTSuite

Our trace-generation tool produced the 500M+ tokens of verified chain-of-thought data behind Qwythos, and SFTSuite orders it into curricula. Data quality — not parameter count — is where our models win, and the data pipeline outlives any single release.

R.03 — Architecture

Microverse — search before you train

An LLM-in-the-loop system that searches architecture space before we commit full training runs — the reason Qwythos ships hybrid linear-attention designs that hold a 1M-token context on consumer hardware.

R.04 — Scaling

One curriculum, every size

The same recipe produced Qwythos-9B and Qwythos-27B-v1, then the Qwen3.8 distill at 9B, 4B and 2B. Same teacher, same curriculum, three capacities — MMLU CoT jumps of +20.5 / +19.9 / +26.5. The pipeline scales with compute, which is exactly what this raise buys.

NEXT — In pretraining now

From distillation to models of our own.

Two next-generation in-house models are in pretraining now, on the infrastructure and data pipelines we already operate. We won't describe them here — no roadmap theatre. We announce models when the weights are ready to download, not before. Every model we've shipped arrived this way; these will too.

2 models · in pretraining
04 Approach Why small · why open

Frontier reasoning,
small enough to own.

Inference is leaving the data centre: a single GPU, an on-prem rack, a laptop. Getting frontier-class reasoning into models that small — models a company can actually own — is the whole of our research.

02.1

Better data, less of it

Qwythos is distilled, not fine-tuned: over 500 million tokens of reasoning traces from rethink, our in-house chain-of-thought tool. When the 9B shipped with a repetition loop, our FTPO method fixed it — 6.7% to 0.0% — with about 2,000 preference pairs, not a retrain.

02.2

Open is our distribution

We have no sales team and no ad budget. Qwythos passed one million downloads because engineers found it, used it and told other engineers — the v2 GGUF alone accounts for over 500,000. Open weights get the work tested, and adopted, at a scale we could not buy.

02.3

Data that cannot leave the building

European enterprises increasingly cannot send regulated data to a foreign API. Qwythos is Apache-2.0, runs on-prem on a single GPU, and can be audited down to the weights. We build in Germany because that constraint is our market — and it is not going away.

05 In numbers All of it public
1,081,832
Total model downloads
Hugging Face · all-time
4,300+
Hugging Face likes
2.6K on the flagship alone
25
Open models published
weights · GGUF · datasets
8
Open-source repositories
github.com/empero-org

Every number on this page is public. Check them: huggingface.co/empero-ai · github.com/empero-org.

— Dispatch

Follow the build.

A short letter, every other Tuesday: progress on Qwen3.8 and Qwythos, new Abacus releases, and the one thing we got wrong that week. No hype, no roadmap theatre. Cancel from any line.

We never share addresses.
06 Contact Open by default

Everything we ship
is a download away.

The weights are on Hugging Face, the code is on GitHub, and the eval transcripts sit in the model cards. If you're running Qwythos and something breaks, if you need a model your auditors can actually read, or if you're wiring Abacus into your team's workflow — write to us.

For developers
Qwythos runs on llama.cpp, Ollama and vLLM from a single download. Issues and PRs get read by the people who trained the models.
For enterprises
Apache-2.0, self-hosted, auditable down to the weights. If regulated data keeps you off hosted APIs, that is the point.
For researchers
Datasets, eval transcripts and methods are published with each release. Replication is encouraged; corrections even more so.

We are independent and intend to stay that way. Research at this pace takes capital — conversations about that reach the same address.