Essay · Jul 03, 2026 · 6 min read

Why we'll never train a model larger than 27B

Our ceiling isn't a limitation we're waiting to outgrow. It's a promise about who gets to hold the model.

Every few months, someone asks us when we'll “go big.” The question is fair. The industry's story so far has been a story of scale: more parameters, more data, more compute, and with them, more capability. Against that backdrop, a company that caps itself at 27 billion parameters looks like it's either short on money or short on ambition.

We'd like to explain why it's neither.

The line is about ownership

A 27B model, quantized to four bits, fits in roughly 16 GB of memory. That is the size of a single consumer graphics card, or a laptop you can buy at an ordinary store, with enough room left over for a long conversation. Cross that line and the model stops being something a person owns. It becomes something a person rents.

Renting isn't evil. But it changes the relationship. When the model lives on someone else's server, your prompts travel there too. The terms can change overnight. The service can disappear, get more expensive, or quietly start behaving differently. We think a large share of what people want from AI — help writing, thinking, coding, learning — should not depend on that kind of trust.

If it can't run on hardware a person can reasonably own, we don't ship it.

Constraints make you careful

When you can't buy your way out of a problem with scale, you have to understand it. A small model punishes sloppy data, so we spend most of our time on data: licensing it, cleaning it, and throwing away far more than we keep. A small model has less room for knowledge, so we teach it to be precise about the edges of what it knows instead of bluffing past them.

That last part matters more than any benchmark. A model that says “I'm not sure” at the right moments is more useful than a slightly smarter one that is confidently wrong. Calibration is a release metric at emjv for exactly this reason.

Small models are easier to keep safe

Safety work scales with the thing you are trying to make safe. Every evaluation, every red-team pass, every interpretability probe costs compute in proportion to the model's size. Staying small means we can afford to run our full suite of 412 safety checks on every release candidate, not just the final one, and to repeat them after every quantization step.

It also means more people can check our work. Independent researchers can study a 27B model on a single machine. We would rather be audited by a thousand people with a desktop than by a handful with a data center.

Local means the safety has to travel

We should be honest about the trade-off. A model running offline has no moderation layer, no usage monitoring, and no remote off-switch. Whatever safety it has, it carries inside its weights. That is a higher bar, not a lower one.

So we train alignment into the model rather than bolting it on around it. We run capability evaluations in high-risk domains before every release, and if a model crosses a threshold we've published in advance, it doesn't ship. Not later, not in a limited preview. It doesn't ship.

What we give up

We will lose on some leaderboards. There will be tasks the largest frontier models can do that ours can't, and when that's the case we'll say so plainly in our model cards. Knowing when to send you elsewhere is part of being trustworthy.

What we gain is a model that is yours: one that works on a plane, in a hospital basement, in a country with unreliable internet, or on a laptop you never want connected to anything at all.

Will the number ever change?

Hardware improves. If the memory in an ordinary computer grows, we may let our ceiling grow with it, and we'll announce that publicly before we do. But the rule underneath the number won't change. We build models for the machine on your desk. That's the whole idea.

— The emjv team