fbpx

Maple Preview: Fast, Free, Local — And Not For Coding

WANT TO BOOST YOUR SEO TRAFFIC, RANK #1 & Get More CUSTOMERS?

Get free, instant access to our SEO video course, 120 SEO Tips, ChatGPT SEO Course, 999+ make money online ideas and get a 30 minute SEO consultation!

Just Enter Your Email Address Below To Get FREE, Instant Access!

Maple Preview is a new open-source reasoning model with ternary weights, an MIT licence and genuinely impressive speed on local hardware. Here’s what it does well, what it doesn’t, and my honest take after testing it.

Short answer

  • 20B model with roughly 1B running cost, using ternary weights — minus, zero or plus.
  • Shrinks around 38GB down to about 5GB, with 256 experts and 8 active per token.
  • 128k context, MIT licence, free to use commercially.
  • Honestly not great at coding — its real story is speed and running on phones.

What Maple Preview is

Maple Preview is a new open-source reasoning model. On paper it’s a 20B model with roughly the running cost of a 1B one, and it uses ternary weights — which is the trick that makes it so small and so quick.

I plugged it straight into the local engine in my agent operating system and asked it to build a snake game. It responded fast. Genuinely fast, on a Mac Studio, running locally and free, with the Wi-Fi switched off.

That matters because most local models I’ve tested are either terrible, or slow enough to bog down my whole setup. This one isn’t.

The specs that actually matter

Spec Detail
Size 20B total, around 1B active at runtime
Weights Ternary — every weight is minus, zero or plus
Experts 256 small networks, 8 activated per token
Context 128k tokens
Licence MIT, so you can use it commercially
Type An actual reasoning model, not just chat

Ternary weights, explained without the jargon

A normal model stores every connection as a precise number with a lot of decimal places. Maple stores each one as just minus, zero or plus — three symbols.

Writing the brain with three symbols makes the file tiny and the maths fast. A model that would normally need around 38GB fits into roughly 5GB, and your chip chews through it far quicker.

The second trick is the expert system. Maple holds 256 small networks inside it, and for each word it generates, a router wakes up only the eight most useful ones. So in theory you get the knowledge of a 20B model at the running cost of a 1B one.

My honest take on quality

I’ll say this for free: it is not the best model in the world at coding.

They’ve compared it against Claude Sonnet 5 in their materials. Having tested it on coding tasks myself, it is nowhere near that level, and I’m not going to pretend otherwise. It also wasn’t as good at coding as Gemma 4, GLM 4.7 Flash or GPT-OSS when I ran them side by side.

What it is, is fast. The whole game with local models is the speed-quality frontier — usually the better a model gets, the slower it runs. Maple sits high on speed while holding decent average performance, and that’s a real trade worth making for some jobs.

I also like how it responds. It gives more detailed answers than a lot of the small models I’ve tested, and it feels pleasant to use, which sounds soft but matters when you’re in it all day.

The mobile angle is the actual story

The thing that makes Maple interesting isn’t beating Sonnet. It’s that they’re targeting phones.

Compare it to something like Qwen 3.6 27B, which is the model I hear about most. On my Mac Studio that’s a lot slower, and depending where you run it a simple question can take minutes. Running that class of model on a phone isn’t realistic.

Free local models running properly on a mobile device is a future that’s clearly coming, and this is one of the first ones built for it rather than shrunk down to fit.

They’ve also shown it running more autonomously on a MacBook Pro, where it decides on its own to remember details. Adapting and learning as it goes is a big advantage for agent work.

Where it fits in a real setup

I keep a local section in my agent OS where I can swap models in and out. Maple slots in there as the fast local builder, and I run something else alongside it for agent work — LFM 2.5 2.6B pairs nicely, since that one’s trained for the Hermes harness.

You can genuinely power a whole agentic OS off a couple of local models: one as the local builder, one for the agent tasks, and a frontier model on top only when you need it.

If you want the wider picture, see the local model setup guide and the best local models for Hermes.

How to try it

  1. Test it in the developer’s own chat demo if you just want a quick feel for it.
  2. Grab the open weights from Hugging Face if you want it running on your own machine.
  3. Point your local engine or agent OS at it, and keep another model configured so you can switch when the task needs more depth.
  4. Judge it on your own tasks. Benchmarks are a starting point, not an answer.

Want the local engine setup? The Agent OS — with the local model section, the swap-in-swap-out engine, the memory system and full mission control — is inside the AI Profit Boardroom, along with training on running the whole thing free on local models.

FAQ

What is Maple Preview?

An open-source ternary-weight reasoning model, around 20B total with roughly 1B active at runtime, with a 128k context window and an MIT licence.

What are ternary weights?

Instead of storing each connection as a precise decimal number, every weight is minus, zero or plus. Three symbols make the file tiny and the maths fast.

How small is it?

A model that would normally need around 38GB fits into roughly 5GB, which is what lets it run quickly on consumer hardware.

Is Maple Preview good at coding?

Honestly, no — not compared to Gemma 4, GLM 4.7 Flash, GPT-OSS or Claude Sonnet 5. In my own tests it was clearly behind on coding quality. Its strength is speed.

So what is it actually good for?

Fast local work where speed matters more than maximum quality, and especially anything targeting smaller devices. That’s where it’s genuinely ahead.

Can it run on a phone?

That’s the design goal, and it’s the most interesting thing about it. Models in the 27B class simply aren’t practical on mobile.

Is it free?

Yes. It’s MIT licensed, so it’s free to use commercially, and running it locally costs you nothing beyond your own hardware.

Do I need Wi-Fi?

No. Once it’s local you can switch the Wi-Fi off and keep working, which is the underrated benefit of the whole local model approach.

The bottom line

Maple Preview is fast, free, local, MIT licensed and architecturally clever — and it is not a coding model. Judge it on speed and on where it’s headed, which is proper local AI on small devices. Test it yourself and see what you think.

About Julian Goldie

I run Goldie Agency, a 7-figure SEO agency, and teach this stuff daily on a 394K+ subscriber YouTube channel. I’ve delivered 240+ client projects on Upwork at a 100% job-success score over 10+ years of ranking sites through every major Google update. The systems I actually run are inside the AI Profit Boardroom, and my link building book is free here.

Picture of Julian Goldie

Julian Goldie

Hey, I'm Julian Goldie! I'm an SEO link builder and founder of Goldie Agency. My mission is to help website owners like you grow your business with SEO!

Leave a Comment

WANT TO BOOST YOUR SEO TRAFFIC, RANK #1 & GET MORE CUSTOMERS?

Get free, instant access to our SEO video course, 120 SEO Tips, ChatGPT SEO Course, 999+ make money online ideas and get a 30 minute SEO consultation!

Just Enter Your Email Address Below To Get FREE, Instant Access!