LongCat-2.0 is free in the Nous Portal for one week — a 1.6 trillion parameter model with a 1M context that can swallow a whole codebase in one pass. It’s also the model that was quietly powering Owl Alpha. Here are the real specs, the benchmarks, and where it actually sits.
Short answer
- 1.6T parameters, ~48B active, 1M context — built for agentic coding.
- Free in Nous Portal for one week, so the window to test it at no cost is short.
- It’s the full model behind Owl Alpha on OpenRouter — mystery solved.
- Strong on search and browsing; behind DeepSeek V4 Pro and Fable 5 on Terminal-Bench.
What LongCat-2.0 is
LongCat-2.0 comes from Meituan’s LongCat team, and it’s free in the Nous Portal for one week — which is why it matters to Hermes users right now rather than at some point in the future.
It’s a 1.6 trillion parameter mixture-of-experts model with around 48 billion active, a 1 million token context window, and it was built for agentic coding from the ground up. It can ingest an entire codebase in one pass.
And here’s the detail that reframes it for anyone who’s been following along: LongCat-2.0 is the full model behind Owl Alpha on OpenRouter. If you tested Owl Alpha and wondered what was actually under the stealth badge, this is it. I covered Owl Alpha in this breakdown when nobody knew what it was.
The architecture, in plain terms
| Piece | What it does |
|---|---|
| LongCat Sparse Attention (LSA) | Scales efficiently across the full 1M-token context instead of falling apart as it fills |
| Zero-Compute Experts | Dynamic activation between 33B and 56B per token — no wasted compute on tokens that don’t need it |
| MOPD | Three specialised expert groups — Agent, Reasoning and Interaction — gate-routed depending on the task |
That third one is the interesting design choice. Most MoE models route to whichever experts happen to fit. LongCat splits its experts by kind of work, so agentic tasks, reasoning tasks and conversational tasks each get routed to a group trained for them.
The Zero-Compute Experts idea is why the active parameter count moves rather than sitting fixed. Easy tokens cost less than hard ones, which is a different approach to efficiency than simply shrinking the model.
🔥 Want this set up without the guesswork? A new frontier model every few days is only useful if you can actually swap one in without rebuilding your setup. Inside the AI Profit Boardroom you get the Agent OS as a ready-to-install file, a 30-day roadmap, daily tutorials the same day new tools ship, and four live coaching calls a week where you share your screen and get unstuck. → Get access here
The benchmarks
| Benchmark | LongCat-2.0 |
|---|---|
| Terminal-Bench 2.1 | 70.8 |
| SWE-bench Pro | 59.5 (vs GPT-5.5 at 58.6) |
| SWE-bench Multilingual | 77.3 |
| FORTE | 73.2 |
| RWSearch | 78.8 |
| BrowseComp | 79.9 |
These are the vendor’s own published numbers. Same discipline as every other launch: take them seriously, treat them as claims until independent testers run the same suites.
One honest comparison worth making, because I have the numbers side by side. On Terminal-Bench 2.1 — which measures how well a model actually works inside a computer and finishes real tasks — LongCat-2.0 scores 70.8. DeepSeek V4 Pro scores 87.9 and Fable 5 scores 88.0 on the same benchmark.
So LongCat is meaningfully behind the current frontier on that specific measure, despite the headline specs. Where it looks strongest is search and browsing — RWSearch at 78.8 and BrowseComp at 79.9 are genuinely good numbers, and those are research tasks, not coding ones.
Full three-way comparison in DeepSeek V4 Pro vs Fable 5 vs Grok 4.6.
Why the 1M context matters with Hermes
A million tokens of context and whole-codebase-in-one-pass is a specific kind of useful, and it isn’t about coding for most people.
- Point it at an entire project folder and ask questions across all of it, rather than feeding it file by file.
- Hand it a full year of documents and get analysis that actually sees the whole set.
- Give an agent a large body of context up front instead of building retrieval plumbing to fetch it piecemeal.
The catch with big context windows is that most of them stop working once they’re actually full — a problem covered in the Pokee-Isaac breakdown. LSA is LongCat’s answer to that, and it’s the part worth testing yourself rather than taking on faith.
How to try it
- Nous Portal — free for one week from launch. That’s the fastest route and it costs nothing.
- OpenRouter — it’s the model behind Owl Alpha, so you may already have used it.
- Point Hermes at it like any other model, then test it on real tool calls rather than chat. See the model setup guide.
- Compare it properly — run the same task through it and through whatever you use now, in separate profiles, and judge on your own work.
The free week is the point. Free-for-a-limited-period offers expire, and this one is short. If you want to evaluate it at zero cost, that window is now rather than whenever you get round to it.
Want to swap models without rebuilding anything? The Agent OS in the AI Profit Boardroom lets you plug in Claude, Hermes and OpenClaw and change the brain underneath from a dropdown — so a free week on a new model is something you test in minutes. Start free with the free AI course and community.
FAQ
What is LongCat-2.0?
A 1.6 trillion parameter mixture-of-experts model from Meituan’s LongCat team with ~48B active and a 1M token context, built specifically for agentic coding.
Is LongCat free in Nous Portal?
It was made free in the Nous Portal for one week from launch. After that window it goes back to normal access, so test it while it’s free.
Is LongCat the model behind Owl Alpha?
Yes — LongCat-2.0 is the full model behind Owl Alpha on OpenRouter. The stealth model people were testing was this.
How does it compare to DeepSeek V4 Pro?
On Terminal-Bench 2.1 LongCat scores 70.8 against DeepSeek V4 Pro’s 87.9 and Fable 5’s 88.0. It looks stronger on search and browsing benchmarks than on terminal work.
What are Zero-Compute Experts?
A design where active parameters flex between 33B and 56B per token depending on difficulty, so easy tokens don’t burn compute they don’t need.
What is MOPD?
Three specialised expert groups — Agent, Reasoning and Interaction — gate-routed by task type, rather than routing purely on fit.
Can I use it with Hermes?
Yes. Point Hermes at it like any other model, ideally in its own profile so you can compare it against what you currently run.
Are the benchmarks independent?
No — they’re published by the LongCat team. Treat them as claims until outside testers confirm.
The bottom line
Hermes LongCat is worth a look this week specifically because it’s free in the Nous Portal, and because the mystery is solved — this is the model that was powering Owl Alpha. 1.6T parameters, 1M context and a whole codebase in one pass are real advantages, but it sits behind DeepSeek V4 Pro and Fable 5 on terminal work. Test it on your own tasks while the free window is open.
About Julian Goldie
I run Goldie Agency, a 7-figure SEO agency, and teach this daily on a 400K+ subscriber YouTube channel. 240+ client projects on Upwork at a 100% job-success score, 10+ years through every major Google update. My systems are in the AI Profit Boardroom; my link building book is free here.
