DeepSeek V4 is an API model with barely any wrapper of its own — it expects you to bring a harness. Here are the three decisions to make before you connect it, and the one setting that actually decides your bill.
Short answer
- V4 has no real consumer app — your agent setup is the harness.
- Choose your host deliberately: DeepSeek’s own API terms allow training on what you send.
- Cache reads cost ~276x less than Fable 5 — check caching is actually being hit.
- A price rise has been announced, so keep the model swappable in its own profile.
V4 doesn’t come with a body
DeepSeek V4 is an API model. There’s barely a consumer wrapper — it expects you to bring your own harness.
For most people that’s a genuine problem: a brilliant cheap brain you can’t actually use. If you already run an agent setup it’s the opposite, because your setup is the harness and V4 is just an extremely cheap brain you drop inside it.
This page is about that wiring — what you have to decide, and the settings that actually change your bill. If you want the comparison of which harness, that’s in the best harness for DeepSeek V4.
Three decisions before you connect anything
1. Which provider hosts it
This one matters more than people expect. DeepSeek’s own API terms let them train on what you send through it. Other providers host the same model without that condition, and OpenRouter has said more are coming online.
For hobby projects and public content, fine. For client work, contracts or anything with customer data, pick your host deliberately.
2. Which variant
V4 Pro is the flagship built heavily for agent work — 1M token context, up to 384,000 tokens of output in one go. V4 Flash is the lighter, faster one. Most serious setups use both, which is covered in the Flash guide.
3. Whether the model is swappable
DeepSeek have posted notice of a significant price increase across the whole API, with no date and no amount given. Today’s pricing is real; it isn’t permanent.
So wire it up in a harness where the model is a dropdown rather than a rebuild. That single choice is what turns a price rise from a crisis into an afternoon.
🔥 Want this set up without the guesswork? Getting a cheap brain wired into a harness properly is the difference between a low bill and a broken agent. Inside the AI Profit Boardroom you get the Agent OS as a ready-to-install file, a 30-day roadmap, daily tutorials the same day new tools ship, and four live coaching calls a week where you share your screen and get unstuck. → Get access here
The setting that actually decides your bill
Agents don’t read your instructions once. They reread them at every step — picture a worker opening the full job folder before every single action. That folder gets read hundreds of times per task.
Each of those rereads is a cache hit, and this is where V4 is unusually cheap:
| Measure | DeepSeek V4 Pro vs Fable 5 |
|---|---|
| Input | ~23x cheaper |
| Output | ~57x cheaper |
| Cache reads | ~276x cheaper |
| Cache hit rate on OpenRouter | ~92% |
So the headline number understates it for exactly the workload a harness exists to run. When you’re configuring, the thing to check is that caching is actually enabled and being hit — a setup that misses cache turns a 276x advantage into a 57x one.
Wiring it in, step by step
- Pick your host and get a key from them, not necessarily from DeepSeek.
- Create a dedicated profile for it in your harness rather than switching your main one. Hermes profiles are ideal for this.
- Point that profile at the endpoint and select the V4 variant you want.
- Test tool calls, not chat. Plenty of models converse well and fall over the moment they have to call a tool. This is the only test that tells you anything.
- Run the same task through your existing setup in a second profile and compare properly.
- Watch the first bill closely and confirm the cache behaviour matches expectations.
Keeping it in its own profile is the habit that pays off. When the price rises, or a better model lands, you’re changing one profile rather than untangling your whole setup.
What V4 can’t do
- It can’t see images. If your workflow depends on screenshots or reading documents as pictures, no harness fixes that. DeepSeek have reportedly said vision doesn’t advance the research goals they care about, so don’t wait for it.
- It isn’t the smartest model available. Fable 5 leads it by around seven points on hard software engineering work. On genuinely difficult multi-step tasks that shows up as more mistakes and more restarts.
- It has no polished home of its own. The harness is on you.
Which is why the sensible pattern is a split: a frontier model plans and reviews, V4 does the volume. Full picture in the three-way comparison.
Want the wiring done once, properly? The Agent OS in the AI Profit Boardroom runs Hermes, Claude and OpenClaw with model swapping built in, so pointing any of them at V4 — or off it when the price moves — is a dropdown. Install file, 30-day roadmap, daily tutorials and four coaching calls a week. Start free with the free AI course and community.
FAQ
How do I run DeepSeek V4 in a harness?
Pick a host, create a dedicated profile in your harness, point it at the endpoint, select the variant, then test tool calls rather than chat.
Does DeepSeek V4 come with its own app?
Barely. It’s an API model that expects you to bring your own harness — a problem if you have no setup, an advantage if you do.
Which provider should I use?
Not automatically DeepSeek’s own. Their API terms allow training on what you send; other hosts run the same model without that condition.
Why is it so cheap for agent work?
Agents reread instructions constantly, and cache reads on V4 Pro cost roughly 276x less than Fable 5, with a ~92% cache hit rate on OpenRouter.
Will the price stay this low?
No. A significant increase has been announced with no date or amount, which is why you should wire it into a harness where the model is swappable.
Can it handle images?
No. V4 is text only, and DeepSeek have indicated vision isn’t a priority for them.
Should I switch everything to V4?
No. Use it for volume and keep a frontier model for the hard reasoning — that split is where the saving actually comes from.
Why use a separate profile?
So you can compare it fairly against your current setup, and so a price change or a better model is a one-profile edit rather than a rebuild.
The bottom line
Running DeepSeek V4 in a harness comes down to three decisions — who hosts it, which variant, and whether the model is swappable — plus one setting worth checking, because the 276x cache advantage is where the savings actually live. Keep it in its own profile, test tool calls rather than chat, and don’t hand it anything that needs eyes or the hardest reasoning.
About Julian Goldie
I run Goldie Agency, a 7-figure SEO agency, and teach this daily on a 400K+ subscriber YouTube channel. 240+ client projects on Upwork at a 100% job-success score, 10+ years through every major Google update. My systems are in the AI Profit Boardroom; my link building book is free here.
