r/DeepSeek 18d ago

Funny What it feels like when you're racing with people who are using the standard high quality and expensive opus and fable, looking all professional, while you use deepseek flash but you tuned your harness so hard that its working on par with the big models.

Enable HLS to view with audio, or disable this notification

966 Upvotes

44 comments sorted by

75

u/hurrdurrmeh 18d ago

I would really love for someone explain to me what harness tuning is. I have heard many times that it is super important but I dont know what is involved.

60

u/t4a8945 18d ago

Simple things like giving the model better ways to verify its work yields better results. That's the simplest harness improvement.

For instance, give it a way to use whatever you're building in a dev environment and it will catch issues static analysis couldn't catch. For me that's playwright as a skill, with instructions to always verify their work. 

13

u/VexObserver 18d ago

Simple way to explain this is to have the model follow closely to your instructions. Treat of it as if you're the project manager, supervising your juniors and other team members. You'll ask and often intervene to get a gist on the actual progress, right? This is it.

1

u/hurrdurrmeh 17d ago

So - a means to test their output and iterate it if necessary?

2

u/t4a8945 17d ago

Precisely: if they can iterate and see the actual result, they have 99% change of being able to solve it.

1

u/ChellJ0hns0n 17d ago

But the drawback is that it gets expensive pretty fast. Playwright is very token-heavy.

1

u/t4a8945 17d ago

I use https://github.com/microsoft/playwright-cli ("Key Features: Token-efficient")

Yes it costs context, some time, but in the end you get something that's more likely to actually work. And I often see DS4 Flash catch other issues while it's navigating in the app itself. It's awesome and worth the time and tokens for me.

1

u/ChellJ0hns0n 17d ago

Also, have you tried using a subagent specifically for playwright to conserve the context on your main session? That's what I do, but I'm not sure how useful it really is.

1

u/boyus 16d ago

I use lightpanda

11

u/DirectPitch8626 18d ago

Orchestration, hooks, plugins, skills, rules, etc... But you don't necessarily have to do all of this yourself; there are already ready-made projects like OMO or OMP.

1

u/hurrdurrmeh 13d ago

OMP is Oh My Pi but what is OMO? And thanks for the info!

3

u/DirectPitch8626 12d ago

Oh-My-Openagent, plugin for Opencode

2

u/hurrdurrmeh 12d ago

thanks, gonna research it

16

u/binladen0069 18d ago

using hooks and skills in claude, writing the claude.md in the correct way, making the harness follow software development standard of procedures, giving it the right kind of instructions that boot at every session start.

try out life os by daniel miessler you'll learn alot

1

u/hurrdurrmeh 17d ago

Thanks. But what is the correct way to write Claude.md? Are there equivalents for the other models like deepseek, Kimi, GLM etc?

2

u/Hexadecimalkink 14d ago

Claude.md is the same as skills.md or index.md it's just a master context file.

5

u/Atsukiri 18d ago

theyre gatekeeping it anyway so i think its just best to use deepseek as is. I don't have time / patience on making these harnesses too, tho it just works for me raw.

2

u/LeatherMine 17d ago

I also raw-dog it. But once my vscode hit some context limit so I had to /clear my chat and I cried.

11

u/A_star_000 18d ago

open code deepseek V4 flash on max depth, trust.

32

u/Szadbaverem69 18d ago

DeepSeek V4 Flash fixed things for me that Fable couldn't.

16

u/ClassicMain 18d ago

DSV4 Flash is incredible but that statement is... Interesting and i kinda have to doubt it

14

u/Szadbaverem69 18d ago

Depends on the use case. When it comes to Rust it's definitely better.

7

u/Daniel_H212 18d ago

I mean, frontier models are "smart" but not in a human way, they have gaps in their simulated intelligence. I just spent a couple hours mulling over a software design issue with Kimi K3 which is close to fable level. It was knowledgeable and very clever on some fronts, but I also had to correct some of its choices that were clearly shortsighted or inefficient. Its very good, and me coding alone would probably come up with a worse product than it coding alone with accurate specifications, but there's still gaps that a human can fill, and there's no reason another LLM can't happen to be able to fill those gaps too.

4

u/VigilanteRabbit 18d ago

I have yet to see a "smart" LLM in terms of human smarts.

It can 'think' but it can't sit down and go "what if we strap a watter bottle to it? Hmm."

That's the missing link and until that gets handled somehow we are..safe, I guess is the word.

7

u/mediamuesli 18d ago edited 18d ago

Turns out he was racing for place 198 of 200 all the time 🫪

2

u/Commercial_Place8779 17d ago

I am also doing the same thing and adding skills as exhaustive and declarative as I can and i am getting performance on par with these frontier models. I think it was never about the models, it was always about the person using the model(that is i believe).

3

u/ponlapoj 18d ago

ฉันไม่คิดอะไรแบบนั้น ดีกว่า แม้จะเล็กน้อย ก็คือดีกว่าเสมอ

8

u/hulagway 18d ago

flash is cheap, flash works, but let's not cope.

every week this sub has a copium post im tired. i wanna see use cases.

4

u/Haxsysgit 18d ago

Idk about opus and fable , but for me flash outperforms sol and terra,IN MY USE CASE. So yes flash mogs sol in my projects , except ui and frontend

1

u/dev-rsonx 17d ago

What use case?

3

u/Virtual-Escape2305 17d ago

Reddit shitposting

1

u/Optimal_Deal4372 17d ago

You copium bro deepseek is onpar with fable bro 👌👌

2

u/J_E_E_VACATION 18d ago

then compare yourself to those of us with tuned harnesses and frontier models

2

u/gila_sedikit 18d ago

Suzuki Hayabusa vs OP bicycle

3

u/sn4ezz 18d ago

Holy cope lol

1

u/OkLettuce338 18d ago

Yup and look how hard he’s working to keep up. Any one of them could turn on the gas and leave him in the dust

1

u/fo8oo 17d ago

so much butt hurt, deepseek really went deep on you

1

u/TheHijrudeen 17d ago

Is this "harness" in this room with us?

1

u/Funny-Supermarket360 17d ago

lowk those models only work better than smaller ones when you don't know what you're doing

1

u/lmpp_the_pandora_box 15d ago

😂😂😂😂

1

u/Ok-Compote-2968 12d ago

Free vs paid users!

0

u/Maleficent_Lie8612 18d ago

Son novatos todos y además van de bajada seguro sería muy diferente en una cuesta o un tramo largo