I can't verify whether or not the LLM is arriving at the conclusion from cheating, or if it's fudging or making stuff up.
Opus 4.6 remains the best model because of this.
I can't verify whether or not the LLM is arriving at the conclusion from cheating, or if it's fudging or making stuff up.
Opus 4.6 remains the best model because of this.
7 comments
Only recent quirk in longer chats with Deepseek, watching the the thought process and output; it sometimes becomes Chinese. Also not sure if this is to intentionally hide the thought process, or the actual information needs to be looked up and explained in Chinese.
Huh? Are you not trying other open-weight models that stream thinking traces?
The new DeepSeek-V4-Flash-0731 should clobber Opus 4.6 in a lot of tasks. Kimi K3 and GLM 5.2 feel like they stand toe-to-toe with Opus 4.8 in my experience. This probably isn't the last stupid decision that Anthropic stands on, you might as well hedge your bet and put some money into another inference provider and see how it goes.
Honestly Anthropic keeps clubbing themselves. They're so worried about their competition they're no longer innovating.