LOADING DATASHEET
LOADING DATASHEET
by DeepSeek
RANKED #185 OF 271 ASSISTANTS · OVERALL #297 OF 6,567 · VIBE SCORE 5.6 · 28 VOICES
Qwen instruction model for multilingual chat, reasoning, and tool use
3 mentions
3 mentions
26 weeks · 20 voices
QUIETLY DEGRADING
Recent vibe and chatter are both sliding.
Closed model. DeepSeek hosts it, so you use it through an app or an API.
DeepSeek R1 Distill Qwen 1.5B is closed, so you rent it. These open-weight models can genuinely stand in for it. Download once, own forever.
Not enough discussion about DeepSeek R1 Distill Qwen 1.5B yet to call it. We found 28 posts but people did not say much either way.
26 POSTS · 2 COMMENTS · STACK OVERFLOW · HACKER NEWS · LEMMY · REDDIT · GITHUB
Other models the crowd has fully reviewed, starting with text models like this one.
1.5B did WHAT?
[R] AutoThink: Adaptive reasoning technique that improves local LLM performance by 43% on GPQA-Diamond
DeepSeek-R1-Distill-Qwen-1.5B Surpasses GPT-4o in certain benchmarks
Context window length table
I fine-tuned DeepSeek-R1-1.5B for alignment and measured the results using Anthropic's new Bloom framework.
Very cool of ChatGPT to walk me through finding and installing the best version of DeepSeek for my system! Even gave me a video to watch :)
I made an almost universal LLM Creator/Trainer
DeepSeek-R1 on iPhone? (DeepSeek-R1-Distill-Qwen-1.5B-GGUF)
I fine-tuned DeepSeek-R1-1.5B for alignment and measured the results using Anthropic's new Bloom framework
Fix AutoTokenizer returning TokenizersBackend for DeepSeek-R1-Distill-Qwen models
Add credential-free DeepSeek open-weight dialogue fallback
[rollout] Continuous Token: DeepSeek-R1 distills cannot construct a builder since #6804 (Qwen builder requires < im_end >)