LOADING DATASHEET
LOADING DATASHEET
by Stepfun
RANKED #223 OF 300 ASSISTANTS · OVERALL #352 OF 6,837 · VIBE SCORE 5.5 · 25 VOICES
Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token....
7 mentions
0 mentions
26 weeks · 20 voices
HOLDING UP
Recent vibe is steady. No clear fade in the crowd.
Open weights. You can use a hosted service, or download it and run it yourself, free.
The internet is split on Step 3.5 Flash. Biggest praise: help with code. Biggest gripe: getting facts right.
23 POSTS · 2 COMMENTS · LEMMY · BLOG · REDDIT · X · GITHUB
3 thumbs up · 1 thumbs down
0 thumbs up · 3 thumbs down
Other models the crowd has fully reviewed, starting with multimodal models like this one.
you can use step 3.5 flash by stepfun for FREE for 15 days😳 stepfun one of chinas ai tigers is running a limited promo
Last week in Image & Video Generation
Why DeepSeek V4 doesn't need more parameters to win
As some of you know, I have the long-running habit of keeping a running list of research papers I want to read, revisit, or cite in future articles and projects. Last year, I shared two organized paper lists, one covering January to June and another one…
I had originally planned to write about DeepSeek V4. Since it still hasn’t been released, I used the time to work on something that had been on my list for a while, namely, collecting, organizing, and refining the different LLM architectures I have covered…
If you have struggled a bit to keep up with open-weight model releases this month, this article should catch you up on the main themes. In this article, I will walk you through the ten main releases in chronological order, with a focus on the architecture…
HalBench: 29 OSS models tested on a custom built Sycophancy and Hallucination Benchmark, Qwen 3.6 and Gemma 4 scoring far above their weight! (While Meta keeps proving they forgot how to spend their money...)
GLM-5 is officially on NVIDIA NIM, and you can now use it to power Claude Code for FREE 🚀
When are we gonna get more 1-Bit models(Medium & Large size)?
Step-3.5-flash Unlosth dynamic ggufs?
Why is Step-3.5-Flash (196B-A11B) much cheaper to run than Qwen3.6-35B-A3B?
Is deepseek v4 flash the best model as subagent?