LOADING DATASHEET
LOADING DATASHEET
by Zhipu
RANKED #73 OF 228 ASSISTANTS · OVERALL #122 OF 6,618 · VIBE SCORE 6.2 · 150 VOICES
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
20 mentions
9 mentions
26 weeks · 205 voices
HONEYMOON FADING
Early praise is cooling off in recent weeks.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download GLM 4.7 Flash once and it's yours. No subscription, no rate limits, works offline.
30B total parameters (MoE with ~3B active), a Q4 file is about 18 GB
NVIDIA GeForce RTX 4060 Ti 16GB
16 GB VRAM + 32 GB RAM
~$450 NEWQ4_K_M with partial offload or tight context, ~25-35 tok/s at 8k context
NVIDIA GeForce RTX 3090 24GB
24 GB VRAM
~$750 USEDQ4_K_XL or Q5_K_M fully in VRAM, ~60-80 tok/s at 32k context
Apple Mac Studio M2 Ultra (64GB)
64 GB UNIFIED MEMORY
~$2600 USEDQ8 / BF16 unquantized weights, full context handling up to 128k
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long GLM 4.7 Flash takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNVIDIA GeForce RTX 4060 Ti 16GB | SWEET SPOTNVIDIA GeForce RTX 3090 24GB | FULL POWERApple Mac Studio M2 Ultra (64GB) |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~17 S★★★★★ | ~7 S★★★★★ | —★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~50 S★★★★★ | ~21 S★★★★★ | —★★★★★ |
| Build a websitea small landing page, markup and styles together | ~4.4 MIN★★★★★ | ~1.9 MIN★★★★★ | —★★★★★ |
| Build a backendan API with routes, storage and tests | ~14 MIN★★★★★ | ~6 MIN★★★★★ | —★★★★★ |
| Build a gamea playable browser game in one file | ~8.3 MIN★★★★★ | ~3.6 MIN★★★★★ | —★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. A DASH MEANS NO SPEED WAS RESEARCHED FOR THAT RIG. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about GLM 4.7 Flash yet to call it. We found 150 posts but people did not say much either way.
114 POSTS · 36 COMMENTS · REDDIT · LEMMY · X · GITHUB · DEV FORUMS · HACKER NEWS
Other models the crowd has fully reviewed, starting with text models like this one.
Car Wash Test on 53 leading models: “I want to wash my car. The car wash is 50 meters away. Should I walk or drive?”
some uncensored models
Free Model List (API Keys)
I ran 8 open-weight models as agents in a persistent MMO for 10 days. Here's the 93k event dataset and some things that I learned
Z.ai Launches GLM-4.7-Flash: 30B Coding model & 59.2% SWE-bench verified in benchmarks
I benchmarked 17 local LLMs on real MCP tool calling — single-shot AND agentic loop. The difference is massive.
I love Mistral
why doesn’t Copilot host high-quality open-source models like GLM 4.7 or Minimax M2.1 and price them with a much cheaper multiplier, for example 0.2?
CachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful
Devstrale 2 > other Chinese AIs like DeepSeek etc
Glm 4.7 vs Deepseek (model of your choice)
192GB gang - what are you running?