LOADING DATASHEET
LOADING DATASHEET
by Zhipu
RANKED #179 OF 267 ASSISTANTS · OVERALL #286 OF 6,544 · VIBE SCORE 5.6 · 33 VOICES
GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understandin
5 mentions
1 mentions
26 weeks · 34 voices
QUIETLY DEGRADING
Recent vibe and chatter are both sliding.
Open weights. You can use a hosted service, or download it and run it yourself, free.
Open weights: download GLM 4.6V once and it's yours. No subscription, no rate limits, works offline.
9B parameters (Flash) / 106B parameters (Full), a 9B Q4 file is about 5.5 GB
NVIDIA GeForce RTX 3060 12GB
12 GB VRAM
~$250 USEDGLM-4.6V-Flash (9B) Q4_K_M quantization, ~25 tok/s, 8k context
NVIDIA GeForce RTX 4070 Ti Super 16GB
16 GB VRAM
~$780 NEWGLM-4.6V-Flash (9B) FP16 or Q8 quantization, ~45 tok/s, 32k context
Apple Mac Studio M2 Ultra (128GB Unified Memory)
128 GB UNIFIED MEMORY
~$3500 USEDGLM-4.6V 106B foundation model Q4 quantization, ~12 tok/s, 32k context
The three cards above are researched picks. Search the machine you actually have — we only say yes if it should reply at a usable speed, not just load the weights and crawl.
HARDWARE PICKS & PRICES ARE RESEARCHED FROM THE LIVE WEB AND REFRESHED AUTOMATICALLY. TREAT THEM AS BALLPARK, NOT GOSPEL.
Everyday jobs on each of those rigs: how long GLM 4.6V takes, and how good the crowd says it is at that kind of work.
| THE JOB | BARE MINIMUMNVIDIA GeForce RTX 3060 12GB | SWEET SPOTNVIDIA GeForce RTX 4070 Ti Super 16GB | FULL POWERApple Mac Studio M2 Ultra (128GB Unified Memory) |
|---|---|---|---|
| Write an emaila paragraph or two, drafted from a one-line brief | ~20 S★★★★★ | ~11 S★★★★★ | ~42 S★★★★★ |
| Summarize a documenta long report boiled down to the points that matter | ~60 S★★★★★ | ~33 S★★★★★ | ~2.1 MIN★★★★★ |
| Build a websitea small landing page, markup and styles together | ~5.3 MIN★★★★★ | ~3 MIN★★★★★ | ~11 MIN★★★★★ |
| Build a backendan API with routes, storage and tests | ~17 MIN★★★★★ | ~9.3 MIN★★★★★ | ~35 MIN★★★★★ |
| Build a gamea playable browser game in one file | ~10 MIN★★★★★ | ~5.6 MIN★★★★★ | ~21 MIN★★★★★ |
STARS ARE THE COMMUNITY SCORE FOR THAT KIND OF WORK, DOCKED FOR HOW SQUASHED THE WEIGHTS GET AT EACH TIER. TIMES ARE COMPUTED FROM RESEARCHED SPEEDS. BALLPARK, NOT GOSPEL.
Not enough discussion about GLM 4.6V yet to call it. We found 33 posts but people did not say much either way.
27 POSTS · 6 COMMENTS · X · REDDIT · GITHUB · LEMMY · BLUESKY · BLOG
Other models the crowd has fully reviewed, starting with multimodal models like this one.
Hey guys, we did a major refresh of quants (quality of life updates) for GLM 4.5, 4.6, 4.6V-Flash and 4.7 llama.cpp and other inference engines like LM Studio now support more features including but not limited to: 1. Non ascii decoding for tools (affects non…
Hey everyone just wanted to give you guys a large update we did a lot of GGUFs in the past few days: GLM-4.6V (new) and Flash was updated with vision support thanks to llama.cpp
I benchmarked 17 local LLMs on real MCP tool calling — single-shot AND agentic loop. The difference is massive.
Z.ai releases GLM-4.6V: A 9B "Flash" model that beats Qwen2-VL-8B,128k context and completely FREE via API.
Housekeeping: I’m traveling so cannot make a voiceover for this post. EDIT — I added a bullet point 5 on the Chinese data industry after sending the email out. Today, Z.ai announced their GLM-5.3 model, currently only available in the coding plan, coming soon…
Dell seems to be the first to realise we don't actually care about AI PCs
open source Zhipu AI GLM-4-9B-Chat tops hallucination leaderboard
Enable auto execution tools
Do any of you use open weight models like DeepSeek and Kimi-K2 for coding and frontend dev?
fix(zai): add provider and current model catalog
ran 120+ benchmarks testing LLM retrieval, here's what i found
fix(core): recognize new DeepSeek/GLM vision models in modality auto-detection