- Status
- Verified
- Trending
- #41
- Downloads, 30 days
- 1.3k
- Weights
- 328 GB
- Sources
- 2
- Revision
- Manifest
GLM-5.3 with 34% of its experts removed, at INT4 — 328 GB, built to serve on 4× H200 (Hopper) in vLLM. Altar-1 was calibrated on cybersecurity traces, coding, tool calling, reasoning, and English.
At a glance
- Task
- Text generation
- Input
- text
- Output
- text
- Parameters
- 501B, 168 experts, 8 active
- Architecture
- GLM MoE Dsa
- Context
- 1M tokens
- Precision
- INT32 95%, BF16 5%
- Format
- Safetensors
- License
- other
- Base model
- Quantized from cyankiwi/GLM-5.3-AWQ-INT4
- Released
- Sep 2026
- Updated
- Sep 2026
- Likes
- 102
- Downloads, all time
- 484
Architecture
- Layers
- 78
- Hidden size
- 6,144
- Attention
- 64 heads
- Experts
- 168 total, 8 active per token
- Vocabulary
- 154,880
- Positions
- 1,048,576
- Tied embeddings
- No
- Quantization
- compressed-tensors
Run it
Pinned to the indexed revision.
hf download AikidoSec/altar-1 --revision 5d591cd98958e4e7e429517887e80ca3bb7ce2c4Through the hub: the same tools, each file from a source that is up (Hugging Face, ModelScope, IPFS), at this revision. The second line checks every file against its address.
export HF_ENDPOINT=https://gethologram.ai
cd "$(hf download AikidoSec/altar-1 --quiet)" && curl -s $HF_ENDPOINT/AikidoSec/altar-1/resolve/main/SHA256SUMS | sha256sum -c --quietPapers
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compressionMike Lasby et al., 2025
Read the full model card
Altar-1 — a 504B parameter Prune of GLM-5.3
GLM-5.3 with 34% of its experts removed, at INT4 — 328 GB, built to serve on 4× H200 (Hopper) in vLLM. Altar-1 was calibrated on cybersecurity traces, coding, tool calling, reasoning, and English. Additionally we used multi-lingual wikipedia articles.
What this is
GLM-5.3 is a 753B mixture-of-experts model: each token uses 8 of 256 expert sub-networks per layer (~40B active). REAP (Router-weighted Expert Activation Pruning) scores each expert’s real contribution and deletes the least useful ones — no retraining. This cut keeps 168 of 256 experts per layer.
The experts are then INT4 W4A16 (compressed-tensors, AWQ), taken from the cyankiwi/GLM-5.3-AWQ-INT4 base. Only the routed experts are 4-bit; attention, the shared expert, the dense layers, and the head stay BF16. vLLM auto-selects the Marlin MoE kernel. Routing is untouched: 8 experts per token out of the 168 that remain, ~40B active parameters, same as the unpruned model.
How close to the original is it?
KL divergence vs full BF16: 0.506 nats (sealed 25-prompt panel, full 154k vocabulary). KL is the standard “how differently do these two models predict” score — 0 = identical, lower = closer. For reference, an EXL3 build of the same 168-expert cut measures 0.511 — at this bit-width the quantization format barely moves the result. Full numbers: fidelity study.
Why these experts
Instead of keeping the globally most-frequent experts (which deletes a domain’s specialists), each expert is scored by its largest share of any single domain’s routed work, so every domain — code, rare languages, structured output — keeps its specialists. Head-to-head vs frequency pruning: fidelity study.
Serving (vLLM, 4× H200)
vllm serve aikido/altar-1 --tensor-parallel-size 4 --trust-remote-code --max-model-len 131072
Requires Hopper (H100/H200). 328 GB of weights across 4× H200 leaves room for a 128k-context KV cache at production batch sizes; vLLM selects the Marlin MoE kernel automatically.
Credits
- Z.AI / zai-org — GLM-5.3, the base model.
- cyankiwi — the GLM-5.3-AWQ-INT4 W4A16 base this prune is built on.
- Cerebras Research — REAP (arXiv:2510.13999).
- 0xSero — performed the REAP prune, and released the 569B and EXL3 builds of the same cut.
Observations: glm-5.3-reap-observations-v1 · Fidelity study: glm-5.3-reap-fidelity-study · Built on 8× NVIDIA RTX PRO 6000 Blackwell.
Deploying Altar
To get help deploying this model to your organization, contact yannick@aikido.dev
If you want to put this model to the test, some of Aikido's products are already powered by Altar, try them today:
License
Inherits the GLM-5.3 license.
Derived on Sep 22, 2026 from Hugging Face at revision 5d591cd9, README.md , config.json .
48 files, 328 GB. Every download is checked against its address.