Catalog
Model by z-ai
glm-5-3-flash
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.
NVIDIA modelchatMultimodalReasoningChatText-to-TextImage-to-Text
Model by z-ai
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.