Tag
9 articles
This article explains the advanced concepts behind Alibaba's Qwen3.8-Max, a 2.4 trillion-parameter multimodal model, including multimodal capabilities, mixture-of-experts architecture, and parameter scaling effects.
This article explains Mixture of Experts (MoE) AI models, how they work like teams of specialists, and why they're important for efficient AI performance.
Soofi Consortium releases Soofi S 30B-A3B, an open hybrid Mamba-Transformer MoE model for German and English. The model leverages 3.2 billion active parameters out of 31.6 billion for efficient multilingual processing.
NVIDIA introduces Nemotron-Labs-3-Puzzle-75B-A9B, a compressed hybrid MoE LLM delivering 2.03x server throughput, leveraging hardware-aware compression and knowledge distillation.
JetBrains has released Mellum2, a 12-billion parameter MoE model trained on 10.6 trillion tokens, designed to accelerate specialized AI tasks in multi-model pipelines.
This explainer explores the advanced Mixture of Experts (MoE) architecture used in Liquid AI's LFM2.5-8B-A1B model, examining how sparse parameter activation enables powerful on-device AI capabilities.
This article explains the advanced AI concepts behind Qwen 3.6-35B-A3B, a multimodal model that combines MoE routing, RAG, and session persistence for intelligent, context-aware AI applications.
Alibaba's Qwen team open-sources Qwen3.6-35B-A3B, a sparse MoE vision-language model with 3B active parameters and agentic coding capabilities.
This explainer article dives into NVIDIA's Nemotron-Cascade 2, an advanced Mixture-of-Experts (MoE) model that demonstrates how strategic parameter allocation can enhance reasoning capabilities while maintaining computational efficiency.