close

DEV Community

#moe

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Your Intel Laptop Can Run 30B Models Now. No NVIDIA. No Cloud. No Problem.

Your Intel Laptop Can Run 30B Models Now. No NVIDIA. No Cloud. No Problem.

Comments
11 min read
How a 176 KB C Binary Runs a 2.78-Trillion-Parameter Model on One CPU with 8 GB of RAM

How a 176 KB C Binary Runs a 2.78-Trillion-Parameter Model on One CPU with 8 GB of RAM

Comments
4 min read
NVIDIA Unveils Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4: A Deployment-Optimized Hybrid MoE LLM

NVIDIA Unveils Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4: A Deployment-Optimized Hybrid MoE LLM

Comments
4 min read
2 TB of Ukrainian Law + DeepSeek V3 860B on GCP: What We'd Get

2 TB of Ukrainian Law + DeepSeek V3 860B on GCP: What We'd Get

Comments
7 min read
Scaling MoE Models with LongCat-2.0: A Deep Dive into 1.6T Parameter Architecture Design

Scaling MoE Models with LongCat-2.0: A Deep Dive into 1.6T Parameter Architecture Design

Comments
3 min read
Step 3.7 Flash is a drop-in — except for one endpoint detail

Step 3.7 Flash is a drop-in — except for one endpoint detail

1
Comments
9 min read
Kimi K2.6 for Local AI in 2026: What VRAM and System RAM You Need to Actually Run the 1T-Parameter MoE Coding Leader

Kimi K2.6 for Local AI in 2026: What VRAM and System RAM You Need to Actually Run the 1T-Parameter MoE Coding Leader

Comments
6 min read
I built a Rust inference engine that streams MoE expert weights from NVMe SSDs, no GPU required

I built a Rust inference engine that streams MoE expert weights from NVMe SSDs, no GPU required

Comments
2 min read
ZAYA1-8B: a 760M-active MoE trained on AMD MI300x

ZAYA1-8B: a 760M-active MoE trained on AMD MI300x

Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.