close

deepseek-ai/DeepSeek-V3.2

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:4Model Size:685BQuant:FP8Context Size:128kTool Calling:SupportedPublished:Dec 1, 2025License:mitArchitecture:Transformer1.5K Open Weights Warm

DeepSeek-V3.2 is a 685 billion parameter language model developed by DeepSeek-AI, featuring a 32,768 token context length. It integrates DeepSeek Sparse Attention (DSA) for efficient long-context processing and a scalable reinforcement learning framework. The model excels in complex reasoning and agentic tasks, with a specialized variant, DeepSeek-V3.2-Speciale, demonstrating performance comparable to or surpassing GPT-5 and Gemini-3.0-Pro in mathematical and informatics olympiads.

Loading preview...

DeepSeek-V3.2: Efficient Reasoning & Agentic AI

DeepSeek-V3.2, developed by DeepSeek-AI, is a 685 billion parameter model designed for high computational efficiency and superior performance in reasoning and agentic tasks. It incorporates several technical innovations to achieve its capabilities, particularly in long-context scenarios and complex problem-solving.

Key Capabilities

  • DeepSeek Sparse Attention (DSA): An efficient attention mechanism that reduces computational complexity while maintaining performance, especially optimized for long contexts up to 32,768 tokens.
  • Scalable Reinforcement Learning: Utilizes a robust RL framework and scaled post-training compute, enabling its high-compute variant, DeepSeek-V3.2-Speciale, to achieve reasoning proficiency on par with or exceeding models like GPT-5 and Gemini-3.0-Pro.
  • Agentic Task Synthesis: Features a novel pipeline for generating large-scale training data, improving the model's compliance and generalization in tool-use and interactive agent environments.
  • Exceptional Reasoning: The DeepSeek-V3.2-Speciale variant has demonstrated gold-medal performance in the 2025 International Mathematical Olympiad (IMO) and International Olympiad in Informatics (IOI).
  • Updated Chat Template: Introduces a revised chat template with enhanced tool calling and a "thinking with tools" capability, supported by provided Python scripts for encoding and parsing messages.

Good For

  • Complex Reasoning Tasks: Ideal for applications requiring advanced logical deduction, problem-solving, and mathematical reasoning, particularly with the "Speciale" variant.
  • Agentic AI Development: Suitable for building sophisticated AI agents that require robust tool-use integration and interaction within complex environments.
  • Long-Context Applications: Benefits from its optimized sparse attention mechanism for processing and understanding extensive textual inputs.
  • Research and Verification: The release includes final submissions for major olympiads, allowing the community to conduct secondary verification of its reasoning capabilities.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p