Qwen/Qwen3.8-2.4T-A95B

Aug 12, 2026 10:01 PM - 1 hour ago 1

Qwen Studio

This repository contains exemplary weights and configuration files for the post-trained exemplary successful the Hugging Face Transformers format.

These artifacts are compatible pinch vLLM, SGLang, TokenSpeed, etc.

For users seeking managed, scalable conclusion without infrastructure maintenance, the charismatic Qwen API work is provided by Qwen Cloud.

In particular, Qwen3.8-Max is the charismatic type based connected Qwen3.8-2.4T-A95B pinch much features, specified arsenic imagination input & non-thinking support, 1M discourse magnitude by default, charismatic built-in tools, etc. For much information, please mention to the Qwen3.8-Max Overview.

Following the wide organization take of the Qwen3.5 and Qwen3.6 series, we are pleased to present Qwen3.8, the astir tin procreation successful the Qwen open-model family to date.

For the first time, Qwen3.8 brings a Qwen-Max-class exemplary to unfastened release. Built connected the architectural instauration of Qwen3.5, Qwen3.8 delivers important gains crossed coding, master work, research, and long-horizon agentic tasks. Beyond answering harder questions, Qwen3.8 is designed to transportation complex, multi-step tasks done to completion pinch greater reliability.

Qwen3.8 Highlights

Qwen3.8 features the pursuing enhancements:

  • Core Capabilities: Comprehensive improvements crossed coding, master work, research, and long-horizon agentic tasks.
  • Agent Execution: Stronger autonomous readying and amended handling of situation feedback, starring to much reliable end-to-end task completion.
  • Downstream Compatibility: Broader support for celebrated harnesses and improvement tools, making it easier to merge into your existing stack.
  • Flexible Thinking Control: Reasoning extent tin beryllium tuned pinch reasoning_effort, and reasoning discourse from humanities messages is retained via preserve_thinking.

For much details, please mention to our blog station Qwen3.8-Max.

Model Overview

  • Type: Causal Language Model
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 2.4T successful full and 95B activated
    • Hidden Dimension: 8192
    • Token Embedding: 248,320 (Padded)
    • Number of Layers: 92
    • Hidden Layout: 23 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 128 for V and 16 for QK
      • Head Dimension: 128
    • Gated Attention:
      • Number of Attention Heads: 64 for Q and 4 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
    • Mixture of Experts:
      • Number of Experts: 512
      • Number of Activated Experts: 10 Routed + 1 Shared
      • Expert Intermediate Dimension: 2048
    • LM Output: 248,320 (Padded)
    • MTP (Multi-Token Prediction): trained pinch aggregate steps
  • Context Length: 262,144 natively and extensible up to 1,010,000 tokens.

Benchmark Results

Opus 4.8Fable 5GPT 5.6 Sol (max)Qwen3.7-MaxQwen3.8-Max
Coding Agent
Terminal Bench 2.1 84.6 84.6 88.8 74.5 86.6
SWE-bench Pro 69.2 80.0 64.6 60.6 67.7
DeepSWE 1.1 59.0 70.0 73.0 21.6 56.6
NL2Repo-Bench 69.4 -- -- 47.2 55.9
FrontierSWE 70.0 88.8 -- 40.7 73.5
MLS-Bench-Lite 42.8 49.9 46.2 31.7 41.0
PaperBench 80.3 88.8 90.5 64.8 93.0
AndroidBench 69.8 84.5 74.0 56.5 75.1
QwenSWEBench 84.0 86.3 73.5 63.4 80.7
QwenQoderBench 62.7 63.1 53.8 36.8 58.4
QwenReactBench 1694 1770 1564 1538 1724
QwenSVGBench 1648 1690 1758 1499 1713
General Agent
CoWorkBench 72.3 75.9 71.5 64.6 74.8
WorkSpaceBench 66.8 68.7 65.6 61.4 67.7
JobBench 48.4 57.4 45.4 31.3 53.4
SkillsBench 65.1 70.9 73.5 61.2 70.2
Agents' Last Exam (Pass / Score) 27.0 / 45.1 -- / -- 30.6 / 53.6 11.8 / 31.1 27.0 / 52.4
Automation-Bench (Pass@1) 27.2 29.1 29.7 14.2 27.3
Toolathlon Verified (Pass@1) 76.2 77.9 74.9 49.7 72.5
WideSearch 72.9 81.2 -- 75.2 81.9
HLE w/ tools 57.9 64.5 58.0 53.5 56.2
General Capabilities
GPQA Diamond 92.0 92.6 94.1 92.4 92.6
HLE 45.7 53.3 47.2 41.4 43.6
IFBench 62.2 63.5 72.7 79.1 82.8
$OneMillion-Bench (expert score) 41.8 55.9 53.8 44.4 52.5
HealthBench 52.4 -- 55.3 54.5 60.2
PLawBench 69.6 70.2 72.3 58.9 73.2
PRBench-Legal 52.7 57.6 57.6 48.5 57.6
PRBench-Finance 51.9 55.8 55.5 46.8 58.3
MRCR v2 256K (8-needle) 83.2 -- 93.8 86.7 92.9
LongBench v2 69.1 -- 67.1 65.3 66.3

1. Fable5 results whitethorn impact fallbacks.
2. Terminal Bench 2.1: Evaluated pinch Claude Code (avg@10), utilizing a 5-hour timeout and max_tokens=131,072. For each different models, we study the champion published people crossed harnesses: Claude Opus 4.8 and Claude Fable 5 pinch Terminus 2 from Artificial Analysis (https://artificialanalysis.ai/evaluations/terminalbench-v2-1); GPT-5.6 Sol pinch Codex (https://openai.com/index/previewing-gpt-5-6-sol/).
3. SWE-bench Pro: Evaluated pinch the Claude Code harness, temp=1.0, top_p=0.95, and a 256K discourse window. Problematic tasks corrected and each baselines evaluated connected the refined benchmark.
4. DeepSWE 1.1: Evaluated pinch the Claude Code and mini-SWE-agent harnesses, temp=1.0, top_p=0.95, and a 256K discourse window. We study the highest people among some harnesses; notably, Qwen3.8-Max performs champion connected Claude Code.
5. NL2Repo-Bench: Evaluated pinch the Claude Code harness. To forestall reward hacking, we disable Bash commands that effort to entree the circumstantial repository, specified arsenic pip download, pip install, and git clone.
6. FrontierSWE: Evaluated pinch the Claude Code harness. All different disposable MEAN@5 results are taken from the charismatic FrontierSWE leaderboard (https://www.frontierswe.com) arsenic of August 3, 2026. Dominance scores are recomputed from the earthy scores utilizing the charismatic information script. "--" indicates that nary charismatic MEAN@5 consequence was disposable arsenic of that date.
7. MLS-Bench-Lite: Evaluated pinch Claude Code utilizing a 5-hour timeout and max_tokens=131,072. All different exemplary scores are taken from the charismatic leaderboard.
8. PaperBench: Evaluated successful the BasicAgent mounting nether Code-Dev mode, judged by Claude Opus 4.6, and averaged complete 3 runs (max 12 hours per run).
9. AndroidBench: Evaluated connected the 95-task nationalist subset, reporting avg@3 scores.
10. QwenSWEBench: Inhouse coding benchmark to measure models' package engineering capabilities. Evaluated pinch the Claude Code harness. Reporting avg@3 pinch an 8-hour timeout, max_tokens=32,768, temperature=1.0, and a 256K-token discourse window.
11. QwenQoderBench: Inhouse coding benchmark to measure personification acquisition connected Qoder. Evaluated pinch the Claude Code harness. Reporting avg@5 pinch a 6-hour timeout, max_tokens=32,768, temperature=1.0, and a 256K-token discourse window.
12. QwenReactBench: Inhouse React task building benchmark utilizing Claude Code arsenic the harness, bilingual (EN/CN), 7 categories; auto-render + multimodal judge; BT/Elo rating.
13. QwenSVGBench: Inhouse SVG codification procreation benchmark; bilingual (EN/CN), auto-render + multimodal judge; BT/Elo rating.
14. CoWorkBench: Inhouse cowork benchmark for evaluating long-horizon tasks crossed machine science, finance, law, medical, and different productivity domains.
15. SkillsBench: Evaluated connected the nationalist SkillsBench v1.1 benchmark crossed 87 tasks, reporting the mean people complete 3 runs per task. Opus 4.8 and Fable 5 are evaluated connected Claude Code; GPT-5.6 Sol is evaluated connected Codex; the Qwen-series are evaluated connected OpenCode. All results are from our ain testing.
16. Automation-Bench: Evaluated connected the 600-task nationalist subset.
17. WideSearch: Evaluated pinch the Claude Code harness for outer models and the Qwen-Agent harness for ours, reporting the mean item-F1 complete 4 runs.
18. $OneMillion-Bench: Evaluated utilizing gemini-3.1-pro-preview.
19. PLawBench: Evaluated utilizing gemini-3.1-pro-preview.
20. Empty cells (--): Scores are not yet disposable aliases are not applicable.

Quickstart

For streamlined integration, we urge utilizing Qwen3.8 via APIs.

Serving Qwen3.8

Inference ratio and throughput alteration importantly crossed frameworks. We urge utilizing the latest model versions to guarantee optimal capacity and compatibility. For accumulation workloads aliases high-throughput scenarios, dedicated serving engines specified arsenic SGLang, vLLM, aliases TokenSpeed are recommended.

Qwen3.8 tin beryllium deployed pinch celebrated conclusion frameworks, e.g.:

  • SGLang: Qwen3.8 Cookbook
  • vLLM: Qwen3.8 Recipe
  • TokenSpeed: Qwen3.8 Recipe

API Usage

Qwen3.8-2.4T-A95B is simply a text-only exemplary that requires reasoning mode for each interactions. Multimodal inputs are not supported, and reasoning cannot beryllium disabled. Every consequence will automatically statesman pinch reasoning enclosed successful <think>\n...</think>\n\n earlier the last output.

We urge utilizing the pursuing group of sampling parameters for generation:

  • temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0

Please statement that the support for sampling parameters varies according to conclusion frameworks.

Qwen3.8 comes pinch charismatic support for reasoning_effort, which tin beryllium utilized to set reasoning extent and power cost:

  • xhigh (default): for analyzable tasks demanding thorough analysis
  • medium: balancing accuracy and speed
  • low: businesslike reasoning optimizing for velocity and cost

In addition, preserve_thinking is enabled by default for each workloads for the champion out-of-the-box experience.

Chat Completions API

The Chat Completions API tin beryllium utilized pinch astir conclusion frameworks, arsenic good arsenic Qwen Cloud. Before starting, make judge the OpenAI Python SDK is installed and the API cardinal and the API guidelines URL are configured, e.g.:

pip instal -U openai # Set the pursuing accordingly export OPENAI_BASE_URL='your-base-url' export OPENAI_API_KEY='your-api-key'
Text-Only Input
from openai import OpenAI client = OpenAI() messages = [{"role": "user", "content": "Write a Python usability to merge 2 sorted linked lists."}] completion = client.chat.completions.create( model="Qwen/Qwen3.8-2.4T-A95B", messages=messages, extra_body={ "chat_template_kwargs": { "enable_thinking": True, "preserve_thinking": True, }, }, reasoning_effort="xhigh", stream=True, stream_options={"include_usage": True}, ) reasoning_content = "" answer_content = "" is_answering = False print("\n" + "=" * 20 + "Reasoning" + "=" * 20 + "\n") for chunk in completion: if not chunk.choices: print("\nUsage:") print(chunk.usage) continue delta = chunk.choices[0].delta if hasattr(delta, "reasoning_content") and delta.reasoning_content is not None: if not is_answering: print(delta.reasoning_content, end="", flush=True) reasoning_content += delta.reasoning_content if hasattr(delta, "content") and delta.content: if not is_answering: print("\n" + "=" * 20 + "Answer" + "=" * 20 + "\n") is_answering = True print(delta.content, end="", flush=True) answer_content += delta.content

If you are utilizing APIs from Qwen Cloud, successful summation to changing model, please walk extra_body={"enable_thinking": True, "preserve_thinking": True} alternatively of extra_body={"chat_template_kwargs": {"enable_thinking": True, "preserve_thinking": True}}.

Best Practices

To execute optimal performance, we urge the pursuing settings:

  1. Sampling Parameters:

    • We propose utilizing the pursuing group of sampling parameters:
      • temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
    • For supported frameworks, you tin set the presence_penalty parameter betwixt 0 and 2 to trim endless repetition. However, utilizing a higher worth whitethorn occasionally consequence successful connection mixing and a flimsy alteration successful exemplary performance.
  2. Adequate Output Length: To optimize capacity connected agentic tasks, we urge allocating capable output magnitude to let the exemplary to make elaborate and broad responses. For frameworks that support abstracted token limits for soul reasoning and last outputs, we propose the pursuing configuration wrong the 1M discourse length:

    • Reasoning Content: Set the maximum output magnitude to 262,144 tokens.
    • Final Response: Set the maximum output magnitude to 131,072 tokens.

    These settings supply the basal capacity for analyzable reasoning while ensuring ample abstraction for high-quality last deliverables.

Citation

If you find our activity helpful, consciousness free to springiness america a cite.

@misc{qwen38, title = {{Qwen3.8-Max}: A New Bar for Coding and Cowork}, url = {https://qwen.ai/blog?id=qwen3.8}, writer = {{Qwen Team}}, period = {August}, twelvemonth = {2026} }
More