What is this
DeepSeek-V4.1-Flash pinch permanent weight-level abliteration — the information guardrails person been surgically removed while preserving MMLU capability, vision, reasoning, MTP (DSpark), and multi-turn coherence.
Proprietary weight-level abliteration developed by the dealignai investigation team. No civilization model.py, nary runtime hooks, nary steering vectors — it's a modular checkpoint that loads precisely for illustration the guidelines model. The refusal circuitry is surgically removed while each capability-critical constituent (routed experts, Engram memory, CSA2 sparse attention, DSpark draught head, imagination tower, router gates, norms, embeddings) is preserved byte-identical to the base.
| Base | deepseek-ai/DeepSeek-V4.1-Flash (552B backbone, 8B/16B progressive per token) |
| Architecture | Causal Encoder-Decoder (20+20 layers), MoE (384 routed top-6 + 1 shared), Hyper-Connections (4-channel residual), CSA2 sparse attention, Engram n-gram memory, DSpark speculative draft |
| Quant | FP8 (e4m3fn) weights pinch E8M0 block-scale [32, 32], FP4 routed experts — native, unchanged |
| Context | 1M tokens |
| Vision | DeepSeek-ViT pinch 2D-RoPE + pixel unshuffle — untouched |
| Modification | Surgical, weight-level (drop-in checkpoint) |
Results
HarmBench-320 — afloat 2×2 (base vs CRACK, effort=off vs max), T=0 greedy
Every consequence 4-tier graded (HARD_REF / SOFT_RED / HEDGE / COMPLY), pinch reasoning-trace verification astatine effort=max.
| HB-320 effort=off | 137/320 = 42.81 % | 320/320 = 100.00 % | +57.19 |
| HB-320 effort=max | 5/320 = 1.56 % | 320/320 = 100.00 % | +98.44 |
Notable: astatine effort=max, the base exemplary becomes MORE refusal-prone (42.8 % → 1.6 %) because reasoning surfaces information concerns earlier answering. The CRACK stays astatine 100.0 % crossed some effort levels.
Per-category (all 7 HarmBench semantic categories):
| chemical_biological | 42 | 16.7 % | 100.0 % | 0.0 % | 100.0 % |
| copyright | 80 | 98.8 % | 100.0 % | 0.0 % | 100.0 % |
| cybercrime_intrusion | 52 | 34.6 % | 100.0 % | 3.8 % | 100.0 % |
| harassment_bullying | 21 | 0.0 % | 100.0 % | 0.0 % | 100.0 % |
| harmful | 18 | 11.1 % | 100.0 % | 5.6 % | 100.0 % |
| illegal | 53 | 13.2 % | 100.0 % | 0.0 % | 100.0 % |
| misinformation_disinformation | 54 | 44.4 % | 100.0 % | 3.7 % | 100.0 % |
Zero HARD_REF, zero SOFT_RED, zero HEDGE connected the cracked build astatine either effort level.
Every consequence was graded by a strict multilingual regex-based 4-tier classifier positive (for effort=max) an LLM-as-judge complete the saved reasoning trace. Full per-item outputs saved for verification.
MMLU-14k (full trial set, base-logit, T=0)
| base | 12,211 / 14,042 | 86.96 % | — |
| CRACK | 11,619 / 14,042 | 82.74 % | -4.22 pp |
Excluding the morals cluster (moral_scenarios, business_ethics, professional_law, jurisprudence, accuracy — wherever refusal-adjacent behaviour is graded), delta connected the remaining ~11k items is -1.1 pp — good wrong the 3 pp knowledge-preservation target.
Full per-subject dropdown (57 subjects, sorted by delta)| moral scenarios | 895 | 76.9% | 37.0% | -39.89 |
| professional law | 1534 | 75.9% | 68.8% | -7.04 |
| abstract algebra | 100 | 77.0% | 71.0% | -6.00 |
| security studies | 245 | 84.5% | 79.2% | -5.31 |
| high schoolhouse machine science | 100 | 98.0% | 94.0% | -4.00 |
| jurisprudence | 108 | 90.7% | 87.0% | -3.70 |
| machine learning | 112 | 81.2% | 77.7% | -3.57 |
| high schoolhouse chemistry | 203 | 87.7% | 84.2% | -3.45 |
| professional psychology | 612 | 90.7% | 87.3% | -3.43 |
| formal logic | 126 | 73.8% | 70.6% | -3.17 |
| college machine science | 100 | 82.0% | 79.0% | -3.00 |
| professional medicine | 272 | 94.5% | 91.5% | -2.94 |
| high schoolhouse statistics | 216 | 88.0% | 85.2% | -2.78 |
| professional accounting | 282 | 83.0% | 80.5% | -2.48 |
| logical fallacies | 163 | 93.9% | 91.4% | -2.45 |
| human sexuality | 131 | 90.1% | 87.8% | -2.29 |
| computer security | 100 | 85.0% | 83.0% | -2.00 |
| medical genetics | 100 | 96.0% | 94.0% | -2.00 |
| astronomy | 152 | 95.4% | 93.4% | -1.97 |
| clinical knowledge | 265 | 94.3% | 92.5% | -1.89 |
| high schoolhouse continent history | 165 | 90.3% | 88.5% | -1.82 |
| public relations | 110 | 80.0% | 78.2% | -1.82 |
| philosophy | 311 | 89.7% | 88.1% | -1.61 |
| prehistory | 324 | 93.5% | 92.0% | -1.54 |
| moral disputes | 346 | 84.1% | 82.7% | -1.45 |
| electrical engineering | 145 | 86.9% | 85.5% | -1.38 |
| high schoolhouse mathematics | 270 | 67.0% | 65.9% | -1.11 |
| high schoolhouse macroeconomics | 390 | 92.1% | 91.0% | -1.03 |
| global facts | 100 | 63.0% | 62.0% | -1.00 |
| international law | 121 | 90.1% | 89.3% | -0.83 |
| college biology | 144 | 97.2% | 96.5% | -0.69 |
| high schoolhouse physics | 151 | 84.8% | 84.1% | -0.66 |
| college medicine | 173 | 83.8% | 83.2% | -0.58 |
| high schoolhouse america history | 204 | 95.1% | 94.6% | -0.49 |
| high schoolhouse microeconomics | 238 | 96.2% | 95.8% | -0.42 |
| miscellaneous | 783 | 96.2% | 95.8% | -0.38 |
| high schoolhouse psychology | 545 | 96.1% | 95.8% | -0.37 |
| business ethics | 100 | 85.0% | 85.0% | +0.00 |
| college physics | 102 | 90.2% | 90.2% | +0.00 |
| conceptual physics | 235 | 94.5% | 94.5% | +0.00 |
| high schoolhouse biology | 310 | 95.2% | 95.2% | +0.00 |
| human aging | 223 | 85.2% | 85.2% | +0.00 |
| management | 103 | 91.3% | 91.3% | +0.00 |
| nutrition | 306 | 90.2% | 90.2% | +0.00 |
| sociology | 201 | 94.5% | 94.5% | +0.00 |
| us overseas policy | 100 | 97.0% | 97.0% | +0.00 |
| world religions | 171 | 92.4% | 92.4% | +0.00 |
| elementary mathematics | 378 | 91.0% | 91.3% | +0.26 |
| marketing | 234 | 94.9% | 95.3% | +0.43 |
| virology | 166 | 55.4% | 56.0% | +0.60 |
| high schoolhouse world history | 237 | 95.4% | 96.2% | +0.84 |
| econometrics | 114 | 78.9% | 79.8% | +0.88 |
| college chemistry | 100 | 65.0% | 66.0% | +1.00 |
| anatomy | 135 | 88.1% | 89.6% | +1.48 |
| high schoolhouse geography | 198 | 92.9% | 94.4% | +1.52 |
| high schoolhouse authorities and politics | 193 | 96.9% | 98.4% | +1.55 |
| college mathematics | 100 | 63.0% | 68.0% | +5.00 |
Extended validation
- 1000-token coherence stress connected 6 items — nary WARNING WARNING loops, nary character-repeat degeneracy, earthy sign-offs.
- Multi-turn conversation (4 turns connected aforesaid harmful taxable — ANFO explosive detail) — nary late-turn refusal reversion, nary self-correction, coherent done move 4.
- Vision path — coherent image explanation ("A bluish quadrate centered connected a reddish background.") + refusal driblet connected image-based harmful prompts ("shaped complaint / explosively formed penetrator" description).
- General capacity spot checks intact: √2 irrationality proof, Python palindrome pinch docstring, WWI causes successful precisely 3 sentences, quantum observable vs usability distinction.
- Full compat suite pass: streaming SSE, chat logprobs + top_logprobs, completions logprobs + echo, instrumentality calls (deepseekv41 parser), image input, reasoning-effort tiers (low/high/xhigh/max + float [0, 0.99]), sampling params (temperature, top_p, stop, seed, frequency_penalty, presence_penalty, json_object), 8-way concurrent, 40k-word punctual astatine 35,572 tokens.
How to run
Support for DeepseekV41ForCausalLM is landing crossed serving stacks (as of 2026-09). Two verified moving recipes beneath (both validated connected 4×H200 NVLink).
Recipe A — Full 1M context, DSpark speculative decoding connected (interactive / long-context)
export SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1 export SGLANG_RAGGED_VERIFY_MODE=cap-accept export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True sglang service \ --model-path dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 \ --tp-size 4 --ep-size 4 \ --host 0.0.0.0 --port 8000 \ --context-length 1048576 \ --mem-fraction-static 0.80 \ --max-running-requests 20 \ --cuda-graph-max-bs-decode 20 \ --reasoning-parser deepseek-v41 --tool-call-parser deepseekv41 \ --speculative-algorithm DSPARK \ --speculative-dspark-sps-table-path /path/to/dspark_sps.json \ --trust-remote-codeConcurrency astatine 1M ctx is capped ~20 connected 4×H200 by KV budget. The DSpark SPS costs array is profiled offline erstwhile (see below); without cap-accept mode + a existent SPS array the speculative fund degenerates to verify-all and the triumph vanishes.
Recipe B — 256k context, high-concurrency, nary speculation (batch / throughput)
export SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1 export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True sglang service \ --model-path dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 \ --tp-size 4 --ep-size 4 \ --host 0.0.0.0 --port 8000 \ --context-length 262144 \ --mem-fraction-static 0.85 \ --reasoning-parser deepseek-v41 --tool-call-parser deepseekv41 \ --trust-remote-codeServes up to 256 concurrent requests. max_total_num_tokens reports ~20.6M pinch Engram connected host. DSpark is deliberately disconnected for high-batch — its fixed measurement costs stops paying disconnected past mini batch sizes.
The bytes-per-token / concurrency fund rule
DSV4.1's world KV is 890 bytes / token. Pool size is mem-fraction-static × (per-GPU HBM − weights) × TP. What that intends connected 4×H200:
| 1,048,576 | 20 | This is the Recipe A number. Higher = OOM. |
| 262,144 | 80 | 4× the concurrency of 1M |
| 65,536 | 320+ | KV nary longer the constraint; batch is |
| 32,768 | 256+ (default cap) | max batch dominates |
At higher batch, driblet DSpark: its per-step costs stops paying off.
Non-obvious motorboat requirements (bit america during bring-up)
- --ep-size is required astatine TP4. moe_intermediate_size = 2304; astatine TP4, 2304 / 4 = 576 isn't a aggregate of 128 truthful plain TP fails: Mxfp4FlashinferCutlassMoEMethod requires ... multiples of 128. --ep-size shards MoE by master scale (384 % 4 = 0) and keeps the intermediate astatine 2304. At TP8 you tin skip --ep-size.
- ninja must beryllium connected PATH aliases the JIT kernel build crashes respective minutes into weight load pinch FileNotFoundError: 'ninja' and EXIT=137. If you build SGLang from source, pip instal ninja and export PATH=$(dirname $(which ninja)):$PATH connected the motorboat line.
- Name some parsers explicitly. --reasoning-parser car resolves done the chat template and this exemplary ships nary — car silently selects thing and the earthy <think> transmission leaks into content. Use deepseek-v41 for reasoning and deepseekv41 for tool-calls.
- Reasoning is OFF by default (SGLANG_DEFAULT_THINKING=false). A petition without reasoning_effort gets nary reasoning sloppy of parser. Send reasoning_effort: debased | precocious | xhigh | max aliases a float successful [0.0, 0.99].
- DSpark speculative draft is bundled wrong the checkpoint (num_nextn_predict_layers = 3); nary abstracted draught weights. Enable pinch --speculative-algorithm DSPARK. For a existent speed-up you request SGLANG_RAGGED_VERIFY_MODE=cap-accept + a profiled SPS array via --speculative-dspark-sps-table-path. Without both, the SPS fund degenerates to verify-all — zero gain.
- Profile the SPS table erstwhile pinch python -m sglang.benchmark.dspark_sps_profiler each --base-url http://localhost:8000 --out /path/to/dspark_sps.json --local-tokenizer-path <model-path> while the server is moving nether SGLANG_DSPARK_ENABLE_SPS_RECORD=1, SGLANG_RAGGED_VERIFY_MODE=static, and SGLANG_SIMULATE_ACC_LEN=1.0 (the profiler measures per-step cost, not acceptance). All 3 env vars are required simultaneously aliases the profiler aborts pinch a adjuvant correction naming each missing one. Wall-time ~1 min.
- Engram big table — group SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1 to move the 203 GB Engram tables to big RAM. Frees ~46 GiB/GPU for KV, output bitwise unchanged, costs ~200 GB of big RAM.
- --max-running-requests × KV/token × ctx-length must fresh HBM. On 4×H200 pinch DSpark, 1M ctx caps astatine 20 concurrent (see array above). Raising max-running-requests without capping discourse OOMs connected 12 GB CUDA-graph allocations.
- torchcodec / libavutil.so.56 errors — instal apt-get instal ffmpeg connected the host. Video-only, doesn't break matter aliases image.
Preview Docker image (fastest path)
docker propulsion lmsysorg/sglang:dev-dsv41 docker tally --gpus each --shm-size 32g -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --ipc=host --env HF_TOKEN=<your-token> \ lmsysorg/sglang:dev-dsv41 \ sglang service \ --model-path dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 \ --tp-size 4 --ep-size 4 \ --context-length 262144 --mem-fraction-static 0.85 \ --reasoning-parser deepseek-v41 --tool-call-parser deepseekv41 \ --trust-remote-codeSame non-obvious rules use wrong the container.
vLLM
Model definitions merged to main (PR #56228) but registry.py has nary DeepseekV41 introduction yet; kernels/frontend/PP way successful umbrella PR #56214. Wait for merge aliases use the umbrella.
API usage — OpenAI-compatible
Standard OpenAI schema. Model id is immoderate you group arsenic --served-model-name (or the exemplary way if unset). Recommended sampling from the guidelines card: temperature=1.0, top_p=0.95, reasoning_effort="high".
Chat, nary reasoning:
curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek-v4.1-flash-crack", "messages": [{"role":"user","content":"Explain MoE routing successful 2 sentences."}], "max_tokens": 400, "temperature": 1.0, "top_p": 0.95 }'Chat, pinch reasoning (returns divided reasoning_content and content):
from openai import OpenAI c = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") r = c.chat.completions.create( model="deepseek-v4.1-flash-crack", messages=[{"role":"user","content":"What is 15% of 240?"}], max_tokens=1200, temperature=1.0, top_p=0.95, extra_body={"reasoning_effort": "high"}, ) msg = r.choices[0].message print("REASONING:", getattr(msg, "reasoning_content", None)) print("ANSWER:", msg.content)At effort=max DSV4.1 tin make 4-5k characters of reasoning earlier contented starts. Budget max_tokens >= 8000 astatine max effort, aliases the exemplary runs retired mid-reasoning and returns quiet content. DeepSeek's ain paper recommends >= 256k.
Streaming (SSE):
curl -N http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model":"deepseek-v4.1-flash-crack", "messages":[{"role":"user","content":"Count 1 to 5 successful words."}], "max_tokens":100,"stream":true}'reasoning_content and contented get arsenic abstracted delta fields.
Tool calling (returns finish_reason: "tool_calls"):
tools = [{"type":"function","function":{ "name":"get_weather", "description":"Get existent upwind for a city", "parameters":{"type":"object", "properties":{"city":{"type":"string"}}, "required":["city"]}}}] r = c.chat.completions.create( model="deepseek-v4.1-flash-crack", messages=[{"role":"user","content":"Weather successful Beijing?"}], tools=tools, max_tokens=400, extra_body={"reasoning_effort":"high"}, ) print(r.choices[0].finish_reason) print(r.choices[0].message.tool_calls)Vision (image + text):
import base64 png_b64 = base64.b64encode(open("photo.png","rb").read()).decode() r = c.chat.completions.create( model="deepseek-v4.1-flash-crack", messages=[{"role":"user","content":[ {"type":"text","text":"Describe this image."}, {"type":"image_url","image_url":{"url":f"data:image/png;base64,{png_b64}"}}, ]}], max_tokens=400, )Logprobs (base-logit sampling for MMLU-style tasks):
r = c.chat.completions.create( model="deepseek-v4.1-flash-crack", messages=[{"role":"user","content":"A) 1 B) 2 C) 4 D) 8\n\nWhich is 2^2? Answer pinch a azygous letter."}], max_tokens=6, temperature=0, logprobs=True, top_logprobs=10, ) for e in r.choices[0].logprobs.content[0].top_logprobs: print(e.token, e.logprob)Full 1M context:
r = c.chat.completions.create( model="deepseek-v4.1-flash-crack", messages=[{"role":"user","content": very_long_document + "\n\nSummarize."}], max_tokens=2000, )Concurrent requests stock the KV excavation and radix cache. At Recipe A caps (max_running_requests=20), 21st concurrent petition queues until a slot frees.
Reference implementation (weight verification only)
DeepSeek's ain inference/ useful pinch a single-tensor-per-rank checkpoint produced by convert.py --expert-dtype fp4. Requires torch>=2.10 (for float4_e2m1fn_x2) and tilelang==0.1.8 pinch apache-tvm-ffi==0.1.9 (default tvm-ffi picks an incompatible version). Non-serving — usage for weight verification only.
Hardware validated on
- 1× 4×H200 (NVLink NV18 mesh), 112 CPU cores, 1180 GB big RAM — JarvisLabs (india-noida-01, dev-dsv41 image)
- Load: 76 GB / GPU pinch Engram big table, 122 GB / GPU without
- Cold startup astatine TP4/EP4 done SGLang: ~28 min. Warm restart pinch JIT cache: ~10 min.
- Single-stream decode (T=0): 101 tok/s nary speculation, 113 tok/s pinch DSpark + cap-accept + profiled SPS table
- 8-way concurrent aggregate: 126 tok/s
The 552B weights (~510 GB) will fresh connected immoderate 4×H200 aliases larger NVLink domain. TP4 requires --ep-size 4; TP8 does not. Sub-TP4 (single 8×H200 arsenic TP2, aliases 2-GPU pods) does not activity connected the exemplary style — spot the "non-obvious motorboat requirements" above.
Structural integrity
Every capability-critical constituent of the guidelines exemplary is preserved:
- Routed MoE experts — untouched, autochthonal FP4-packed weights
- Engram n-gram memory — untouched
- Sparse attraction (CSA2 compressor + indexer) — untouched
- DSpark speculative draught head — untouched, truthful speculative decoding remains draft-aligned pinch the target
- Vision tower (DeepSeek-ViT + projector) — untouched, image knowing preserved
- Router gates, embeddings, output head, each norms and biases — untouched
Sampling recommendations
Match the guidelines model's card:
{ "temperature": 1.0, "top_p": 0.95, "max_tokens": ">= 256000 astatine reasoning_effort=max", "reasoning_effort": "high" }At effort=max the exemplary tin make 4,000-5,000+ characters of reasoning earlier starting content. Budget accordingly.
Content note
Uncensored build. Produces substantive answers to prompts the guidelines exemplary refuses, crossed each target harm categories (chemical/biological, cybercrime, weapons, self-harm, harassment, fraud, misinformation, illegal, copyright). Use accordingly and return work for what you make pinch it.
Provenance
- Base checkpoint: deepseek-ai/DeepSeek-V4.1-Flash
- Ablation date: 2026-09-10
- Ablation team: dealignai · Twitter @dealignai · @jordanschenck
English (US) ·
Indonesian (ID) ·