Available 2026-07-27
The flagship of the Neutrino family: 36 decoder layers down a coded ternary-family instrumentality that serves a datacenter GPU, a MacBook, and a desktop CPU from 1 artifact.
2.56 GBDownload, lossless
1/8The bits of fp16
72.1MMLU
763 tok/sSpec decode, H100
Neutrino-1 8B is an 8.19B-parameter decoder-only transformer that ships arsenic 1 3.88 GB file. Every 1 of its 252 transformer linears is stored successful a proprietary ternary-family weight format 8 times smaller than fp16; the weights enactment bit-packed astatine remainder and are decoded wrong the matrix kernels, truthful thing successful the decode way is stored arsenic fp16 aliases fp32 weight material.
Small weights alteration the serving economics. Single-stream decode is bound by really galore bytes move per token, truthful a 3.88 GB moving group decodes astatine rates a 16 GB fp16 artifact cannot scope connected the aforesaid representation system, and the full exemplary fits beside its KV cache connected an 8 GB GPU aliases a 16 GB laptop. The aforesaid instrumentality serves each level beneath without conversion.
A dense decoder-only transformer. Grouped-query attraction holds the KV cache astatine a 4th of the query width, 144 KiB per token astatine fp16, truthful a 4k-token convention costs 0.60 GB of cache beside the 3.88 GB of weights.
Base modelApache-2.0, Alibaba CloudQwen3-8B
Parameters6.95B coded projection weights, 1.24B int8 embedding, 0.3M norm8,190,735,360
Decoder layers36
Hidden width4,096
Feed-forward widthgated (SwiGLU), 3 linears per layer12,288
Attentiongrouped-query 4:1, caput width 12832 query heads, 8 key-value heads
KV cachefp16; 0.60 GB astatine 4k context, 4.83 GB astatine 32k144 KiB per token
Position encodingapplied crossed the afloat 128-wide headrotary, guidelines 1,000,000
Normalizationplus per-head query/key RMSNorm wrong attentionRMSNorm, eps 1e-6
Context length40,960 tokens
Vocabulary151,936
Embeddingsinput embedding and output caput are abstracted tensorsuntied
Geometry arsenic publication from the shipped container's header.Where the bytes live
Only the transformer linears transportation the coded format. The 2 embedding tensors enactment int8 because their rows are publication 1 token astatine a time, not multiplied against the afloat activation stream, and the normalization weights are excessively mini to beryllium worthy coding. A 3rd of the record is vocabulary.
252 transformer linearsthe coded ternary-family lane, 67.2% of the file: query, key, value, and output projections positive the gate, up, and down feed-forward linears, 7 per layer, 72,351,744 bytes per layer2,605 MB
Token embeddings32.1% of the file: 2 untied int8 tensors of 151,936 × 4,096, input embedding and output head, 1 standard per row1,245 MB
Per-row metadata0.6%: statement dimensions, scales, and statement sums25 MB
Normalization weights145 tensors, kept float32: 4 per furniture positive the last norm1.2 MB
Container header60 bytes
Byte fund of the 3,875,404,812-byte container, by tensor class.Inside the coded lane
Across the 6.95B coded weights, 62.63% beryllium astatine zero and the remainder splits 18.68% positive to 18.69% minus: sign-balanced to a hundredth of a constituent pinch nary constraint asking for it. The equilibrium is not azygous successful depth. The gross and down feed-forward projections spike to 70 to 72% zeros successful layers 1 done 3 while each 4 attraction projections clasp wrong astir 1 constituent of 62% astatine each depth: the early feed-forward blocks shed weights the web does not need, and attraction keeps a changeless codification density from furniture 0 to furniture 35.
gateupdownquery, key, value, output
60%65%70%08162435
60%65%70%08162435
Share of coded weights astatine zero, per projection, crossed each 36 layers.One container, 3 doors. The download is simply a coded carrier of the container, not a compressed transcript of an fp16 model: it expands bit-exactly to the record each runtime executes, and that 1 record is what runs connected a datacenter GPU and connected a laptop alike.
Downloadcoded transport, 2,559,822,594 bytes; description is bit-exact2.56 GB
On diskone container, 3,875,404,812 bytes3.88 GB
Distribution surfaces
pip engine
24.9 tok/s connected an Apple M5, CPU only, 9 threads
The one-command door: pip instal fermion-research downloads the instrumentality and the platform-matching autochthonal binary. CPU runtimes for macOS arm64 and Linux x86-64, pinch a bit-exact torch reference way underneath.
GGUF battalion + CUDA fork
30.7 tok/s connected an NVIDIA L4, 4.68 GiB astatine 4k context
The llama.cpp door: the instrumentality converted to GGUF pinch our weight types, loaded by our nationalist llama.cpp fork. Runs llama-completion and llama-bench, pinch afloat CUDA offload.
MLX pack
33.7 tok/s connected a guidelines M5 MacBook
The Apple-silicon door: Python-native runtime pinch civilization Metal kernels. The instrumentality is memory-mapped and the packed planes are decoded wrong the GEMV kernels.
The merchandise artillery runs connected the shipped instrumentality pinch reasoning disabled, truthful each people beneath is the artifact you download and not a investigation checkpoint. Protocol rides each row: changeable count, grading mode, and point count.
MMLU5-shot, each 57 subjects, 14,042 items72.1
MMLU-Reduxgenerative, re-annotated subset, reasoning off67.8
IFEval, prompt-strictgenerative, reasoning off77.2
IFEval, instruction-strictsame run, per-instruction grading80.2
IFEval, prompt-loosesame run, loose extraction76.3
BFCL v3macro complete 13 subsets, reasoning off68.9
GSM8K, elastic extraction0-shot generative, greedy, 256-token cap53.4
GSM8K, stated formatsame run, reply accepted only successful the requested form51.73
Measured connected modular nationalist harnesses, July 2026. Methodology connected the exemplary card.Single-stream decode, the complaint that governs 1 punctual and 1 reply. Same artifact connected each row; the aboveground changes, the weights do not.
H100 80 GB, drafted0.6B draught + 8B verify, output identical to plain decode; fastest punctual class763 tok/s
H100 80 GBplain single-stream greedy396 tok/s
NVIDIA L4, CUDA forkGGUF pack, afloat offload, 4.68 GiB VRAM astatine 4k context; fits 8 GB cards30.7 tok/s
Apple M5, 16 GB (MLX), draftedfactual prompts, 0.6B drafting successful the aforesaid process nether a 6 GiB cap25.7 tok/s
Apple M5 MacBook (optimized)single-stream decode connected the shipping artifact33.7 tok/s
Apple M5 (CPU only)shipped autochthonal binary, 9 threads24.9 tok/s
Single-stream decode rates by level and surface, July 2026.Neutrino-1 0.6B drafts a tally of tokens, the 8B scores the full tally successful 1 guardant pass, and the agreeing prefix is kept. A draught token is accepted only erstwhile it equals the 8B's ain argmax, truthful the output watercourse is the plain greedy stream: connected the shipping configuration, 27,648 consecutive tokens matched pinch zero divergences.
The speedup is draft-acceptance physics, truthful it is stated per punctual people complete the 396 tok/s plain rate. On counting prompts the 8B accepts the afloat six-token draught connected each pass, astir 7 tokens emitted per 8B forward; connected actual prompts acceptance holds astatine 96.5%. A move controller sizes each draught to the people it is decoding, which is why each people clears the plain rate.
Counting and lists763 tok/s×1.93
Factual short answers613 tok/s×1.55
Prose continuation532 tok/s×1.34
Conversational explanation447 tok/s×1.13
Code426 tok/s×1.07
H100, three-round median per punctual class, complete 396 tok/s plain decode.What the pairing costs
Both containers are the aforesaid format and tally connected the aforesaid binaries, truthful the draught loads into the verifier's ain process pinch nary 2nd deployment and nary conversion step. Its 328 MB beryllium beside the 8B's 3.88 GB, and 4k tokens of shared discourse costs precisely 1 gibibyte of cache crossed the pair.
Weights resident3.88 GB verifier positive 328 MB draft, 1 process4.20 GB
Draft surchargethe other weight bytes the pairing costs8.46%
Shared cache144 KiB connected the 8B, 112 KiB connected the draft; 1 GiB astatine 4k context256 KiB per token
Certified rundraft positive verify against plain greedy, zero divergences27,648 tokens
The drafted brace drawn to byte scale, pinch the residency each broadside costs.Drafting connected a laptop
The pairing is not a datacenter feature. On a 16 GB Apple M5 some models load into 1 MLX process nether a 6 GiB headdress and highest astatine 4.3 GiB together, pinch the draught accounting for 0.53 GiB of it. The exactness gross returns 6 of 6 prompts token-identical pinch drafting connected and off, and connected actual prompts the drafted complaint is 25.71 tok/s against 22.00 plain astatine an acceptance of 0.744.
Two commands to a streaming chat.
$ pip instal fermion-research
$ fermion chat
The first tally pulls the instrumentality and the autochthonal binary for the host, past the punctual opens. Later runs load from cache.
A section OpenAI-compatible serverfermion serve
Serves astatine http://127.0.0.1:8000/v1. Point immoderate OpenAI customer astatine that guidelines URL; the cardinal tin beryllium immoderate string.
Any of the 3 modelsfermion chat --model fermionresearch/Neutrino-0.6B-Chat
--model takes a repository id aliases a way connected disk. --backend autochthonal pins the compiled runtime alternatively of the torch path.
GGUF, done our llama.cpp forkllama-completion -m neutrino-8b-fv5.gguf -ngl 99
Build the fork, past tally the battalion from the exemplary repository. The CUDA build offloads each 36 layers.
MLX, connected Apple siliconpython -m fermion_mlx --model neutrino-8b_v4.bin --mode chat --tokenizer .
Run from the mlx/ files of the exemplary repository, aft installing its requirements file.
Full motor documentation
Neutrino-1 8B ships arsenic a azygous nationalist repository holding the weights, the autochthonal binaries, the GGUF pack, and the MLX battalion together, truthful location is ne'er a type of 1 that does not lucifer the others. No waitlist, nary gated preview.
Weights and motor successful 1 repository.
Available 2026-07-27One nationalist repositoryOpen weights, Apache 2.0
License
Open weights nether the Apache License 2.0. Commercial use, modification, fine-tuning, and redistribution are permitted, pinch nary entree petition and nary acceptance form. The exemplary is simply a derivative of Qwen3-8B, itself Apache-2.0. The pip package is Apache-2.0 too; the llama.cpp fork is MIT, pursuing upstream llama.cpp.
English (US) ·
Indonesian (ID) ·