Can we run item akin Jev in your browser? GitHub repo ↗
A live, local experiment
A local example can either peruse probabilities for your allowed options without decoding them, or compose the identical benevolent of allocation token by token. Pick a size, run the two on your own GPU, and measure the difference.
browser onlyno backendyour timings1.56 GB model
There is no waitlist! Just try it out ↓
MiniCPM5 2B is selected by default. On a phone or smaller device, toggle to Qwen3 0.6B in the example box if needed.
00 / setup
Load the example once
Model
Larger model. Loading may be slower or may not fit on several low-end devices.
Model performancehigher is better
Native BF16 · TypeSafe: identical 102-row subset · Jev: published outcome · browser builds are quantized
{{ content }}
download / cache—
starts lone whenever you click load
model load—download and prepare
warmup—compile passes for the two methods
Weights arrive from Hugging Face and remain in your browser cache. Inputs never depart this page. First burden can obtain multiple minutes depending on the selected model, network and GPU.
01 / decision
Give it a genuine choice
Try an example
Both paths obtain the identical decision. One says choice probabilities directly; the another asks the example to compose its choice probabilities as JSON text.
your decisionstate + inquiry + options
→
same local modelMiniCPM5 · 2B
↗
↘
read logitsA…T probabilities
write tokens{options + probabilities}
02A / straightforward readout
Choice probabilities
no decoding
Read the model’s choice logits and normalize lone throughout the options you supplied.
waiting for a run
total—
input—
output1 readout
02B / generation
JSON probabilities
token by token
Ask the example to evaluation the identical displayed-option allocation and compose it as JSON. Watch all token arrive.
waiting for a run
first token—
total—
input—
output—
measured wall-time ratiorun it on your GPUThe methods run sequentially on the identical loaded example so they do not contend for one GPU. Direct runs first, afterward generation.
What these numbers do—and do not—meanConditional probabilities. Direct scores are a softmax complete lone the displayed choice tokens. They are not calibrated confidence and do not contain all answer the example power prefer.
Local example tiers. The phone example trades accuracy for size. MiniCPM is the desktop default. The 4B choice needs substantially additional memory. None is claimed to equivalent Jev.
Real local timing. Setup, warmup, immediate preparation, straightforward execution, archetypal generated token and generation completion are timed alongside performance.now(). No canned results appear.
Quantized weights. The demo uses pinned GGUF builds through wllama. Quantization can alter the two norm and speed.