OpenJev

Hacker News by 3 min read 526x views
OpenJev

Share Post

openjev

Can we run item akin Jev in your browser? GitHub repo ↗

A live, local experiment

A local example can either peruse probabilities for your allowed options without decoding them, or compose the identical benevolent of allocation token by token. Pick a size, run the two on your own GPU, and measure the difference.

browser onlyno backendyour timings1.56 GB model

There is no waitlist! Just try it out ↓

MiniCPM5 2B is selected by default. On a phone or smaller device, toggle to Qwen3 0.6B in the example box if needed.

00 / setup

Load the example once

Model

Larger model. Loading may be slower or may not fit on several low-end devices.

Model performancehigher is better

Native BF16 · TypeSafe: identical 102-row subset · Jev: published outcome · browser builds are quantized

{{ content }}

download / cache—

starts lone whenever you click load

model load—download and prepare

warmup—compile passes for the two methods

Weights arrive from Hugging Face and remain in your browser cache. Inputs never depart this page. First burden can obtain multiple minutes depending on the selected model, network and GPU.

01 / decision

Give it a genuine choice

Try an example

Both paths obtain the identical decision. One says choice probabilities directly; the another asks the example to compose its choice probabilities as JSON text.

your decisionstate + inquiry + options

→

same local modelMiniCPM5 · 2B

↗
↘

read logitsA…T probabilities

write tokens{options + probabilities}

02A / straightforward readout

Choice probabilities

no decoding

Read the model’s choice logits and normalize lone throughout the options you supplied.

waiting for a run

total—

input—

output1 readout

02B / generation

JSON probabilities

token by token

Ask the example to evaluation the identical displayed-option allocation and compose it as JSON. Watch all token arrive.

waiting for a run

first token—

total—

input—

output—

measured wall-time ratiorun it on your GPU

The methods run sequentially on the identical loaded example so they do not contend for one GPU. Direct runs first, afterward generation.

What these numbers do—and do not—mean

Conditional probabilities. Direct scores are a softmax complete lone the displayed choice tokens. They are not calibrated confidence and do not contain all answer the example power prefer.

Local example tiers. The phone example trades accuracy for size. MiniCPM is the desktop default. The 4B choice needs substantially additional memory. None is claimed to equivalent Jev.

Real local timing. Setup, warmup, immediate preparation, straightforward execution, archetypal generated token and generation completion are timed alongside performance.now(). No canned results appear.

Quantized weights. The demo uses pinned GGUF builds through wllama. Quantization can alter the two norm and speed.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads