OpenJev in your browser

๐Ÿš€ Read this insightful post from Hacker News ๐Ÿ“–

๐Ÿ“‚ **Category**:

โœ… **What Youโ€™ll Learn**:

openjev

A live, local experiment

A local model can either read probabilities for your allowed options without decoding them, or write the same kind of distribution token by token. Pick a size, run both on your own GPU, and measure the difference.

browser onlyno backendyour timings1.56 GB model

There is no waitlist! Just try it out โ†“

MiniCPM5 2B is selected by default. On a phone or smaller device, switch to Qwen3 0.6B in the model box if needed.

00 / setup

Load the model once


Larger model. Loading may be slower or may not fit on some low-end devices.

Model performancehigher is better

Native BF16 ยท TypeSafe: same 102-row subset ยท Jev: published result ยท browser builds are quantized

download / cacheโ€”starts only when you click load

model loadโ€”download and prepare

warmupโ€”compile passes for both methods

Weights come from Hugging Face and remain in your browser cache. Inputs never leave this page. First load can take several minutes depending on the selected model, network and GPU.

01 / decision

Give it a real choice

Try an example

Both paths receive the same decision. One reads option probabilities directly; the other asks the model to write its option probabilities as JSON text.

your decisionstate + question + options

โ†’

same local modelMiniCPM5 ยท 2B

โ†—
โ†˜

read logitsAโ€ฆT probabilities

write tokensโšก

02A / direct readout

Choice probabilities

no decoding

Read the modelโ€™s choice logits and normalize only across the options you supplied.

waiting for a run

total
โ€”

input
โ€”

output
1 readout

02B / generation

JSON probabilities

token by token

Ask the model to estimate the same displayed-option distribution and write it as JSON. Watch every token arrive.

waiting for a run

first token
โ€”

total
โ€”

input
โ€”

output
โ€”

measured wall-time ratiorun it on your GPU

The methods run sequentially on the same loaded model so they do not contend for one GPU. Direct runs first, then generation.

What these numbers doโ€”and do notโ€”mean

Conditional probabilities. Direct scores are a softmax over only the displayed option tokens. They are not calibrated confidence and do not include every answer the model might prefer.

Local model tiers. The phone model trades accuracy for size. MiniCPM is the desktop default. The 4B option needs substantially more memory. None is claimed to match Jev.

Real local timing. Setup, warmup, prompt preparation, direct execution, first generated token and generation completion are timed with performance.now(). No canned results appear.

Quantized weights. The demo uses pinned GGUF builds through wllama. Quantization can change both quality and speed.

OpenJev / browser labGitHub repo ยท model ยท wllama

๐Ÿ’ฌ **Whatโ€™s your take?**
Share your thoughts in the comments below!

#๏ธโƒฃ **#OpenJev #browser**

๐Ÿ•’ **Posted on**: 1789730293

๐ŸŒŸ **Want more?** Click here for more info! ๐ŸŒŸ

By

Leave a Reply

Your email address will not be published. Required fields are marked *