Works on turbowarp! https://turbowarp.org/1370277845 Warning: Might be resource iNtEnSiVe Note: No animals were killed in the making of this project A real language model — 491,040 parameters — running entirely in Scratch. No extensions. No TurboWarp (if you want lag). No internet. Just blocks and lists. Click the green flag and type in the ask box. IT IS SLOW. Every reply is millions of multiply-adds done in Scratch blocks, one at a time. Give it a moment — the PROGRESS line at the top shows what it's doing. Shift+click the green flag for turbo mode; it helps a lot. Newest messages appear at the TOP of the list, so you never have to scroll. TRY THIS hello what's your name? tell me a joke I'm having a rough day what's the capital of Chad? IT REMEMBERS THINGS YOU TELL IT /mem name: Jamie; then ask: what's my name? /mem favorite_color: blue; then ask: what's my favorite color? Facts live outside the weights and get fed in as a prefix, so a 491K model only has to turn a supplied fact into a sentence. Other fact types are hit and miss — these two are the reliable ones. COMMANDS /clear start over /mem <fact> set the remembered fact /nomem forget it /temp 0.4 lower = safer and more repetitive, higher = wilder (0.7 default) /new 24 fewer tokens per reply = faster /toktest check the tokenizer against known-good output
Fold = simplifies user responses THE MODEL Weights are PocketLM 500k by Greninja9257 (MIT licensed), trained in PyTorch. https://github.com/Greninja9257/PocketLM Scripted using Goboscript 1024-token vocabulary | 5 layers | 6 attention heads | 256-token context RMSNorm, RoPE, grouped-query attention, SwiGLU, tied embeddings, bias-free. Character-level BPE tokenizer, ported exactly — same word split, same merge order. Sampling is temperature 0.7, top-p 0.9, repetition penalty 1.1. NOTHING IS QUANTIZED The weights are the checkpoint's own float16 bit patterns, three packed into eight base-64 characters and unpacked back into sign/exponent/mantissa at runtime. Not rounded, not approximated — bit-identical. I checked the whole forward pass against the original Python implementation and it agrees to 3.9e-07, which is just floating-point noise. That costs space, so: project.json is 1.78 MB against Scratch's 5 MB limit, about 36%. It fits with room to spare. 2,403 blocks. Character codes come from 95 costumes named after ASCII characters, which is how a string of packed weights becomes numbers. WHAT TO EXPECT It holds short conversations, tells real jokes, and declines what it doesn't know. It does NOT know facts about the world — that isn't a bug, it's a 491,040-parameter model. For scale, that's roughly 0.0003% of GPT-3.