TURBOWARP RECOMMENDED!! Type anything, and Cud will do its best to answer. Cud (version 2) is a transformer language model, which has similar architecture to modern tools like ChatGPT, but shrunk down to a much more simple scale and entirely in Scratch. Every response is computed LIVE from 365,386 numbers stored in two lists! Unfortunately, the "no-internet" claim no longer holds, as I use the translate blocks to translate text when Cud is prompted to do so... but everywhere else, it could apply! Here are some conversation starters: "hi", "what color is the sky", "say purple monkey window", "tell me a joke", "i had a bad day", "what is the capital of france" (Verity EBK) What can it do? > Math (Try "what is 1+1+1+1+50") > Mimicry/Repeating (Try "say _" or "repeat _") > Spelling/Typo Tolerant (Try "wat colr is teh sky?") > Translate (Try "what is hello in French?") > Feelings - sort of... (Tell it about your day!) > Memory - usually remembers around 5 of the last messages > Honesty - the biggest one in my book, it knows when to accept what it, well, doesnt know. It has only 365,386 parameters, while modern AI has billions, so it's not the most capable. It doesn't learn while you talk either, and it gets things wrong (MANY things wrong...) Sign in if you want it to know your name!
Up/down arrow: Scroll, Left arrow: delete, "clearthechat" clears the chat Cud v2 is a decoder-based transformer trained from scratch in Python's PyTorch, then quantised into integers and exported as .txt lists that Scratch reads at runtime via, well, lists! No responses are pre-written. There are no if/else answers anywhere in this project (well, there are, just not for responses!) and every sentence is generated one word piece at a time from its learned weights. Here are some details, if you're wondering. > Parameters: 365,386 > Layers: 2, 2 attention heads each > Model width: 96 (head width of 48, feedforward of 256) > Context: 96 tokens (v1 had 8, so a 12x increase!) > Vocabulary: 1,898 pieces... not exactly words! > Weights in lists: 182,208 embeddings + 183,178 others > Precision: integers (so limited), scaled by 20,000 The tokenizer has words, subword chunks, and singular characters as a really really final fallback. Unknown words are broken down into manageable chunks, which is why they work, unlike v1. The model uses routing, deciding what general semantic path to go down before typing. These include greet, smalltalk, fact, self, math, unknown, nonsense, translate, echo, and support, and that choice is important to what it decides to say. Some tokens are replaceable, like <num>, <trans>, and <name>, so it gets certain info correct every time, and not down to chance. Training was done on a corpus of around 68 million tokens that generated unique patterns to enforce learning. It didn't use any data centers, just my MacBook, so you don't need to worry about that. :) The final perplexity was 0.933, with 30k training steps covering around 7k different sentence patterns/shapes. The UI was custom built from the ground up by myself, so feel free to use it (with credit, please!) Credits: > Pytorch - training weights > Scratch - everything in the runtime :P > UI and all art - myself :D > Name and Logo - Claude/Anthropic inspiration > And of course, thank you to the thousands of people who gave attention to v1 and gave me critical information to shape this new version. :)