TURBOWARP EVEN MORE RECOMMENDED THAN V2!! > https://turbowarp.org/1386471146 ---------------------------------------------------------------- > Enter to send message, left-arrow to delete (since Scratch has no backspace detection!) > Type a query, and Cud will try to answer to the best of its abilities, based on how it was trained! > Cud v3 is a transformer language model, with similar architecture to modern tools like ChatGPT, but shrunk down and built entirely in Scratch. Every response is computed LIVE with its weights! > It can answer certain types of questions much better than others! This model specializes in more things than v2, such as jokes, world knowledge, and multi-message threads. > It has a much better memory, remembering around 7-10 of your previous messages. > Sign in if you want it to know your username :D
Scroll or use up/down arrows to navigate the chat. Use the left-arrow to delete your last typed character. > "cleanup" to clear all cache lists, for if you want to mod my project! (exports under the 5MB limit if you do this)... this stops the functionality of the chat. > "clearthechat" to clear your chat without stopping functionality Displays 1 and 2 allow you to switch between the classic chat interface and a new robot character display mode! Model Details: Parameters: 6,727,488 (v2 had only 365,386) Layers: 15 (v2 had 2) Attention: 6 heads that share 2 key/value heads (grouped-query attention, GQA) Model width: 192 (head width of 32, feedforward of 448 using SwiGLU) Positions: ALiBi, which tells the model how far back each word is, direct upgrade from v2 Context: 256 tokens (v2 had 96, and v1 had 8 (!!)) Vocabulary: 1,024 pieces, plus two hashed tables for common 2- and 3-piece combos Precision: 3 bits per weight, learned during training (quantization-aware training, aka QAT) This model does a lot more math so Turbowarp makes it a lot less laggy! Training was done on around 300 million tokens and 18,310 training steps. It had multi-message threads, so it learns to keep up with a real conversation rather than just single questions, a jump from v2. PyTorch: training weights Scratch: everything in the runtime :P UI and all art: myself :D Name and Logo: Claude/Anthropic inspiration And again, thank you to EVERYONE who tried both v1 and v2 and told me what screwed up. It shaped this version. :) *Important note: This project comes with safeguards installed against profanity and explicit/sensitive topics. If discussed, the project will do its best to divert you away from the topic. Update v3.0.0: release Update v3.5.0: • 5,213,376 --> 6,727,488 parameters (retrained) • 8 layers --> 15 layers • Reduced 4-bit precision to 3 bits • Says when it doesn't know something instead of guessing, and handles topic switches better • Replies now appear word by word as they are written instead of having a loading anim and then streaming with a uniform time