Click in the field: place a training point 1 / 2: which class you are placing (cyan / amber) SPACE: train / pause S: run exactly one epoch R: re-initialise the weights, keep the data P: next dataset: XOR, concentric, moons, spiral, linear C: clear all points ↑ / ↓: learning rate, 0.1 – 4.0 ← / →: epochs per frame, 1 – 64 A red ring around a point means the network currently gets it wrong, so you can watch the last stubborn points flip one at a time. The right-hand panel is the loss curve above and the live weight matrix below: every edge is one weight, cyan positive, amber negative, thickness by magnitude.
A 2-8-6-1 feedforward network trained by backpropagation inside the project — watch the decision boundary form, the weights change and the loss fall, in real time, and place your own training points with the mouse to make it start over. THIS IS REAL GRADIENT DESCENT Nothing is precomputed and nothing is animated. Weights start random, and each frame runs N full-batch epochs — forward pass, backward pass, weight update — over whatever points are on the field. Add a point mid-training and the very next epoch includes it. The maths, in full: 2 -> 8 tanh -> 6 tanh -> 1 sigmoid 105 weights and biases loss = binary cross-entropy dL/dz3 = a3 - y (BCE and sigmoid derivatives cancel) dL/dz2 = (W3 . d3) * (1 - a2^2) (tanh' = 1 - a^2) Scratch has neither tanh nor exp, only antiln (eˣ). So tanh(z) = 2/(1 + e^-2z) - 1, which saturates correctly for free: a large negative z makes e^-2z overflow to Infinity, and 2/Infinity is 0, giving −1. No clamping anywhere. tanh is worth its exponential precisely because 1 - a² gives the derivative for one multiply once you already have the activation. VERIFIED NUMERICALLY The compiled .sb3 was run under the headless harness and its final weights, points and loss history dumped, then the whole network was **recomputed from those weights in Python** and compared: * XOR, 36 points, lr 0.5, 723 epochs: loss 0.6886 → 0.0018, monotone at every one of 91 sampled points, 36/36 correct. * Python's independent recomputation from the dumped weights: loss 0.001793, 36/36. The project's own loss history sample agrees to four decimal places. * Mouse placement checked the same way: clicking twice in the field took the point count from 60 to 62 and training continued from there. The banner is a second, independent implementation — the same architecture, the same initialisation scheme, the same dataset formula, vectorised in numpy — and it converges the same way: on the project's own spiral preset, loss 0.6686 → 0.0003 and 68/68. Everything in that image is the output of that run. TWO DECISIONS WORTH EXPLAINING Full batch, not per-sample SGD. The loss curve is part of the exhibit, and batch descent gives a curve you can actually read; SGD on 40 points gives noise at the same average. It also makes "epoch" a meaningful unit for the counter. Drawing runs on a slower clock than training. Painting the decision surface means a forward pass per cell — 600 of them — which cannot happen every frame. So the field and the weight diagram repaint every fifth frame while the loss curve and the counters repaint every frame, and training never waits for drawing. The whole epoch loop lives in a func, which compiles to "run without screen refresh"; in an ordinary script the repeat over training points would yield per point and the network would manage about half an epoch a second. The loss history keeps 160 samples, one per pixel column of the plot. When it fills, every second sample is dropped and the sampling interval doubles (the plot labels it X2, X4, …), so the curve always spans the whole run rather than scrolling and hiding the initial collapse — which is the interesting part. TWO BUGS THE HARNESS CAUGHT THAT READING THE CODE DID NOT Both were invisible in the source and obvious in a screenshot: 1. The right-hand panel cleared itself with a pen stroke spanning the full stage width. All original - code, art and sound. See Inside is open.