left rail: pick a kernel — identity, box blur, Gaussian, sharpen, emboss, outline, Sobel X, Sobel Y, Sobel magnitude SPLIT / A: before-and-after split view; drag inside the image to move the divider IMAGE / I: cycle the three source images RESET / S: back to identity, sliders zeroed BRIGHTNESS / CONTRAST / THRESHOLD: drag the tracks GRAY / INVERT: toggles The matrix panel on the right always shows the 3×3 that is currently running, with its divisor and bias. The histogram under the image is the output luminance in 64 bins, with the source distribution drawn behind it as an outline, so a brightness or contrast change visibly moves the data rather than merely being asserted.
An image processing lab in Scratch: a 48×36 bitmap held in lists, convolved with real 3×3 kernels, and drawn with the pen — with the nine numbers that produced the picture shown next to the picture. THE INTERESTING DECISION The pipeline is split in two, and the split is the design: imgs.txt -> p_r/p_g/p_b + lum -> [3x3 CONVOLUTION] -> m_r/m_g/m_b -> [POINT OPS via a 256-entry LUT] -> pen A convolution is 1564 interior pixels × 9 taps × 3 channels = **42,228 multiply-adds** and takes a second or two. A brightness change is a table lookup. Caching the convolution result in mid means dragging a slider re-runs only the second half of the pipeline, so the expensive stage happens exactly when the kernel changes and never once per frame of a drag. The point ops go through a 256-entry lookup table rebuilt only when a slider moves, which turns three multiplies and two clamps per channel per pixel into three list reads. Three smaller things that are all Scratch constraints rather than choices: The tap loop is unrolled. The nine weights and the divisor are read into locals before the loop rather than indexed per tap — that alone removes 15,552 list reads from a single pass. A row of pixels is one pen stroke, not 48 dots. Scratch's pen has round caps and nothing else, so a dot of exactly the pixel pitch leaves the four corners of each cell unpainted and the image reads as a beehive; oversizing the dot to cover the corners makes every pixel bleed into its neighbours. Drawing each row as a polyline at pen size = pitch is exact instead: vertically a stroke covers precisely one row, and horizontally each segment's trailing cap is overwritten by the next segment, so every pixel ends up a clean 5×5 cell. The only artefact is a half-cell phase shift, which the origin cancels — and it is 2 blocks per pixel instead of 4. There is no timing readout, deliberately. timer() in Scratch only advances between yields and a func never yields, so sampling it either side of the convolution returns the same value: a "took 0 ms" display would have been a lie. The status bar reports the multiply-add count instead, which is the honest number and the one that explains why the magnitude pass costs what it does. VERIFICATION The Sobel magnitude pass and the Gaussian pass were both dumped out of a headless scratch-vm run and compared against an independent Python implementation, pixel by pixel: | pass | pixels compared | mismatches | |---|---|---| | Sobel magnitude | 1564 | 0 | | Gaussian, all three channels | 4692 | 0 | The Gaussian comparison initially showed 211 mismatches with a maximum difference of 1, which turned out to be the reference being wrong: Python's round() is banker's rounding and Scratch's rounds half away from zero, and sum/16 lands exactly on .5 whenever the weighted sum is 8 mod 16. With floor(v + 0.5) the two agree exactly. HONEST LIMITATIONS - 48 × 36 pixels. 1728 of them, drawn as 5px cells. That is the size at which a 3×3 convolution finishes in about a second in vanilla Scratch; four times the pixels would be four times the wait, every time you click a kernel. - The one-pixel border is not convolved. Clamping the sample coordinates would put four comparisons inside the hottest loop in the project for the sake of 164 of 1728 pixels. All original - code, art and sound. See Inside is open.