## Concept explanation A **convolution** uses a small set of numbers called a **filter** or **kernel** to scan across an image one local patch at a time. At each position, the filter multiplies its weights by the nearby pixel values and adds the results into one output number. Because the **same filter** is reused everywhere, it can detect the same kind of pattern—such as a vertical or horizontal edge—no matter where that pattern appears in the image. This combination of **local computation** and **weight sharing** is what turns an image into a **feature map**. ## What you see You’re looking at a grayscale input image on the left, the active `3×3` filter in the center, and the output feature map on the right. The teal box shows the current input window being examined. As you move the window, the matching output cell is highlighted and filled with the filter’s response value, while the response card summarizes the local products that were combined. This makes it easier to see that each output cell depends only on one small input region, even though the filter itself stays unchanged. ## Try it yourself - **Drag the horizontal slider** and watch the teal window move across one row while the output cell updates. - **Drag the vertical slider** to compare how the same filter responds in different parts of the image. - **Switch the filter dropdown** between vertical and horizontal edge detectors to see how the feature map changes. - **Pause on a bright-dark boundary** and notice that the response becomes stronger when the local patch matches the filter’s preferred pattern. - **Press the reset button** and trace how the output map starts filling again from the top-left position. - **Compare neighboring positions** and notice that each new output value comes from a different local patch, but always uses the same `3×3` weights. ## Concept explanation A **convolution filter** is a small grid of weights that slides across an image and scores how well each local patch matches a pattern. Different filters become sensitive to different structures because their positive and negative weights reward some arrangements of pixels and suppress others. A **horizontal edge** filter responds where brightness changes from dark to light across rows, a **vertical edge** filter responds to side-to-side changes, and other filters can respond to **corners** or **blobs**. The resulting **feature map** is bright wherever the selected filter finds a strong match. ## What you see You can compare three linked views. On the left is a tiny input image made of lines, corners, and filled regions. In the middle is the currently selected `3×3` filter, showing the weights that define what pattern it is looking for. On the right is the feature map, where brighter cells mark locations where that pattern appears strongly in the input. When you hover over a valid patch in the input, the matching location in the feature map is emphasized so you can connect the local image structure to the filter response. ## Try it yourself - **Choose a different filter** from the menu and watch how the bright regions move to places where that pattern exists. - **Click Next filter** and compare how the same image creates very different feature maps for horizontal, vertical, corner, and blob detectors. - **Hover over a line segment** in the input and see whether the highlighted activation is strong or weak for the current filter. - **Hover over a corner or filled region** and notice that some filters respond strongly while others barely activate. - **Toggle Show receptive field** to see the exact `3×3` image patch being compared against the filter weights. ## Concept explanation **Max pooling** is a downsampling step that looks at a small neighborhood, such as a `2×2` patch, and keeps only the **largest activation** from that patch. This reduces the size of the feature map while making the representation less sensitive to tiny shifts: the exact brightest cell inside a neighborhood can move, but the pooled value often stays similar because the strongest local response is still preserved. ## What you see You can compare three stages side by side: the left panel shows a simple underlying image with a movable object, the next panel shows the resulting **feature map**, and the right panel shows the smaller **pooled output** after max pooling. The center panel isolates one highlighted neighborhood so you can watch which cell wins locally and how that single maximum becomes one value in the pooled grid. ## Try it yourself - **Drag the small dot** inside the left grid and watch nearby feature-map cells brighten and dim as the activation shifts. - **Use the arrow keys** to nudge the object by small amounts and compare how the detailed feature map changes more than the pooled output. - **Adjust the shift step** to make each keyboard move tiny or larger, then test how much movement still leaves the pooled response stable. - **Change the activation spread** to see how a tighter or broader response affects which local cell becomes the maximum. - **Switch the highlighted region** to inspect a different `2×2` neighborhood and follow its winner into the pooled grid. - **Turn on snap to cell centers** to compare neat cell-by-cell jumps with smoother sub-cell motion. - **Press Reset object** and repeat the experiment, noticing that max pooling keeps the strongest local response even when the exact peak position slides within a neighborhood. ## Concept explanation A **convolutional network** builds understanding in stages. Early layers respond to simple local signals such as **edges**, because each filter only sees a small patch of the input. Deeper layers combine those earlier responses across larger effective receptive fields, so they can detect **motifs** like corners or junctions, and then even larger **object parts** made from those motifs. The key idea is that stacked layers do not just repeat the same work — they compose simpler features into more abstract ones. ## What you see You are looking at one input image on the left and up to three feature layers to its right. The first layer shows edge-like activations, the second shows corner or motif responses formed from multiple edges, and the third shows a larger object-part pattern. When you select a cell in the input, teal outlines mark the downstream activations whose receptive fields include that selected region, so you can trace how one small patch influences broader and more abstract features as depth increases. ## Try it yourself - **Click a filled part of the input shape** and compare how many highlighted activations appear in the edge layer versus the deeper layers. - **Move the depth slider from `1` to `3`** to reveal how the same selected region affects progressively more abstract representations. - **Switch between `T shape`, `Corner`, and `Box`** to see how different input structures create different motif and object-part responses. - **Click an empty background cell** and notice how much weaker the downstream highlights become when the selected region contains little signal. - **Use `Center selection`** to reset to a central patch and study how receptive-field composition spreads outward through the stack. ## Concept explanation A convolution slides a **kernel** across an input grid and produces an output grid. **Stride** controls how far the kernel jumps each time, so larger stride samples the input more sparsely and usually makes the output smaller. **Padding** adds extra border cells around the input, which helps the kernel still reach edge information instead of ignoring it. Together, stride and padding determine output size, how much spatial detail is preserved, and how much of the border contributes to the result. ## What you see You can compare the padded input grid on the left with the output grid on the right. The highlighted output cell is the one you are inspecting, and the teal box on the input shows that cell’s receptive field: the exact input region that contributes to it. Faint extra windows show how the kernel is sampled across the input, so when stride increases you can see the windows spread out, and when padding increases you can see the kernel keep touching border regions. ## Try it yourself - **Move the stride slider** from `1` to `2` or `3` and watch the output grid shrink as the sampling windows skip farther across the input. - **Increase the padding slider** and notice how the receptive field can extend beyond the original input while still producing output cells near the edges. - **Click different output cells** to see how each one maps to a different covered region in the input. - **Hover over output cells** to preview receptive fields before selecting one. - **Switch the kernel size** between `3 × 3` and `5 × 5` to see how a larger window increases spatial coverage and changes the output dimensions. - **Pick a corner output cell**, then **raise padding** and compare how much real input data versus padded border is used. - **Press Reset selection** to jump back to the top-left output cell and compare configurations again from the same reference point. ## Concept explanation A **convolutional layer** applies many learned **filters** to the same input at the same time. Each filter has separate weights for the red, green, and blue channels, so it can combine color information differently to detect specific patterns such as edges, blur, or texture. Every filter produces one **feature map**, so when you learn more filters, the output gets deeper: width and height stay tied to image position, while depth grows because you now have more pattern detectors responding in parallel. ## What you see On the left, you can see the RGB image split into its three channels, plus a small combined preview with the center patch highlighted. In the middle is a bank of filters, where each card shows the three channel-specific slices of one filter. On the right, the selected filter’s center-patch response is broken into per-channel contributions, and below that the full stack of output feature maps appears as layered cards. The highlighted map corresponds to the currently selected filter, showing that one chosen detector contributes one slice of the output depth. ## Try it yourself - **Click different filter cards** or **use the filter menu** to switch between detectors and compare how the highlighted output map changes. - **Turn off the R, G, or B channel checkboxes** and watch how the center-patch contribution and selected feature map update when that color information is removed. - **Select an edge-like filter** and then **disable one channel at a time** to see which color channel was helping that detector respond most strongly. - **Pick several different filters in sequence** and notice that the stack keeps the same number of layers overall: **each learned filter adds one feature map**. - **Press Reset channels** after experimenting so you can restore the full RGB input and compare the complete response again. ## Concept explanation In a convolutional neural network, a late **feature tensor** stores several high-level activation maps, where each map responds to a different visual pattern such as **edges**, **texture**, or **symmetry**. To make a prediction, the network often aggregates each map into a single summary value using **global average pooling**, then combines those pooled values with class-specific weights to produce **class scores**. This means a class confidence does not usually come from one pixel or one map alone; it comes from distributed evidence across many spatial locations and feature channels. ## What you see On the left, you can inspect a compact stack of activation maps that represents the late CNN feature tensor. In the center, each map is compressed into one pooled value, showing the aggregation step before classification. On the right, those pooled features are transformed into class score bars for example categories. The teal outlines mark spatial regions that are currently most influential for the selected class view, so you can see which parts of the feature tensor are contributing most strongly. ## Try it yourself - **Click a region in the front activation map** and watch nearby evidence strengthen across the tensor while the class score bars update. - **Move the Edge, Texture, Shape, and Symmetry sliders** to increase or decrease the contribution of different high-level features. - **Change the Evidence view dropdown** to inspect how the same pooled tensor supports a different class. - **Compare the pooled values in the middle with the bar lengths on the right** to see how class-specific weights turn shared features into different predictions. - **Press Reset evidence** to return to the starting state and try a new evidence pattern.