## Concept explanation A **diffusion model** learns to reverse a gradual corruption process. You can think of training as repeatedly taking a clean image, adding a little more **noise**, and asking the model to predict how to remove that noise step by step. When the image is only slightly corrupted, much of the original structure is still present; as the noise grows, the signal becomes harder to recognize. The key idea is that if information can be destroyed in many tiny increments, then a model can learn to recover it in many tiny denoising increments too. ## What you see You are comparing the same image in two states: the left side stays clean as the reference signal, while the right side shows that image after a chosen amount of corruption. As you move the slider, the noisy image updates continuously, so you can watch edges, shapes, and colors fade into randomness. The meter below gives a rough visual cue for how much recognizable structure remains, helping you connect “more noise” with “less recoverable information.” ## Try it yourself - **Drag the noise slider slowly upward** and notice which features disappear first: fine lines, sharp edges, or large colored regions. - **Slide back downward** to imagine the reverse process a diffusion model tries to learn: removing a little noise at a time instead of all at once. - **Switch the corruption style** to compare smooth grain with chunkier damage, and see that different noise patterns can hide the same underlying image. - **Click `New sample`** to keep the same noise level but change the random corruption pattern, showing that many noisy images can come from the same clean source. - **Pause near very high noise levels** and ask yourself how much of the original image is still visibly present, versus how much would need to be inferred by a learned denoising model. ## Concept explanation A **diffusion model** does not just look at a noisy image and guess what to remove. It also needs the **timestep** (or noise level) to know *how corrupted* that image is. The same visible pattern can require very different denoising behavior depending on whether the model is at an early, lightly noisy stage or a late, heavily noisy stage. Without `t`, the model would not know whether to make tiny detail-preserving corrections or large structure-recovering ones. ## What you see You are looking at one image sample in the center, with its noise amount controlled by the selected timestep on the timeline below. As you move through time, the nearby input panel shows that the model receives both `x_t` and `t`. That pairing is the key idea: the noisy image alone is ambiguous, but the noisy image plus its timestep tells the model what kind of denoising step to apply. ## Try it yourself - **Drag the timeline handle** left and right to move between cleaner and noisier stages of the diffusion process. - **Watch the center image** as the corruption grows or shrinks, and notice that the same underlying shape is still there. - **Look at the model input panel** and compare how the message changes at low vs. high `t`. - **Use the timestep slider** in the control panel to set an exact stage and see that it matches the draggable handle. - **Switch the base shape** to test whether the idea holds for different image content. - **Press `New sample`** to change the random noise pattern while keeping the same lesson: the model still needs both the noisy image and the timestep to denoise correctly. ## Concept explanation A **Diffusion Transformer (DiT)** does not usually process an image as one huge full-resolution grid of pixels all at once. Instead, it splits the image into small square **patches**, and each patch becomes a **token** in a sequence, similar to how words become tokens in a language model. This creates a trade-off: **smaller patches** preserve more local detail but produce a longer token sequence, while **larger patches** reduce sequence length but compress more of the image into each token. ## What you see You are looking at a stylized image on the left with a patch grid drawn over it. Each square in that grid corresponds to one token card in the sequence on the right, arranged in reading order from top-left to bottom-right. As the patch size changes, the grid becomes finer or coarser, and the token cards update to show how the same image can be represented with many small tokens or fewer large ones. ## Try it yourself - **Move the patch size slider** to make the squares smaller, and notice how the token sequence grows because the image is being broken into more pieces. - **Move the patch size slider** the other way to make the squares larger, and watch the sequence shorten as each token covers more of the image. - **Compare the grid and the token count** to see that `tokens = patches across × patches down`. - **Switch the scene dropdown** to test the same tokenization idea on different image content. - **Press Reset** to return to a middle patch size and compare from a neutral starting point. ## Concept explanation In a transformer, **attention** lets each patch token compare itself with every other token and pull in useful information from patches that look related, even if they are far apart in the image. That matters for denoising because a noisy patch does not have to guess from only its local pixels: it can borrow **global context** from distant patches with similar structure, helping the model reconstruct a cleaner image. ## What you see On the left, each node is a patch token from a noisy image, and the faint web of lines represents possible attention links across the whole image. When you click one token, its strongest connections brighten, showing which other patches it relies on most. On the right, the reconstructed image is assembled from the patch tokens, and the highlighted squares show which image regions are sharing information most strongly for the current focus token. ## Try it yourself - **Click different tokens** in the graph and watch how the strongest attention links jump to distant patches with similar structure. - **Increase the `Noisy input` slider** to make the patches less reliable, then notice how the highlighted links become more important for interpreting the image. - **Adjust the `Attention blend` slider** to see how strongly each patch is pulled toward information from its connected neighbors. - **Switch the `Underlying pattern`** between `Face-like`, `Stripes`, and `Corners` to compare how attention finds long-range relationships in different image structures. - **Press `Reset focus`** to return to the central patch and compare its context-sharing pattern with edge or corner patches. ## Concept explanation A **Diffusion Transformer (DiT)** starts from noisy latent tokens and learns a **denoising direction** for each step. With **external conditioning** such as a class label or text concept, the model does not change the starting noise itself — instead, it changes how the transformer interprets that same noise. The condition acts like a semantic hint that biases attention and feature mixing, so the next denoising move points toward “cat,” “car,” or “house” structure even when the initial token sequence is unchanged. ## What you see On the left, you see one shared noisy token sequence. The colored condition token feeds into the transformer block and strengthens some tokens more than others, which is why certain tokens glow more strongly and receive stronger steering arrows. On the right, the preview panel shows the kind of structure that the current denoising step is being pulled toward, along with a compact readout of the most influenced tokens. Switching conditions keeps the noisy start fixed while changing the semantic direction of the update. ## Try it yourself - **Switch the condition** between `cat`, `car`, and `house` and notice how the same token row gets reweighted differently. - **Watch which tokens brighten most** and compare them with the ranked influence bars in the preview panel. - **Increase the guidance strength** to see the denoising direction become more decisive and structured. - **Raise the noise level** and observe how the preview becomes wobblier when the condition is not strongly enforced. - **Adjust the attention pulse** to make the token-to-block activity feel more or less active over time. - **Press `Reset noise`** to generate a fresh noisy starting point, then test how each condition steers that new sequence. ## Concept explanation A **Diffusion Transformer (DiT)** is trained to predict a target from a noisy sample at a chosen **timestep**. A common target is the **added noise** itself: if the model can estimate what noise was injected, you can subtract that estimate and make the sample cleaner. You can also ask the model to predict the **clean signal**, but noise prediction is often a more stable training objective because the target stays well defined across many noise levels. ## What you see You are looking at three linked panels. The left panel shows the noisy input the model receives, the middle panel shows what the model is trained to predict, and the right panel shows the result after using that prediction to denoise. As you change the timestep, the amount of corruption changes; when you switch the prediction target, the middle and right panels update so you can compare how well each target supports cleaning at early versus late denoising stages. ## Try it yourself - **Move the timestep slider** toward a high value and notice how the noisy input becomes heavily corrupted while the noise-prediction target still leads to a strong cleaned result. - **Switch the target to `Predict clean signal`** and compare the cleaned panel with the noise-target case at the same timestep. - **Drag the timestep slider back to a low value** and see that both targets become easier, but the noise target still gives a clear subtract-and-clean interpretation. - **Toggle between the two targets several times** while watching the middle panel, and notice that predicting noise matches the quantity you want to remove from the input. - **Click `New example`** to test the same training idea on a different underlying pattern. ## Concept explanation In a **Diffusion Transformer (DiT)**, generation starts from random noise and improves the sample through repeated **denoising steps**. Each step gives the model another chance to remove noise, sharpen structure, and stabilize the image, so using more steps usually improves **quality**. But every extra step also adds computation, which increases **sampling time**. That creates a trade-off: fewer steps make generation faster, while more steps often produce cleaner and more reliable results. ## What you see You’re looking at a synthetic image preview that evolves from noisy blur toward a more coherent result as sampling progresses. The timeline shows where you are in the denoising sequence, the snapshot panel gives quick views of early, middle, and late stages, and the trade-off card summarizes the current outcome. On the right, the step slider changes the total denoising budget, while the quality and time meters update so you can compare speed against refinement. ## Try it yourself - **Press `Play sampling`** and watch the sample move from noise toward structure across the denoising timeline. - **Lower the total-step slider** to a small value and replay the animation to see a faster run that leaves the image blurrier or less stable. - **Raise the total-step slider** and replay again to notice how extra denoising passes improve fine detail and consistency. - **Compare the blue quality meter and gold time meter** after each replay to see why better samples usually cost more compute. - **Use `Show final frame`** to jump straight to the end state for the current step budget, then change the slider and compare the final result immediately.