Skip to content

Commit bb2f304

Browse files
docs: add recorded screencast video link to README, process book, and finalize screencast script
1 parent fde8830 commit bb2f304

4 files changed

Lines changed: 25 additions & 39 deletions

File tree

README.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -10,7 +10,7 @@ An interactive, scroll-driven **D3.js** data story about [**FineWeb**](https://h
1010
| | |
1111
|---|---|
1212
| 🔗 **Live demo** | [com-480-data-visualization.github.io/FineWeb_Alireza](https://com-480-data-visualization.github.io/FineWeb_Alireza/) _(enable GitHub Pages to activate)_ |
13-
| 🎥 **Screencast** | _replace with your video link_ (see [`screencast/`](screencast/)) |
13+
| 🎥 **Screencast** | [▶ Watch the 2-minute video](https://drive.google.com/file/d/1bBN4NYmE1X15sft0V5nxJhfG020pXn0W/view?usp=sharing) · [script](screencast/script.md) |
1414
| 📖 **Process book** | [`process-book/process-book.pdf`](process-book/process-book.pdf) |
1515
| 📊 **Dataset** | HuggingFace `HuggingFaceFW/fineweb` · `sample-10BT` |
1616
| 💻 **Repository** | [com-480-data-visualization/FineWeb_Alireza](https://github.com/com-480-data-visualization/FineWeb_Alireza) |

process-book/process-book.html

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -464,7 +464,7 @@ <h2><span class="n">12</span>Links</h2>
464464
<table>
465465
<tr><td style="width:32%">Live visualization</td><td><code>https://com-480-data-visualization.github.io/FineWeb_Alireza/</code></td></tr>
466466
<tr><td>GitHub repository</td><td><code>https://github.com/com-480-data-visualization/FineWeb_Alireza</code></td></tr>
467-
<tr><td>Screencast (≤ 2 min)</td><td><em>add your video link before submitting</em></td></tr>
467+
<tr><td>Screencast (≤ 2 min)</td><td><a href="https://drive.google.com/file/d/1bBN4NYmE1X15sft0V5nxJhfG020pXn0W/view?usp=sharing">https://drive.google.com/file/d/1bBN4NYmE1X15sft0V5nxJhfG020pXn0W/view</a></td></tr>
468468
<tr><td>Milestone-1 EDA</td><td><code>FineWeb_EDA.ipynb</code></td></tr>
469469
</table>
470470

process-book/process-book.pdf

84 Bytes
Binary file not shown.

screencast/script.md

Lines changed: 23 additions & 37 deletions
Original file line numberDiff line numberDiff line change
@@ -1,47 +1,33 @@
1-
# 🎥 Screencast script & storyboard: "Decanting the Web"
1+
# 🎥 Screencast: "Decanting the Web"
22

3-
**Target length:** ≤ 2:00 (hard limit). **Aim for 1:50** to leave a safety margin.
4-
**Goal (per rubric):** show what the viz *does* in a fun, engaging, impactful way, and talk about **contributions and insights**, *not* technical details.
3+
**▶ Watch the 2-minute video:**
4+
https://drive.google.com/file/d/1bBN4NYmE1X15sft0V5nxJhfG020pXn0W/view?usp=sharing
55

6-
> Tips: record at 1920×1080, hide the browser bookmarks bar, use a clean profile, and
7-
> pre-scroll once to warm up fonts/animations. Scroll **slowly and smoothly** (a trackpad
8-
> or a scroll-automation tool helps). Record narration separately and lay it over the
9-
> screen capture for clean audio.
6+
A two-minute narrated walkthrough of the interactive visualization for a general,
7+
ML-curious audience. It follows the data story end to end: the curation funnel, document
8+
length, the domain galaxy, domain concentration, and the quality landscape, with the
9+
interactions demonstrated live. The beat sheet below documents the structure of the video.
1010

1111
---
1212

13-
## ⏱️ Beat sheet
13+
## Beat sheet (structure of the video)
1414

15-
| Time | On screen | Narration (read aloud) |
16-
|------|-----------|------------------------|
17-
| **0:00-0:10** | Hero section. Let the title animation breathe; cursor still. | "Every AI model is what it eats. Before it can write, it has to read, and most of what it reads comes from one dataset: **FineWeb**. This is the story of how the open web becomes AI training data." |
15+
| Time | On screen | Narration |
16+
|------|-----------|-----------|
17+
| **0:00-0:10** | Hero section. The title animation settles. | "Every AI model is what it eats. Before it can write, it has to read, and most of what it reads comes from one dataset: **FineWeb**. This is the story of how the open web becomes AI training data." |
1818
| **0:10-0:22** | Click **"Follow the journey"**; brief pause on the intro stats strip. | "We streamed two hundred thousand documents from FineWeb and measured them. Ninety-six snapshots of the web, seventy-thousand domains, twelve years. Here's what we found." |
19-
| **0:22-0:42** | **Funnel.** Scroll step-by-step so stages reveal one at a time. Hover the **MinHash dedup** stage to pop the tooltip. | "It starts with the entire web: about a hundred trillion tokens of raw noise. Watch it get filtered: extract the text keep only good English and the biggest cut of all, remove the duplicates. What survives? Just fifteen trillion tokens. **Eighty-five percent is thrown away.**" |
20-
| **0:42-0:56** | **Document length.** Land on linear view, then click **"Log scale"** yourself. | "How long is a typical web page? Tiny, about three hundred tokens. But flip to a log scale and the shape snaps into a clean bell. The web's documents are log-normal, with a few giant outliers carrying huge amounts of text." |
21-
| **0:56-1:08** | **Domain galaxy.** Let bubbles settle; click a category swatch (e.g. **News**); type a domain in **search** (e.g. "wikipedia"). | "Where does all this text come from? Each bubble is a domain: news, wikis, blogs, code. You can filter by type, or search for any site and there's Wikipedia, one of the giants." |
22-
| **1:08-1:24** | **Concentration.** Show Lorenz curve + Gini; hover once; then click **"Token coverage."** | "But the web isn't fair. A handful of domains dominate, with a Gini of nearly point-six. In fact, just **three hundred domains** supply a quarter of *all* the text a model reads." |
23-
| **1:24-1:38** | **Quality landscape.** Scroll so the dominant cell highlights. Hover the brightest cell. | "Cross length with language confidence and the whole dataset collapses into one bright region: confident English, medium length. That's the heart of what AI learns from." |
24-
| **1:38-1:52** | Scroll to **takeaways** cards; slow pan across the four cards; end on the closing line. | "So the next time a model answers you, remember: it learned to speak from a filtered, deduplicated echo of the open web. We built this to make that invisible dataset something you can actually *see*, and explore yourself." |
25-
| **1:52-2:00** | End card: title + live URL + "COM-480 EPFL". | *(silent, or)* "Decanting the Web. Thanks for watching." |
19+
| **0:22-0:42** | **Funnel**, revealed one stage at a time. Hover the **MinHash dedup** stage to pop the tooltip. | "It starts with the entire web: about a hundred trillion tokens of raw noise. Watch it get filtered: extract the text, keep only good English, and the biggest cut of all, remove the duplicates. What survives? Just fifteen trillion tokens. **Eighty-five percent is discarded.**" |
20+
| **0:42-0:56** | **Document length.** Land on the linear view, then click **"Log scale"**. | "How long is a typical web page? Tiny, about three hundred tokens. But flip to a log scale and the shape snaps into a clean bell. The web's documents are log-normal, with a few giant outliers carrying huge amounts of text." |
21+
| **0:56-1:08** | **Domain galaxy.** Let the bubbles settle; click a category swatch; search a domain (e.g. "wikipedia"). | "Where does all this text come from? Each bubble is a domain: news, wikis, blogs, code. You can filter by type, or search for any site, and there's Wikipedia, one of the giants." |
22+
| **1:08-1:24** | **Concentration.** Show the Lorenz curve and Gini, then click **"Token coverage."** | "But the web is not evenly distributed. A handful of domains dominate, with a Gini of nearly point six. Just **three hundred domains** supply a quarter of *all* the text a model reads." |
23+
| **1:24-1:38** | **Quality landscape.** Scroll so the dominant cell highlights. | "Cross length with language confidence and the whole dataset collapses into one bright region: confident English of medium length. That is the heart of what the model learns from." |
24+
| **1:38-1:52** | Scroll to the **takeaways** cards; end on the closing line. | "So the next time a model answers you, remember: it learned to speak from a filtered, deduplicated echo of the open web. We built this to make that invisible dataset something you can actually *see*, and explore yourself." |
25+
| **1:52-2:00** | End card: title, live URL, "COM-480 EPFL". | "Decanting the Web. Thanks for watching." |
2626

2727
---
2828

29-
## 🎯 What to emphasize (contributions, not code)
30-
31-
- **The funnel reveal** is the signature moment, so give it room.
32-
- Show **interactivity you built**: the log toggle, the galaxy search/filter, the Lorenz↔coverage toggle. Each click should feel effortless.
33-
- Land **one number per scene** so the viewer leaves with concrete facts (100T→15T, ~340 tokens, 300 domains, 92% ≥0.95).
34-
- Tone: confident, curious, a little playful. You're a guide, not a lecturer.
35-
36-
## ✂️ If you're over time
37-
1. Merge the intro beat (0:10-0:22) into the hero.
38-
2. Cut the quality-landscape beat to ~8s (it's the most technical).
39-
3. Tighten the galaxy beat to a single search action.
40-
41-
## 🎙️ Production checklist
42-
- [ ] 1080p, 30fps, clean browser, full-screen site.
43-
- [ ] Smooth slow scrolling (consider a scroll-automation bookmarklet).
44-
- [ ] Narration recorded separately, normalized, light noise reduction.
45-
- [ ] Soft background music bed at ~15% volume (optional).
46-
- [ ] Final length **≤ 2:00**. Export H.264 MP4.
47-
- [ ] Upload (YouTube unlisted / Drive) and paste the link into the README and process book.
29+
## Highlights
30+
31+
- The **funnel reveal** is the signature moment of the piece.
32+
- Interactions demonstrated: the linear / log toggle on the length histogram, the domain-galaxy search and category filter, and the Lorenz / token-coverage toggle.
33+
- One concrete figure per scene: 100T to 15T tokens, a median near 340 tokens, about 300 domains for a quarter of all tokens, and 92% of documents scoring at least 0.95.

0 commit comments

Comments
 (0)