| Student's name | SCIPER |
|---|---|
| Maëlys Bühler | 387090 |
| Nima Dekhli | 326426 |
| Alexandre Maquet | 394705 |
Milestone 1 • Milestone 2 • Milestone 3
10% of the final grade
This is a preliminary milestone to let you set up goals for your final project and assess the feasibility of your ideas. Please, fill the following sections about your project.
(max. 2000 characters per section)
Find a dataset (or multiple) that you will explore. Assess the quality of the data it contains and how much preprocessing / data-cleaning it will require before tackling visualization. We recommend using a standard dataset as this course is not about scraping nor data processing.
Hint: some good pointers for finding quality publicly available datasets (Google dataset search, Kaggle, OpenSwissData, SNAP and FiveThirtyEight).
The dataset we chose for our project is the GTFS timetable of the Swiss train network of this year. It is made available on the opentransportdata.swiss platform. This provides us with a comprehensive overview of the train routes, schedules, connections and station location in Switzerland.
The GTFS (General Transit Feed Specification) format is widely used for public transportation data. It consists of several text files, each containing specific information about the transit system. Some of the key files include all the stops, routes, trips, and stop times. The files are linked together through unique identifiers, allowing us to reconstruct the entire transit network and its schedule.
Since the dataset is provided by the official swiss public transportation authority and since it is being used by many applications and services, we can assume that the data is of very high quality. It is updated twice a week.
As is, the GTFS dataset is clean and well structured, but it is raw and not directly usable for our visualization. Moreover, the dataset is quite large and contains many details and informations that are not relevant for our project. We will need to process the data to extract the relevant information and create a more compact and efficient representation of the train network. Moreover, since all the app rendering will be executed on the client side, we need to make sure to optimize the data for performance and avoid unnecessary data transfers.
Frame the general topic of your visualization and the main axis that you want to develop.
- What am I trying to show with my visualization?
- Think of an overview for the project, your motivation, and the target audience.
This visualization aims to create an interactive, gamified experience exploring the Swiss train network. The objective is to expand users' understanding beyond their usual routes, encouraging discovery of the entire national network.
The concept is a game where users start at a random Swiss station with a specific destination objective. They must navigate the network by selecting the correct trains to reach their target. Players are provided with the distance (in km) and the fastest possible travel time to the destination, but the target station's location remains hidden on the map.
Users must easily identify available trains from their current location, understand their routes, and see all intermediate stops. With each move, the remaining distance to the objective updates dynamically, guiding the player toward their goal.
The visualization features an interactive map displaying visited stations and accessible next stops. Zoom functionality allows users to contextualize their general location within Switzerland while examining local connections from their current station.
The primary target audience includes train enthusiasts eager to explore the network through gameplay. The design ensures accessibility for users with varying levels of knowledge, offering an engaging learning experience for beginners while providing a strategic advantage to those already familiar with the Swiss network.
Pre-processing of the data set you chose
- Show some basic statistics and get insights about the data
In the Jupyter Notebook located in src/gtfs.ipynb, we did a first basic processing of
the dataset. In particular, we managed to dynamically create the departure board
for a given train station in Switzerland at a particular date and hour. For instance,
for the Renens VD train station at 09:00 AM, we get the following trains:
09:04 IR95 Brig
09:04 R3 Vallorbe
09:06 RE33 Annemasse
09:08 IC1 St. Gallen
09:09 R9 Allaman
09:11 IC51 Basel SBB
09:16 R1 Grandson
09:16 R2 Bex
09:20 IC5 Lausanne
09:21 R8 Lausanne
...
- What others have already done with the data?
This dataset is used in both the scientific and commercial domains. For instance, the smart ticketing company Fairtiq uses timetables and transport networks data to optimize journey detection within their automatic fare billing application. On a more academic side, the Institute for Transport Planning and Systems at ETH Zürich frequently uses this dataset to conduct studies on the Swiss railway network.
Regarding its previous use within the Data Visualization course, a group of students carried out in 2025 a project called Cartarail aiming to visualize the impact of terrain geography and relief on the travel times of various railways connections.
- Why is your approach original?
GTFS datasets are primarily used in an utilitarian manner, such as in applications that provide services, schedules and itineraries to passengers, or analytically for research purposes, like optimizing train networks and understanding transportation flows.
Our approach is original because it repurposes this data to create a playful experience. The goal is to encourage the user to explore the rail network in an interactive and gamified way, rather than presenting them with static information through traditional charts and metrics.
- What source of inspiration do you take? Visualizations that you found on other websites or magazines (might be unrelated to your data).
Our first source is GeoGuessr, a browser-based game where the user is dropped at a random location on the planet. Using Google Maps Street View only, they must successfully orient themselves within their environment to deduce their precise location on Earth. This idea of initial disorientation and spatial investigation perfectly matches the starting mechanic of our project.
Our second inspiration is Wiki Game, which requires for the user, starting from a random Wikipedia article, to navigate to another article on a precise subject as a destination using only hyperlinks. This node-based navigation constraint directly inspires the core mechanic of our project, which is to connect two train stations navigating by exclusively through the connections of the railway network.
- In case you are using a dataset that you have already explored in another context (ML or ADA course, semester project...), you are required to share the report of that work to outline the differences with the submission for this class.
We haven't used this dataset in a previous project
10% of the final grade
The milestone 2 document can be found here.
The deployed prototype can be found at https://com-480-data-visualization.github.io/via/.
Ensure you have NPM installed on your machine. Then run the following commands:
cd via-ui
npm install
npm run devThe website is accessible on http://localhost:5173.
80% of the final grade
Live demo: https://youtu.be/WjtaBsId2Pg
Process book: processbook.pdf
Website: https://com-480-data-visualization.github.io/via/.
Ensure you have NPM installed on your machine. Then run the following commands:
cd via-ui
npm install
npm run devThe website is accessible on http://localhost:5173.
pfaedle -x switzerland-260518.osm.pbf -i via/data/gtfs_fp2026_20260311.zip -m rail -c pfaedle/pfaedle.cfg
python3 via/src/filter.py ./gtfs-out ./via/via-ui/public/data
- < 24h: 80% of the grade for the milestone
- < 48h: 70% of the grade for the milestone