One row per metro system on Earth, operational or under construction: 251 systems across 67 countries, from Wikipedia's list of metro systems enriched by each system's own article. Every system carries its city and country, opened and last-expanded years, station and line counts, route length in km, the latest annual ridership in millions with the year it was measured, daily ridership and operator from the article infobox, and the derived efficiency columns the list only implies: riders per km and riders per station. The 17 systems under construction ride along with a Status column, construction-start and projected-opening years, so China's pipeline is visible next to its operating boom: 47 Chinese systems, nearly all opened since 2000, including both giants (Shanghai the busiest at 3.77 billion annual riders, Beijing the longest at 909 km, the crown having changed hands mid-decade). London Underground anchors the old end at 1890, dated by electric metro service. This dataset began life as a ridership-by-year panel until a probe showed Wikidata holds a median of two dated ridership points per system; it ships as what the data honestly supports, a snapshot done properly. Built with a 15-test acceptance suite that learned Beijing had overtaken Shanghai and that Pyongyang's latest ridership estimate is from 2009.
251 rows in one CSV file, with 19 columns in total: System, City, Country, Status, Opened_Year, System_Age_Years, Last_Expanded_Year, Construction_Started_Year, Projected_Opening_Year, Stations, Lines, Length_KM, Annual_Riders_Millions, Ridership_Year, Riders_Per_KM_Millions, Riders_Per_Station_Millions, Daily_Riders, Operator, System_URL.
It is built by programmatically scraping en.wikipedia.org, last pulled on 2026-09-12. Datasets are versioned; older versions stay downloadable.
Yes. Every CodeSights dataset is completely free as a CSV download; a free account is all it takes. Anyone can preview the data without signing in.
Yes. The exact Python scraper that built it is viewable on the dataset page by any signed-in member, so every number is reproducible.
Automated scraping leaves room for error and the underlying sources change over time, so no version is guaranteed accurate or complete. If a number matters, verify it against the original source.