Every dinosaur genus ever coined: all 1,881 entries on Wikipedia's list, with the list's own verdicts kept as a Status column (1,428 valid, plus junior synonyms, nomina nuda, preoccupied names, and the 111 that turned out not to be dinosaurs at all, one famously a piece of petrified wood). Each article's taxobox supplies the full classification, one column per Linnaean rank plus every Clade row comma-joined in order, the type species, the naming authority, and the years the genus and species were established. The temporal range ships as written plus numeric From/To Mya; length as written plus low/high/mid meter twins; the discovery year parses from the Discovery and History sections (never postdating the naming); diet is captured only where the article states it. Synonym entries that link to another genus's article keep only name, status and URL so no row wears borrowed taxonomy. Built with a 9-test known-facts suite.
1,881 rows in one CSV file, with 26 columns in total: Name, Status, Kingdom, Phylum, Class, Clades, Order, Suborder, Superfamily, Family, Subfamily, Tribe, Type_Species, Named_By, Year_Genus_Established, Year_Species_Established, Discovery_Year, Temporal_Range, From_Mya, To_Mya, Length, Length_Low_M, Length_High_M, Length_Mid_M, Diet, Genus_URL.
It is built by programmatically scraping en.wikipedia.org, last pulled on 2026-09-07. Datasets are versioned; older versions stay downloadable.
Yes. Every CodeSights dataset is completely free as a CSV download; a free account is all it takes. Anyone can preview the data without signing in.
Yes. The exact Python scraper that built it is viewable on the dataset page by any signed-in member, so every number is reproducible.
Automated scraping leaves room for error and the underlying sources change over time, so no version is guaranteed accurate or complete. If a number matters, verify it against the original source.