One row per star with an IAU-recognized proper name: 653 names from Wikipedia's list of proper names of stars, the WGSN catalog plus the NameExoWorlds additions, and the dataset where astronomy becomes social science. The Name_Origin column is dense on every row and answers the famous question with receipts: Arabic leads with 209 names (Betelgeuse, Aldebaran, the whole medieval translation chain), then Chinese (67), Latin (63), Greek (40) and Sumerian (16), the night sky as a history of who was doing astronomy when. Each star's own article Starbox fills the physics on 354 rows: distance in light-years (Proxima Centauri nearest at 4.2, with Rigil Kentaurus and Toliman right behind, which is the Alpha Centauri system agreeing with itself), apparent magnitude (Sirius brightest at -1.46, and the parser learned that the sky writes minus as U+2212), and spectral type (Betelgeuse's M means red). Constellation and designation complete each row; physics coverage is the articles' coverage, honestly reported. Built with a 9-test acceptance suite that pins Sirius as the brightest row in the file, Proxima as the nearest, Betelgeuse's Arabic name and red class, and Arabic's overall dominance.
653 rows in one CSV file, with 9 columns in total: Star, Designation, Constellation, Name_Origin, Distance_ly, Apparent_Magnitude, Spectral_Type, Notes, Star_URL.
It is built by programmatically scraping en.wikipedia.org, last pulled on 2026-09-22. Datasets are versioned; older versions stay downloadable.
Yes. Every CodeSights dataset is completely free as a CSV download; a free account is all it takes. Anyone can preview the data without signing in.
Yes. The exact Python scraper that built it is viewable on the dataset page by any signed-in member, so every number is reproducible.
Automated scraping leaves room for error and the underlying sources change over time, so no version is guaranteed accurate or complete. If a number matters, verify it against the original source.