← All datasets

The Named Stars

One row per star with an IAU-recognized proper name: 653 names from Wikipedia's list of proper names of stars, the WGSN catalog plus the NameExoWorlds additions, and the dataset where astronomy becomes social science. The Name_Origin column is dense on every row and answers the famous question with receipts: Arabic leads with 209 names (Betelgeuse, Aldebaran, the whole medieval translation chain), then Chinese (67), Latin (63), Greek (40) and Sumerian (16), the night sky as a history of who was doing astronomy when. Each star's own article Starbox fills the physics on 354 rows: distance in light-years (Proxima Centauri nearest at 4.2, with Rigil Kentaurus and Toliman right behind, which is the Alpha Centauri system agreeing with itself), apparent magnitude (Sirius brightest at -1.46, and the parser learned that the sky writes minus as U+2212), and spectral type (Betelgeuse's M means red). Constellation and designation complete each row; physics coverage is the articles' coverage, honestly reported. Built with a 9-test acceptance suite that pins Sirius as the brightest row in the file, Proxima as the nearest, Betelgeuse's Arabic name and red class, and Arabic's overall dominance.
653 rows 1 file last pulled 2026-09-22
Stars 653 rows × 9 columns
StarDesignationConstellationName_OriginDistance_lyApparent_MagnitudeSpectral_TypeNotesStar_URL

Questions and answers

What is the The Named Stars dataset?

One row per star with an IAU-recognized proper name: 653 names from Wikipedia's list of proper names of stars, the WGSN catalog plus the NameExoWorlds additions, and the dataset where astronomy becomes social science. The Name_Origin column is dense on every row and answers the famous question with receipts: Arabic leads with 209 names (Betelgeuse, Aldebaran, the whole medieval translation chain), then Chinese (67), Latin (63), Greek (40) and Sumerian (16), the night sky as a history of who was doing astronomy when. Each star's own article Starbox fills the physics on 354 rows: distance in light-years (Proxima Centauri nearest at 4.2, with Rigil Kentaurus and Toliman right behind, which is the Alpha Centauri system agreeing with itself), apparent magnitude (Sirius brightest at -1.46, and the parser learned that the sky writes minus as U+2212), and spectral type (Betelgeuse's M means red). Constellation and designation complete each row; physics coverage is the articles' coverage, honestly reported. Built with a 9-test acceptance suite that pins Sirius as the brightest row in the file, Proxima as the nearest, Betelgeuse's Arabic name and red class, and Arabic's overall dominance.

How big is the The Named Stars dataset?

653 rows in one CSV file, with 9 columns in total: Star, Designation, Constellation, Name_Origin, Distance_ly, Apparent_Magnitude, Spectral_Type, Notes, Star_URL.

Where does the data come from?

It is built by programmatically scraping en.wikipedia.org, last pulled on 2026-09-22. Datasets are versioned; older versions stay downloadable.

Is the dataset free to download?

Yes. Every CodeSights dataset is completely free as a CSV download; a free account is all it takes. Anyone can preview the data without signing in.

Can I see the code that built this dataset?

Yes. The exact Python scraper that built it is viewable on the dataset page by any signed-in member, so every number is reproducible.

Can the data contain errors?

Automated scraping leaves room for error and the underlying sources change over time, so no version is guaranteed accurate or complete. If a number matters, verify it against the original source.

Browse every free dataset on CodeSights