All 617 dog breeds on Wikipedia's list, the 62 extinct or critically endangered ones flagged, one row each: origin, coat and colour from the breed infobox, height and weight as written plus low/high/mid numeric twins in both cm/in and kg/lb (male and female figures merged into the breed's full range, the missing unit system converted), and life expectancy extracted from article prose since Wikipedia's infobox dropped the field, with the exact source phrase kept alongside the numeric twins so study-specific figures keep their context.
617 rows in one CSV file, with 24 columns in total: Breed, Status, Origin, Coat, Colour, Height, Height_Low_Cm, Height_High_Cm, Height_Mid_Cm, Height_Low_In, Height_High_In, Height_Mid_In, Weight, Weight_Low_Kg, Weight_High_Kg, Weight_Mid_Kg, Weight_Low_Lb, Weight_High_Lb, Weight_Mid_Lb, Life_Span, Life_Span_Low_Years, Life_Span_High_Years, Life_Span_Mid_Years, Breed_URL.
It is built by programmatically scraping en.wikipedia.org, last pulled on 2026-09-02. Datasets are versioned; older versions stay downloadable.
Yes. Every CodeSights dataset is completely free as a CSV download; a free account is all it takes. Anyone can preview the data without signing in.
Yes. The exact Python scraper that built it is viewable on the dataset page by any signed-in member, so every number is reproducible.
Automated scraping leaves room for error and the underlying sources change over time, so no version is guaranteed accurate or complete. If a number matters, verify it against the original source.