Two joinable datasets covering every American film from the 2000-2025 Wikipedia year lists: 6,842 films (opening date, production company, director, stars, runtime, country, budget and box office in millions, and an Academy Award winner flag derived from the Academy Award dataset by strict title-plus-year matching) and 41,078 star-per-film credits covering 12,027 stars, each enriched from their own article: ages at release, career spans, birthplaces, marriages and children. The largest dataset on the platform.
47,920 rows across 2 joinable files, with 31 columns in total: Title, List Year, Release Date, Production Company, Director, Stars, Run Time (min), Country, Budget (millions), Budget Range, Box Office (millions), Academy Award Winner, Wikipedia Link, Movie Title, Movie Release Date, Star, Star Wikipedia Link, Birth Date, Age at Release, Death Date, Years From Release to Death, Year Started Acting, Year Ended Acting, Years Acting at Release, Birth City, Birth State, Birth Country, Number of Marriages, Avg Marriage Length (years), Number of Children, Age Started Acting.
It is built by programmatically scraping en.wikipedia.org, last pulled on 2026-08-31. Datasets are versioned; older versions stay downloadable.
Yes. Every CodeSights dataset is completely free as a CSV download; a free account is all it takes. Anyone can preview the data without signing in.
Yes. The exact Python scraper that built it is viewable on the dataset page by any signed-in member, so every number is reproducible.
Join the files on movie title + release date, the family key shared with the studio and Academy Award movie datasets.
Automated scraping leaves room for error and the underlying sources change over time, so no version is guaranteed accurate or complete. If a number matters, verify it against the original source.