← All datasets

Large Language Models

One row per model on Wikipedia's list of large language models: 167 models from GPT-1's 117 million parameters (June 11, 2018) to the current frontier, the arms race this platform's own audience is living through. Everything ships as written PLUS parsed numeric twins: Parameters_B from the stated count (mixture-of-experts entries multiply out, disclosed; 99 models state a count and the frontier labs' silence after 2022 is itself a finding), Corpus_Tokens_B where the corpus is stated in tokens, and Training_Cost_M_USD only where the table gives dollars, 25 honest values, because the 2025 tables dropped the column and a naive parse once read Apache 2.0 as two million dollars. Release dates are ISO where the day is known, including the short dates whose year lives in the section heading (DeepSeek-R1's Jan 20 means 2025-01-20), with a year twin on every row. Open_Weights reads the license: 93 of 167 are open, and the openness-by-year and openness-by-lab splits are the first charts anyone will make. Anthropic (17), OpenAI (14) and Google DeepMind (12) lead the developer column; the parameter explosion runs three orders of magnitude in four years and the 10-test acceptance suite pins GPT-1, GPT-3's 175B and DeepSeek-R1's MIT license. This dataset dates fast by nature; repulls will version it.
167 rows 1 file last pulled 2026-09-22
Models 167 rows × 13 columns
ModelDeveloperRelease_DateRelease_YearParameters_BParameters_As_WrittenCorpusCorpus_Tokens_BTraining_Cost_M_USDLicenseOpen_WeightsNotesModel_URL

Questions and answers

What is the Large Language Models dataset?

One row per model on Wikipedia's list of large language models: 167 models from GPT-1's 117 million parameters (June 11, 2018) to the current frontier, the arms race this platform's own audience is living through. Everything ships as written PLUS parsed numeric twins: Parameters_B from the stated count (mixture-of-experts entries multiply out, disclosed; 99 models state a count and the frontier labs' silence after 2022 is itself a finding), Corpus_Tokens_B where the corpus is stated in tokens, and Training_Cost_M_USD only where the table gives dollars, 25 honest values, because the 2025 tables dropped the column and a naive parse once read Apache 2.0 as two million dollars. Release dates are ISO where the day is known, including the short dates whose year lives in the section heading (DeepSeek-R1's Jan 20 means 2025-01-20), with a year twin on every row. Open_Weights reads the license: 93 of 167 are open, and the openness-by-year and openness-by-lab splits are the first charts anyone will make. Anthropic (17), OpenAI (14) and Google DeepMind (12) lead the developer column; the parameter explosion runs three orders of magnitude in four years and the 10-test acceptance suite pins GPT-1, GPT-3's 175B and DeepSeek-R1's MIT license. This dataset dates fast by nature; repulls will version it.

How big is the Large Language Models dataset?

167 rows in one CSV file, with 13 columns in total: Model, Developer, Release_Date, Release_Year, Parameters_B, Parameters_As_Written, Corpus, Corpus_Tokens_B, Training_Cost_M_USD, License, Open_Weights, Notes, Model_URL.

Where does the data come from?

It is built by programmatically scraping en.wikipedia.org, last pulled on 2026-09-22. Datasets are versioned; older versions stay downloadable.

Is the dataset free to download?

Yes. Every CodeSights dataset is completely free as a CSV download; a free account is all it takes. Anyone can preview the data without signing in.

Can I see the code that built this dataset?

Yes. The exact Python scraper that built it is viewable on the dataset page by any signed-in member, so every number is reproducible.

Can the data contain errors?

Automated scraping leaves room for error and the underlying sources change over time, so no version is guaranteed accurate or complete. If a number matters, verify it against the original source.

Browse every free dataset on CodeSights