Skip to sources
Time Machine

Statisticians

Scrub through 197years of this role's history, from when it first emerged, through every wave of technology that reshaped it, to the cited projections for where it's heading next.

2026drag to travel through time
1850187519001925195019752000now
2026
Known today as Statisticians (BLS SOC 15-2041; parallel emergence of the Data Scientist label for the ML sub-specialty)
Latest actual · 2024
32K
BLS OEWS May 2024, sourced from O*NET which reflects the same BLS establishment-survey figure. Employment has grown 69% from the 2000 SOC-introduction baseline (19,000 in 2000), driven by pharmaceutical clinical-trial expansion, the AI-validation demand surge of the 2020s, and the growing use of survey-based official statistics across government programs. The 32,200 figure applies to the narrow BLS SOC 15-2041 definition; the broader workforce of people doing statistician-level work under other titles (biostatistician, clinical data scientist, survey methodologist) is estimated at two to three times this figure.
Latest actual · 2024
$103,300
BLS OEWS May 2024 median annual wage, sourced from O*NET which reflects the same BLS establishment-survey figure. At $103,300, statisticians earn 2.3x the 2024 all-occupations median ($48,060). The profession's real wages have grown substantially since 2000: the 2024 nominal wage is 99% higher than 2000, while the CPI increased roughly 70% over the same period, implying meaningful real-wage growth. The top 10% of statisticians earned over $170,000 in May 2024, reflecting the premium commanded by clinical statisticians with FDA-submission experience and senior government survey methodologists.
Each dot is a cited figure over time; the dotted line only links them (values between aren't measured). Hollow dots are estimates.
Beat · 2025

The American Statistical Association publishes "AI and the Future of Statistics," a policy white paper positioning the profession as the natural validator of AI outputs. The white paper identifies four irreducible human-statistician functions: (1) authoring statistical analysis plans for FDA-regulated clinical trials, where legal accountability cannot be delegated; (2) designing government survey sampling frames, where political defensibility requires human domain knowledge; (3) expert witness testimony in legal and regulatory proceedings; and (4) translating domain-science questions into valid statistical designs. The ASA's framing is explicitly optimistic: AI tools augment the routine code-writing and exploration work, freeing statisticians to concentrate on the judgment-intensive tasks that define the profession's value.

Tools of the era

The tools that defined the work

Select an era to see how it reshaped the work.

  • Ledgers, mortality tables, and mechanical calculators (Babbage-era statistics)

    The statistician of the mid-19th century worked with pen, paper, and actuarial tables. The primary instrument was the published mortality table and the census ledger. Logarithm tables (published since the 1600s) accelerated multiplication; mechanical calculators such as the Comptometer (patented 1887) and later the Marchant (1910s) could add, subtract, and multiply faster than a clerk could do by hand. The actual statistical methods available were almost entirely descriptive: means, proportions, frequency distributions. The inferential revolution of Fisher and Pearson lay decades ahead. This era's statistician was above all a compiler and summarizer of public facts, not a hypothesis tester.

    Ledger workPaper recordkeeping
  • Fisher's inference framework: significance tests, ANOVA, experimental design (1925-1935 publications)

    Ronald Fisher's "Statistical Methods for Research Workers" (1925) and "The Design of Experiments" (1935) gave statisticians their core intellectual toolkit: the F-test, analysis of variance, randomized controlled trials, and the concept of a null hypothesis. Egon Pearson and Jerzy Neyman formalized hypothesis testing in the 1930s, introducing the framework of Type I and Type II errors and statistical power. These were not just theoretical advances: they defined what a statistician did for the next seven decades. The statistician's job became designing experiments, computing test statistics by hand (or later with mechanical calculators), and advising scientists on whether their results met the significance threshold. The institutionalization of this framework in university statistics departments (Iowa State 1927, North Carolina State 1941, Stanford 1948) created a reproducible pipeline from graduate education to professional employment.

    Effect on the work

    The Fisher-Neyman-Pearson framework created a sustainable professional niche for statisticians that no other scientific role could easily fill: the authority to design and validate quantitative studies. This framework anchored statistician employment in academia and government research for the following 50 years.

    Work toolChanging equipment
  • Wartime statistical mobilization: Sequential Analysis, operations research, sampling theory

    The Statistical Research Group at Columbia University (1942-1945) brought 18 of America's leading statisticians together to solve military problems: Abraham Wald developed sequential analysis, which allowed military inspectors to make accept/reject decisions on manufactured items with fewer samples than fixed-sample plans required. Wald also solved the survivorship-bias problem for aircraft armor placement. W. Allen Wallis and Frederick Mosteller applied sampling and experimental design to ammunition testing. The war demonstrated, at institutional scale, that statisticians could save lives and resources in ways that no other professional category could replicate. The postwar expansion of the National Science Foundation (1950), NIH research funding, and the FDA's increasing reliance on clinical trials was a direct institutional consequence of this wartime proof-of-concept.

    Effect on the work

    WWII established statistics as a strategically important profession in the eyes of federal agencies and defense contractors, driving a sustained postwar expansion of federal statistical employment and academic statistics programs.

    Work toolChanging equipment
  • Mainframe statistical software: SPSS (1968), SAS (1972/1976), and the programming-language era

    SPSS (Statistical Package for the Social Sciences) was first released in 1968 at Stanford, commercialized in 1975, and had 600 organizational users by 1975 including NASA and Procter and Gamble. SAS was conceived in 1966, first released in 1968 at North Carolina State, and incorporated as SAS Institute in 1976. Both packages moved statistical computation from hand calculation and mechanical calculators onto mainframes and, by the mid-1980s, onto personal computers. This transformation had two effects: it dramatically expanded what a single statistician could compute (enabling regression models with hundreds of variables, bootstrapping, and simulation), and it made basic statistical analysis accessible to researchers who were not professional statisticians. The number of people performing statistical analyses grew much faster than the number of credentialed statisticians, increasing both demand for statistical expertise and competition from adjacent roles.

    Effect on the work

    SAS and SPSS created a new skills premium for statisticians who could translate between methodological validity (what the analysis should do) and software implementation (what the code actually computes). The gap between a statistician and a researcher who happened to use SAS became the primary professional moat for the occupation from the 1970s through the 1990s.

    Mainframe processingComputerized records
  • R language and the open-source statistics revolution (S language 1976, R 1993)

    Bell Labs's S language (1976) and its open-source successor R (first released 1993, CRAN repository launched 1997) transformed the economics of statistical software. Where SAS and SPSS charged institutional licenses measured in thousands of dollars per seat, R was free. By the mid-2000s, R had become the dominant tool for academic statistics and a growing force in pharmaceutical biostatistics. The CRAN package ecosystem grew from a few dozen packages in 1997 to over 2,000 by 2007. R enabled a generation of statisticians to build and share novel methods at a pace that proprietary software could not match. The Bayesian revival of the 1990s-2000s (BUGS, MCMC samplers, hierarchical models) was powered largely by R and free-standing Bayesian packages. Python entered the statistical toolkit seriously around 2008-2010 with NumPy, SciPy, and eventually pandas.

    Effect on the work

    R commoditized the computation side of statistics while raising the ceiling on methodological sophistication. The statistician's professional moat shifted further from "can run SAS" toward "can design the study and interpret the results," because any researcher with internet access could now run the models.

    Work toolChanging equipment
  • Big Data and the Data Scientist divergence (Hadoop 2006, scikit-learn 2007, the "sexiest job" era)

    In 2008, DJ Patil and Jeff Hammerbacher began using the title "data scientist" at LinkedIn and Facebook respectively. In 2012, the Harvard Business Review labeled data scientist "the sexiest job of the 21st century." The data scientist role drew heavily on statistics (regression, hypothesis testing, experimental design for A/B tests) but combined it with large-scale data engineering (SQL, Hadoop, Spark) and predictive machine learning (random forests, gradient boosting, neural networks) in ways that the traditional BLS 15-2041 statistician category did not capture. A separate BLS occupation code (15-2051 for Data Scientists) was created in the 2018 SOC revision and populated in OEWS from 2022. The effect on 15-2041 statisticians was complex: the headline data-science boom drew talent away from the statistician pipeline, but simultaneously generated enormous demand for statisticians to validate AI and ML outputs, design the experiments that tested algorithmic interventions, and serve as the inferential backbone for clinical and government programs that the data-science framing did not cover.

    Effect on the work

    The data-science divergence split the quantitative analyst workforce into two tracks: a rapid-growth ML-and-engineering track (15-2051 Data Scientists) and a steady-growth inferential-and-regulatory track (15-2041 Statisticians). BLS employment for 15-2041 continued growing through the 2010s despite the data-science boom, reaching approximately 28,000-33,000 by the late 2010s.

    Work toolChanging equipment
  • AI-augmented statistics: GitHub Copilot, Jupyter AI, Julius AI, and the LLM-validation demand surge

    GitHub Copilot (launched 2022), Jupyter AI/Jupyternaut (2024), Julius AI (2023), and SAS Viya's embedded AI assistants transformed routine statistical coding from a manual task into an AI-accelerated workflow. Statisticians can now generate boilerplate R, Python, and Stata analysis code from natural-language prompts, accelerating the exploration and table-generation phases of analysis substantially. The ASA's 2025 white paper "AI and the Future of Statistics" explicitly named the human statistician's irreplaceable role as "validating the question before answering it" and positioned the profession as the natural counterparty to AI outputs that require inferential defensibility. The FDA's 2025 guidance on AI in drug development explicitly preserves the statistical reviewer role for analysis-plan authorship and results validation. The net effect: AI handles more of the code-writing and preliminary exploration work, while the study design, regulatory documentation, and inferential-validity functions remain firmly human.

    Effect on the work

    BLS projects +8.5% employment growth for statisticians from 2024 to 2034, reaching 34,900 positions, driven by AI expanding demand for defensible statistical validation rather than reducing it. The LinkedIn 2026 Skills on the Rise report identified AI-augmented statistical work as a category with a 56% wage premium.

    AI audit toolsPattern detection
Projection cone · present → 2034

What credible sources project

Scrub the slider past now to anchor each scenario on the scrubber. The spread is the range of futures credible sources project for this role.

Employment outlook
Projected change in the number of people doing this work.
BLS National Employment Matrix 2024-34
2034
+8.5%
BLS Employment Projections program industry-occupation matrix. For 2024-34, BLS projects 15-2041 Statisticians employment to grow from 32,200 (2024) to 34,900 (2034), a gain of 2,700 positions. This is classified as "much faster than average" growth against the all-occupations average of approximately 3-4%. The BLS methodology models growing demand for statistical expertise across three primary channels: (1) pharmaceutical and biomedical research expanding clinical-trial capacity; (2) AI adoption creating institutional demand for statistical validation of model outputs; and (3) federal and state government survey programs increasing their analytical workforce. The projection does not model the possibility that AI tools could absorb a portion of entry-level statistical coding work, which could moderate actual growth below the projection.
AI task exposure
Share of the role’s tasks that researchers estimate AI can do. This is a measure of task exposure, not a forecast of jobs lost.
Eloundou et al. — "GPTs are GPTs" (2023/2024)
2028
35%
of tasks
GPT-4 task-by-task LLM exposure scoring on O*NET tasks for Statisticians (15-2041). Eloundou et al. place statisticians in a moderate-to-high LLM exposure range: tasks involving code writing (R, Python, Stata scripts), data preparation, and output formatting are substantially exposed to LLM augmentation. However, tasks involving study design, regulatory documentation authorship, survey sampling methodology, and cross-disciplinary consulting score low on exposure, because they require irreducible human judgment, legal accountability, and domain translation. The 35% aggregate exposure estimate reflects this split: roughly a third of a statistician's task hours can be materially accelerated by LLMs, while the two-thirds involving inferential judgment and regulatory responsibility cannot. This is a task-exposure estimate, not an employment projection: Eloundou measures how many hours of work are LLM-accessible, not how many jobs will be eliminated.
Goldman Sachs — "The Jobs AI Is Likely to Boost and Those It May Disrupt" (2025)
2030
28%
of tasks
Goldman Sachs task-automation analysis on the professional and technical services sector. Goldman's 2025 update identifies statisticians as a profession where AI is more likely to boost productivity than eliminate positions, because the regulatory and inferential core of the role is not automatable under current LLM architectures. The 28% figure reflects the share of statistician task-hours that Goldman identifies as automatable (primarily code generation, literature search, and descriptive analysis), with the remaining 72% requiring human judgment in study design, regulatory submission, and expert consultation. Goldman explicitly flags statistical methods validation as a growing function as AI outputs require independent quantitative review across finance, pharmaceutical, and government contexts.
Today, in this role

What's shifting in the work right now

The historical view above shows how this role has moved. This is the present-day detail: which AI tools are picking up which tasks, where the edge still is, and the natural directions this work can grow.

What's changing in your day

Three parts of your work where AI is already doing real lifting, and what stays yours.

AI is sitting alongside you hereRun rapid exploratory and descriptive analyses using Julius AI or JMP Pro AI: submit uploaded datasets to natural-language prompts to generate histograms, correlation matrices, summary tables, and initial regression outputs

Run rapid exploratory and descriptive analyses using Julius AI or JMP Pro AI: submit uploaded datasets to natural-language prompts to generate histograms, correlation matrices, summary tables, and initial regression outputs; then direct the follow-up inferential analysis based on patterns identified in the AI-generated outputs.[8],[9]

Where your edge is

Tools like Julius AI and JMP Pro AI dramatically accelerate the exploratory phase but frequently misidentify the appropriate test family when the statistician does not constrain the prompt. Develop disciplined prompting habits: specify the distributional assumptions, the comparison structure, and the inferential goal explicitly so the AI output aligns with the pre-registered analysis plan rather than data-dredging.

AI is sitting alongside you hereWrite R, Python, or Stata analysis scripts using AI-assisted coding tools: use GitHub Copilot or Jupyter AI (Jupyternaut) to accelerate boilerplate code for data cleaning, merging, imputation, and table generation

Write R, Python, or Stata analysis scripts using AI-assisted coding tools: use GitHub Copilot or Jupyter AI (Jupyternaut) to accelerate boilerplate code for data cleaning, merging, imputation, and table generation; then review, validate, and adapt the generated code against the pre-registered SAP before running any inferential analyses.[5],[10]

Where your edge is

AI code generation for statistics is fast but error-prone on domain-specific conventions: SAS-style output formatting, regulatory-compliant CDISC dataset structures, and survey-weighted regression syntax are all areas where copilots hallucinate plausible-but-wrong code. Build a systematic code-review practice: always test generated code on a simulated dataset with known answers before applying it to study data.

AI is sitting alongside you hereBuild automated reporting and dashboard pipelines in SAS Viya or Jupyter notebooks: configure SAS Viya's automated model building and scoring pipelines or author parameterized Jupyter notebooks that regenerate standardized statistical output tables (frequency distributions, regression summaries, survival curves) from updated data without manual re-execution.

Build automated reporting and dashboard pipelines in SAS Viya or Jupyter notebooks: configure SAS Viya's automated model building and scoring pipelines or author parameterized Jupyter notebooks that regenerate standardized statistical output tables (frequency distributions, regression summaries, survival curves) from updated data without manual re-execution.[11],[12]

Where your edge is

Automated pipelines eliminate the manual table-regeneration cycle that can consume 20-40% of a production statistician's time on recurring reports. Invest time in pipeline design and parameterization so the automated output is validation-ready: build in automated assumption checks, flagging logic for data anomalies, and audit-trail logging so the pipeline output can be traced to input data without manual reconstruction.

Where this role is heading

Natural next steps for someone with your foundation: not exits, evolutions.

A direction you could grow

Computer and Information Research Scientists

Statisticians with strong mathematical foundations and a research publication track record can transition into Computer and Information Research Scientist roles — particularly in AI/ML methods research, statistical learning theory, or computational statistics. As AI proliferates across science and industry, statisticians who develop algorithmic depth become valuable contributors to the teams designing and validating next-generation AI systems. The ASA 2025 white paper notes that "statisticians are uniquely positioned to solve the inferential validity problem in AI" — a research agenda that is explicitly addressed in academia and industrial AI labs. The pivot requires deeper programming fluency and typically benefits from a PhD or equivalent research portfolio.

What you'd add
· Probabilistic programming: Stan, PyMC, Pyro for Bayesian methods research
What it takesA real upskill, but a natural one
Share this year
Drops anyone you send it to straight into 2026.
Preview card
Part of Tech & Data · see all 28roles →
Different role?

See the same long-arc view for your own profession.

Browse the directory by industry, or search by title or SOC code. New roles ship every few weeks. Every profile cites every claim.

Browse all roles

The data behind this timeline

On record since1839
Latest tracked employment32,200 (US, 2024)
Latest median pay$103,300 (2024)
Outlook+8.5% by 2034 (BLS National Employment Matrix 2024-34)
View all 29 cited data points
YearUS employmentMedian annual paySource
18701,200n/aESTIMATE
19023,500n/aESTIMATE
194412,000n/aESTIMATE
196022,000$7,500ESTIMATE
1980n/a$24,000ESTIMATE
199035,000n/aESTIMATE
200019,000$51,990BLS-OEWS
200318,370$59,560BLS-OEWS
200417,030$58,620BLS-OEWS
200517,480$62,450BLS-OEWS
200619,660$65,720BLS-OEWS
200720,270$69,900BLS-OEWS
200820,680$72,610BLS-OEWS
200921,370$72,820BLS-OEWS
201022,830$72,830BLS-OEWS
201123,770$73,880BLS-OEWS
201225,570$75,560BLS-OEWS
201324,950$79,290BLS-OEWS
201426,970$79,990BLS-OEWS
201529,870$80,110BLS-OEWS
201633,440$80,500BLS-OEWS
201736,540$84,060BLS-OEWS
201839,920$87,780BLS-OEWS
201939,090$91,160BLS-OEWS
202038,860$92,270BLS-OEWS
202131,370$95,570BLS-OEWS
202230,780$98,920BLS-OEWS
202329,950$104,110BLS-OEWS
202432,200$103,300BLS-OEWS
Embed this timeline on your site

Free for any site. Paste this where the timeline should appear; it stays interactive, every datapoint stays cited, and it sets no cookies on your page. How embedding works

<iframe src="https://futurehistory.earth/embed/15-2041"
  width="100%" height="430" style="border:0"
  title="Statisticians, a Future History timeline"
  loading="lazy"></iframe>

See all roles in Tech & Data