Data Scientists
Scrub through 74years of this role's history, from when it first emerged, through every wave of technology that reshaped it, to the cited projections for where it's heading next.
The tools that defined the work
Select an era to see how it reshaped the work.
Mainframe statistical packages (FORTRAN, SPSS 1968, SAS 1976)
The predecessor practitioners: statisticians, operations research analysts, and quantitative economists relied on mainframe batch computing to run regressions, ANOVA, and time-series analyses. Tukey's "Exploratory Data Analysis" was largely pencil-and-paper work, supplemented by mainframe runs that could take hours to return results. SPSS (Statistical Package for the Social Sciences) was released in 1968 by Norman Nie, Dale Bent, and Hadlai Hull at Stanford, making multivariate statistics accessible to researchers without programming skills. SAS followed in 1976 from North Carolina State University. These tools shaped a generation of analysts who treated data analysis as a slow, deliberate, batch-oriented craft.
Effect on the workThe mainframe era created a small, highly credentialed analyst class, primarily in academia, government agencies, and large manufacturers. Estimates of US statisticians and quantitative analysts in the 1970s run to roughly 20,000-30,000 professionals. The computing barrier kept the field elite and the headcount low.
Mainframe processingComputerized records Personal workstation + S/R + early relational databases (SQL, S-Plus, R 1993)
The arrival of affordable Unix workstations and personal computers transformed statistical computing from a shared-mainframe activity into individual, interactive work. S (developed at Bell Labs by John Chambers beginning in 1975, commercialised as S-Plus in 1988) enabled statistical programming at the command line. R, the open-source implementation of S, was first released by Ross Ihaka and Robert Gentleman at the University of Auckland in 1993 and became the dominant academic statistics language by the early 2000s. SQL relational databases gave analysts direct access to business transaction data for the first time. The combination of R/S-Plus for analysis and SQL for data access defined the workflow of what would become the data scientist predecessor role.
Effect on the workR and SQL democratised quantitative analysis in corporate settings. Analyst headcounts in finance (credit risk, quant trading), insurance (actuarial), and market research grew substantially through the 1990s. BLS tracked statisticians and operations research analysts separately; combined they numbered roughly 60,000-80,000 in the late 1990s according to available OEWS estimates.
Work toolChanging equipment Python + MapReduce / Hadoop (big data era: Python 2.4, Hadoop 2006, scikit-learn 2007)
The mid-2000s brought two developments that reshaped the analyst role into what would be called data science. First, Python gained its scientific stack: NumPy (2005), SciPy (2006), and the beginnings of scikit-learn (2007) made Python a credible alternative to R for machine learning. Second, Google's 2004 MapReduce paper and Doug Cutting's Apache Hadoop (2006) gave analysts a framework for processing datasets too large for a single machine. Facebook, Google, LinkedIn, and Amazon were accumulating user-behaviour datasets of a scale no existing statistical toolkit had been designed to handle. The combination of Python fluency, machine learning libraries, and distributed computing was exactly the triangular skillset Patil and Hammerbacher were reaching for when they coined "data scientist" in 2008.
Effect on the workThe Hadoop ecosystem and Python ML stack created a new type of practitioner who was neither a pure statistician nor a pure software engineer. By 2012, "data scientist" job listings on Indeed had grown 15,000 percent from their 2008 baseline. LinkedIn data showed that the number of professionals adding "Data Scientist" to their profiles between 2012 and 2015 grew at a rate consistently 50 percent higher than for software engineers and data analysts.
Work toolChanging equipment Deep learning + cloud ML platforms (TensorFlow 2015, AWS SageMaker 2017, Jupyter notebooks)
AlexNet's win at the 2012 ImageNet competition announced deep learning's arrival as a practical engineering tool, not just an academic curiosity. Within three years, TensorFlow (Google, 2015) and PyTorch (Facebook, 2016) gave data scientists production-grade deep learning frameworks. AWS SageMaker (2017), Azure ML, and Google AI Platform began abstracting away the infrastructure layer, letting data scientists focus on model design rather than server management. The Jupyter notebook, originally IPython (Fernando Perez, UC Berkeley, 2001), became the universal data science workspace: a live computational document that combined code, output, and narrative in a single shareable file. This era produced the first formalised data-science education infrastructure: fast.ai (2016), Coursera machine learning specialisations, and dozens of university master's programmes.
Effect on the workThe cloud ML era dramatically lowered the barrier to data science practice, expanding the practitioner base far beyond the original Python-Hadoop specialists. Industry surveys (ODSC, KDnuggets, Burtch Works) tracked the median base salary rising from roughly $91,000 at entry level in 2015 toward $100,000+ by 2018, as demand consistently outpaced supply. The 2018 LinkedIn Workforce Report reported a national skills shortage of 151,717 data science positions.
Work toolChanging equipment MLOps + AutoML + feature stores (MLflow 2018, DataBricks Delta Lake 2019, DVC)
The model-deployment gap became the defining problem of this era. Studies found that 87% of machine learning models never made it to production (VentureBeat, 2019). MLflow (Databricks, 2018), DVC, Weights and Biases, and eventually AWS SageMaker Pipelines addressed the lifecycle gap between building a model in a Jupyter notebook and running it reliably in production. AutoML platforms (Google AutoML, H2O.ai, DataRobot) began automating the model-selection and hyperparameter-tuning work that had previously occupied junior data scientists. The data scientist role bifurcated: routine model training was increasingly automated; senior practitioners concentrated on problem framing, feature engineering, and the production-readiness work that AutoML could not yet handle.
Effect on the workThe AutoML era sparked the first serious debate about whether "data scientists will be automated away." In practice, employment continued to grow, because each wave of automation expanded what was computationally possible and thus expanded the scope of problems organisations could tackle, generating more demand for the humans who could direct the automation.
Work toolChanging equipment LLM-powered tools + agentic notebooks (ChatGPT Nov 2022, GitHub Copilot, Hex Magic, Databricks Genie Code)
The release of ChatGPT in November 2022 changed the data scientist's workbench within months. GitHub Copilot (general availability June 2022) began generating Python and SQL cells in Jupyter notebooks. Hex Magic (2023) and Databricks Genie Code (2025) went further, translating natural-language data questions directly into multi-cell analytical workflows. The practical effect was to accelerate the routine EDA and boilerplate-code layer of the job significantly. Senior practitioners reported spending less time writing standard cleaning and transformation code and more time on hypothesis formation, model evaluation, stakeholder communication, and the architectural decisions that AI tools could not yet make. The post-ChatGPT demand surge for people who could build and govern LLM-based systems drove BLS-measured employment from 169,000 (2022) to 246,000 (2024).
Effect on the workRather than contracting the field, the LLM era expanded it. Job postings requiring LLM and GenAI skills in data science grew from near-zero in 2022 to 31% of data science listings by 2026 (Medium/AI Analytics Diaries survey of 500 postings). NLP skills in DS postings rose from 5% to 19% of openings in a single year. GenAI skills commanded a 15-25% wage premium over baseline data science compensation. BLS projects 33.5% employment growth in 15-2051 from 2024 to 2034.
AI audit toolsPattern detection
What credible sources project
Scrub the slider past now to anchor each scenario on the scrubber. The spread is the range of futures credible sources project for this role.
What's shifting in the work right now
The historical view above shows how this role has moved. This is the present-day detail: which AI tools are picking up which tasks, where the edge still is, and the natural directions this work can grow.
What's changing in your day
Three parts of your work where AI is already doing real lifting, and what stays yours.
AI is sitting alongside you hereRun natural-language queries against cloud data warehouses using Cortex AISQL or Snowflake Cortex Analyst, then validate AI-generated SQL for correctness, index efficiency, and schema compliance before sharing results.
Run natural-language queries against cloud data warehouses using Cortex AISQL or Snowflake Cortex Analyst, then validate AI-generated SQL for correctness, index efficiency, and schema compliance before sharing results.[8],[9]
Cortex AISQL (GA November 2025) and similar tools can generate complex SQL from plain English. The human job shifts to specifying the business question precisely and auditing the output — not writing SQL from scratch.
AI is sitting alongside you hereRun exploratory data analysis (EDA) in AI-assisted notebooks: use Hex Magic or Databricks Genie Code to generate multi-cell SQL+Python analysis from natural-language prompts, then interpret the results and direct follow-up hypotheses.
Run exploratory data analysis (EDA) in AI-assisted notebooks: use Hex Magic or Databricks Genie Code to generate multi-cell SQL+Python analysis from natural-language prompts, then interpret the results and direct follow-up hypotheses.[10],[11]
Shift time from writing boilerplate EDA code to evaluating AI-generated outputs, catching schema errors, and translating patterns into business hypotheses. Notebook fluency with AI co-pilots is now table-stakes.
AI is sitting alongside you hereBuild, tune, and compare predictive and generative models: select architectures, design feature-engineering pipelines, run hyperparameter sweeps, and evaluate against holdout sets using loss functions and business-relevant metrics.
Build, tune, and compare predictive and generative models: select architectures, design feature-engineering pipelines, run hyperparameter sweeps, and evaluate against holdout sets using loss functions and business-relevant metrics.[4],[3]
Use AutoML to benchmark baselines quickly; reserve human effort for custom architectures, domain-specific feature engineering, and the judgment calls AutoML cannot make (e.g., bias audits, business-constraint integration).
Where this role is heading
Natural next steps for someone with your foundation: not exits, evolutions.
Computer and Information Research Scientists
The AI Research Scientist path deepens theoretical foundations (deep learning, reinforcement learning, mathematical optimization) and moves into published-research or R&D environments at labs, academia, or corporate AI divisions. This pivot requires stronger academic credentials (typically a PhD) but offers the highest long-term resilience as AI reshapes applied roles.
See the same long-arc view for your own profession.
Browse the directory by industry, or search by title or SOC code. New roles ship every few weeks. Every profile cites every claim.
Browse all roles