Research

OpenAI Releases GABRIEL, an Open-Source Toolkit for Measuring Qualitative Data with GPT

OpenAI's Economic Research Team released GABRIEL, an open-source GPT-powered Python toolkit that converts unstructured text and images into quantitative measurements for researchers.

Scaling social science research
Scaling social science researchAI-generated
By Sophie Lindqvist5 min read

Updated

Why it matters

  • OpenAI's Economic Research Team released GABRIEL, an open-source Python toolkit that uses GPT to turn unstructured text and images into quantitative measurements.
  • Researchers describe what to measure in plain language (e.g., "how family-friendly is this job listing?") and GABRIEL applies the question consistently across thousands or millions of documents, returning a score for each.
  • The toolkit includes a paper, "GPT as a measurement tool," benchmarking GPT on labeling qualitative data across many use cases; OpenAI reports high accuracy.
  • GABRIEL bundles dataset merging with mismatched columns, smart deduplication, passage coding, theory ideation, and deidentification of personal information from text.
  • GABRIEL is available now as an open-source library on GitHub with a Colab tutorial notebook, and OpenAI says it requires minimal technical background.

OpenAI's Economic Research Team released GABRIEL today, an open-source Python toolkit that uses GPT to convert unstructured text and images into quantitative measurements. The library targets economists, social scientists, and data scientists who need to study qualitative data at scale, and it is available now on GitHub with a tutorial notebook for getting started.

The release addresses a specific bottleneck in social science research. Qualitative data — what OpenAI describes as "the richest stories about the world" — spans syllabi, interviews, social media posts, and photographs, and it exists in enormous volumes. But converting that raw material into rigorous evidence is, in the company's words, "incredibly time-consuming." Often it isn't feasible at all. The result, according to OpenAI, is that "social scientists are forced to forego important avenues of research, not because the data doesn't exist, but because it's impossible to analyze."

That constraint shapes what research questions the field can pursue. Text measurement has traditionally required armies of human coders, months of labeling time, or narrow keyword approaches that miss meaning. GABRIEL positions GPT as a general-purpose substitute: a model that can read a document the way a trained research assistant would and return a consistent numeric score.

How the toolkit works

GABRIEL's core mechanism is simple to describe. A researcher states what they want to measure in plain language — OpenAI's example is "how family-friendly is this job listing?" — and the toolkit applies that same question consistently across thousands or millions of documents, returning a score for each one. The measurement question, once written, becomes reusable across an entire corpus.

The design shifts labor rather than eliminating it. In OpenAI's framing, the tool lets researchers "spend less time on repetitive data labeling and more time on the work that actually requires expertise: choosing what to measure, validating results, and drawing careful conclusions." The human contribution moves upstream to question design and downstream to validation; the model handles the middle.

The use cases OpenAI lists sketch the breadth of the approach. GABRIEL can analyze a large collection of scientific papers to see which specific methods appear and how they evolve over time. It can examine course curricula to measure how much attention goes to different subjects or skills. It can extract structured historical details for every small town across Europe, or process a trove of customer reviews and surface patterns in what people value most.

Accuracy claims and supporting paper

Alongside the toolkit, OpenAI has published a paper, "GPT as a measurement tool," in which the team benchmarks GPT at labeling qualitative data across many use cases. The company reports that the model is "highly accurate" on these tasks. The paper provides the methodological backing for researchers who need to evaluate whether GPT-based measurement is defensible for their own work — the validation step the team itself flags as a core part of the researcher's job.

The stakes for measurement quality are real. If GPT labels qualitative data as reliably as OpenAI claims, it changes the economics of research designs that were previously impractical: panel studies over decades of documents, cross-country comparisons of curricula, large-scale analysis of historical archives. If the accuracy degrades on harder or more ambiguous constructs, researchers will need to catch that in validation — which makes the benchmark paper, rather than the toolkit alone, the document to scrutinize.

Utilities beyond measurement

GABRIEL bundles practical tooling that measurement projects typically require. According to OpenAI, the toolkit supports merging datasets even when the columns don't match, smart deduplication, passage coding, ideating new scientific theories, and deidentifying personal information from text to preserve privacy.

The deidentification feature matters for compliance. Qualitative social science data frequently contains personal information, and processing it through an external model raises both ethical review and data protection concerns. OpenAI does not detail in the announcement how the deidentification works or whether processing can be done locally, but the feature's inclusion signals that the team considered research-ethics workflows in the design.

The dataset-merging capability points at a second, less glamorous bottleneck: much research time goes into wrangling inconsistent data files rather than analyzing them. Automated column matching and deduplication attack that problem directly.

Availability and accessibility

GABRIEL ships as an open-source Python library, hosted on OpenAI's GitHub, with a tutorial notebook designed to run in Google Colab. OpenAI says the toolkit "is designed to require minimal technical background," widening the potential user base beyond data scientists to economists and social scientists without programming-heavy training.

The choice to open-source the code is notable. Researchers can inspect, audit, and extend the measurement pipeline rather than treating it as a black box — a meaningful property for a field where methodological transparency is a publishing requirement.

OpenAI frames the release as part of a broader mission: "A core part of our work at OpenAI is enabling scientists to move faster and solve harder problems." The company says it will "keep improving GABRIEL over time based on feedback from the academic community."

That feedback loop will determine the tool's real impact. GABRIEL is the latest entry in a fast-moving competition to make LLMs standard instruments for social science measurement, and OpenAI is betting that the pairing of an open-source library, a tutorial notebook, and a public benchmark paper will make GABRIEL the default starting point for researchers who want to turn human stories into numbers.

Original: cdn.openai.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

115 articles

Related articles

  1. OpenAI Puts ChatGPT Inside Excel With GPT-5.4 Built for Finance
  2. OpenAI Launches ChatGPT for Financial Services With Built-In Data
  3. OpenAI Says a Quarter of U.S. Workers Now Use ChatGPT on the Job
  4. OpenAI Says Over a Quarter of U.S. Workers Now Use ChatGPT on the Job
  5. OpenAI Launches ChatGPT for Financial Services With Built-In Market Data

« Previous articleNext article »