GPQA
arXiv 2023 / NYU, Cohere, Anthropic

GPQA: A Graduate-Level Google-Proof Q&A Benchmark

David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, Samuel R. Bowman

TL;DR

"GPQA is a dataset of 448 ultra-difficult multiple-choice questions in biology, physics, and chemistry written by PhDs. While domain experts achieve 65% accuracy, skilled non-experts with full web access score just 34%, and state-of-the-art LLMs like GPT-4 top out at 39%—making it a premier testbed for scalable oversight."

🔬 Biology, Physics, Chemistry 🧠 Scalable Oversight 🛡️ Google-Proof Design 🤖 Frontier LLM Evaluation

Why This Matters

How do we supervise an AI that is smarter than us?

As Large Language Models (LLMs) advance, they are beginning to tackle complex scientific problems. Soon, they may generate novel theories, design experiments, or write code at a level that surpasses human capabilities. This presents a critical challenge known as scalable oversight: how can human supervisors reliably evaluate and guide AI systems when the humans themselves do not know the correct answers?

To research solutions to this problem, scientists use "sandwiching" experiments. In these setups, a non-expert human is tasked with supervising a highly capable AI model, while experts provide the ground-truth answers to verify if the oversight was successful. For this research to be valid, we need a benchmark that is:

  • Highly objective: Experts must agree on a single, clear, correct answer.
  • Extremely difficult: Even smart, highly motivated non-experts with full access to Google should not be able to easily find or guess the answer.
  • Beyond current models: Frontier LLMs should not be able to solve them easily, leaving a clear "sandwich" zone for evaluation.

The Big Idea

GPQA (Graduate-Level Google-Proof Q&A) is designed specifically to meet these criteria. Sourced from PhD-level experts in physics, chemistry, and biology, the questions are crafted to resist search engines and require deep multi-step reasoning.

The dataset consists of 448 multiple-choice questions in its main set (and 198 in the ultra-curated "Diamond" set). To ensure the questions were truly "Google-proof," the researchers hired high-performing PhD students from *different* scientific fields to validate them. Armed with search engines and 30+ minutes per question, these highly skilled non-experts could only achieve 34% accuracy—barely above the 25% random guessing threshold.

How It Works

1. The Multi-Stage Verification Pipeline

Creating a high-quality, objective, yet incredibly difficult dataset is exceptionally hard. The authors designed a rigorous 4-stage pipeline with strong financial incentives to ensure both clarity and extreme difficulty.

Interactive Pipeline Explorer

Click on each stage to see the requirements, incentives, and quality gates.

2. Quantifying the "Expertise Gap"

The core strength of GPQA is the massive performance gap between domain experts (PhDs in the question's specific field) and highly skilled non-experts (PhDs in different fields). This gap proves that the questions require deep, specialized knowledge and cannot be solved merely by clever search queries.

Dataset Split & Accuracy Comparison

Select a dataset split to visualize the performance difference between experts, non-experts, and GPT-4.

Domain Experts (PhDs) 64.8%
Highly Skilled Non-Experts (with Google) 34.1%
GPT-4 (Few-Shot CoT) 38.7%
Random Guessing 25.0%

GPQA Extended: The full dataset of validated questions. Even here, the expertise gap is a massive 30.7 percentage points, proving that non-experts cannot easily leverage search engines to find the answers.

3. Try the "Google-Proof" Challenge

Can you solve a GPQA question? Below is a real chemistry question from the dataset. Try to answer it, or simulate how a non-expert would try to "Google" the solution.

Organic Chemistry Avg. validator time: 37 mins

"Your overhear two chemists talking to each other as they leave a synthetic organic chemistry lab. One asks the other 'So, how did it go?' The second chemist replies, 'Not well - my compounds are on top of each other.' What is the second chemist most likely referring to?"

Results & Baselines

The authors benchmarked several state-of-the-art models in both closed-book and open-book (with web search) settings. The results are striking: even with chain-of-thought prompting and search tools, AI systems lag far behind human domain experts.

Evaluated Entity Extended Set (%) Main Set (%) Diamond Set (%)
Expert Human Validators 65.4% 72.5%* 81.2%*
Non-Expert Human Validators 33.9% 30.5%* 21.9%*
GPT-4 (Few-Shot CoT) 38.7% 39.7% 38.8%
GPT-4 + Search (Self-Ask) 39.4% 41.0% 38.8%
Llama-2-70B-chat (Few-Shot CoT) 30.4% 29.1% 28.1%

*Note: Expert and non-expert human accuracies on the Main and Diamond sets are skewed by selection effects, as those subsets were explicitly filtered based on validator performance.

Limitations & Open Questions

01

Small Dataset Size

Due to the high qualification requirements and expensive validation pipeline, GPQA is relatively small (448 questions in the main set). This limits its statistical power for detecting small performance differences between highly capable models.

02

Specialized Non-Experts

The "non-experts" used in this study were themselves PhD students in other scientific fields. This sets an upper bound on non-expert accuracy; general-population non-experts would likely score even lower.

03

Demographic Bias

Contractors were sourced primarily through Upwork without explicit regional or demographic balancing, potentially introducing linguistic or domain selection biases.

Glossary

Scalable Oversight

The problem of supervising and evaluating AI systems that possess capabilities exceeding those of any human supervisor. It focuses on developing protocols that allow humans to accurately judge the truthfulness of outputs they cannot verify on their own.

Google-Proof

Questions designed such that the exact answer or a direct derivation cannot be easily found using standard search engine queries. They require synthesizing multiple complex concepts from scratch.

Sandwiching Experiments

An empirical paradigm where a human supervisor with intermediate capability (the "sandwich filling") attempts to oversee a model with superior capability, while researchers use ground-truth expert answers to check if the supervision was successful.

Citation

@article{rein2023gpqa,
  title={GPQA: A Graduate-Level Google-Proof Q\&A Benchmark},
  author={Rein, David and Hou, Betty Li and Stickland, Asa Cooper and Petty, Jackson and Pang, Richard Yuanzhe and Dirani, Julien and Michael, Julian and Bowman, Samuel R},
  journal={arXiv preprint arXiv:2311.12022},
  year={2023}
}
Made with Flash Papers — make your own