Benchmarks / LAKEQA

LAKEQA

Released
Exploratory question answering over data lakes

Evaluating whether deep-research agents can discover, integrate, and reason over relevant datasets at million-table scale.

Updated June 2026
1M+ Data lake scale
ICML 2026 Publication

Background

Deep-research agents typically search the web while ignoring the structured data stored in public and enterprise data lakes. LAKEQA targets analytic questions that require discovering relevant datasets, combining heterogeneous evidence, and showing where an answer came from.

What it evaluates

  • Dataset discovery at million-table scale
  • Cross-source integration and analytic reasoning
  • Enumeration, aggregation, and evidence synthesis
  • Answer provenance

Dataset and tasks

Tasks are exploratory rather than simple fact lookup: the relevant sources are not handed to the agent in advance, and a successful answer may require several structured and unstructured datasets.

Evaluation methodology

Submissions are assessed for answer quality and provenance. The benchmark is designed to expose failures at each stage—from finding the right datasets to integrating them and producing a supported final answer.

Representative task

Answer an analytic question whose evidence is distributed across a large data lake, identify the relevant sources, and return a verifiable answer with provenance.

Results

The paper and public leaderboard provide current benchmark results and evaluation details.