LAKEQA
ReleasedEvaluating whether deep-research agents can discover, integrate, and reason over relevant datasets at million-table scale.
Background
Deep-research agents typically search the web while ignoring the structured data stored in public and enterprise data lakes. LAKEQA targets analytic questions that require discovering relevant datasets, combining heterogeneous evidence, and showing where an answer came from.
What it evaluates
- Dataset discovery at million-table scale
- Cross-source integration and analytic reasoning
- Enumeration, aggregation, and evidence synthesis
- Answer provenance
Dataset and tasks
Tasks are exploratory rather than simple fact lookup: the relevant sources are not handed to the agent in advance, and a successful answer may require several structured and unstructured datasets.
Evaluation methodology
Submissions are assessed for answer quality and provenance. The benchmark is designed to expose failures at each stage—from finding the right datasets to integrating them and producing a supported final answer.
Representative task
Answer an analytic question whose evidence is distributed across a large data lake, identify the relevant sources, and return a verifiable answer with provenance.
Results
The paper and public leaderboard provide current benchmark results and evaluation details.