Trustworthy, Safe, Efficient and Reliable Agents
The Data, Agents, and Processes Lab (DAPLab) at Columbia University develops systems, infrastructure, and interaction principles for AI agents to safely and reliably automate real work.
Agents turn every layer of the stack into a semantic quality problem. As a consequence, trustworthy agent automation cannot be solved at a single layer of the stack. DAPLab vertically integrates expertise across operating systems, data systems, AI, HCI, security, and enterprise workflows to rethink what agent development, evaluation, and interaction, along with the computing infrastructure should look like. We work closely with industry partners to ground this research in real organizational needs and ensure it delivers practical impact.
For more information about the lab, contact ewu@cs.columbia.edu.
Why vertical integration matters
AI agents fail across the entire stack: models hallucinate, retrieval misses critical context, execution environments lack isolation, workflows leak data, and human oversight breaks under scale. Fixing only one layer is not enough.
DAPLab brings together researchers across systems, databases, AI, HCI, security, and organizational workflows because trustworthy automation requires coordinated advances across the full agent stack — from infrastructure and state management to evaluation, safety, and human interaction.
- Human Interaction
- Agents + Models
- Processes + Assurance
- Data + Execution Systems
- OS + Infrastructure
News
All news →Asaf Cidon and coauthors win SOSP 2026 Best Paper and Distinguished Artifact
The paper “It’s the Kernel’s Fault! Custom Page Fault Handling With bpf_fault” by Tal Zussman, Riju Dey, Hasan Zengin, Yiming Fang, David Hildenbrand, and Asaf Cidon received the SOSP 2026 Best Paper and Distinguished Artifact awards.
SceniX acquired by World Labs; AMD to acquire World Labs
Huge congrats to DAPLab faculty Yunzhu Li and Changxi Zheng, whose robotics startup, SceniX, was acquired by World Labs in July 2026. In September a few months later, AMD announced an agreement to acquire World Labs, bringing the SceniX team’s work on robotics, simulation, and spatial intelligence into AMD.
Fall 2026 DAPLab Research Seminar
The DAPLab’s Tuesday 12PM research seminar continues in CSB 453 (CS Conference Room). We invite internal and external speakers that can share cutting-edge agent-systems research or can talk about processes in their organizations and how they are trying to automate them.
MortarBench featured in Realtor.com
Realtor.com covered our MortarBench research on AI mortgage origination agents. Top models got nearly 1 in 4 answers wrong under realistic conditions — and showed systematic bias, flagging non-English names as “foreign origin” at 5× the rate of English names. Matthew Toles and Zhou Yu are quoted in the piece.
Haonan Wang receives inaugural Workday AI PhD Fellowship
Haonan Wang has been awarded the inaugural Workday AI PhD Fellowship to pursue his work on Data Lake Agents. The core question: how do you teach an enterprise agent to gather sufficient, traceable evidence from an organization’s heterogeneous data sources and correctly execute the resulting workflow — while satisfying regulatory requirements, organizational policies, role-based permissions, approval hierarchies, fairness constraints, and operational budgets?
Events
All events →Upcoming
-
From Multi-Agent Collaboration to Emergent Communication in LLM Agents
-
Towards Recursive Self-Improvement: Lessons from Self-Evolving Reinforcement Learning
-
What would it cost to end extreme poverty?
-
Agent Evaluation Science Fall 2026
Recent
-
OpenAgents: The Next-Gen Collaboration OS for Teams of Humans and Agents
-
Selective interventions for information acquisition: Incentives and audits for biodiversity
-
Columbia AI Night
-
Fall 2026 Personal Health Assistant Course
Publications
All →-
StateFork: Branchable Infrastructure for Agent Exploration
-
SANA: What Matters for QA Agents over Massive Data Lakes?
-
Prove2Me: An Open Collaborative Platform for Scaling Math Formalization
-
The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark
-
OpenForgeRL: Train Harness-native Agents in Any Environment
-
TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization
-
fog: Expressing Motion and Emotion through Function Composition of AI-Generated Code
-
PersonaJudge: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data
Thanks to our industry partners