From questions to reproducible insights
Consider a question that comes up regularly in drug development. A trial missed its primary endpoint, but a minority of patients showed a strong response. Can we identify and target the subgroup that benefits? Answering often means bringing internal trial data together with external clinical and genomic cohorts. Here’s how that unfolds in the agentic era.
The researcher creates a “clean room” and requests access to both sources. Each data steward’s policies apply automatically: which methods the agent can use, which individual and aggregate data it can reach, what can leave. The researcher activates their preferred agent and asks the question in plain language; the agent proposes a stratification, writes the code, and runs it.
An association appears. The agent flags that it may not be real — the responders are concentrated in one ancestry group, so the signal could reflect population structure rather than the variant. It proposes an adjustment and documents the rationale. The researcher approves and runs the revised approach. The signal holds in one subgroup and disappears in another.
The researcher asks the agent to prepare a downloadable summary with the supporting evidence. The agent does so, flags that data steward policies require explicit approval because of re-identification risk, and drafts an egress justification for the researcher to review. The researcher reviews, the agent submits, the data steward approves, the results are downloaded.
The final work product contains a reproducible summary of inputs, methods, and results. All activity is audited, blocked actions are flagged, and no protected data leaves the environment.




