Key takeaways
- Wearable devices capture continuous physiological signals at population scale.
- These systems optimize for predictive performance while overlooking statistical validity, leading to spurious correlations, leakage, and…
- By combining hypothesis generation, parallel statistical analysis, model training, adversarial validation, and literature-grounded…
What happened
Wearable devices capture continuous physiological signals at population scale. These streams, ranging from heart rate dynamics to sleep patterns, can reveal early physiological changes before symptoms appear. The bottleneck is no longer data collection, but turning these signals into reliable, clinically meaningful biomarkers. Existing language model-based agent systems automate parts of the scientific workflow, but can often break down on physiological time-series data.
The Biomarker Discovery Framework did not simply select existing variables; it constructed novel composite features. For instance, in the mental health domain, it identified sleep duration variability and sleep onset variability as top correlates of depression severity.
In the metabolic domain, it derived a cardiovascular fitness index (steps divided by resting heart rate) as a non-invasive correlate of insulin resistance, linking it to prior work on glucose regulation and cardiometabolic fitness. We deployed the Biomarker Discovery Framework across three distinct large-scale cohorts totaling 9,279 participant-observations, spanning both mental health and metabolic disease domains.
Across the two depression domains, Biomarker Discovery Framework prioritized different operationalizations of a related circadian-instability construct. 252, p To assess manuscript quality, 15 experts in medicine, biomedical data science, machine learning, bioinformatics, and digital health reviewed blinded reports from the Biomarker Discovery Framework and three contemporary AI research systems (Google DeepMind’s AI co-scientist, Biomni, and Google ADK’s Data Science Agent).
Why it matters
These systems optimize for predictive performance while overlooking statistical validity, leading to spurious correlations, leakage, and brittle features. To this end, we introduce the Biomarker Discovery Framework, a multi-agent system that structures candidate biomarker prioritization as an iterative research loop under human supervision.
By combining hypothesis generation, parallel statistical analysis, model training, adversarial validation, and literature-grounded reasoning, Biomarker Discovery Framework accelerates the discovery process while maintaining strict statistical rigor and preserving human oversight. Across three cohorts (N = 9,279 participant-observations), Biomarker Discovery Framework recovered known clinical signals, identified convergent biomarkers across independent datasets, and improved downstream prediction when combined with demographic features.
Biomarker Discovery Framework combines deterministic computation for numerical analysis with generative reasoning for hypothesis formation and interpretation. An Orchestrator agent decomposes natural-language research directives into execution plans and guides specialized agents through a six-phase process. 252). The workflow then checked stability, leakage, subgroup consistency, and alternative explanations before framing the result as a literature-grounded circadian-instability hypothesis for human review.
To assess the Biomarker Discovery Framework's capability to extract plausible physiological insights from noisy data, we applied it independently across three large-scale cohorts totaling 9,279 participant-observations, spanning both mental health (DWB and GLOBEM) and metabolic disease (WEAR-ME) domains. The pipeline autonomously identified 41 candidate digital biomarkers for mental health and 25 for metabolic outcomes. The table below shows a curated sample of candidate associations.
Spearman’s ρ summarizes the direction and strength of an association. The 95% confidence interval quantifies uncertainty, and the adjusted p-value accounts for multiple comparisons. The mechanism presents a literature-grounded hypothesis rather than a causal conclusion. Importantly, the final column describes the strength of prior evidence — not clinical validation in this study — and the stars denote evidence-tier markers rather than statistical-significance codes.
What to watch
Biomarker Discovery Framework, Biomni, and the Data Science Agent were scored together in 21 sessions, and Biomarker Discovery Framework was scored in a separate 13-session set using the same evaluation instrument. In the blinded evaluation, the Biomarker Discovery Framework received the highest mean scores across all seven quality dimensions.
Under the study’s simulated editorial rubric, it was the only system to receive any “Accept” or “Minor Revision” recommendations: 2 Accept, 8 Minor Revision, 8 Major Revision, and 3 Reject. 4% for the baselines, and ranked the Biomarker Discovery Framework first in 9 of 13 four-system ranking sessions.


