Key takeaways
- A large proportion of clinical diagnoses can be derived from language-based interviews alone.
- Current language models (LMs) have demonstrated strong differential diagnosis assessment capabilities when evaluated on
- These evaluations do not capture how everyday patients report their health symptoms, for example with varying levels of
What happened
A large proportion of clinical diagnoses can be derived from language-based interviews alone. These diagnostic interviews are typically conducted by clinicians through doctor-patient interactions during in-person or remote visits. While these interactions are the gold standard for symptom assessment, they can often suffer from financial, geographic, and systemic barriers that limit their accessibility.
After assessing the accuracy of SymptomAI’s differential diagnoses (DDx), we further compare SymptomAI’s diagnoses against biosignals from participants’ Fitbit wearable devices in the time leading up to their conversation with SymptomAI. We show that SymptomAI conversations that led to diagnosis with an infectious disease etiology coincide with physiological trends that may indicate an immune response, suggesting further evidence of SymptomAI’s performance.
We found that the clinicians preferred the DDx generated by SymptomAI over those provided by the other clinicians in over 50% of the cases. This indicates that SymptomAI DDx aligned with our clinicians’ medical assessments just as often or more often than that of other clinicians. We found that SymptomAI’s performance above clinical baselines was greatest for cases where the clinician’s felt least confident in their own DDx.
Why it matters
Current language models (LMs) have demonstrated strong differential diagnosis assessment capabilities when evaluated on curated medical case-studies, highlighting their potential to support the diagnostic process. However, existing evaluations have largely relied on curated, highly detailed and sometimes synthetic patient vignettes, which may not reflect real world experience and clinical presentation variability.
These evaluations do not capture how everyday patients report their health symptoms, for example with varying levels of medical literacy, incomplete information, and other complexities that arise through natural conversation. This represents a key gap, leading to uncertainty of how LMs might perform in real-world contexts.
To address this gap, we conduct an in-situ comparative research study of a set of experimental conversational prototype AI agents designed to explore how conversational AI might conduct end-to-end symptom interviews and differential diagnostic assessment for research benchmarking purposes. 0 SymptomAI agents. All diagnoses, labels, and disease associations generated during the study were for research analysis only and did not constitute confirmed clinical diagnoses or official medical assessments.
Two weeks after their interaction with the AI agents, we asked research participants to report any diagnoses they received from a visit with a healthcare provider. Using this data, we conducted a clinical expert annotation study comparing SymptomAI’s diagnostic performance relative to real clinicians' medical assessments.
What to watch
Given SymptomAI's accuracy against clinical baselines, we can also explore its potential at scale. Currently, the cost of clinical labels prohibits real-world analyses of population-scale datasets. Accurate symptom checking systems like SymptomAI have the potential to enable automated reference labeling of clinical quality diagnosis, which can open up large-scale analyses of physiological data — a task that is currently impossible at scale.
One such example is correlating wearable biosignals with different categories of illness. The most notable changes in wearable biosignals are observed for acute respiratory infections. To study this at population scale, we collected daily biometric data from our consenting participants for up to 30 days prior to their interaction with SymptomAI. We find clear biosignal shifts indicating symptom onset in the days appr



