AI Can Spot ADHD Years Before Diagnosis, Duke Study Finds | Executive Function News
Independent Journalism on Executive Function and Neuroscience

Executive Function News

Reporting on the Science and Practice of Executive Function
Neuroscience

AI Can Spot ADHD Years Before Diagnosis, Duke Study Finds

An artificial intelligence model trained on the electronic health records of more than 140,000 children can identify which five-year-olds are likely to be diagnosed with ADHD years later. The Duke Health team that built it sees a path to earlier support. The deployment questions, given a long history of algorithmic medical bias, are larger.

A pediatric examination room with examination table and medical equipment, illustrating the routine healthcare settings where electronic health records used by the Duke AI model are generated.
An April 2026 Duke Health study published in Nature Mental Health demonstrated that AI models analyzing routine pediatric electronic health records can identify children likely to develop ADHD years before typical diagnosis. Photo: Executive Function News.

A team at Duke Health has trained an artificial intelligence model to read the kind of medical records a pediatrician produces during a routine well-child visit, the developmental notes, the growth charts, the brief observations entered into the electronic record after a parent describes their toddler’s tantrums or sleep problems or motor coordination, and predict, with measurable accuracy, which children in that ordinary clinical stream will eventually be diagnosed with ADHD. The prediction can be made by age five. ADHD is typically diagnosed around age seven, and often considerably later for girls and for children whose presentations do not match the textbook hyperactive profile. The gap is what makes the technology consequential. The gap is also what makes it complicated.

The research was published in Nature Mental Health on April 27, 2026, by a group at Duke University School of Medicine led by Elliot Hill, a data scientist in the Department of Biostatistics and Bioinformatics, with senior author Matthew Engelhard, a physician-scientist in the same department. The team included Naomi Davis of Duke’s Department of Psychiatry and Behavioral Sciences, autism researcher Geraldine Dawson, and biostatisticians Benjamin Goldstein and De Rong Loh. The work was supported by the National Institute of Mental Health.

What the Duke team built is, in plain terms, a pattern-matching system. The model was first pretrained on the electronic health records of more than 720,000 patients, learning the general structure of how medical events appear together over time. It was then fine-tuned on records from more than 140,000 children, some of whom went on to be diagnosed with ADHD and some of whom did not. The model learned to recognize the early combinations of developmental notes, behavioral observations, and clinical findings that distinguished one trajectory from the other.

By age five, the model could classify a child’s ADHD risk with accuracy the authors describe as clinically meaningful. The same accuracy held, the paper reports, across demographic groups defined by sex, race, and insurance status, an unusual claim for an AI medical screening tool and one that future independent replications will be watched closely on.

“This is not an AI doctor,” Engelhard said in a Duke Health statement accompanying the paper. “It’s a tool to help clinicians focus their time and resources, so kids who need help don’t fall through the cracks or wait years for answers.”

Key Findings
  • An AI model trained on electronic health records from over 140,000 children, with pretraining on records from 720,000 patients, can predict future ADHD diagnosis by age five.
  • ADHD is typically diagnosed around age seven and often years later for girls, who are systematically underdiagnosed under current screening practices.
  • The Duke team reports the model’s accuracy holds across demographic groups defined by sex, race, and insurance status within the Duke Health system.
  • The technology surfaces predictive patterns hidden in routine medical data that human clinicians do not consistently see, raising both the promise of earlier intervention and questions about deployment, consent, and downstream effects.
  • The history of algorithmic medicine, including a widely-deployed 2019 healthcare algorithm later shown to systematically deprioritize Black patients, is a sobering backdrop to claims of demographic fairness in any new clinical AI tool.
  • ADHD evaluation systems in many regions are already overwhelmed by demand. A screening tool that flags more children for evaluation will arrive into that bottleneck rather than around it.

The Recent History of Medical AI

The Duke study does not arrive on a blank page. It arrives on a page that has been actively written, over the past decade, by a series of medical AI deployments whose track record ranges from genuinely useful to catastrophically harmful. Understanding the Duke work requires understanding what came before it.

On the useful end of the ledger sit applications like AI-assisted detection of diabetic retinopathy, which has been deployed in primary care settings to flag patients whose eye scans suggest emerging vision-threatening disease. Sepsis prediction algorithms running in hospital electronic health records have, in some implementations, demonstrably reduced mortality by alerting clinicians to deteriorating patients earlier than human pattern recognition would have. Cancer image classification tools, particularly in dermatology and radiology, have matched or in some cases exceeded specialist accuracy on specific tasks. These successes share several features: a narrow, well-specified diagnostic question, a high-quality ground truth, and deployment as a flag for clinician attention rather than as an autonomous decision-maker.

On the other end of the ledger sits the case study that every researcher in clinical AI has now read: a 2019 paper in Science by Ziad Obermeyer and colleagues at the University of California, Berkeley, which dissected the algorithm used by Optum, a subsidiary of UnitedHealth Group, to identify high-need patients for enrollment in care management programs. The algorithm was used at scale, affecting decisions about extra care for an estimated 200 million Americans each year. It was, the researchers showed, systematically deprioritizing Black patients.

The mechanism was instructive. The Optum algorithm was not asked to predict illness directly. It was asked to predict future healthcare spending, on the assumption that patients who would spend more in the coming year were patients who needed more care now. The proxy seemed reasonable. The problem was that Black patients, due to structural barriers, unequal access to healthcare, and a long history of discriminatory care patterns, systematically spent less on healthcare than equally sick white patients. The algorithm, taught to predict spending, learned to predict access. Black patients with the same illness severity as white patients were given lower risk scores. Fewer Black patients were enrolled in the high-touch care programs the algorithm controlled access to. The bias in the underlying data became bias in the algorithm’s output, applied at scale to hundreds of millions of people.

Obermeyer’s paper is now the canonical reference for what can go wrong when AI is deployed in clinical settings. It is not the only example. IBM Watson Health’s oncology product, marketed for years as a tool to help cancer specialists choose treatments, was quietly shut down after evidence emerged that it had been giving unsafe and sometimes outright wrong recommendations. The technology was sold before it was reliable. The careers and reputations damaged in its failure are still being repaired.

The Duke researchers know this history. They are working in a field that has learned, painfully, what proxy variables can do, what unrepresentative training data can do, what deployment without external validation can do. The careful framing of the Duke team’s claims, this is not an AI doctor, the model flags children for clinician attention rather than diagnosing them, the equity claim is preliminary and will need external validation, reflects a field that has internalized the lessons of the previous decade. Whether that internalization is sufficient to prevent the next round of failures is the question deployment will answer.

What the Study Did

The Duke researchers worked with a body of clinical data that has, until recently, been difficult to study at scale: the longitudinal electronic health record. Every time a child sees a pediatrician, the visit produces a record. Some of what is recorded is structured data, height, weight, blood pressure, immunizations, billing codes. Some is unstructured text, the pediatrician’s brief notes about the visit, parental concerns reported in the room, observations of the child’s behavior. Over the first years of a child’s life, these records accumulate into a dense, time-stamped picture of how the child has developed, what concerns have come up, what referrals have been made, what diagnoses have been recorded.

For the Duke study, the team assembled records from more than 140,000 children seen in the Duke Health system, some of whom received an eventual ADHD diagnosis and some of whom did not. Before fine-tuning on this dataset, the model was pretrained on a much larger set of more than 720,000 patient records, a step that helped the system learn the underlying grammar of how clinical events relate to one another before being specialized for the ADHD prediction task.

The model’s task was straightforward to state and difficult to do: given a child’s records up to a particular age, predict whether that child will eventually be diagnosed with ADHD. The researchers tested its accuracy at multiple ages, with the headline result being that meaningful prediction was possible by age five.

The patterns the model identified were not single events. They were combinations: of developmental milestones reached or delayed, of behavioral observations, of clinical findings, of patterns in how often a child was brought in for unscheduled visits. The patterns are difficult to articulate in human terms because that is precisely what the model is doing that a human clinician cannot easily do, integrating many small signals across years of data into a single risk estimate.

“We have this incredibly rich source of information sitting in electronic health records,” Hill said. “The idea was to see whether patterns hidden in that data could help us predict which children might later be diagnosed with ADHD, well before that diagnosis usually happens.”

What the Model Was Actually Looking At

The model’s predictions were not opaque. When the Duke team analyzed which record events most influenced its predictions, certain categories of information consistently rose to the top: developmental concerns, behavioral observations, language or speech delays, learning concerns, emotional symptoms, and patterns of repeated visits about attention or behavior. The model also identified co-occurring psychiatric conditions as predictive markers, meaning that other mental health patterns frequently traveled alongside later ADHD diagnoses.

This list, on its own, would not surprise an experienced pediatrician. The same kinds of signs are what a thoughtful clinician would already be tracking in a child whose development they were concerned about. The model’s contribution is not the recognition of any one pattern but the integration of many patterns across years of data into a single risk estimate. A human pediatrician sees a child for fifteen minutes at a time, weeks or months apart, and is asked to remember subtle threads of concern across visits. The model has access to every entry, every note, every code, and can recognize combinations that no single clinician would carry in working memory.

This is also where the model’s most interesting limitation lives. What it learned to predict is not “true ADHD” in some objective sense. It is “ADHD as it is diagnosed within the Duke Health system.” The signals it uses to predict are the signals that, in this particular data environment, preceded the diagnostic label being applied. If diagnostic practice in another system is different, by being more aggressive or more conservative, by drawing the boundary in different places, by including or excluding particular comorbidities, the model trained on Duke data might predict less well. This is not a fatal flaw, but it is a reason that external validation across multiple healthcare systems is the next test the work needs to pass.

The Underdiagnosis of Girls

If algorithmic prediction has a strongest single use case in pediatric ADHD, it is the systematic underdiagnosis of girls. The case is worth understanding in detail, because it illustrates both why earlier detection matters and why doing it well is harder than it looks.

Girls are diagnosed with ADHD at roughly half the rate of boys. Researchers do not believe the prevalence is actually that different. What is different is how the condition presents and how the existing detection apparatus responds to those presentations. Girls are more likely than boys to present with the inattentive subtype of ADHD, which produces internal struggles, daydreaming, difficulty completing work, organizational chaos, rather than the externalizing behaviors, fidgeting, calling out, leaving the seat, that are most likely to trigger teacher complaints and referrals for evaluation.

The detection bias is, in this sense, a function of the detection apparatus. Current ADHD screening is largely complaint-driven. A child whose behavior disrupts the classroom gets referred. A child whose attention is failing internally, who is quietly anxious about her inability to keep up, who learns to mask her difficulty through extra effort or social camouflage, does not. By the time the diagnosis finally arrives, often in adolescence or adulthood, the original ADHD has typically been joined by anxiety, depression, low self-esteem, and the cumulative damage of years of misattributed struggle.

The longitudinal evidence on this is striking. Studies of girls diagnosed with ADHD in childhood, followed prospectively into early adulthood, show substantial elevated rates of anxiety, mood disorders, and self-harm relative to their non-ADHD peers. A 2025 meta-analysis published in the Journal of Attention Disorders found that depression rates in girls with ADHD reached roughly 21 percent, more than double the 9 percent rate observed in boys with ADHD. The mechanism most researchers point to is precisely the late-diagnosis pathway: girls with undetected ADHD are not somehow producing different brains; they are producing different consequences of having ADHD that no one has identified.

This is the case for which earlier algorithmic detection has its strongest theoretical appeal. A model that surfaces ADHD-trajectory patterns directly from medical records, without requiring a teacher complaint or a parent’s articulate request for evaluation, could in principle identify girls who would otherwise spend a decade slipping below the radar of complaint-driven screening. The qualification matters. The same algorithmic detection could also be biased in the same way the data is biased. If girls with ADHD show up in their medical records in less visible ways than boys do, because they are also masking in healthcare interactions, the model trained on those records may inherit the same blind spots that human clinicians have. Whether the Duke team’s claim of cross-sex accuracy holds up in external validation will be the empirical test.

Researchers studying the female-ADHD phenotype have, over the past five years, increasingly argued that current diagnostic criteria themselves need adjustment to better capture the inattentive and internalizing presentations. The Duke tool does not change the underlying criteria. It changes how the existing criteria are applied at the point of entry to evaluation. If it works equitably, it does so by surfacing the children whose presentations would otherwise be filtered out of the referral pipeline.

The Promise

To understand the appeal of earlier detection, it helps to understand the current trajectory of an undiagnosed child with ADHD.

Most children with ADHD are not diagnosed in early childhood. The typical pattern is that the child enters kindergarten or first grade, struggles in ways that are visible to teachers, is referred for evaluation, waits months or years to be seen, eventually receives a diagnosis, and then begins receiving support. The total delay between when symptoms first emerge and when support arrives is often four or more years. During that delay, the child is failing in school. The child is being told, implicitly or explicitly, that they are lazy, undisciplined, careless, distracted by choice. The child’s relationships with teachers and peers are deteriorating. The child’s self-concept is forming around a series of failures that the adults around them have not yet learned to attribute to a treatable cognitive difference.

By the time the diagnosis arrives, the secondary harms of years of misattributed struggle are often as substantial as the primary cognitive symptoms. The child has internalized shame. Their academic confidence is damaged. The trajectory their teachers and family expect for them has narrowed.

The case for earlier identification is the case for shortening that delay. A child flagged at age five for evaluation can receive supports, classroom accommodations, behavioral interventions, sometimes medication, before the secondary harms accumulate. The interventions that are well-supported by evidence work better when they are delivered early. The relational dynamics with teachers and peers can be set on a different trajectory if the child’s needs are understood before the child’s identity has formed around being the kid who cannot do what other kids can.

“Children with ADHD can really struggle when their needs aren’t understood and adequate supports are not in place,” Davis said in the Duke statement. “Connecting families with timely, evidence-based interventions is essential for helping them achieve their goals and laying a foundation for future success.”

The promise of the Duke tool is precisely this kind of compression of the delay. A pediatrician using such a system would not be told that a specific child has ADHD. They would be told that a specific child shows patterns that resemble the early trajectories of children who were later diagnosed with ADHD, and that further evaluation may be warranted. The clinician would still be the decision-maker. The tool would be focusing attention, not making diagnoses.

The Deployment Questions

The case for earlier detection runs into several substantive questions when the technology actually has to be deployed.

False Positives

ADHD is already one of the most commonly diagnosed conditions in American children. The most recent CDC data finds that roughly 11 percent of US children carry an ADHD diagnosis, with rates of 15 percent in boys and 8 percent in girls. A screening tool that flags additional children for evaluation will, by basic arithmetic, generate some number of false positives, children identified as showing patterns consistent with ADHD trajectories who would not actually develop the condition. The cost of a false positive is not zero. A child flagged by an algorithm for ADHD evaluation enters a clinical pathway with real consequences: parent worry, school awareness, possible labeling, possible early treatment for a condition the child does not have. The Duke paper reports the model’s accuracy in technical terms, but the operational question of false positive rates at the population level, what proportion of children flagged would have been correctly flagged and what proportion would not, is the question that will determine whether the tool helps or harms the affected population on net.

Labeling Effects

Decades of research on the diagnostic cascade suggests that medical labels in childhood are not inert. They affect how parents perceive their child. They affect how teachers calibrate their expectations. They affect what the child is told about themselves. For children who genuinely have a condition, an accurate label is often beneficial because it organizes intervention and removes the misattribution of cognitive differences to character flaws. For children who do not have the condition the label suggests, an inaccurate label can produce real harm: lowered expectations, premature medicalization of ordinary developmental variation, an identity built around a diagnosis that does not actually describe the child. Algorithmic flagging that funnels children into evaluation may produce both kinds of labels at scale, and the ratio of help to harm depends on the precision of the system in real-world deployment. Children who carry diagnoses they do not need can lose, through that diagnosis alone, opportunities the system would have given them if they had been seen as unlabeled and ordinary.

Consent

The model the Duke team built was trained on existing electronic health records, almost certainly with appropriate institutional review board oversight for that specific research use. The deployment question is different. If primary care systems begin using such tools as routine screening, families will be having their children’s medical records processed by an algorithm to generate predictions about future mental health diagnoses. Whether this counts as a use of medical data that requires specific informed consent, or as a routine quality-improvement use that is covered by general consent to be a patient in the system, is a question that current health privacy frameworks are not well-equipped to answer.

The gap between research-stage IRB oversight and deployment-stage routine use is one of the most contested areas in current healthcare AI policy. When a researcher analyzes existing records to develop a model, the use is bounded, the population is consented to research participation through their general patient agreements, and the output is published rather than deployed. When the same model is then used to make decisions about individual patients in care, the analogies change. The patient whose record is now being processed by the algorithm did not necessarily consent to algorithmic risk-stratification as a condition of being a patient. They may not know it is happening. They may have specific objections if they did know.

The legal frameworks governing this are still being written. HIPAA, the U.S. health privacy law, does not contain a specific category for algorithmic processing of records to generate risk predictions about future diagnoses. Most institutional consent forms do not address it. The Duke researchers have not, in their published work, made claims about deployment that would require resolving these questions. The questions will need to be resolved by the institutions that choose to deploy the tool.

Supply

ADHD evaluation in many regions is already characterized by long waitlists. Specialty clinics in pediatric mental health are oversubscribed. Some states report waitlists of twelve to eighteen months for an initial developmental evaluation. The bottleneck on faster ADHD identification is not, in most places, that children are not being flagged. It is that flagged children cannot be evaluated, and evaluated children cannot be connected to evidence-based treatment, in any reasonable timeframe.

A predictive tool that surfaces more children at younger ages will arrive into a system that is already overwhelmed. The most concerning version of this scenario is one in which the additional flagged children sit on waitlists during the years that the supports were supposed to be delivered, and arrive at evaluation, eventually, at the same age they would have arrived without the algorithmic flag, but with the additional burden of having been algorithmically tagged for a condition they may or may not have. Whether the Duke tool improves outcomes will depend at least as much on what happens after the flag as on the flag itself.

Where This Goes

The fifth question is the broader one of where this kind of algorithmic mental health screening goes. ADHD is the first application because it is common, because the research team had access to a useful dataset, and because the case for earlier identification is reasonably strong. The same approach, predicting future mental health diagnoses from patterns in routine pediatric records, could be applied to depression, anxiety, autism, eating disorders, and conduct disorders. Each application will have its own promise-question ratio. The framework being established with ADHD will shape how the entire category of pediatric mental health prediction develops.

This is not an AI doctor. It’s a tool to help clinicians focus their time and resources, so kids who need help don’t fall through the cracks or wait years for answers. Dr. Matthew Engelhard, Duke University

The Autism Detection Precedent

The history of pediatric mental health is not without precedent for the dynamic the Duke work is initiating. The clearest parallel is the autism early-detection movement of the early 2000s and 2010s, which made similar claims about the value of compressed identification delays and produced a similar mix of clear benefits and complicated downstream effects.

The Modified Checklist for Autism in Toddlers, or M-CHAT, was developed and refined in the early 2000s as a brief parent-administered screening instrument that pediatricians could deliver during routine 18-month and 24-month visits. The American Academy of Pediatrics began recommending universal autism screening in 2007. The result was a significant earlier identification of children on the autism spectrum, more access to early intervention services, and a meaningful improvement in some outcomes for children who would otherwise have been identified years later.

It also produced effects the original advocates did not predict. Identification was not equally distributed. Children whose parents were comfortable reporting concerns on the screening tool were more likely to be flagged. Children of color, children in families that did not present concerns to pediatricians comfortably, children whose presentations did not match the screening instrument’s underlying assumptions, were not. The earlier-detection promise was disproportionately delivered to families who were already navigating the healthcare system well. The diagnostic system bent toward earlier identification, but in ways that, in the absence of deliberate effort, reproduced existing inequities in who got identified.

The autism precedent is not a reason to oppose algorithmic ADHD detection. It is a reason to take seriously the question of who benefits from the earlier-detection promise and who does not. A tool deployed without attention to differential data availability, differential clinician deployment, and differential downstream service access will produce a similar pattern: better outcomes for the populations already best served, less change for the populations the technology was theoretically going to help most.

The Identity Question

The Duke tool prevents one kind of harm and creates the possibility of another. The harm it prevents is the one its developers are focused on: years of misattributed struggle, identity damage from being treated as lazy or careless when one is actually contending with a treatable cognitive difference. The harm it creates, if deployed widely, is a different kind of identity formation: the five-year-old whose parents are told by an algorithm that their child shows patterns consistent with future ADHD diagnosis.

What does it do to a young child to grow up under the shadow of an algorithmic prediction? The honest answer is that no one knows yet. The literature on diagnostic labeling effects is rich, but it concerns the effects of actual diagnoses applied after evaluation. The effects of pre-diagnostic algorithmic risk stratification on identity formation, parent-child relationships, teacher expectations, and the child’s own self-narrative are not yet established.

The neuroaffirming community, which has argued for years that neurodivergent identities should be supported rather than corrected, has views on this that are worth engaging with seriously. The case from neuroaffirming perspectives is that early identification, when it leads to earlier support and accommodation, is genuinely valuable. The concern is when identification leads not to support but to surveillance, to normalization pressure, to early medicalization of patterns that, in another framing, would be features of a particular brain rather than symptoms of a disorder. Algorithmic detection that funnels children into a primarily medical pathway, with medication and clinical treatment as default interventions, may not deliver the kind of support neurodivergent children actually benefit from most. Support that affirms difference rather than corrects it requires a particular kind of clinical and educational infrastructure. Whether the systems being built around algorithmic detection are oriented toward that kind of support, or toward older models of medicalized intervention, is one of the substantive questions the field will need to answer.

The framing parents are given about an algorithmic risk score also matters. A child whose parents are told “your child shows patterns that resemble children who were later diagnosed with ADHD; here are some resources for understanding what that might mean and how to support your child if it becomes relevant” is a different child than one whose parents are told “your child has an 80 percent risk of ADHD and we recommend early evaluation and treatment.” The first framing treats the child as a developmental being whose trajectory is open. The second treats the child as already a patient. The choice between these framings is not made by the algorithm. It is made by the clinical infrastructure built around the algorithm.

The Methodology and Its Limits

The Duke team has made some choices that are worth understanding for readers trying to evaluate the strength of the finding.

The model was built using methods now standard in healthcare AI: a foundation model approach, with pretraining on a large general dataset followed by fine-tuning on a task-specific dataset. This is the same architectural approach that powers most modern AI systems. The technical contribution of the Duke work is not a new kind of model but a careful application of existing methods to a specific problem with a specific dataset.

The validation was done within the Duke Health system, which has its own particular patient population, clinical practices, and documentation standards. Models that perform well on one health system’s records do not always transfer to other systems. The next test for this work will be external validation: can the same model, or a model trained the same way, predict ADHD trajectories in a different healthcare environment with comparable accuracy? Until external validation is done, the equity claim, that the model performs comparably across demographic groups, is preliminary. It is an observation about how the model performs within the Duke validation set, not yet an established property of the model in general clinical deployment.

The ground truth in the study, what counts as an “ADHD case” for purposes of training the model, is itself a clinical artifact. The children identified as having ADHD in the Duke records are children who received an ADHD diagnosis from a Duke clinician. The model is therefore learning to predict not “true ADHD” in some objective sense but “ADHD as it is diagnosed in this particular healthcare system.” If diagnostic practices vary across systems, populations, or eras, models built on one set of practices may not generalize to others.

The model was not designed to diagnose. It was designed to flag children for further evaluation. The accuracy figures the paper reports are about flagging, not diagnosis. This is an important distinction: a model that performs accurately at flagging children for evaluation is not the same thing as a model that is accurate at saying a child has ADHD. The downstream evaluation by a clinician is what determines whether a flag becomes a diagnosis.

What This Means for Families and Clinicians

For families with a young child showing developmental concerns or behavioral patterns that worry them, the Duke study is most usefully understood as a signal about where the field is heading, not as a tool currently available to them. The model is not deployed for general clinical use. Pediatricians in most practices do not have access to it. The question of whether to seek evaluation for a young child remains, as it has been, a question to discuss with a primary care provider based on observed concerns.

For pediatricians and family medicine clinicians, the research is an indication that the technological landscape of pediatric mental health screening is changing. Whether or not the Duke model itself reaches their practice, similar tools are likely to be developed and offered for use. The clinical and ethical questions about how to integrate algorithmic risk scores into pediatric practice, when to share them with families, how to weigh them against clinician judgment, how to handle false positives, are questions the field will need to work through before deployment becomes routine.

For the ADHD community more broadly, the research surfaces a real tension. Earlier identification is a goal almost everyone supports in principle. Algorithmic identification, with all the questions it raises about consent, labeling, and inequity, is a deployment strategy that is more contested. The conversation about how to get the benefits of earlier identification without the costs of algorithmic surveillance is one the field will be having for years.

For the broader healthcare system, the Duke work is one entry in what will be a long line of AI-based mental health prediction tools. The framework that develops around this first generation of tools, who gets to deploy them, under what conditions, with what consent, with what oversight, will shape what is possible and what is permitted in the second generation. The decisions being made about ADHD prediction now will be the precedents for prediction tools targeting other conditions later.

The Study at a Glance

The Duke Study
Early Attention Deficit Hyperactivity Disorder Prediction from Longitudinal Electronic Health Records
Authors: Elliot D. Hill, De Rong Loh, Naomi O. Davis, Benjamin A. Goldstein, Geraldine Dawson, Matthew Engelhard • Institution: Duke University School of Medicine, Departments of Biostatistics and Bioinformatics and Psychiatry and Behavioral Sciences • Journal: Nature Mental Health, April 27, 2026 • DOI: 10.1038/s44220-026-00628-2Sample: Electronic health records from more than 140,000 children at Duke Health, with pretraining on records from over 720,000 patients • Method: Foundation-model approach with pretraining and fine-tuning, validated within the Duke Health system across demographic subgroups • Key finding: The model can predict eventual ADHD diagnosis by age five, with accuracy reported as comparable across sex, race, and insurance status within the validation system. • Funding: National Institute of Mental Health.

Sources and Further Reading

  1. Hill, E. D., Loh, D. R., Davis, N. O., Goldstein, B. A., Dawson, G., & Engelhard, M. (2026). Early attention deficit hyperactivity disorder prediction from longitudinal electronic health records. Nature Mental Health. DOI: 10.1038/s44220-026-00628-2.
  2. Duke Health press release: AI Tool May Spot ADHD Years Before Children Are Diagnosed, April 27, 2026.
  3. Duke Department of Psychiatry and Behavioral Sciences: AI Tool May Spot ADHD Years Before Children Are Diagnosed.
  4. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447-453.
  5. Centers for Disease Control and Prevention: ADHD in Children: Data and Statistics.
  6. Wang, S., Stewart, T. M., Ozen, I., Mukherjee, A., & Rhodes, S. M. (2025). Rates of Depression in Children and Adolescents With ADHD: A Systematic Review and Meta-Analysis. Journal of Attention Disorders.
  7. EurekAlert release: AI tool may spot ADHD years before children are diagnosed, April 27, 2026.

About This Publication

Executive Function News is a research and analysis publication of NBEFC®, the National Board for Executive Function Certification. NBEFC offers board certification for executive function coaches at nbefc.org. Our editorial process applies independent journalistic standards to research coverage, regardless of the topic’s relationship to NBEFC’s programs.

Scroll to Top