Executive Function News
Adults Who Overreported Symptoms Scored Higher on Executive Function Questionnaires. That Still Did Not Explain Why the Questionnaires Disagree With the Tests.
There is a tidy explanation for why executive function rating scales and objective tests keep failing to agree: some people exaggerate. Researchers tested that explanation in 99 adults undergoing ADHD evaluations, using validity measures to sort credible from noncredible reporters. Overreporting turned out to be real and detectable. It also turned out not to be the answer. What the ratings tracked instead was psychological distress, and that link was strongest among the people reporting credibly.
Executive function rating scales and objective executive function tests have been failing to agree with each other for well over a decade, across age groups, instruments, and study designs. One explanation has always been available and has the advantage of being simple: in the settings where this matters most, some proportion of people are overstating their difficulties. Adult ADHD evaluation is exactly such a setting. The assessment leans heavily on self-report, rates of noncredible presentation are documented as high, and the incentives are concrete, since a diagnosis can produce access to stimulant medication and to academic accommodations. If exaggeration were driving the divergence, this is where it would show.
Julie Suhr, Adrienne Jankowski, Taylor Lambertus and Grace Lee set out to test that directly, and published the result in Psychological Assessment. They took an archival sample of 99 adults who had presented for ADHD evaluation, sorted them by performance on symptom validity measures, and asked whether validity group status changed the relationship between rating scales and tests.
It did not. The explanation failed, and the way it failed is more useful than a confirmation would have been.
Why Adult ADHD Evaluation Is the Hard Case
The authors chose this population deliberately, and the reasoning is worth spelling out, because it is what makes the null result meaningful.
Adult ADHD assessment sits in an unusual position among psychological evaluations. The diagnostic criteria are behavioural, and the behaviours are ones only the person themselves can reliably report on. There is no blood test, no scan that settles it, and no adult equivalent of the teacher who watched the child for six hours a day. Childhood evaluation draws on parents and teachers as independent informants. Adults typically arrive alone, and retrospective childhood information is often unavailable or supplied by the same person being assessed.
That places enormous weight on self-report, and executive function rating scales carry much of it. They are quick, standardised, normed, and they ask about precisely the everyday difficulties an adult seeking evaluation wants to describe.
Then there are the incentives. A diagnosis can unlock stimulant medication, which is a controlled substance with recreational and performance-enhancing demand, and formal academic accommodations, including extended time on examinations. These are real, concrete benefits attached to a particular outcome of the assessment. Nothing about that implies most people seeking evaluation are pursuing them dishonestly, and the great majority are not. It does mean this is a setting where a validity question is legitimate rather than paranoid, and where clinicians have been trained to ask it.
So if overreporting were going to account for the gap between ratings and tests, this is the population and the setting where the effect should be largest and easiest to detect. That is what makes its failure to explain the gap informative rather than merely inconclusive.
What Symptom Validity Testing Actually Establishes
Before the findings, a word about the terminology, because it is easy to read more into it than it carries.
Symptom validity tests are measures designed to detect response patterns that are improbable if a person is reporting accurately. Some work by including items describing symptoms that almost nobody genuinely experiences, or symptom combinations that rarely co-occur. Someone endorsing many of these is producing a pattern that does not fit how the condition typically presents.
What such a test establishes is that a response profile falls outside expected bounds. It does not establish intent, and the technical label of noncredible reporting is not a synonym for lying. A person can produce an invalid profile through catastrophic thinking about their own functioning, through severe distress that colours every judgment, through misreading the questions, or through a genuine belief that their difficulties are more extreme than an outside observer would judge. Classification is also probabilistic rather than certain, and any cutoff will misclassify some people in both directions.
This matters for reading the study fairly. The group labelled noncredible is a group whose responses met a statistical threshold, not a group proven to have deceived anyone.
What the Study Found
Three results, in ascending order of interest.
First, the rating scale was vulnerable. Adults classified as noncredible reporters scored higher on the executive function rating scale than adults classified as credible reporters. That is a straightforward finding and it confirms something clinicians have long suspected: a commonly used self-report measure of executive function can be pushed upward, and the direction of that push is toward reporting more impairment.
Second, and this is the null that carries the paper, rating scale scores were largely unrelated to executive function test scores across the full sample, and validity group status had no moderating effect. In plainer terms: the ratings and the tests disagreed just as thoroughly among the credible reporters as among the noncredible ones. Removing the people whose reporting looked suspect did not repair the relationship. The divergence was still there.
Third, the ratings did relate to something. Higher executive function rating scores were associated with greater psychological distress. And here is the detail that reframes the whole study: for depression scores, the effect sizes were larger in the credible reporting group.
Rating scale scores were also related to self-reported functional impairment, and that held regardless of validity group.
- The sample was 99 adults, drawn from archival records of people who presented for ADHD evaluation.
- The setting was chosen deliberately: adult ADHD assessment relies heavily on self-report and carries external incentives including stimulant access and academic accommodations.
- Adults classified as noncredible reporters scored higher on the executive function rating scale than credible reporters.
- Across the full sample, rating scale scores were largely unrelated to executive function test scores.
- Validity group status had no moderating effect on that relationship. The gap persisted among credible reporters.
- Higher rating scale scores were associated with greater psychological distress.
- For depression specifically, the effect sizes were larger in the credible reporting group.
- Ratings were related to self-reported functional impairment regardless of validity group.
- The authors conclude that a commonly used executive function rating measure is vulnerable to noncredible reporting, while noncredible reporting is not a major contributor to the ratings and tests divergence.
Two Findings That Pull in Opposite Directions
The results are easy to misreport in either direction, and both misreadings are already predictable.
One camp will take the first finding and conclude that executive function rating scales are compromised in high-incentive settings and should be discounted. The other will take the second finding and conclude that concerns about overreporting have been overblown, since removing the noncredible group changed nothing.
Both are reading half the paper. What it establishes is that the rating scale is susceptible to overreporting, and that overreporting is not the reason the scale disagrees with objective tests. Those are separate claims and both are supported.
Holding them together produces the useful conclusion. If exaggeration were the whole story, screening it out would restore agreement between measures. It did not. Something more fundamental separates what a rating scale captures from what a test captures, and the growing evidence is that they are measuring different things rather than measuring the same thing with different amounts of error.
That interpretation now has support across the developmental span. A companion study in the same journal, following 110 children from age three, found a median correlation of 0.07 between performance-based executive function tasks and parent ratings, with only the tasks registering developmental improvement and only the ratings predicting behavioural outcomes. Preschoolers being rated by their parents have no incentive to exaggerate and no capacity to coordinate it. The divergence appears there too.
What a Well-Aimed Null Result Is Worth
Studies reporting that an expected effect did not appear are published less often than studies reporting that one did, and they are read less carefully when they are. That is a mistake in a case like this one.
A null result carries weight in proportion to how well the study was positioned to detect the effect. A poorly designed study in an unlikely population that finds nothing has told you very little, because the absence could easily be the design’s fault. This study was aimed at the place the effect was most likely to be: a clinical sample with documented rates of noncredible presentation, real external incentives, and heavy reliance on the instrument in question. The researchers had a validity measure to split the sample with, and the split worked, since the two groups did differ on rating scale scores exactly as predicted.
Having established that their grouping variable was detecting something real, they then found that it did not moderate the relationship of interest. The tool worked; the hypothesis did not survive it. That is a considerably stronger form of null result than a study that simply failed to find anything.
The practical consequence is that a common reassurance no longer holds. It has been possible to argue that rating scales are trustworthy once you have screened for overreporting, and that the poor agreement with tests is mostly a contamination problem. This study screened for overreporting and the poor agreement remained.
The Depression Finding, and Why Its Direction Matters
The most practically consequential result is the association between rating scores and distress, particularly depression, and the reason it matters is the direction of the effect size difference.
If elevated executive function ratings mainly reflected exaggeration, the association with depression should have been strongest among the noncredible reporters, since general symptom amplification would inflate both scales at once. Instead the depression association was stronger in the credible group. Among people whose reporting profile looked valid, the more depressed they were, the more executive difficulty they reported.
There are at least three ways to read that, and the study does not adjudicate between them.
Depression may genuinely impair executive functioning. This is well established clinically, and if so the ratings are picking up something real, just not something specific to ADHD.
Depression may alter self-perception without altering capacity. A depressed person may judge their own functioning more harshly, recall failures more readily than successes, and report accordingly. The rating would then reflect the appraisal rather than the performance.
Or executive difficulty may contribute to depression. Years of missed deadlines, disorganisation, and unmet intentions are demoralising, and the causal arrow could run in that direction, or in both.
A cross-sectional design cannot separate these. What it can establish is that a high score on an executive function rating scale, taken alone, does not identify why the score is high. An adult reporting substantial executive difficulty may have ADHD. They may be depressed. They may be both, which is common. The rating scale by itself cannot tell you which.
What the Ratings Did Predict
It would be a mistake to leave this thinking the rating scale measures nothing useful.
Rating scale scores related to self-reported functional impairment, and that held in both validity groups. Whatever the ratings capture, it tracks the experience of things going badly in daily life. That is not a trivial thing to measure, and it is arguably closer to what brings a person to an evaluation in the first place than a test score is. Nobody seeks an assessment because they performed poorly on a card-sorting task. They seek one because their life is not working the way they want it to.
The limitation is that functional impairment was itself self-reported, so part of the association reflects the same person answering both sets of questions. This is the same shared-method problem that constrains the preschool study, where parents supplied both the executive function ratings and the behavioural outcomes. It is a structural feature of self-report research rather than a flaw specific to this paper, and it means the association should be read as suggestive rather than as independent confirmation.
What the Study Cannot Tell You
Several constraints deserve stating plainly.
The sample is 99 adults from archival clinical records, which means the researchers worked with data collected for clinical rather than research purposes, with whatever measures those clinicians happened to use. Archival designs are efficient and they trade away control over what was administered and how.
The design is cross-sectional, so no causal claim about depression and executive function follows from it in either direction.
The sample consists of people who sought ADHD evaluation, which is a self-selected group with a particular relationship to the question being asked. Findings here do not automatically extend to adults being screened in other contexts.
And symptom validity classification, as discussed, is a probabilistic judgment. Some people in the noncredible group were probably reporting accurately, and some in the credible group were probably not.
What Follows for Assessment and for Coaching
The implications are specific and, for anyone using these instruments, fairly demanding.
A high executive function self-report score is a starting point, not a finding. It establishes that a person reports substantial difficulty. It does not establish ADHD, objective cognitive impairment, or any particular neurological mechanism. This study demonstrates two separate reasons why: the score can be elevated by overreporting, and it can be elevated by distress in people reporting entirely credibly.
Emotional and contextual factors belong in the interpretation, not in a footnote. Given the depression association, an elevated executive function score in a person who is also struggling emotionally has more than one plausible reading. Treating the score as a clean measure of executive capacity discards that.
Concrete functional examples do work a score cannot. A number establishes that something is difficult. Which specific situations break down, at what time of day, under what conditions, and what happens immediately before, gives a picture that a scale total cannot. In a domain where the standard instruments disagree with each other this comprehensively, specific functional detail may be the most reliable information available.
Screeners and diagnostic assessments are different objects. A screener can identify who might benefit from a closer look. Diagnosis, particularly where medication or formal accommodations are at stake, requires a qualified evaluator using multiple sources, and this study is part of why. In a setting with external incentives, the authors warn specifically against relying on executive function rating scales without accounting for noncredible reporting and psychological distress.
The larger point is one the assessment literature has been circling for years and this study sharpens. The reflex when two measures disagree is to find the flaw in one of them. Suhr and colleagues went looking for the flaw in the obvious place, in a setting practically designed to produce it, and found that it was there and that it was not sufficient. The measures still disagreed. The most reasonable conclusion is that they were never measuring the same thing, and that using either one alone means seeing only part of the person in front of you.
Sources and Further Reading
- Suhr, J. A., Jankowski, A., Lambertus, T., & Lee, G. J. (2026). Does noncredible reporting moderate the relationship of executive functioning ratings to executive functioning tests, psychological symptoms, or functional impairment in adults seeking assessment for attention-deficit/hyperactivity disorder? Psychological Assessment.
- Contreras Solorzano, A., Kim, Y., Graves, A. R., Garcia-Barrera, M. A., & Müller, U. (2026). Measuring preschoolers’ executive function through performance-based measures and parent reports: A longitudinal study. Psychological Assessment, 38(8), 532-546.
- Toplak, M. E., West, R. F., & Stanovich, K. E. (2013). Practitioner review: Do performance-based measures and ratings of executive function assess the same construct? Journal of Child Psychology and Psychiatry, 54(2), 131-143.
- Buchanan, T. (2016). Self-report measures of executive function problems correlate with personality, not performance-based executive function measures, in nonclinical samples. Psychological Assessment, 28(4), 372-385.
- Hlutkowsky, C. O., All, K. E., Roule, A. L., Warner, T. A., & Huang-Pollock, C. (2026). A comparison of commercially available parent and teacher rating forms in the concurrent prediction of executive functioning performance in children. Journal of Attention Disorders.
Executive Function News is a research and analysis publication of NBEFC®, the National Board for Executive Function Certification. NBEFC offers board certification for executive function coaches at nbefc.org. Our editorial process applies independent journalistic standards to research coverage, regardless of the topic’s relationship to NBEFC’s programs.