Preschoolers’ Test Scores and Their Parents’ Ratings of Executive Function Barely Agreed. The Median Correlation Was 0.07. | Executive Function News
Independent Journalism on Executive Function and Neuroscience

Executive Function News

Reporting on the Science and Practice of Executive Function
Executive Function

Preschoolers’ Test Scores and Their Parents’ Ratings of Executive Function Barely Agreed. The Median Correlation Was 0.07.

Researchers followed 110 preschoolers through three assessments, six months apart, measuring executive function two ways: with performance tasks administered to the child, and with a standard questionnaire completed by the parent. The two produced almost unrelated pictures of the same children. That result is not an anomaly, and treating it as a problem to be fixed misunderstands what each instrument is doing.

A preschool-age child sorting colored shapes at a table.
Performance-based executive function assessment in early childhood typically looks like this: a structured task, administered by an adult, with a clear instruction and a scoreable outcome. What it captures and what a parent observes at home turn out to overlap far less than most people assume. Photo: Executive Function News.

A correlation of 0.07 is, for practical purposes, no relationship at all. It means that knowing one number tells you essentially nothing about the other. That is the median figure reported in the August issue of Psychological Assessment by Alejandra Contreras Solorzano, Yaewon Kim, Abigail Reid Graves, Mauricio Garcia-Barrera and Ulrich Müller, describing the relationship between preschoolers’ performance on executive function tasks and their parents’ ratings of the same children on the BRIEF-P. Two instruments, both widely used, both named for the same construct, agreeing at a level indistinguishable from chance.

The study followed 110 children who were 36 to 47 months old at the start, with a mean age just over 40 months. Each completed five performance-based tasks at three time points, six months apart. Parents completed the BRIEF-P at each point and, at the final assessment, also rated their children’s socioemotional and behavioural functioning. Rather than relying on simple correlations alone, the researchers used growth models, which track individual trajectories over time rather than comparing group averages at each stop.

The usual reaction to a result like this is to ask which measure is wrong. That is the wrong question, and the reason it is wrong is the most useful thing in the study.

What Each Instrument Is Actually Asking

A performance-based executive function task puts a child in a controlled situation, gives an explicit instruction, removes competing demands, and records what happens. Sort these cards by color, now sort them by shape. Say night when you see the sun. Remember this sequence and repeat it back. The examiner has the child’s attention, the task is bounded, the motivation is supplied by the situation, and the score reflects what the child can do under those conditions.

The BRIEF-P asks something else entirely. It is a 63-item questionnaire for children aged 2 through 5, developed by Gioia, Espy and Isquith, covering inhibition, shifting, emotional control, working memory, and planning and organizing. A parent rates how often specific everyday problems occur. Does the child have trouble with transitions? Get overwhelmed by tasks with several steps? Become upset by small changes in routine?

Read those two descriptions side by side and the surprise is not that they correlate at 0.07. The surprise is that anyone expected otherwise. One measures capacity under optimal conditions. The other measures how often things go wrong under ordinary ones. A child can have the capacity and still have things go wrong constantly, because home has no examiner, no single task, and no removal of competing demands.

What the Tasks Actually Look Like

It helps to be concrete, because “performance-based executive function task” sounds more technical than what happens in the room.

In the Day-Night task, a child is shown a card with a sun and has to say “night,” then a card with a moon and has to say “day.” The correct answer is the one the picture is not. Success requires holding an arbitrary rule in mind while suppressing the obvious response, which is inhibition and working memory operating together.

In dimensional change card sorting, a child sorts cards by one dimension, colour, and then is told to switch and sort the same cards by another, shape. Younger preschoolers characteristically keep sorting by the first rule even while correctly reciting the second one, which is a failure of shifting rather than of understanding.

Susan Carlson’s work on developmentally sensitive measures for this age group catalogued the design problem these tasks face. A preschool executive function task has to be hard enough to produce variation and simple enough that a four-year-old can hold the instruction, and the window between those is narrow. Push slightly too hard and everyone fails. Ease off slightly and everyone succeeds. Either outcome destroys the measurement.

Held against that, a parent answering whether their child has trouble with transitions is doing something not remotely similar. There is no instruction to hold, no bounded trial, no scoring. They are summarising months of accumulated impression.

Why Preschool Is the Revealing Case

This divergence has been documented before, but preschool is where you would most expect the measures to converge, which makes the failure to converge more informative.

Older children and adults have compensatory strategies. They build routines, use reminders, arrange their environments, and lean on other people. Those strategies can mask an underlying difficulty on a rating scale while a lab task still detects it, or conversely can leave a person functioning poorly in daily life despite intact test performance. Preschoolers have far fewer of these. Less scaffolding sits between the underlying capacity and the visible behavior.

And yet the correlation here is lower than the figure from the broader literature. When Maggie Toplak, Richard West and Keith Stanovich reviewed the question across 20 studies covering children and adults, they found a median correlation of 0.19, with only about a quarter of the 286 correlations they examined reaching statistical significance. Their conclusion was that performance-based measures and rating measures assess different underlying constructs. This preschool study lands well below even that modest figure.

One point about the numbers is worth pausing on, because it affects how the range should be read. On the BRIEF-P, a higher score means more executive difficulty, while on a performance task a higher score means better performance. Agreement between the two would therefore appear as a negative correlation. The reported range runs from -0.30, which represents modest agreement, up to 0.03, which represents none at all. Even the strongest association in the set is weak.

There is also a measurement problem specific to this age group. A systematic review of BRIEF-P studies in preschoolers with ADHD-compatible symptoms found floor effects on tasks tapping emotionally laden executive function and ceiling effects on the more purely cognitive ones. Instruments in early childhood are working near the edges of their usable range, which compresses variance and weakens any correlation that depends on it. That does not explain away a 0.07, but it belongs in the accounting.

By the Numbers
  • 110 children, 36 to 47 months old at baseline, were assessed three times at six-month intervals. Each completed five performance-based tasks per wave.
  • Correlations between the performance composites and parent BRIEF-P ratings ranged from -0.30 to 0.03, with a median reported as 0.07.
  • Only performance scores improved over time in the growth models, at a highly significant level.
  • Only BRIEF-P scores predicted behavioural difficulties, lower prosocial behaviour, and other executive-related outcomes.
  • Those outcomes were also parent-reported, so the authors flag shared method variance as a possible explanation.
  • The sample was 48 percent female and 85 percent White, which limits how far the findings generalise.
  • An earlier review across 20 studies found a median correlation of 0.19 between the two measure types, with about a quarter of correlations reaching significance.
  • The BRIEF-P contains 63 items for ages 2 through 5, covering inhibition, shifting, emotional control, working memory, and planning.
  • Reviews of BRIEF-P use in preschoolers have documented floor effects on emotionally laden executive tasks and ceiling effects on cognitive ones.
  • A separate preregistered study found prefrontal connectivity measured by fNIRS predicted both task performance and behaviour ratings, despite the two measures not converging with each other.

The Age Signal Only One Measure Picked Up

The longitudinal design produced something a single snapshot could not, and it is the most quietly interesting result in the study.

Performance-test scores improved with age. That is exactly what should happen. Executive function develops rapidly across the preschool years, and a well-constructed task should register that development. The task is the same at each assessment; the child gets better at it.

Parent ratings did not carry that developmental signal in the same way. On its face that looks like a failure of the questionnaire. It is better understood as a feature of what a rating scale is.

A parent is not scoring against a fixed standard. They are scoring against expectations, and expectations move as the child ages. A three-year-old who cannot wait his turn is a three-year-old. A five-year-old who cannot wait his turn is a problem. The same behavior, unchanged, can generate a worse rating a year later, because the yardstick has shifted. Meanwhile the demands placed on the child have escalated too, at home and increasingly at preschool, so the child’s improving capacity is being measured against a rising bar.

The child is getting better and the demands are getting harder at the same time. A performance task sees only the first. A parent sees the difference between them. On why the two measures diverge over time

This has a direct consequence for anyone tracking change. If you measure progress only with repeated questionnaires, you are measuring the gap between capacity and demand, not capacity itself. A child could be improving substantially while their ratings stay flat or worsen, simply because expectations rose faster than the child did. That is not a scoring error. It is what the instrument measures.

The Shared-Rater Problem

The study also found that only the parent ratings predicted behavioural difficulties, lower prosocial behaviour, and other executive-related outcomes. The performance tasks predicted none of them. That asymmetry is the finding most likely to be quoted, and it needs careful handling, because it points two directions at once.

The optimistic reading is that rating scales have ecological validity that lab tasks lack. Parents see the child across hundreds of unstructured situations. They observe the thing that actually matters, which is functioning, rather than a proxy for it. If parent ratings predict real-world difficulty better, that is a point in their favor.

The problem is that the difficulty outcomes were also reported by parents. So what the study established is that a parent’s rating of a child’s executive function predicts the same parent’s rating of that child’s behavior and social functioning. That correlation could reflect a real underlying pattern that this parent is well positioned to observe. It could equally reflect something about the rater: a parent under strain, or with a particular threshold for what counts as a problem, or holding a general impression of their child that colors every item they answer.

The study cannot distinguish between those. This is not a flaw specific to this paper; it is a structural feature of any design where one informant supplies both predictor and outcome. It is also why teacher ratings and direct observation matter as independent sources. Research comparing parent and teacher BRIEF-P ratings has found that they diverge too, with parents in one study attributing more executive difficulty than teachers did to the same children.

The uncomfortable implication is that a rating scale carries information about the rater as well as the child. Work by Buchanan, published in this same journal, found that adult self-report measures of executive function problems correlated with personality rather than with performance-based measures. Whether something analogous operates in parent reports of children is not settled, but the possibility is not exotic.

A Third Informant Produces a Third Picture

If parents and tasks disagree, an obvious move is to ask a teacher. That produces its own complication.

A 2023 study of 130 children aged roughly three to six had both parents and teachers complete the BRIEF-P on the same children, alongside direct measures of emotion comprehension, language, and non-verbal reasoning. Parents and teachers differed systematically, with parents attributing more executive difficulty to the children than teachers did. More usefully, the associations between rated executive difficulty and the directly measured skills were stronger when the ratings came from teachers.

There are several plausible reasons and the study does not adjudicate among them. Teachers see many children of the same age and have a calibrated sense of normal that a parent with one or two children cannot have. Teachers also see the child in a structured environment with clear expectations, which is closer to a task than a living room is. Parents, on the other hand, see the unstructured hours where executive demands are heaviest and support is thinnest, and they see the child tired.

The practical point is not that teacher ratings are better. It is that three sources produce three pictures, and each source is positioned to see something the others cannot. Treating any single one as the truth against which the others are checked is the error the whole literature keeps pointing at.

Both Measures May Be Valid

A low correlation between two instruments is often read as evidence that at least one of them is failing. There is reason to think that inference is wrong here.

A preregistered study published in Developmental Science in 2022 took 41 children aged 4 and 5 and measured resting-state functional connectivity in the prefrontal cortex using functional near-infrared spectroscopy, alongside a task-based executive function measure and teacher BRIEF-P ratings. Patterns of prefrontal connectivity predicted both the task performance and the behaviour ratings at one month and again at four months, controlling for age and verbal ability.

That is a meaningful result for this question. A neural measure related to both instruments, even though the two instruments barely relate to each other. Two measures can each carry real signal about the developing brain while capturing different portions of it. Low convergence between them is not proof that one is noise.

There is also a technical constraint worth naming. The correlation between any two measures is capped by how reliable each of them is on its own, and preschool performance tasks are not highly reliable instruments. Young children are variable across sessions in ways that have nothing to do with the construct: tired, shy, distracted by the room, differently motivated by a particular examiner. Some portion of a 0.07 reflects that ceiling rather than genuine construct separation. It does not account for the whole gap, since Toplak and colleagues found the same pattern in adult samples with more reliable measures, but an honest reading includes it.

One further note on the number itself. A median correlation summarises many pairings between individual tasks and individual rating subscales. Some of those pairs will have run higher than 0.07 and some will have sat at or below zero. The median is the right summary for describing a body of relationships, but it is not a single relationship, and no individual pairing should be assumed to equal it.

A Finding That Keeps Being Found

What makes this study worth attention is not novelty. It is the accumulation.

Miranda and colleagues examined this in 209 preschoolers in 2015, comparing working memory and inhibition tasks against parent and teacher BRIEF ratings, with ADHD symptoms and word reading as outcomes. Toplak, West and Stanovich synthesized 20 studies in 2013 and concluded the two measure types assess different constructs. Buchanan reported the personality finding in 2016. In 2026, Hlutkowsky, All, Roule, Warner and Huang-Pollock compared commercially available parent and teacher rating forms against actual executive function performance, opening with the observation that it is often argued the two measure the same construct at different levels of analysis.

Often argued, and repeatedly not supported. Across more than a decade, with different age groups, instruments, and designs, the answer keeps coming back the same. The 0.07 is the sharpest version of it, in the youngest sample, with the longitudinal design that rules out a one-off measurement artifact.

At some point a finding this durable stops being a methodological curiosity and becomes a fact about the instruments that anyone using them needs to know.

What Follows for Anyone Assessing a Child

Several things follow, and they are more practical than the abstract framing suggests.

Disagreement between measures is information, not error. When a child performs well on tasks and rates poorly on a questionnaire, the reasonable interpretation is that capacity exists and something in the environment is preventing it from being deployed. When the pattern reverses, the child may be functioning on scaffolding that a bounded task strips away. Both patterns tell you something specific about where to intervene. Neither is resolved by picking a winner.

Do not average them into a single score. Combining two measures that correlate at 0.07 produces a composite that represents neither. If the instruments measure different constructs, a summary number discards the distinction that made collecting both worthwhile.

Ask what a score is for before choosing an instrument. To know whether a capacity has developed, a performance task is the appropriate tool. To know whether daily life is going badly, a rating scale gets closer. To know why, neither is sufficient without specifics about which situations break down and what surrounds them.

Be cautious about tracking change with questionnaires alone. This is the implication with the sharpest edge. Repeated ratings by a single informant are vulnerable to shifting expectations, rising demands, and whatever else is happening to that informant across the measurement period. Performance data, observation, and specific functional examples give the change estimate something to stand on.

Collect functional detail alongside scores. A number tells you that something is difficult. Knowing which transitions break down, at what time of day, with which adults present, and what happens immediately before, tells you what to actually do. That detail is not a supplement to assessment. In a domain where two validated instruments agree at 0.07, it may be the most reliable part of it.

The authors’ own stated conclusion is exactly this: that the findings underscore the importance of using multiple methods to assess executive function. They also name shared method variance themselves as a candidate explanation for the parent-rating result, which is the objection a critic would raise, volunteered by the researchers who found it.

It is worth adding that the stakes here are not academic. A 2024 review and meta-analysis by Stucke and Doebel in Psychological Bulletin confirmed that early childhood executive function predicts concurrent and later social and behavioural outcomes. Executive function in the preschool years matters for how a child’s life goes. That is precisely why measuring it badly, or measuring it with one instrument and believing you have the whole picture, has consequences.

None of this means either instrument should be discarded. It means the field has spent a long time treating agreement between them as the standard for validity, and a decade of results now says that was the wrong standard. Two measures of the same name that do not converge are not a problem to be solved. They are two different questions, and both are worth asking.

Sources and Further Reading

  1. Contreras Solorzano, A., Kim, Y., Graves, A. R., Garcia-Barrera, M. A., & Müller, U. (2026). Measuring preschoolers’ executive function through performance-based measures and parent reports: A longitudinal study. Psychological Assessment, 38(8), 532-546.
  2. Stucke, N. J., & Doebel, S. (2024). Early childhood executive function predicts concurrent and later social and behavioral outcomes: A review and meta-analysis. Psychological Bulletin, 150(10), 1178-1206.
  3. Garon, N. M., Piccinin, C., & Smith, I. M. (2016). Does the BRIEF-P predict specific executive function components in preschoolers? Applied Neuropsychology: Child, 5(2), 110-118.
  4. Toplak, M. E., West, R. F., & Stanovich, K. E. (2013). Practitioner review: Do performance-based measures and ratings of executive function assess the same construct? Journal of Child Psychology and Psychiatry, 54(2), 131-143.
  5. Buchanan, T. (2016). Self-report measures of executive function problems correlate with personality, not performance-based executive function measures, in nonclinical samples. Psychological Assessment, 28(4), 372-385.
  6. Miranda, A., Colomer, C., Mercader, J., Fernández, M. I., & Presentación, M. J. (2015). Performance-based tests versus behavioral ratings in the assessment of executive functioning in preschoolers: Associations with ADHD symptoms and reading achievement. Frontiers in Psychology, 6, 545.
  7. Hlutkowsky, C. O., All, K. E., Roule, A. L., Warner, T. A., & Huang-Pollock, C. (2026). A comparison of commercially available parent and teacher rating forms in the concurrent prediction of executive functioning performance in children. Journal of Attention Disorders.
  8. Eng, C. M., et al. (2022). Longitudinal investigation of executive function development employing task-based, teacher reports, and fNIRS multimethodology in 4- to 5-year-old children. Developmental Science.
  9. Carlson, S. M. (2005). Developmentally sensitive measures of executive function in preschool children. Developmental Neuropsychology, 28(2), 595-616.
  10. Gioia, G. A., Espy, K. A., & Isquith, P. K. (2003). BRIEF-P: Behavior Rating Inventory of Executive Function, Preschool Version. Psychological Assessment Resources.

About This Publication

Executive Function News is a research and analysis publication of NBEFC®, the National Board for Executive Function Certification. NBEFC offers board certification for executive function coaches at nbefc.org. Our editorial process applies independent journalistic standards to research coverage, regardless of the topic’s relationship to NBEFC’s programs.

Scroll to Top