Pre-K assessment metadata may identify executive function struggles, MDRC finds
A nine-week pilot involving 428 children found response time and accuracy data from Khan Academy Kids and Kibeam showed the strongest links with established measures of attention, impulse control and other executive functioning skills
MDRC researchers examined whether metadata automatically generated by technology-based pre-K assessments could provide signals of children’s executive functioning skills
Data already generated when young children complete digital literacy, language and math assessments could offer schools and assessment developers another way to identify children struggling with executive functioning, according to new research from MDRC.
The study examined whether information such as response times, accuracy and task completion could provide a useful signal of executive functioning skills in children aged three to five, without requiring a separate assessment.
Researchers Emily Hanno, Sophie Litschwartz, Emily Swinth, Victor Porcelli and Amy Taub tested the approach using data from Khan Academy Kids and Kibeam. The nine-week pilot, carried out in fall 2024 through MDRC’s Measures for Early Success Initiative, involved 428 children across 29 classrooms, 18 pre-K programs and five states.
The clearest result was not simply that speed could reveal how well a child was concentrating.
Measures combining how quickly children responded with whether they answered correctly had the strongest and most consistent positive associations with established executive functioning assessments. Across the two assessment platforms, correlations ranged from 0.31 to 0.67, with seven of the 12 correlations statistically significant.
By contrast, very fast answers did not appear to be a useful signal. The researchers found that their “not too fast” measure had correlations ranging from -0.24 to -0.01 with direct executive functioning assessments.
Slower responses proved more informative. The findings suggest children taking unusually long to answer may be showing inattention or disengagement, while quick responses can simply reflect the relatively straightforward format of early years assessment tasks.
Digital assessment data put to a second use
Executive functioning covers mental processes including attention, working memory, impulse control and the ability to shift focus.
These skills can currently be assessed through dedicated activities or educator observations, but MDRC explored whether information collected automatically during existing digital assessments could reduce the need for additional testing.
Khan Academy Kids was used to assess literacy and math through short tablet-based activities, while Kibeam used a handheld sensor-equipped device to record children’s responses to questions about physical books.
Researchers developed six potential executive functioning measures from the data.
Three focused on whether children responded unusually quickly or slowly. Two combined response speed with answer accuracy, while another tracked the proportion of assessment items children completed.
For most of the metadata measures, reliability reached or was consistent with the 0.6 threshold used by the researchers for exploratory work.
A smaller group of children also completed established executive functioning assessments in person, including Head-Toes-Knees-Shoulders Revised and Pencil Tap, allowing the team to compare the digital signals with more conventional measures.
The strongest results came from average response time on correctly answered questions and a two-vector measure combining speed and accuracy.
The two-vector score came with a complication. Because academic accuracy is built directly into the calculation, it was also strongly related to the literacy, language and math skills being assessed, with correlations between 0.7 and 0.9. That makes it harder to separate executive functioning from academic performance when using that measure.
Most of the other metadata measures had much weaker relationships with academic scores, generally below 0.3.
Better at spotting difficulties than separating stronger performers
The data also suggest the approach may be more useful at the lower end of the executive functioning range.
When researchers compared children’s positions on the metadata measures with their scores on direct assessments, the measures were better at differentiating children with lower executive functioning scores.
They were less precise when separating children with average and stronger skills.
For Khan Academy Kids, the researchers found signs of a ceiling effect, with average and high performers producing similar metadata scores.
That points toward a more limited potential use for the technology. Rather than generating a finely graded measure of executive functioning across every child, metadata could eventually help identify children who may need additional attention or more detailed assessment.
The research also found early indications that reliability was broadly similar when Khan Academy Kids results were compared by age and by program type. However, the samples were too small to conduct more extensive fairness testing, including whether the measures worked equally well for different groups of children.
Evidence remains preliminary
MDRC is clear that the measures are not ready for wider use. The study relied on relatively small samples, particularly when metadata scores were compared with established executive functioning assessments. The researchers say that limits the strength of conclusions about validity and fairness.
Both Khan Academy Kids and Kibeam were also still in development during the 2024 pilot and have since undergone substantial revisions.
Metadata itself can capture more than one behavior. A slower response could reflect attention, but children’s motivation, motor skills, language processing or the way an assessment is administered may also influence the result.
For example, the researchers note that dual-language learners may take longer to process an English-language question without that necessarily indicating weaker executive functioning.
Further work could examine additional signals already generated by digital assessments, such as changes in response time when children encounter a new type of question, the number of screen taps before choosing an answer or how often a child needs to be reminded what to do next.
For teachers, the researchers stop short of recommending automated executive functioning scores. They instead suggest that children’s behavior during technology-based assessments may already provide useful observational information, particularly when a child repeatedly struggles to remain engaged.
Larger studies will now be needed to establish whether the approach can be replicated across different children, assessment products and classroom settings before metadata-based executive functioning measures can move beyond exploratory use.