Inside the Team
All Episodes
Why 4.8 Survey Scores Don’t Mean Learning

Why 4.8 Survey Scores Don’t Mean Learning

0:00|0:00

Traditional smile-sheet surveys may look impressive, but they often tell you almost nothing about real learning or performance. This episode breaks down why the old 1-to-5 rating model fails and how a performance-focused approach can reveal whether people can actually apply what they learned on the job.


Chapter 1

The 0.09 Illusion: Why Your High Survey Scores Mean Nothing

Jordan Avery

So, Priya, I- I- I have to tell you about this- this kickoff meeting I was in last week. The, uh, the program lead walks in, literally beaming, right? They had just wrapped up this massive three-day leadership offsite, and they- they- they put up this slide with a giant, gold 4.8 out of 5. Everybody is high-fiving, talking about how the- the interactive Lego building game was a huge hit, and how the catering was "next-level" delicious.

Priya Nair

Oh, let me guess. The- the catering gets its own paragraph in the review.

Jordan Avery

Oh, absolutely! "Best artisanal sandwiches ever!" But, okay, fast-forward exactly three months. The 360-degree feedback data starts rolling in for these exact same managers. And guess what? Zero change. Like, literally, no measurable change in how they actually manage their teams. Not- not a single needle budged.

Priya Nair

It is the classic L&D tragedy. We- we mistake a great party for great learning. We look at those smiley-face surveys—what we call "smile sheets"—and we think, "Hey, they loved it, so it worked!" But the data on this is- is actually pretty brutal, Jordan. There's a famous meta-analysis—well, a couple of them actually, by researchers like Sitzmann and Alliger, looking at over 150 scientific studies—and they found that the correlation between traditional smile sheet ratings and actual learning results is... wait for it... 0.09.

Jordan Avery

Zero-point-zero-nine? That's... I mean, statistically, that is essentially zero. That's like saying a learner's rating of a class has about as much to do with what they learned as, I don't know, the weather outside that day.

Priya Nair

Exactly! It is uncorrelated. And the reason is those standard 1-to-5 Likert scales. You know the ones: "The trainer was engaging," or "I found this training valuable." Strongly Agree to Strongly Disagree. Those questions don't actually measure skill acquisition. They- they measure facilitator charisma. They measure comfort. It's- it's measuring learner compliance, basically. "Did you have a nice time in this air-conditioned room?"

Jordan Avery

Right, "Was the chair comfortable?" But- but wait, let me play devil's advocate here. I- I mean, surely, if a learner is miserable, if they hate the facilitator, they're not going to learn anything either, right? Like, a little bit of enjoyment has to matter for engagement.

Priya Nair

Well, sure. Enjoyment is- is a hygiene factor. It's like having a clean classroom. But treating "did you like it" as your primary success metric? That is a massive trap. Think of it like a surgeon. Would you evaluate a surgeon's competence solely by their bedside manner? "Well, she was incredibly friendly, so I'm sure my bypass surgery went great!" No! You- you want to know if the patient survived and if the artery is clear.

Jordan Avery

Okay, that is a terrifying but incredibly accurate analogy. We- we have to stop measuring the bedside manner and start measuring the surgery.

Chapter 2

The Performance-Focused Smile Sheet: Shifting from Liking to Doing

Jordan Avery

So, how do we actually do that? If the 1-to-5 scale is broken, how do we fix it? Well, there's this framework by Dr. Will Thalheimer called the Performance-Focused Smile Sheet, or PFSS. And instead of these vague numerical ratings, he says we should use highly descriptive, behavior-based multiple-choice options. You- you basically force the learner to think about what they can actually do on the job.

Priya Nair

Ooh, okay. Give me an example. What does that actually look like in practice?

Jordan Avery

So, Thalheimer suggests this question, which is basically the holy grail of post-training surveys: "In regard to the course topics taught, HOW ABLE ARE YOU to put what you've learned into practice on the job?" But here is the trick—you don't give them "Strongly Agree." Instead, you give them concrete, diagnostic options. For example, option A might be: "I have a general awareness, but I'll need more guidance to actually do this." Option B is: "I'm able to perform actual job tasks at a fully competent level." And option C could be: "I'm already an expert; I didn't need this training."

Priya Nair

Oh, I love that! Because as a learner, I actually have to stop and think about my own capability, rather than just clicking "5" down the page to get the survey over with.

Jordan Avery

Yes! Exactly. It forces a moment of honest self-assessment.

Priya Nair

But, okay, let's- let's do a quick reality check here. If I go to my VP of HR, and I say, "Hey, we're throwing out our 4.8 average dashboard metric," they might freak out. They like their neat little line graphs, Jordan. How do you pitch this to executives who are, let's be honest, kind of obsessed with those high average numbers?

Jordan Avery

Right, because a 4.8 looks great on an annual report, even if it means nothing. But- but here's how you reframe it. You go to that VP and say, "Look, instead of reporting a meaningless 4.7 out of 5, we can now show you that 85% of our team is fully competent to perform these actual job tasks on Monday morning, while 15% are going to need some follow-up coaching from their managers." Which of those two metrics actually helps them run the business?

Priya Nair

Wow. Yeah. "85% are ready to perform on Monday" is a- is a business metric. "4.7" is just... fluff.

Jordan Avery

Exactly. It changes the whole conversation from "did they like it" to "can they do it."

Priya Nair

Okay, so let's give the listeners something they can actually use tomorrow. A- a weekly experiment. I call it the "One-Question Rewrite."

Jordan Avery

Tell me more.

Priya Nair

For your very next training session, look at your survey. Find that standard, boring question that says "This course was useful" with the 1-to-5 scale. Delete it. Just- just delete it completely. And replace it with Thalheimer's "How Able Are You" multiple-choice question. Just that one change will start giving you predictive, actionable data instead of polite compliance. It's- it's a tiny shift, but the data you get back will be completely different.

Jordan Avery

I love that. Swap the fluff for real capability data. Alright, let's stop there. Go rewrite those surveys, folks. Talk soon, Priya.

Priya Nair

See ya, Jordan.