Cross-individual generalizability of machine learning models for ball speed prediction in baseball pitching
This study evaluates the cross-individual generalizability of machine learning models for predicting baseball pitch speed across 50 pitchers, revealing that while within-individual performance is high, generalization drops significantly and varies by expertise level and specific body segments, highlighting the need for cross-individual validation to improve practical applicability in sports science.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you are trying to teach a robot to guess how hard a baseball pitcher is throwing, just by watching their body move. You have a bunch of data from 50 different pitchers, ranging from high school kids to professional stars.
This paper is basically a report card on how well that robot works when it meets a new pitcher it has never seen before.
Here is the story of what they found, broken down into simple concepts:
1. The "Classroom" vs. The "Real World" Test
The researchers ran two types of tests to see how smart their AI model was:
- The Classroom Test (Within-Individual): They taught the robot about Pitcher A, then immediately tested it on Pitcher A again. The robot was a genius here, getting a score of 91% (R² = 0.91). It knew exactly how Pitcher A moved.
- The Real World Test (Cross-Individual): They taught the robot about Pitchers A through Z, but then tested it on a brand new pitcher it had never seen. Suddenly, the robot's score dropped to 38% (R² = 0.38).
The Analogy: Imagine a student who memorizes the answers to a specific practice test. If you give them that exact same test, they get an A. But if you give them a new test with the same types of questions but different numbers, they might struggle. The robot was good at memorizing specific people, but bad at understanding the general rules of pitching that apply to everyone.
2. The "Expert" vs. "Intermediate" Trap
The researchers wondered if the robot got confused because some pitchers were pros and others were amateurs. They split the group into "Experts" (pros) and "Intermediates" (college/industrial players).
- What they found: The robot didn't get the amount of error wrong for either group, but it got the direction wrong for the intermediates.
- The Metaphor: Think of the robot as a coach who assumes that if a player moves their body a certain way, they must be throwing the ball fast.
- Experts are like efficient athletes; they move efficiently and get great results.
- Intermediates might move their bodies in a similar way but lack that "secret sauce" of efficiency.
- The Result: The robot looked at an intermediate pitcher, saw the movement, and said, "Oh, that looks like a pro move! That must be a 90 mph pitch!" But the ball was actually only 80 mph. The robot overestimated the intermediate players because it didn't realize they were less efficient than the pros in its training data.
3. Which Body Parts Were the "Superheroes"?
The researchers tried to figure out which parts of the body helped the robot make better guesses. They covered up different body parts in the video data to see what happened.
- The Heroes: The Trunk (torso) and the Pivot Leg (the back leg that pushes off). Even when the robot only looked at the very beginning of the pitch (when the pitcher starts shifting their weight), the pivot leg gave it enough clues to make a decent guess (a score above 0.25).
- The Strugglers: The Arms (both throwing and leading arms). The robot couldn't guess well based on the arms alone.
- The Lesson: It seems that while every pitcher has their own unique arm style, the way they use their core and push off the ground is more similar across different people. The robot learned that "pushing off the ground" is a universal rule for speed, but "how you swing your arm" is too personal to generalize easily.
4. The Big Takeaway
The main point of this paper is a warning for sports scientists and coaches: Don't trust a machine learning model just because it works great on the athletes you trained it on.
If you build a model to predict ball speed using data from 50 specific players, it might fail miserably when you try to use it on a 51st player. The paper shows that to make these tools useful in the real world, we have to test them on new people, not just the same people over and over.
In short: The robot is good at memorizing individuals, but it's still learning how to understand the general "language" of human movement, especially when it comes to the tricky differences between a pro and a semi-pro. The trunk and the back leg are the most reliable translators in that language.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.