Three per-pitch machine learning grades anchor the site's command-and-stuff stack: Stuff+, DLoc+ v1, and Pitching+, with PLoc+ sitting alongside them as an MLB-only pitcher-season command projection, a pitcher-season deception model (2024+), and a season-level Arsenal Synergy grade. All grades use a 100-based scale where 100 is league average and higher is better. Per-pitch models score each pitch separately against left-handed and right-handed batters.
How nasty is the pitch? Grades based on physical characteristics: velocity, movement, spin, release point, extension, arm angle, tunnel deception, and trajectory.
XGBoost · 24 features
How well was the pitch actually located? DLoc+ is the descriptive command grade. It starts from pitch-type-by-count-by-matchup-side target surfaces, then adjusts those maps for batter chase tendency, walk acceptability by game situation, and promoted shape-routing before grading the pitch’s final location.
Surface-based target maps · Not a forward-looking command projection
What should this season’s command profile become next year? PLoc+ aggregates a pitcher’s season of location behavior and projects the following season’s DLoc+. It is meant to sit next to DLoc+, not replace it: descriptive command tells you what happened, while PLoc+ estimates what command quality should carry forward.
Pitcher-season projection · MLB only in v1 · Projects next-year DLoc+
The full picture. Grades pitch physics and plate location together in one overall pitch effectiveness grade, on one scale for starters and relievers.
Physics + location · MLB 2021+ and Triple-A
How much harder is the pitcher to hit than expected? Trained on bat tracking residuals (bat speed suppression, swing length shortening) to isolate deception independent of Stuff+/DLoc+/Pitching+. Measures delivery deception and timing disruption.
Ridge regression · ~15 features · r = −0.160 with FIP · Available 2024+
How well a pitcher’s pitch types work together as a unit, trained on the gap between actual CSW rate and the CSW predicted by the individual Stuff+ grades. Captures pairwise velocity and movement contrasts, tunnel consistency, and specialist vs. diverse arsenal dynamics: two elite pitches that tunnel well together score higher than three average pitches.
Season-level · RidgeCV · 26 features · r = 0.75 year-over-year stability · Independent of Stuff+ (r = −0.05)
Grade Scale
100 = league average. 110 = good. 120 = elite. 80 = below average. One standard deviation is roughly 10 points.
Signal Strength
Trust season grades (100+ pitches). Question monthly grades. Ignore single-game grades. As sample size shrinks, noise dominates.
DLoc+ vs Stuff+
Separate questions by design. A pitcher can have elite Stuff+ and average DLoc+ — that tells you the pitches are nasty but the executed locations still need work.
Run Value Sign
Raw xRV: negative = runs prevented = good for pitchers. The 100 scale flips this so higher is always better, regardless of the underlying metric.
DLoc+ (Descriptive Location+) surface grade for how well pitches were actually located.
Pitcher-season projection of next-year DLoc+ from this season's pitch-location behavior.
Physics-only model trained on residual xRV after DLoc+.
Combined model that merges command and stuff into one grade.
What drives each pitch type's grade, and how to read a 100-centered score.
DLoc+ measures descriptive command, while PLoc+ projects what that command signal should look like next year at the pitcher-season level. Stuff+ measures physical nastiness on residual xRV after command is removed, and Pitching+ brings descriptive command and stuff back together with sequencing and context.
All pitch grade models target xRV. Negative xRV means runs prevented (good for pitchers). Every score is scaled to a 100 mean so higher is always better.
The pipeline order: DLoc+ -> Stuff+ residuals -> Pitching+.
Why boosting works and how early stopping protects generalization.
A case study in underfitting and why one extra split matters.
What data powers the models and how features are built.