Follow me on Twitter!


Thursday, May 29, 2014

Regional Predictions, Week 4

Although last week's regional competitions had their share of drama, I'll admit it felt like a bit of a letdown after the amazing weekend prior. Thankfully, this fourth and final weekend looks like it has some fantastic stuff in store. I mean, how could Northern California's men's competition not be insane? We have seven former Games competitors vying for three spots, including three men who finished in the top 10 at the Games last year. Like we saw in the Central East, there will be some men not heading to the Games that probably would have gone in virtually any other region in the world.

With that in mind, let's get to some assorted topics before we move onto the predictions for week 4:
  • Taken in a vacuum, I don't have a problem with Dave Castro's statement (speaking for HQ I presume) that there will not be any wild cards given out this year. However, in context of this season, I'm not a fan.
    • Why even announce that wild card spots will be available (which they did earlier this year) if you are going to rule out that possibility before the Regionals have even finished? I cannot conceive of a scenario where wild cards would make more sense than they do for Sam Briggs this year. She is the reigning Fittest Woman on Earth, she had a single bad event in one of the most volatile events ever programmed at Regionals (1-attempt handstand walk) and she still finished fourth in a stacked region. If you're not going to use a wild card in that situation, then you're never going to use it.
    • Talent is so clearly bunched in a few regions (and has been for a few years). I can understand the argument that the regionals are set up with a limited number of spots in each region to increase drama and make things more exciting. However, I find it difficult to accept the argument that this system is ideal for finding the fittest athletes in the world. I get that cross-regional comparisons are not perfect, but I challenge anyone to argue that Graham Holmberg (4th in Central East) is not among the 40 best CrossFitters in the world. As it stands now, he is ranked ahead of the champions from 9 other regions. For Castro to argue that "the right athletes" are going to the Games seems a bit disingenuous. If you're just setting it up this way for drama, that's fine, but let's just call it what it is.
  • Although we have one week to go, the data from across all regions has allowed me to get a sneak peak at some interesting things from this year's regionals.
    • In terms of correlation with success across all Regional and Open events, it appears that events 3 and 7 are the top events at this point. I'll admit when I was wrong, and I was wrong on event 7. The top athletes are all crushing it, and it is damn exciting. Event 3 is a bit surprising, but again, look at the athletes who are doing well there, and they're usually dominating across the board.
    • On the other end of the spectrum, event 5 for the men actually has the lowest correlation with overall success. My guess here is that this is the one event this season that truly favors taller athletes, and so you are seeing some athletes with huge performances who otherwise are struggling. For the women, this event is not so bad, mainly because there are no athletes jumping 10-11 feet in the air and getting to the top of the rope in a couple pulls.
    • Not surprisingly, the two single-modality events (1 and 2) are among the least correlated with overall success for both men and women. Event 2 is slightly worse than event 1, but not by as much as you might think.
    • Events 4 and 6 are kind of middling in this respect. I expected event 6 to really bring out the top all-around athletes, but it might just be so grueling that it heavily favors the endurance specialists.
    • If we look at Open events in this context, 14.3 has the lowest correlation with overall success among Regional athletes (as it did for the entire Open field). On the other side, 14.4 was the highest correlation with overall success among Regional athletes (as it did for the entire Open field). In fact, it is basically neck-and-neck with Regional event 3 for the top spot across all events this season.
    • Some have suggested that results in the handstand walk might be correlated with success in event 4 (which has tons of handstand push-ups). It doesn't appear that way; ranks on those two events are not particularly correlated (52% for men, 44% for women - both of those figures are middle of the road this respect). The only combination of events that really stands out is events 1 and 7, which were 77% correlated for women and 68% correlated for men.
  • Last week I posted a chart and some statistics regarding the accuracy of my predictions (I should note that these are after removing athletes who withdrew prior to event 1). After week 3, the calibration plot looks about the same, but the mean-square error has dropped from 4.38% to 3.93%. For reference, last year's model was 4.43% and a model giving each athlete an equal chance would be about 6.40%. Below is the calibration plot (read last week's post for an explanation):


Alrighty... with all that out of the way. Let's get onto the predictions. This week, the only athlete for whom I made a manual adjustment to the model was Jason Khalipa. This year's events might not really favor him, but the guy has been so freaking consistent over the past 6 years that I felt he warranted special consideration.

With that said, here you go. Enjoy the final week of Regionals, everyone!

[Update 5/31: I've made a couple fixes to account for women's name changes since last year, as well as making the adjustment for Andrea Ager that I suggested in the comments a couple nights ago. I treated her as if she did not compete at Regionals last year, rather than as if she finished very low. Her low finish was due to a DQ in the OHS event, not due to a poor performance overall.]



Note that Africa only has one qualifying spot. All other regions this week have three.

Also note that the pictures look prettier this week because I'm posting from a Mac. Excel is terrible on a Mac, but at least it exports nicely to pictures.

Thursday, May 22, 2014

Regional Predictions, Week 3

I've said before that I have no doubt the CrossFit Games is a totally viable spectator sport. My opinion on the matter hasn't changed, but I'm beginning to think that it's the Regionals that should be on ESPN. And I doubt anyone who followed the coverage on the Games site this past weekend would disagree with me.

There was high drama all weekend, and with 10 different competitions to follow at various time zones across the globe, events were broadcast basically non-stop. We had the reigning fittest woman on Earth fighting just to make the Games in Europe, possibly the most competitive men's competition ever in Central East, a changing of the guard in Australia and the first real challenge to Camille and Michelle Letendre's dominance in Canada East. So before we move on to predictions for this week, let's get to some quick thoughts on week 2:
  • The handstand walk claimed another victim this week in Sam Briggs. I'm sure HQ would never admit that the programming was anything less than perfect, but anyone following that competition knows that Briggs was one of the three fittest women there, and likely the fittest. She dominated events 3-6 and placed decently on the two heavier workouts (1 and 7).
  • That being said, she will almost assuredly get a special invite, and the three women who did qualify certainly deserved it. Along with Southern California, Europe looks to be one of the two strongest women's regions in the world (Australia also looked sneaky-tough last weekend, too).
  • I certainly hope that HQ also extends a special invite to Graham Holmberg. We'll have to wait to see how the rest of the competitions play out, but I'd be he would have qualified in every other region (and likely won many of them). He came through when it counted most, smashing the event record in event 7, but unfortunately, four other men in his region also beat the event record (including third-place Will Moorad). He was also the only man in that region other than Rich Froning to take an outright first place, and he did it twice.
  • Super-impressive performance by Moorad to grab a spot in the Central East. As disappointed as I was to see Graham fall off, it is nice to see someone else break through in that region.
  • If the Games were programmed like the Regionals, Camille Leblanc-Bazinet might be the most dominant athlete in the sport. If the event has barbells and gymnastics, she's basically guaranteed a spot in the top 10 in the world. Hopefully she can fare a bit better with the more unorthodox Games events this year.
  • While I'm not convinced that event 7 is as good a test of fitness as, say event 4 or event 6, but I can't dispute that it is great for the viewers. Those 8 overhead squats have derailed more than one athlete's shot at the Games and it's given several others the chance to make up ground in dramatic fashion.
  • Event 6 looks like an absolute beast of a workout. Unlike prior years, I haven't been able to test out the workouts (due to a back injury), but I don't recall seeing Rich Froning that gassed in a workout since the rope climb/sled push workout at the 2012 Games. 
So how did my predictions do last week? Well, despite a couple of shockers, things actually pretty well. The charts below show how well the predictions were calibrated this year compared to last year. The athletes were bucketed based on their predicted rank, and each bucket was plotted as a poitn on the blue line.  The location on the x-axis represents my predicted chances of qualifying and the location on the y-axis represents the actual chances of qualifying.  The red line represents perfect predictions, so when the blue line is below the red line, my predictions over-estimated the chances of qualifying for those athletes.  As you can see, this year, the blue line tracks much more closely to the perfect predictions.




Of course, calibration only tells half the story. We could easily create a well-calibrated model by predicting an even chance for each athlete. If there are 30 athletes in a region with 3 qualifying spots, we could estimate that each athlete has a 10% chance of qualifying, and indeed, we would be right in some sense. It is sure that 10% of them will qualify. But we are also trying to be accurate. A perfect model would predict 100% chances for the 3 athletes that do qualify and 0% for all others. To measure our accuracy, we can calculate the means square error across all our athletes. In that respect, I did about as well as last year, and considerably better than the perfectly calibrated model with even predictions for each athlete.
  • 2014 Week 2, CFG Analysis Predictions - 4.38%
  • 2013 Weeks 3-4, CFG Analysis Predictions - 4.43%
  • 2013-2014 Equal Chance Predictions - 6.40%
So, with that in mind, I really didn't make any changes to the model this week. The only region where I had to deviate significantly was for the women in Asia. There are no returning Games qualifiers, hardly any returning regional competitors and just a handful of athletes in the top 2000 in the Open worldwide. My default model would have basically given all athletes the same chance, so I modifiied it as best I could, but the predictions are still pretty weak in that region.
[UPDATE 5/23/2014: In the Asia region, I just realized Candice Ford was Candice Howe last year, which meant I originally assumed she did not compete at the regionals last year. The predictions below are now fixed to account for this.]

Anyway, without further ado, here are my week 3 predictions. Enjoy the Regionals everyone!


Note that Asia only has one qualifying spot. All other regions this week have three.

Thursday, May 15, 2014

Regionals Predictions, Week 2

Welcome back, everyone. The first week of the 2014 Regionals is in the books, and in many ways, I think it played out like I expected. The handstand walk derailed one heavy favorite (Stacie Tovar), all of the athletes I picked to fare well did (Lucas Parker, Elizabeth Akinwale, Talayna Fortunato) and the weekend as a whole seemed to favor athletes with strong gymnastic abilities and pure strength, rather than those with the biggest engines. I was pleasantly surprised by event 7, which appeared to be more balanced than I expected and provided some exciting shake-ups in a few regions.

This week looks to be one of the most (if not the most) exciting weeks of the Regional schedule. We all know that the Central East men's region is the toughest in the world, but the European women's region and the Central East women's region are also super-competitive as well this year. It's also a little intriguing to see which "records" will fall now that the bar has been set on all the events. Personally, I think every single one of the men's records will go down (potentially all in the Central East?) and many of the women's records will go down (doubtful that Akinwale's records in events 1 and 2 will fall).

With that in mind, let's get down to the business at hand. Last week, I made a bunch of excuses for why I wasn't able to get formal predictions finished in time, but this week I was able to make it happen. I've been able to estimate the odds of qualifying for each athlete in all 10 competitions, and at the bottom of this post I've shown the odds for the top contenders in each region.

But first, here's a recap of the methodology, which is largely similar to what was done last year:
  • Learn from prior results
    • Separate the 2013 Regional competitors into various categories, based on their performance in the 2013 Open, 2012 Games and 2012 Regionals.
    • See how frequently athletes in each category posted a very high (top 20 worldwide) or relatively high (20-50 worldwide) regional performance in 2013 (based on the cross-regional rankings last year).
    • Repeat the first two steps one year further back. Combine the results with what I came up with in the first two steps. This helped me get a bigger sample size and hopefully improve the predictions.
  • Apply learnings to what we know this year to make predictions
    • For each athlete this year, place them in one of the categories based on their performance in the 2014 Open, 2013 Games and 2013 Regionals.
    • Depending on their category, randomly generate a worldwide ranking for each athlete this year. The category affects this randomized worldwide ranking, i.e. those who had better results in the past year will generally get a better randomized worldwide ranking this year.
    • Re-rank all athletes within a region based on these randomized worldwide rankings.
    • Repeat 200 times and see how often each athlete qualifies for the 2014 Games.
Now, for those so inclined, here are a few details on this process:
  • To get a large enough sample size to build this model, I combined men and women.
  • The process for creating the categories of competitors was not straightforward. There was quite a bit of judgment on my part to make sure that each category had sufficient athletes to be credible and that the categories produced results that made sense with each other. For instance, I wanted to separate out the top 2012 Games competitors (I chose top 15), but that meant I could not further break those athletes down based on 2012 regional rank, because there just would not be enough athletes there to get a credible sample.
  • The process for randomly generating the numbers is as follows:
    • Generate a uniform random number between 0 and 1 (=rand() in Excel). If the first is lower than the athlete's chance of finishing in the top 20, assign him or her to the top 20. If not, then if it is lower than the athlete's chances of finishing in the top 50, assign him or her to be between 20-50. Otherwise, the athlete is assigned to be between 50 and 100.
    • Once we have assigned the athlete to the a range of ranks, generate another uniform random number between 0 and 1. Multiply this by 20 to get the athlete's exact place within the range (multiply by 30 if they are in the 20-50 range or multiply by 50 if they are in the 50-100 range). Generally, you'll need to be in the top 50 worldwide to qualify, but depending on how other athletes fare, it's possible to end up in the 50-100 range and still be in the top 3.
  • Here is a lifting of the categories I used to break down the athletes:
    • Top 15 at prior Games
    • Below 15 at prior Games, top 40 worldwide at prior Regionals
    • Below 15 at prior Games, below 40 worldwide at prior Regionals
    • Did not make prior Games, top 50 worldwide at prior Regionals, top 100 in current Open
    • Did not make prior Games, top 50 worldwide at prior Regionals, below 100 in current Open
    • Did not make prior Games, 50-100 worldwide at prior Regionals
    • Did not make prior Games, below 100 worldwide at prior Regionals, top 250 in current Open
    • Did not make prior Games, below 100 worldwide at prior Regionals, below 250 in current Open
    • Did not make prior Games, did not compete at prior Regionals, top 75 in current Open
    • Did not make prior Games, did not compete at prior Regionals, 75-150 in current Open
    • Did not make prior Games, did not compete at prior Regionals, below 150 in current Open
Last year, my predictions weren't bad, but they generally overestimated the chances for the athletes on the low end and high end, but I underestimated the chances for the athletes in the middle. Consider:
  • Athletes predicted 0-10% - 2.5% expected to qualify, 0.4% qualified
  • Athletes predicted 10-50% - 20.7% expected to qualify, 39.5% qualified
  • Athletes predicted 50-100% - 66.3% expected to qualify, 57.1% qualified
I have more data to train the model this year, which should help to calibrate things a little better, and I was a bit more liberal in applying some manual adjustments to some elite athletes. For instance, Julie Foucher did not compete last year, but I treated her as if she had finished in the top 15 in the Games. As a three-time top 5 athlete, I think this is only fair. Other athletes for whom I made at least some adjustment included Rich Froning, Samantha Briggs, Annie Thorisdottir, Frederick Aegideus and Camille Leblanc-Bazinet.

So with that in mind, below are my predictions for week 2 (athletes with less than 5% chance are not shown). As always, keep in mind that this is all in fun, and it's all simply based on the numbers. I'm not making any sort of judgment about the effort these athletes have put in, I'm simply reflecting how athletes in similar situations have performed in the past. Enjoy week 2, everyone!


*Note that Canada East only has two qualifying spots. All other regions this week have three.

Thursday, May 8, 2014

Quick Regional Programming Thoughts

As much as I was hoping to have my regional predictions set up to go this week, I wasn't able to make that happen. A confluence of events over the last month - a 5-hour actuarial exam, a trip out of town to have my son baptized, my wife's birthday and numerous trips to the chiropractor/ART to try to fix some back issues that flared up after the Open - pretty much made that impossible for me. By next week, I do expect to be able to produce my stochastic regional predictions (like I did for the final two weeks of last year's season). But for this week, your guess is as good as mine about who will claim the first 24 spots at the Games.

That being said, I did have time to do a bit of work assessing the programming for this year's Regionals. Here are my thoughts:
  • Overall, I like the programming. In particular, I think events 3-6 seem to me to be well-balanced workouts that should be fun to watch.
  • Events 1 and 2 introduce a ton of volatility into the situation. Only getting 3 attempts on the hang snatch means we could see some top athletes get burned by taking a gamble on those 2nd and 3rd lifts. And a max handstand walk on a single attempt gives plenty of opportunity for a catastrophic failure that could cost an otherwise fit athlete a shot at the Games. In my opinion, I don't think this event really helps us find the fittest athletes, and it may end up preventing some really stellar athletes from making it.
  • Event 7 is really a wait-and-see event for me. It seems like a really weird design for a workout to have just 8 reps of the overhead squat, since it doesn't seem like it gives enough time for athletes to make up ground on that movement. But hopefully I'm wrong and this event turns out to be a better test of fitness than it appears to be on paper, especially considering it's the finale.
  • This year's regional is in some ways the heaviest regionals to date and in some ways the lightest. When weights are involved, the average relative load (1.55 men, 0.98 women) is the highest in the past four years. However, the programming is only 37% lifting, the smallest percentage of any Regionals. In fact, the 2011 Games is the only HQ competition with a lower percentage (34%).
  • When you combine those two factors, you get an load-based emphasis on lifting (LBEL) of 0.58 for men and 0.36 for women, both slightly lower than 2011 and 2013 and much lower than 2012. The chart below shows the progression of each of these metrics at the Regionals since 2011.

  • All-in-all, I believe this year's Regionals look more like the Games than in any past year. And what that means to me is that the emphasis is on strength, both in terms of weightlifting and very challenging bodyweight movements (like legless rope climbs or strict handstand push-ups). Don't get me wrong, you can't do well here without a high level of conditioning, but you will be punished much harder for lacking in strength.
  • Look for handstand push-ups and legless rope climbs to completely decimate the women's leaderboard. With the legless rope climbs, we saw how much variability there was at the Games last year. As far as handstand push-ups, keep in mind that before kipping became common (think 2011 and earlier), this was an extremely difficult movement for many top women, even at the Games.
  • With that in mind, here are some athletes that should do well this programming: Lucas Parker, Josh Bridges, Lacee Kovacs, Chris Spealler, Matthew Fraser, Elizabeth Akinwale, Talayna Fortunato, Annie Thorisdottir, Camille Leblanc-Bazinet.
  • Beyond those names, all the podium athletes from last year's Games should do fine. I'm just not sure they will fare any better because of this programming.
That's it for today. I hope you all enjoy the opening weekend of Regionals, and I'll see you again next week.

Friday, April 25, 2014

A Look Back at the 2014 Open: Part II

Note: For details on the dataset I am using, including which athletes are and are not included, please see the introduction to Part I. Thanks again to Andrew Havko for pulling the 2011 and 2014 data sets for me.

Welcome back. As I mentioned in Part I, the second part of my look back at the Open will deal more with the results of this year's Open. Much like in Part I, I'm also going to be putting things in perspective by making some comparisons to past years.

So let's start part II by looking at the event correlations, which is one way to judge how effective each workout is for measuring overall fitness (see "Are Certain Events 'Better' Than Others" from 2012 for more explanation). The charts below show, for the 2012-2014 men's* Opens, the correlation between each event and the sum of ranks for all other workouts in that season. Remember, correlations range from -100% to 100%, with 100% meaning that a higher rank on an event always indicates a higher rank across other events and 0% meaning that there was no relationship whatsoever between the rank on an event and the rank across other events.


Not surprisingly, event 14.4 was most correlated with overall success this year (both looking at the entire field and the top 1,000 only). We've seen in the past that events with more movements typically have higher correlations, although there have been several couplets with high correlations (13.4 for instance). Note that two of the weakest correlations were for 12.1 and 12.2, both single-modalities. We also notice that 14.3 had a low correlation, which probably doesn't come as a shock to most of us who follow the sport closely. That event basically boiled down to a heavy deadlift workout, and consequently we saw relative unknowns like Steven Platek end up in the top 10 worldwide.

For those unfamiliar with with what the concept is, below is a visual representation. The first chart shows the relationship between 14.3 rank and the sum of all other ranks, and the second charts shows the relationship between 14.4 rank and the sum of all other ranks. Notice how the bunching is tighter for 14.4 - there are fewer athletes who did well on 14.4 but did not do well otherwise (these would be dots in the top left) and fewer athletes who struggled on 14.4 but did well otherwise (these would be dots in the bottom right).



The correlations were generally lower this year than in 2013, although they were higher than in 2012. Lower correlations aren't necessarily "bad" - the fact that 2014 had lower correlations than 2013 is partly due to the fact that the workouts were more varied, which I personally liked. I think they struck a nice balance this year between not having any events that were too specialized (like 12.1 or 12.2), but not having the same athletes finishing at the top of each event (which occurred to some extent in 2013). I think this was reflected in the point totals needed to qualify for regionals each year: in the Central East for example, the 60th place competitor scored 589 points in 2014, compared with 450 in 2013 (with approximately 65% of the competitors of 2014) and 508 in 2012 (with approximately 30% of the competitors of 2014). More variety in the events means that athletes can afford more points and still reach regionals.

Let's move on to a comparison of performance between new athletes and continuing athletes. The charts below show the average percentile rank (0% being first place, 100% being last place) of athletes in each event, split between athletes who finished all 5 events in 2013 and those that either did not compete in 2013 or did not finish all events.


Like last year, the returning athletes did fare much better than the newcomers, and the difference was consistent across all events (the gap is nearly identical to last year). This shouldn't come as a surprise. In fact, if we extend this further and look at athletes who competed all the way back in 2011, we find that they finished in approximately the top 25% for men and 20% for women.

This leads us to probably the most interesting analysis I have in this post: a comparison of 11.1 and 14.1. Last year, when I looked at 13.3 vs. 12.4, I found that overall, the field performed nearly identically across the two years. However, when we isolated this to athletes who competed in both years, we saw a significant improvement in 13.3.

So what about 14.1 vs. 11.1? First, I compared the results across all athletes who finished all 5 events in either year. Interestingly, the average score in 2014 was approximately 10% lower for both men and women when we look at the entire field. This supports my belief that the Open didn't really become "inclusive" until 2012, when HQ put a lot more effort into convincing the community that everyone could and should participate in the Open. The 2011 Open was intimidating also: event 11.3 required male athletes to be able to squat clean 165 and female athletes to be able to squat clean 110, or else they would get a DNF.

Below is a graph showing the distribution of scores in 11.1 and 14.1 for all women who finished all events in either year. The x-axis is shown in terms of stations completed, not total reps, because this accounts for the fact that each round has twice as many double-unders as snatches. A score here of 10 stations is equal to 5 full rounds or 225 reps. Note that the 11.1 distribution is skewed to the right, indicating more athletes with high scores.


Like last year, however, this only tells part of the story. If we limit our analysis to athletes who finished all events in both years, we see the improvement we expected. Female athletes who competed in both years scored approximately 23% higher in 2014, male athletes who competed in both years scored approximately 14% higher, and both men and women averaged an impressive 283 reps in 2014. Below is a graph showing the distribution of scores in 11.1 and 14.1 for women who finished all 5 events in both years. Now you can see that the 14.1 distribution is skewed to the right.











Another thing I looked into was the percentage of athletes who finished the workout on the double-unders. I used this as a proxy to see how the community has improved on the double-unders. The idea is that for athletes who are competent on double-unders, that station will be much shorter than the snatches, thus we should see fewer athletes finishing on the double-unders if the athletes are stronger in that area. I found the following:
  • Men (entire field) -  2011: 50.7%, 2014: 47.4%
  • Women (entire field) - 2011: 50.0%, 2014: 49.9%
  • Men (competed both years) - 2011: 49.3%, 2014: 41.3%
  • Women (competed both years) - 2011: 47.8%, 2014: 41.2%
Although I was expecting to see the percentage drop more for the entire field, the numbers do support my belief that the community has become more proficient at double-unders in the past three years. This is particularly true for athletes who have been competing each year.

Finally, I took a look at the attrition from week-to-week this season. The chart below shows how the field declined each week over the past four seasons. To reduce clutter, I averaged the results for men and women each season.




Again, we see that things shifted pretty significantly after 2011. As I mentioned earlier, I believe that there was a much smaller percentage of "casual" Open participants in 2011 than we see today. The chart above supports that. Fewer athletes dropped off that season because the athletes who chose to sign up were generally more committed to competing for the long-haul.

Since then, the men's and women's field has finished up with between 59% and 64% of those that completed event 1. The percentage dropping off each week has varied between 4% and 19%, averaging out to approximately 11%**.

Well that's it for today. I've covered a lot in these past two posts, but at the same time I think there is plenty more work that can be done with the Open data, especially now that I have all four years of Open data to play around with and compare. That being said, I think it's shift the focus to Regionals, so I'll see you all again in a few weeks.

*The women's correlations are very similar.
**11% would probably be a good place to start for the attrition estimate needed to do the mid-week overall projections (described in recent posts).

Saturday, April 19, 2014

A Look Back at the 2014 Open: Part I

As I did last year around this time, today I'll be starting my review of the 2014 CrossFit Games Open. Of course there will be limitations to what this analysis will be able to cover, partly due to the data I'm able to get at this point (I am hoping to eventually pull down all the individual stats, like Fran time, max deadlift, etc.). Still, I think there is enough data out there to help us further our understanding of the current state of our sport and where we may be headed. Like last year, I'll be breaking this post up into two parts. To start, here is a list of topics I plan to cover, followed by a list of things I will not be touching on in this post:

Will cover:
  • Breakdown of the programming of this year's Open, much like my "What to Expect from the Open" posts from fall 2012 and fall 2013 (Part I)
  • Correlations between events this year, compared with last year (Part II)
  • Comparison of performance by new competitors vs. returning athletes (Part II)
  • Comparison of 11.1 and 14.1 results (Part II)
  • Attrition in this year's Open, compared with past years (Part II)
Will not cover:
  • Comparison between regions (don't have region information on the data at the moment)
  • Breakdown by age group (don't have age information, either)
  • Predictions for regionals (coming in the next few weeks)
  • Probably lots of other subjects that I simply didn't think of. If you have suggestions for future analysis, by all means, post to comments or email me.
Finally, here are some notes on the data set I am using for any work dealing with the results of the Open (thanks again to Andrew Havko for helping me pull this data, along with the 2011 Open results, which I've been wanting to get my hands on for some time):
  • Excluded any athletes who did not complete all 5 events. This simply makes for fairer comparisons. I did look at all scores in order to calculate the number who dropped off each week, but that is it.
  • Masters competitors (54 and under, since older groups are scaled) are lumped in with everyone else. As mentioned above, I don't have age information this dataset since I pulled it straight off the worldwide leaderboard.
  • I have re-ranked athletes on each event among the athletes in this dataset.
  • Athletes were identified as returning athletes if their full name was in last year's dataset. There are multiple athletes with the same exact name, but I had no way around this without region or age information. I assume any impact here is minor. The one manual fix I made was to make sure the Ben Smith at the top of the leaderboard was matched up with the correct Ben Smith from last year's data. 
So let's get started.

We'll start with the programming this year. I generally liked the programming this year (although it certainly didn't play to my strengths), mainly because HQ finally threw us some curveballs and some workouts that I didn't expect. Among the things we saw this year that hadn't occurred in previous Opens:

  • Rowing
  • A workout for time (rather than AMRAP)
  • A workout with more than 3 movements
  • Weights over 300 pounds for men and 200 pounds for women
  • Pull-ups and thrusters not in the same workout

Having said that, the Open is still the Open, and many things remained the same or similar as prior years. For instance, the loading was still much lower than the Regional and Games level. In fact, by my measurements, this was the lightest Open yet. Below is a basic comparison of the average loading* used each year in the men's competition (the pattern is the same for women).


The average relative weight was down from all prior years and there less than 50% lifting, which meant that the load-based emphasis on lifting (LBEL) was down about 10% from the historical average**. An investigation for another day is whether this Open favored smaller athletes because of that lower LBEL.

Now let's take a look at which movements have been used across the three years, and how they have been valued. This is presented slightly different than last year: each value represents how much that movement was worth as a percentage of the total for that season. The reason for presenting it this way (as opposed to counting total events) is that 2011 had six events and the other years had five, so this accounts for the fact that each event was not worth as much in 2011.


We see that this year, the programming hit almost every movement that has been used at any point in the past (push-ups and jerks were the exceptions) and added one new movement (rowing). Not surprisingly, we see the same movements being emphasized as in prior years: snatches, burpees, thrusters and pull-ups. One interesting note is that this is the first year in which no movement has accounted for more than 10% of the total points. However, the caveat here is that this methodology assumes all movements in a given workout are valued equally. In reality, there are instances where this is not necessarily true: for instance, most people would agree the deadlifts were valued far more than the box jumps in 14.3

As in past years, you'll also notice that the Olympic-style lifts and derivatives (thruster, overhead squat), as well as basic gymnastics movements, were the biggest keys to Open success. However, we did see a bit more value placed on other areas, such as powerlifting (e.g. deadlift) and pure conditioning (double-unders, rowing). Although Castro did surprise us with a few things this year, I still think it's a safe bet that the more advanced movements you might see at the Games (e.g. ring handstand push-ups, heavy medicine ball cleans) are not going to be tested in the Open. That's not to discount the usefulness of these other skills in training; it's just that you're not likely to see that tested until at least the regionals.

Finally, here's a chart I put together showing the relationship between loading, the number of movements, and the length of workout in the past three years of the Open. This chart was shown last year, but I have added the 2014 workouts, which are represented by the red balls. In the chart below, the x-axis represents the time domain, the y-axis represents the number of movements*** and the size of each bubble represents the LBEL of that particular workout (roughly how "heavy" was each workout). A plus-symbol indicates the weight varied during the workout and the arrows indicate the time varied for the workout****.


This year's workouts, although unique in how they were programmed, still didn't stray too far from what we've seen in the past in terms of loading and time domain. Keep in mind that for the above chart, I'm using averages for the variable-weight event (14.3) and variable-time events (14.2 and 14.5). For many CrossFitters, they did see a very long workout in 14.5 (which took many people beyond 20 minutes) and a very short workout in 14.2 (which only lasted 3-6 minutes for a majority of the field).

You can still observe some general trends here, which are often true of CrossFit programming in general. The shorter workouts tend to involve fewer movements and can occasionally go heavy, while longer workouts can potentially involve 3 or more movements but generally have light-to-moderate loading.

That's it for Part I. In Part II, which I expect to be out this week, I'll be focusing more on the results of this season's Open. See you soon.


*For background on these metrics, please see my post "What to Expect from the 2013 Open and Beyond." You may notice that the loading for prior years has changed slightly, which is due to me updating the relativities between lifts as I gather more data.
**For any workout with a variable element, such as the weight in 14.3 or the time in 14.2 or 14.5, I used the average of the top 1000 athletes. This is consistent with prior years.
***I considered the 11.3 a single-modality for this chart even though it technically included cleans and jerks.
****Workout 13.5 is actually hidden from view here because 14.3 is covering it up (same time domain and number of movements). Workout 13.5 was also a variable-time workout and would have had the arrows on the ball.

Tuesday, April 1, 2014

Can Mid-Week Projections Work?

Two weeks ago, I proposed a method to project an athlete's overall ranking before score submissions had closed for the week. To me, it made sense on paper, but it was admittedly untested. So I put out a request for help on testing it in week 4, and thanks to Andrew Havko (among others), I was able to make that happen.

So can it work? It appears that it can. That's not to say the projections are 100% accurate, and they are far from precise very early each week. But I think it's clear that the projections can give an athlete a good sense of where they would likely finish the week if they stick with their current score, which is something that is nearly impossible currently.

I tested these projections at three points during week 4: Friday 8 a.m., Saturday 5:30 p.m. and Sunday 3:30 a.m. (all EDT). The method requires one key assumption, which is the percentage of athletes who will drop off from the prior week, and for this I used 10%. Certainly this would need a bit more careful thought if it were to be implemented by HQ.

For each athlete, I projected their overall worldwide ranking at each of these times. For athletes whose score did not change by the end of the week, I compared my projection to their ultimate ranking. In total, the error of my projections were as follows:
  • Friday 8 a.m. (<1% of field reporting) - 9,575 mean absolute error*, 9,404 mean error
  • Saturday 5:30 p.m. (16% of field reporting) - 1,003 mean absolute error*, -787 mean error
  • Sunday 3:30 a.m (21% of field reporting) - 1,454 mean absolute error*, -1,362 mean error
Interestingly, the projections (at least using this first basic method) got slightly worse overall from Saturday to Sunday. The reason is that the distribution of scores submitted by Saturday 5:30 p.m. was more similar to the ultimate distribution than on Sunday. What I found was that, in general, the scores submitted very early on during the week are well above average, and the quality slowly declines throughout the week.  That is until Monday evening, when a slew of athletes replace their first score with a second improved submission. It turned out in this case that Saturday afternoon was a pretty accurate indication of how the current week's scores will turn out.

However, let's look a little more closely at the errors. Although an error of 1,003 (our best mean absolute error) is pretty small for an athlete finishing, say, 40,000th, it would be a very large error for an athlete finishing 2,000th. Thankfully, the size of the errors generally increased as the ranking increased. Below is a chart showing the percentage error for athletes across the spectrum of rankings, using our Saturday afternoon projections.


So you see that generally, we never really stray further than 3% error at any point. That's not too bad when you consider that there's currently no way to get even a good ballpark estimate until at least mid-day Monday.

Still, maybe we can do better. What if we had actually used the perfect assumption (8% in this case) for the percentage of athletes who would drop off from the prior week?  Well, in total, we improve for our Saturday and Sunday projections, with the mean absolute error going down to 338 for Saturday and 581 on Sunday. Interestingly, though, in this particular case it doesn't necessarily improve the projections across the board for Saturday and Sunday. Below is the same chart as above, but with the perfect assumption for attrition.


Although our error gets a little worse near the top, once we get near the middle of the pack, these projections are nearly spot-on. And even near the top, a 5% error isn't that bad - that's like these projections putting Josh Bridges at 100th overall, whereas he actually finishes 105th.

One way we can theoretically adjust to get even closer is to make an adjustment for the skill level of the atheletes who have submitted scores at a given point. This could involve looking at the average ranking of the athletes from their prior week's scores and comparing that to what we'd expect by week's end. The trouble is, it's challenging to know what the level will be at week's end. You might expect that the field would average out to be at the 50th percentile in prior weeks, but that wasn't actually the case here. The average athlete submitting a score for 14.4 was actually about the 48th percentile in prior weeks, which is due to the fact that the athletes dropping out after 14.3 were generally from the bottom of the pack.

My point is that while such an adjustment is possible, it might not be practical. And considering the projections even with my base 10% attrition assumption weren't too bad, I don't think further adjustments are necessary, beyond refining that attrition assumption to make it as accurate as we can.

Finally, while I think this method would produce reasonable results if implemented by HQ next year, there are some caveats about the testing done here:

  • I've only done testing for one week. There may be more (or less) error if we made these projections in week 2 or week 5.
  • I'm almost certain that the percentage error would increase a bit if we do this for each region. The sample size is much smaller, which means that even if the same principles apply, we're likely to see more variability. For one thing, it's going to take longer each week before the projections are even remotely meaningful, since many regions had less than 100 entries until late each Friday afternoon.
  • I only tested this for the men's field. I don't see any reason why the results would be much different for women, aside from the field being smaller, which would likely increase our percentage error a bit.
All that being said, I feel that implementing this method would provide a realistic glimpse into where an athlete will wind up. As long as athletes understand that this is merely an estimate, the information provided can be quite useful. 

Would this revolutionize the sport? Of course not. But I think it would be yet another improvement to the athlete experience as the largest stage of our sport continues to grow.


*Mean absolute error is the average of our errors, if we ignore the direction of the error. So if we are off by -500 for one athlete and +500 for another, the mean absolute error is 500 but the mean error is 0.